In Part 1, we covered why countries need their own genome project and the end-to-end architecture from raw sequencing to cohort VCFs. This part covers the design principles that keep everything standardized, scalable, and governed — plus a realistic five-phase implementation roadmap.
Part 2 (This Post): Design principles and implementation roadmap
Part 3: Variant calling, cohort analytics, and data organization
Part 4: Infrastructure, talent, and government support: 2018 → 2026
Part 5: Summary, extensions, and the questions ahead
1. Design Principles
1.1. Standardization as the Foundation
Without CI/CD integration, "standardization" is just a claim, not a guarantee.
No CI/CD = manual testing = inconsistent results across developers/environments
With CI/CD = automated testing = identical results every time, on HPC or Cloud
Every sample must go through identical processing—this is non-negotiable for population-scale analyses.
Why standardization matters:
Reproducibility: Same input data + same pipeline = same results every time
Quality assurance: Systematic issues caught early and applied uniformly across all samples
Cross-cohort comparisons: Different batches and timepoints can be compared without systematic bias
Joint analyses: Cohort-wide statistics are only valid when all samples processed identically
Implementation:
Use Nextflow for pipeline versioning and reproducibility
gVCF format enforces standardized output across all samples
Version control (Github, Gitlab, Gitea, etc): Track pipeline versions, reference genome versions, annotation database versions
Figure 1: CI/CD for nf-modules, showing how common patterns (modules, subworkflows) stay identical, reproducible, and reusable across multiple pipelines.
Figure 2: CI/CD for nf-germline-short-read-variant-calling — identical inputs produce identical outputs, and the workflows can be deployed on HPC or cloud.
1.2. Scalability Through Region-Based Parallelization
Joint variant calling works on the entire cohort simultaneously, not sample-by-sample. The key to scaling is splitting the genome into smaller, independent regions (windows) and processing all samples together within each region in parallel.
The Parallelization Strategy:
All samples in each region: For region 1, call variants across ALL samples simultaneously; for region 2, call variants across ALL samples; etc.
Region-based parallelization: Each genomic region runs on a separate SLURM job, enabling horizontal scaling
Failure isolation: If one region fails, you only re-run that region—not the entire genome for all samples
Bed File Strategy (Chromosome-Based with Intervals):
For small-to-medium cohorts (under 10k samples): Use simple chromosome-based BED files (one per chromosome):
chr1.bed covering chr1
chr2.bed covering chr2
... and so on for all 22 autosomes
chrX.bed and chrY.bed for sex chromosomes
For large cohorts (>10k samples): Split each chromosome into 100kb intervals for finer-grained parallelization:
chr1_region_1.bed: chr1
chr1_region_2.bed: chr1
chr1_region_3.bed: chr1
... continuing across all chromosomes with thousands of regions
Why this interval-based approach scales:
Memory efficiency: Each region holds only ~100kb of genomic data per sample, not entire chromosomes
Job parallelization: 22 chromosomes × 24,000 intervals = 528,000 potential SLURM jobs (realistic for 50k+ sample cohorts)
Failure recovery: Failed region only requires re-running 100kb interval, not entire chromosome
SLURM queue optimization: Small jobs (100kb regions) fit better into queue scheduling than large jobs
Example workflow with GLnexus (chromosome-based for <10k):
Terminal window
1
# For each chromosome in parallel:
2
glnexus_cli--configDeepVariant_h37\
3
--bedchr1.bed\
4
sample1.gvcf.gzsample2.gvcf.gz...sampleN.gvcf.gz\
5
> cohort_chr1.bcf
6
7
# For chr2, chr3, etc. (run in parallel across SLURM cluster)
8
9
# Merge all chromosomal BCF files
10
bcftoolsconcatcohort_chr*.bcf-Oz > cohort.vcf.gz
Example for large cohorts (interval-based for >10k):
For data storage, standardize on storing only the data that is needed:
Data should be stored and access should be granted with the appropriate permissions.
BAM/CRAM files should be kept in a form that can be restored to the original format (FASTQ), while saving storage and remaining usable as inputs for tools that extract more information from the data.
Hail MatrixTable, VDS, PLINK, and BGEN are pre-processed, high-quality data ready for analysis.
Users do not need to re-ingest data, repeating the pre-processing and quality-control steps.
Table 1: Pipelines and bioinformatics tools utilized in genomic resources in the national biobank
Compute-to-data: Analysis runs on the secure cluster; results exported, not raw data
Role-Based Access Control:
1
Raw gVCF data
2
├─ Accessible to: Original institution + data stewards
3
└─ Access: read-only, audit-logged
4
5
Cohort VCF (QC'd)
6
├─ Accessible to: Project researchers
7
└─ Access: for analysis with ethics approval
8
9
Public allele frequencies
10
├─ Accessible to: General public
11
└─ Download from website
This design principle enables nations to collaborate across institutions while maintaining strict data governance and participant privacy.
2. Infrastructure Layer & Implementation Roadmap
Building a national genome project requires careful sequencing of phases. This section outlines both the infrastructure requirements and a realistic implementation timeline.
2.1. Infrastructure Requirements Overview
A national genome project infrastructure stack consists of three core layers:
Layer
Component
Technology
Purpose
Compute
HPC Cluster
SLURM + compute nodes
Variant calling, joint genotyping, analytics
Storage
Object Storage
S3-compatible (on-premises)
gVCFs, VCFs, analytics formats
Analytics
Data Platform
Hail, Python, Spark
Population statistics, GWAS, QC
Table 3: Core infrastructure layers for national genome project
Figure 3: Realistic 6–12 month implementation roadmap showing all five phases. Key insight: Phase 2 (Sample Collection) is the bottleneck (3–6 months). Phase 3 (Sequencing) and Phase 4 (Variant Calling + QC) are rapid (1–2 months each). Infrastructure setup runs in parallel with sample collection.
Parallel: Order HPC hardware and establish S3 partnership contracts
Outcomes at End of Phase 1:
IRB/ethics approval obtained
Recruitment strategy finalized
Sequencing capacity confirmed (partner with sequencing center or in-house)
HPC hardware ordered (delivery 4–8 weeks)
S3 partnership contract signed
2.3. Phase 2: Sample Collection (3–6 months)
Goals:
Recruit and enroll 1,000 participants
Perform DNA extraction
Establish data governance protocols
Timeline:
Multi-hospital recruitment: 3–6 months (usually the slowest phase)
DNA extraction & quality control: Parallel to recruitment
Cohort metadata assembly: Concurrent
Key Challenge: Sample collection is typically the bottleneck. Plan for:
Hospital coordination across multiple sites
Participant compliance and scheduling
DNA quality control (ensure >95% pass-through)
Metadata standardization
Outcomes at End of Phase 2:
1,000 samples collected, extracted, and QC'd
Metadata database established
Samples ready for sequencing
2.4. Phase 3: Sequencing (1–2 months)
Figure 4: Sequencing procedure: Blood samples are collected from participants and genomic DNA is extracted using standard laboratory protocols. The DNA is then prepared for sequencing and processed on high-throughput platforms, generating raw sequencing data in FASTQ format. The raw data is uploaded to object storage such as Amazon S3 or an S3-compatible partner. The HPC system downloads the data from S3, performs alignment and variant calling, and uploads the processed outputs back to S3 for long-term storage and downstream population analysis.
2 GPU nodes with NVIDIA GPU (for DeepVariant acceleration)
Reduces per-sample variant calling from 4-6 hours (CPU-only) to 1-2 hours per sample
Optional optimization (not required for 1k-sample target)
Using Ansible for Infrastructure Automation:
Terminal window
1
# Deploy entire SLURM cluster from scratch
2
ansible-playbook-iinventory.inicluster_slurm.yml
3
4
# Validates:
5
-1.Allnodescansubmit/runjobs
6
-2.Jobschedulingpoliciesworking
7
-3.Networkconnectivitystable
Domestic S3 Storage Partnership (Initiated)
Industry-standard storage: Object storage such as Amazon S3, Google Cloud Storage, or Azure Blob Storage or Local compatible S3 provider is widely used across industries, ensuring scalability, compatibility, and long-term sustainability beyond bioinformatics use cases.
High availability requires dedicated expertise: Population genomics projects generate hundreds of TB to PB-scale data. Maintaining high availability, throughput, backup, and security requires a dedicated infrastructure team, which increases operational complexity.
Reduced operational cost and complexity: Partnering with an S3 provider shifts infrastructure management to experienced teams, reducing maintenance overhead and allowing the project to focus on analysis and scientific outcomes.
Using the external S3 provider is a feasible approach
Options:
Commercial domestic cloud provider (with local data residency)
Government-backed infrastructure (if available in your country)
Academic cloud partnership (university data center)
Key requirements:
Data never leaves the country (sovereignty requirement)
On-premise SLURM nodes have direct network access (low latency)
Contract negotiated for 5-10 year commitment
Scalable from this project while opening the possibility to extend the storage volume
Quality Control For Raw Data
After sequencing, the raw FASTQ data should undergo quality control to ensure it is suitable for downstream analysis. This step typically includes checking sequencing quality scores, read length distribution, GC content, adapter contamination, and duplication levels using tools such as FastQC and MultiQC.
Samples that do not meet quality thresholds may require re-sequencing or additional preprocessing such as adapter trimming and filtering. Performing quality control at this stage ensures that the data is reliable and ready for subsequent alignment and variant calling steps.
Outcomes at End of Phase 3:
All 1,000 samples sequenced (FASTQ files delivered)
FASTQ files transferred to HPC cluster and archived in S3
Ready for variant calling in Phase 4
2.5. Phase 4: Variant Calling + QC (1–2 months)
Goals:
Call variants on all 1,000 samples using DeepVariant
Perform joint genotyping with GLnexus
Complete cohort QC and finalization
Variant Calling Pipeline (Parallel to Sequencing):
Per-variant QC: allele frequency distribution, population structure
Rapid queries: filter by allele frequency, population, phenotype
Research enablement:
Export to PLINK for GWAS
Export to BGEN for imputation studies
Direct analysis via Hail Batch for distributed computing
Data Access & Governance:
Domestic S3 Storage (Long-term):
Established contract: 5-10 year commitment
Data residency: All data stays in-country
Pay-as-you-grow model: ~$0.02-0.05/GB/month
Direct SLURM access: low-latency queries
Data access policies:
Researchers access data via controlled analytics platform
Training required before data access
Audit trails for all data access
No raw data downloads (aggregate statistics only)
Outcomes at End of Phase 5:
Production-grade operations established
Research projects supported
Clinical integration pathways clear
Infrastructure reusable for future projects
2.7. Overall Timeline Summary
Total Project Duration: 6–12 months (realistic range accounting for sample collection bottleneck)
Phase
Activity
Duration
Bottleneck
Phase 1
Planning & Ethics
2–4 months
IRB approval
Phase 2
Sample Collection
3–6 months
Multi-hospital recruitment
Phase 3
Sequencing
1–2 months
Sample delivery speed
Phase 4
Variant Calling + QC
1–2 months
Computational throughput
Phase 5
Operations & Research Enablement
Year 1+
Ongoing
Table 4: Project timeline summary with phase durations and bottlenecks
Key Insight: The critical path is Phase 2 (Sample Collection), which typically takes 3-6 months. Phases 3 and 4 are rapid (1–2 months each). Infrastructure (Phase 1 planning and Phase 3 implementation) runs in parallel with sample collection.
2.8. Critical Success Factors
Technical:
Robust variant calling (validated benchmarking)
Reliable joint genotyping (tested thoroughly)
Automated QC (catches errors before publication)
CI/CD integration (ensures reproducibility)
Organizational:
Dedicated team (at least 2-3 full-time staff, stable)
Clear governance (who decides on data access, pipeline changes)
Institutional buy-in (hospital systems, university commitment)
Sustainable funding (multi-year commitment, not grant-dependent)
Strategic:
Population engagement (transparent about data use)
Research partnerships (early collaboration with universities)
Clinical integration (demonstrate direct patient benefit)
International standards (align with 1000 Genomes, GA4GH, etc.)
2.9. Risk Mitigation & Contingencies
Risk: Pipeline bugs causing systematic errors in 1k samples
Mitigation: Comprehensive CI/CD testing; region-based processing with checkpoints; ability to re-run any region
Risk: Data loss or corruption
Mitigation: Multi-site backup; immutable gVCFs archived; version control on all code/configs