Recent national genome initiatives highlight the growing need for population-specific genomic resources. After reading VN1K: a genome graph-based and function-driven multi-omics and phenomics resource for the Vietnamese population and EGP1K: Whole-Genome Sequencing of 1,024 Egyptians Characterizes Population Structure and Genetic Diversity, it becomes clear that building a national-scale 1000-genome project is increasingly important for understanding genetic diversity, improving disease research, and enabling precision medicine.
Special credit goes to the VN1K team at the Vingroup Big Data Institute and their collaborators. Building the first comprehensive resource for the Vietnamese population — 1,011 individuals, nearly 40 million variants with 8.5 million of them novel, and multi-omics layers including long-read methylation — came with enormous challenges, from sample collection and high-depth sequencing at scale to multi-omics integration and data accessibility. Their work paved the way for what this series describes.
In this series, I outline how a country can design and implement a project at a similar scale — from sample collection to cohort-level genomic analysis. This discussion focuses on the technical and architectural aspects of building the 1000-genome resource. It does not cover the downstream data platform or user portal for accessing processed data, which will be explored in a separate article.
The goal is to provide a practical, scalable roadmap that countries — especially those with emerging genomics infrastructure — can adapt to their own national genome initiatives.
Giang Nguyen — Founder @ Vieomics. During my time at DNAnexus, I worked on bioinformatics infrastructure and large-scale genomic data analysis, including workloads involving hundreds of thousands of samples. Here, I'll show how it should be done using open-source tools only — validated on real samples, with small datasets for proof of concepts.
The original 1000 Genomes Project (2008-2015) was a watershed moment in human genomics—it created the first comprehensive catalog of global genetic variation, providing a reference map that democratized variant calling and population genetics. But here's the catch: that data, while invaluable globally, doesn't fully represent your country's population, healthcare challenges, or genetic landscape.
Modern genome analysis relies heavily on population-specific variant frequencies. When you're doing variant interpretation, rare disease diagnosis, or pharmacogenomics, you need to know: How common is this variant in the population I'm treating? The 1000 Genomes Project surveyed ~2,500 individuals across diverse populations, but the genetic variation in your local population—shaped by unique migration histories, founder effects, and evolutionary pressures—may differ significantly.
Clinical implications are real:
Large-scale genomic projects generate not just data—they generate infrastructure, expertise, and economic value. Countries that build their own projects gain:
The good news: many countries have already built or are building large-scale genome projects. Here's what they've taught us:
These programs share common patterns: they all standardized variant calling, implemented joint genotyping strategies, built cohort-wide analytics layers, and invested heavily in data governance. The technical blueprint is proven—now it's about adapting it to your country's resources, population, and healthcare priorities.
Figure 1: Overview of genomic resources in the national biobanks. a Geographical distribution of the biobanks, with the sample numbers representing the total cohort size targeted or achieved by each biobank. Major biobanks possessing large-scale WGS datasets exceeding 10,000 individuals are highlighted. An asterisk (“*”) indicates the targeted cohort size. b Detailed information on WGS sample sizes, ancestry composition, and health conditions of the respective biobank datasets were recently disclosed
Reference: https://link.springer.com/article/10.1186/s44342-025-00040-9
While national and government-funded programs dominate the landscape, private companies and organizations have also built impressive large-scale genomic initiatives. Understanding their approach offers valuable lessons on infrastructure, scaling, and data governance.
Direct-to-Consumer (DTC) Companies
Pharmaceutical & Drug Development
Clinical Genomics & Diagnostic Companies
Global Infrastructure & Sequencing
Key Insight from Private Programs:
Private organizations succeed because they:
For national programs, the private sector model teaches us that data utility drives participation. People contribute samples when they see immediate clinical benefit (diagnosis, health insights) or participate when aligned with personal incentives (ancestry, precision medicine).
Building a 1000 Genomes-scale project requires:
But here's the key insight from existing programs: the software, the pipelines, and the infrastructure are mature enough to make this achievable for any well-resourced national program.
This series walks you through the architecture, design decisions, and implementation strategy for building a national genome project—from variant calling to cohort analytics.
A national genome project is fundamentally a data transformation pipeline: raw sequencing reads → standardized variants → cohort-wide genotypes → research-ready analytics. The architecture must handle massive scale, enforce consistency across thousands of samples, and satisfy strict data governance requirements.
Here's the architecture we'll build around, anchored in real, proven components:
Figure 2: End-to-end architecture showing data flow from raw sequencing through population analytics, with concrete tools and storage at each stage.
Why SLURM?
The foundation is a high-performance computing cluster managed by SLURM.
Figure 3: SLURM HPC cluster and Simple Object Storage S3 architecture
Key setup:
Every genome must go through identical variant calling logic. This is non-negotiable for cohort analyses.
The pipeline: nf-germline-short-read-variant-calling
Why gVCF? It includes not just called variants, but also "confident no-calls" at each genomic position. This is crucial for joint genotyping later—without it, you lose information about coverage depth and genotype quality.
Scale: Processing 100 samples/month at 30x coverage:
Figure 4: The nextflow pipeline for short read variant calling using Illumina platform, optimized for 30X WGS. The pipeline is integrated with nf-modules where it standardizes to share the common components between pipelines
After variant calling, you have thousands of individual gVCF files, each called independently. But independent calling is not enough for population-scale genomics. You must combine them into a cohort VCF using joint variant calling.
Figure 5: Joint genotyping reconciles per-sample calls. At chr1, Sample A's gVCF records an A>G variant while Sample B records no variant at all — which may reflect low coverage or a signal below the calling threshold rather than true homozygosity. Joint calling re-genotypes every site in every sample using the depth and quality evidence each gVCF carries, producing one cohort VCF with genotypes that are comparable across the whole cohort.
Four options at a glance
All four tools take per-sample gVCFs and produce a cohort call set, but they differ in method, license, and — critically — what happens when new samples arrive:
| Tool | Method | Adding new samples | Scale |
|---|---|---|---|
| GLnexus (Apache-2.0) | Joint genotyping | ❌ Full re-run | ~100k samples |
| DPGT (GPL-3.0) | Joint genotyping (Spark) | ❌ Full re-run (resumable) | Millions of samples |
| DRAGEN iGG — Illumina (proprietary) | Joint genotyping | ✅ Incremental batches | Population cohorts |
| Hail VDS combiner (MIT) | Combination only (no re-genotyping) | ✅ Incremental, restartable | 150k+ genomes (gnomAD) |
Only the first three tools re-genotype every site from gVCF evidence. The Hail VDS combiner merges call sets without revisiting per-sample genotypes, so it cannot improve variant calls the way joint genotyping does.
Option A: GLnexus (Efficient, Recommended for under 100k samples)
# GLnexus merges gVCFs into a cohort VCFglnexus_cli --config DeepVariant_h37 --bed <bed file> *.gvcf.gz > cohort.bcfbcftools view cohort.bcf -O z > cohort.vcf.gzOption B: DPGT with Spark (Open Source, Built for Scale)
#!/bin/bash# DPGT runner for joint genotyping cohort VCF
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"BUILD_LIB_PATH="${PROJECT_ROOT}/build/lib"DPGT_JAR="${PROJECT_ROOT}/DPGT/target/dpgt-1.3.2.0.jar"
# Defaults (override with env vars if needed)INPUT_LIST="${INPUT_LIST:-${PROJECT_ROOT}/cohort_vcf/1KGP/gvcf_input.list}"REFERENCE_FASTA="${REFERENCE_FASTA:-${PROJECT_ROOT}/reference/Homo_sapiens_assembly38.fasta}"OUTPUT_DIR="${OUTPUT_DIR:-${PROJECT_ROOT}/cohort_vcf/1KGP/results}"TARGET_REGION="${TARGET_REGION:-chr12:111760000-111763759}"JOBS="${JOBS:-4}"ALLOW_OVERWRITE="${ALLOW_OVERWRITE:-0}"
echo "================================"echo "DPGT Cohort VCF Runner"echo "================================"echo ""
# Check prerequisitesecho "Checking prerequisites..."if [ -z "$DPGT_JAR" ] || [ ! -f "$DPGT_JAR" ]; then echo "ERROR: DPGT JAR not found at $DPGT_JAR" echo "Run 'make build' to compile DPGT first" exit 1fi
if [ ! -f "$BUILD_LIB_PATH/libcdpgt.so" ]; then echo "ERROR: libcdpgt.so not found at $BUILD_LIB_PATH" echo "Run 'make build-cpp' to compile C++ libraries" exit 1fi
if [ ! -f "$INPUT_LIST" ]; then echo "ERROR: input list not found: $INPUT_LIST" echo "Create it with one gVCF path per line (3-sample trio supported)." echo "Example existing list: ${PROJECT_ROOT}/gvcf_input.list" exit 1fi
if [ ! -f "$REFERENCE_FASTA" ]; then echo "ERROR: reference fasta not found: $REFERENCE_FASTA" exit 1fi
if [ -d "$OUTPUT_DIR" ] && [ "$(ls -A "$OUTPUT_DIR" 2>/dev/null || true)" != "" ]; then if [ "$ALLOW_OVERWRITE" = "1" ]; then echo "Output exists. Removing: $OUTPUT_DIR" rm -rf "$OUTPUT_DIR" else echo "ERROR: output directory exists and is not empty: $OUTPUT_DIR" echo "Set ALLOW_OVERWRITE=1 or choose another OUTPUT_DIR" exit 1 fifi
echo "DPGT JAR: $DPGT_JAR"echo "C++ Library: $BUILD_LIB_PATH/libcdpgt.so"echo "Input List: $INPUT_LIST"echo "Reference: $REFERENCE_FASTA"echo "Output Dir: $OUTPUT_DIR"echo "Region: $TARGET_REGION"echo ""echo "Note: default region is a small smoke-test interval."echo ""
# Runtime environmentexport LD_LIBRARY_PATH="$BUILD_LIB_PATH:${LD_LIBRARY_PATH:-}"
echo "Running DPGT joint genotyping..."echo ""
# Some environments need explicit local filesystem implementations for Spark/Hadoopjava \ -Dspark.hadoop.fs.file.impl=org.apache.hadoop.fs.LocalFileSystem \ -Dspark.hadoop.fs.AbstractFileSystem.file.impl=org.apache.hadoop.fs.local.LocalFs \ -cp "$DPGT_JAR" \ org.bgi.flexlab.dpgt.jointcalling.JointCallingSpark \ -i "$INPUT_LIST" \ -r "$REFERENCE_FASTA" \ -o "$OUTPUT_DIR" \ -j "$JOBS" \ -l "$TARGET_REGION" \ --local
echo ""echo "Run complete."echo "Output files in: $OUTPUT_DIR"find "$OUTPUT_DIR" -maxdepth 1 -type f -name "result*.vcf.gz" -print || trueOption C: Illumina DRAGEN Iterative gVCF Genotyper (Proprietary)
Illumina's DRAGEN iterative gVCF Genotyper (iGG) attacks the re-run problem directly: it aggregates gVCFs batch by batch, so adding new samples means running only the new batch through the first step and re-aggregating the cohort census — not redoing the analysis end to end. The trade-offs: it is a licensed, closed-source feature metered per gigabase of input, and the input gVCFs must come from the DRAGEN ecosystem.
Option D: Hail VDS Combiner (Combination Only)
If you only need to combine call sets — not re-genotype them — Hail's VariantDatasetCombiner is the open-source choice: restartable and failure-tolerant, it can incrementally combine GVCFs and existing VDS datasets (gnomAD's 150,000 genomes were built this way). The catch: it combines datasets and preserves each sample's existing calls — it does not run a joint re-genotyping pass, so it cannot improve variant quality across samples the way GLnexus, DPGT, or iGG do.
Output: A single cohort.vcf.gz with all samples and all variants. For a ~1000-sample project, it can reach 200–500 GB (it depends on the diversity of variants and samples)
The raw cohort VCF isn't analysis-ready. It needs QC and must be converted to efficient formats. QC should be adjusted according to the quality of the ingested data
Using Hail's MatrixTable:
# Load cohort VCF into Hailmt = hl.import_vcf('cohort.vcf.gz')
# Sample-level QCmt = mt.filter_cols(hl.agg.count_where(mt.GT.is_non_ref()) > 100)
# Variant-level QCmt = mt.filter_rows(hl.agg.count_where(mt.GT.is_non_ref()) > 0)
# Export to multiple formats for different toolsmt.export('cohort.plink') # PLINK format for association studiesmt.export_bgen('cohort.bgen') # BGEN format (more efficient)mt.write('cohort.vds') # Hail VDS (best for Hail)Export formats for different use cases:
QC filtering removes:
Data Sovereignty Strategy: Rather than building on-premises storage infrastructure, partner with a domestic S3 provider (cloud or government-backed) for long-term data residency. This keeps data in-country while avoiding large capital investment in storage hardware.
Nowadays, the data should be accessed via the data platform where users need to:
Architecture:
The system architecture consists of:
Why this partnership model?
Long-term value: After the 1k-sample project completes, this S3 infrastructure remains available for future extensions (multi-omics storage, disease cohort data, long-term archival)
In Part 2, we cover the design principles — standardization, region-based parallelization, storage tiers, and federated governance — plus the five-phase implementation roadmap.
This is Part 1 of the series on building a national 1000-genome project. Continue to Part 2 for the design principles and roadmap.