In the previous parts, we moved from vision and architecture to design principles, the implementation roadmap, cohort analytics, storage, and Vietnam's 2018 → 2026 journey. This final part stays short: a quick recap of how to build and operate the program, the optional extensions and future roadmap, and the questions ahead.
Series Overview
- Part 1: Vision and end-to-end architecture
- Part 2: Design principles and implementation roadmap
- Part 3: Variant calling, cohort analytics, and data organization
- Part 4: Infrastructure, talent, and government support: 2018 → 2026
- Part 5 (This Post): Summary, extensions, and the questions ahead
1. Summary: How to Build It, and What Comes Next
The stack, in one line: standardized variant calling (Nextflow + nf-core) → region-based joint genotyping (GLnexus/DPGT) → cohort analytics (Hail MatrixTable) → S3-compatible storage that keeps data in-country.
The shortest path from zero to 1,000 genomes:
- Phase 1 - Pilot (months 1-3): start with 100-500 samples on a small footprint — 1 SLURM login node + 4 compute nodes, ~50 TB S3, one Nextflow instance, a 4-node Hail cluster, and a core team of 2-3 FTE. Review every sample manually, establish QC baselines, and document standard operating procedures.
- Phase 2 - Scale to 1k (months 4-12): the same architecture grows to 20-30 compute nodes, 50-100 TB S3, and an 8-12 node Hail cluster. At 1k scale, GLnexus-based joint genotyping means no Spark cluster is needed. Output: 1,000 QC'd samples, a versioned cohort VCF, tested PLINK/BGEN exports, automated QC and monitoring, and a public allele-frequency database.
- Phase 3 - Operate and enable research (year 2+): a sustainable 2-3 FTE team supports GWAS queries, rare-disease diagnosis, clinical-lab partnerships, and feedback to participants.
What it takes: compute capacity worth roughly $150-300k — an on-premise HPC cluster, or the equivalent consumed as cloud services — plus $55-95k per year in running costs (team compensation follows the local labor market), a long-term data-residency contract with a domestic S3 provider, and clear governance for data access.
Why it matters: population-specific allele frequencies, rare-disease diagnosis with cohort context, GWAS powered by local variants, and precision medicine tailored to your population.
Extended work & future roadmap (optional): once the core program succeeds, these are the natural next steps:
- Multi-omics (years 2-4+): add proteomics, then transcriptomics, then metabolomics and imaging to move from variant lists to systems biology — disease-variant-protein links, biomarker panels, and precision medicine.
- EHR & phenotype linkage (year 2+): structured clinical data, then NLP-based phenotyping, then outcomes tracking — enabling rare-disease diagnosis and clinically meaningful GWAS. Roughly +$100-200k capital and +$30-50k annually.
- Interactive allele-frequency portal (~$5-20k): let researchers worldwide query your population frequencies and compare them against gnomAD — where local frequencies often change clinical interpretation.
- Disease-specific sub-cohorts: enrich rare-disease or cancer subsets within the 1k for higher-power discovery; mostly operational cost through clinical partnerships.
- Machine learning & polygenic risk scores (~$20-50k): build population-specific PRS, validate them against external cohorts, and deploy clinically where predictive.
None of these are required for a successful 1k-sample program. They are opportunities that open up when the core succeeds, stakeholders want to expand, funding allows, and international partnerships develop.
The technology is ready. The infrastructure is proven. The question now is: When will your country start building?
2. In 2026, what does a 1,000-genome project cost, and how long does it take?
The economics have changed decisively. By 2026, a 1,000-genome program is not a research moonshot but an industrialized service: whole-genome sequencing, secondary analysis, and cohort aggregation are sold as end-to-end contracts, and the same operating model has already been proven at far larger scales (UK Biobank, All of Us). A country no longer needs to build the entire value chain itself:
- Sequencing — contract it end to end. Illumina and comparable vendors offer end-to-end solutions covering sequencing, analysis, and data aggregation. At 1,000 samples, domestic providers such as Gene Solutions can adapt their existing services, keeping logistics, pricing, and support in-country. Public benchmarks in 2026 put research-grade 30x WGS — including library preparation and basic analysis — at roughly $220-650 per sample (about $221 on Ultima UG100, ~$650 on NovaSeq X Plus), while Illumina lists $200 per genome for flow-cell consumables alone.
- Data and platform — own them locally. The cohort can be stored with a domestic cloud provider (Viettel, FPT, VNPT, or CMC), while the software layer for storage, access control, and cohort management is built domestically. Object storage for 50-100 TB costs on the order of $1,000-3,000 per month, and the country retains ownership of both the data and the platform.
- Turnaround — fast once the pipeline exists. With raw data in hand and validated pipelines, assembling the managed cohort is a matter of days. Under favorable conditions — sufficient cloud compute and pre-validated workflows — completing a 1k run within a single day is not unrealistic.
The indicative first-year budget for 1,000 research-grade genomes is therefore on the order of $0.5M (roughly $0.4-0.8M depending on local pricing), dominated by sequencing and people rather than infrastructure:
- Sequencing and basic analysis: ~$0.3-0.6M
- Cloud storage and compute: ~$30-60k for the year (or a one-time $150-300k for an on-premise cluster, as described in Section 1)
- Core team (2-3 FTE): local salaries — the largest recurring cost line
This range excludes sample collection, consent, ethics review, and clinical phenotyping, which depend on the country's health system and study design. Compared with 2018 — when the same program was constrained by scarce expertise and infrastructure — the entry cost is several times lower, and the dominant investment is execution: a small team, a validated pipeline, and a storage contract. The timeline follows suit: a pilot of 100-500 samples in the first months, the full 1,000 within roughly a year.
3. How does the project stay alive and keep growing after the first 1,000 samples?
Delivering the first cohort is a milestone, not an endpoint. Programs that stall at 1k usually fail for the same reason: the data exists, but nothing sustains its use. Those that reach 50k or 500k tend to share three characteristics:
- Phenotypes and follow-up, not genotypes alone. At scale, value shifts from variants to outcomes — cancer incidence, disease progression, treatment response, and long-term follow-up. This is what allows a 50k or 500k cohort to answer clinical questions a 1k cohort cannot.
- Multi-omics beyond germline variant calling. Methylation, transcriptomics, proteomics, and metabolomics turn a variant catalogue into a systems-level resource, widening both its scientific and its commercial value.
- A revenue model that funds reinvestment. The most durable programs combine public and private support. Approved access to data and compute can be licensed for downstream analysis — precision medicine, disease screening, prognosis — with pharmaceutical and diagnostics partners providing the income that maintains and expands the platform. The governance and cost-recovery mechanisms described in Section 1 (secure analysis environment, quota-based allocation, cost recovery) are the backbone of this model.
This leads to the honest strategic question: is it worth investing in 50,000 samples and additional omics layers when buyers — and therefore success — are not guaranteed? The answer is that success is engineered, not assumed. The 1k program is the de-risking step: it proves the pipeline, the governance, and the first research results, and it builds the industry relationships. Expansion should be stage-gated: scale when demand and co-funding are demonstrated, not before, so that each increase in investment is backed by evidence rather than expectation.
4. What is the next step for VN1K?
VN1K has already delivered the hard part of the foundation: a pangenome-informed, multi-omics and phenomics resource built from 1,011 individuals across 53 provinces, with 42 million variants — 8.5 million of them absent from global databases. In his recent VietNamNet interview, Prof. Vu Ha Van set out what that foundation enables: preventive genetic testing, personalized drug selection, precision treatment, and even VinGenChip for identifying the remains of fallen soldiers. The same interview makes the next constraint explicit: much of the most valuable work — polygenic risk scores (PRS) and pharmacogenomics (PGx) — requires far more than 1,000 healthy genomes.
Why 1,000 healthy samples are not enough.
- PRS needs tens of thousands of participants with disease phenotypes and outcomes to train and validate risk models. A 1k healthy cohort can describe allele frequencies, but it cannot power disease-risk prediction — and scores derived in European populations transfer poorly to East Asian populations, which is exactly why Vietnamese-specific models matter.
- PGx needs genotypes linked to drug response and adverse events, plus enough carriers of actionable variants. Clinically important examples in East Asian populations —
HLA-B*15:02 for carbamazepine, CYP2C19 for clopidogrel, CYP2C9/VKORC1 for warfarin — require both scale and clinical annotation before they can guide prescriptions.
- Both need longitudinal phenotypes, EHR linkage, and consent frameworks that go beyond a one-time sampling study.
A staged path forward.
- Expand the cohort from 1,000 healthy individuals to at least 10,000-50,000, adding disease cohorts and the ethnic groups still underrepresented (sampling covered 53 of 63 provinces).
- Deepen the phenotype layer through hospital partnerships: longitudinal diagnoses, medications, lab values, and follow-up outcomes.
- Develop Vietnamese-specific PRS and PGx panels on the expanded cohort, validated in local clinical settings before deployment.
- Continue the multi-omics work — long-read sequencing, RNA, and methylation are already part of the VN1K resource — to interpret variants and prioritize actionable ones.
- Translate into products, as the interview describes: preventive testing kits, drug-selection guidance, and VinGenChip.
Why government funding fits. The foundational layers — cohort expansion, secure data and compute infrastructure, governance, and clinical validation — are public goods: their returns accrue to the health system over decades rather than to a single company, so market incentives alone will underfund them. This is where state investment is most effective — not as a subsidy for profit, but as funding for future healthcare, as with national biobanks elsewhere (UK Biobank, All of Us). At 2026 sequencing prices from the first question, expanding to 10,000 genomes is a roughly $3-6M sequencing commitment, with phenotyping and follow-up as additional costs; private partners and product revenue can sustain the downstream layers, as the previous question describes. The VN1K results — 8.5 million novel variants and domestic technology that could cost two to three times less than imported alternatives — are already evidence of the return on that investment.
A personal note of respect for the VN1K team. Building a population-scale, pangenome-informed resource from 1,011 individuals across 53 provinces — and identifying 8.5 million variants absent from global databases — took years of meticulous effort, and the milestone stands on its own. The road ahead is not easy: scaling the cohort, building the phenotype layer, and turning the data into PRS and PGx tools that reach patients will require sustained support. But the hardest step — from zero to a reference resource for the Vietnamese population — has already been taken, and it was taken by this team. That work deserves recognition, and the next phase deserves the investment to continue it.
If you are interested in my blog, feel free to discuss with me via:
References
Series repositories
- omicslab-hpc — Ansible automation to build the SLURM HPC cluster. https://github.com/vieomics/omicslab-hpc
- nf-germline-short-read-variant-calling — standardized short-read germline variant-calling pipeline. https://github.com/vieomics/nf-germline-short-read-variant-calling
- nf-modules — Nextflow modules and subworkflows shared across the series pipelines. https://github.com/vieomics/nf-modules
- omicslab-kit — gkit proof-of-concept implementations (GLnexus, DPGT, Spark-on-SLURM, Hail). https://github.com/vieomics/omicslab-kit
Foundational Projects & Research
- 1000 Genomes Project (2015). "A global reference for human genetic variation." Nature, 526(7571), 68-74. https://www.nature.com/articles/nature15393
- UK Biobank (2015). "UK Biobank: An Open Access Resource for Identifying the Causes of a Wide Range of Complex Diseases of Middle and Old Age" PLOS Medicine (now widely cited). https://doi.org/10.1371/journal.pmed.1001779
- All of Us Research Program. National Institutes of Health. https://allofus.nih.gov/
Relevant Genomic Biobanks Review
- Hyeji Lee et al. (2025). "Lessons from national biobank projects utilizing whole-genome sequencing for population-scale genomics." Genomics & Informatics. https://link.springer.com/article/10.1186/s44342-025-00040-9
Workflow Management
- Nextflow: https://www.nextflow.io/
- nf-core: https://nf-co.re/
- nf-modules ecosystem: https://github.com/vieomics/nf-modules
Variant Calling
- GATK (Genome Analysis Toolkit): https://gatk.broadinstitute.org/
- DeepVariant: https://github.com/google/deepvariant
- Our PoC: nf-germline-short-read-variant-calling pipeline, https://github.com/vieomics/nf-germline-short-read-variant-calling
Joint Genotyping
- GLnexus: https://github.com/dnanexus-rnd/GLnexus
- DPGT (Distributed Powerful Genotyping Tool): https://github.com/BGI-flexlab/DPGT
- Our PoC: gkit implementations (GLnexus, DPGT, Spark-on-SLURM, Hail), https://github.com/vieomics/omicslab-kit
Population Analytics
- Hail (Hail is the platform for population-scale genetics): https://hail.is/
- Spark (Distributed computing framework): https://spark.apache.org/
Storage & Infrastructure
- SLURM (Simple Linux Utility for Resource Management): https://slurm.schedmd.com/
- Our PoC: omicslab-hpc (Ansible automation for SLURM clusters), https://github.com/vieomics/omicslab-hpc
Standards & Interoperability
Genomic Data Standards
- GA4GH (Global Alliance for Genomics and Health): https://www.ga4gh.org/
- Beacon API for variant discovery
- Data Repository Service (DRS) for data access
- VCF format specifications
- VCF (Variant Call Format): https://samtools.github.io/hts-specs/
- gVCF (genomic VCF): Genomic variant call format retaining no-call information
Quality Control Standards
- bcftools: VCF manipulation tools, https://samtools.github.io/bcftools/
Variant Annotation & Interpretation
- VEP (Variant Effect Predictor): https://www.ensembl.org/info/docs/tools/vep/
- ClinVar: Curated variant-phenotype database, https://www.ncbi.nlm.nih.gov/clinvar/
- gnomAD: Population-scale variant database (55k genomes), https://gnomad.broadinstitute.org/
Privacy & Data Governance
- Data Protection Regulations: GDPR (EU), HIPAA (US), and regional equivalents
- GA4GH Data Use Ontology: https://www.ga4gh.org/post/data-use-ontology
Clinical Genomics & Disease Interpretation
- OMIM (Online Mendelian Inheritance in Man): https://omim.org/
- ClinGen (Clinical Genome Resource): https://clinicalgenome.org/ - Standards for variant interpretation
- Rare Disease & Undiagnosed Disease Networks: Examples from US, EU, Australia, etc.
Population Genetics & GWAS
- PLINK: Whole-genome association analysis toolset, https://www.cog-genomics.org/plink/
- GCTA (Genome-wide Complex Trait Analysis): https://cnsgenomics.com/software/gcta/
- Hardy-Weinberg Equilibrium testing standards
- Allele frequency calculations and stratification by ancestry
Vietnam & VN1K
- VN1K is a pangenome-informed multi-omics and phenomics resource for the Vietnamese population (2026). Nature Communications. https://www.nature.com/articles/s41467-026-75375-0
- Thanh Hung (2026). "Vietnamese genome project has multiple practical applications: Prof Vu Ha Van." VietNamNet. https://vietnamnet.vn/en/vietnamese-genome-project-has-multiple-practical-applications-prof-vu-ha-van-2541109.html
Cost References
- Illumina. "$200 genome on NovaSeq X Series." https://www.illumina.com/systems/sequencing-platforms/novaseq-x-plus/applications/broad-sequencing.html
- University of Minnesota Genomics Center. "Human Whole Genome Sequencing with Analysis." https://genomics.umn.edu/service/human-whole-genome-sequencing
This concludes the series on building a national 1000-genome project. Thank you for following along!