In Part 3, we covered variant calling, cohort analytics, and the storage architecture. This part steps back from the technical stack to look at the bigger picture: Vietnam's journey from 2018, when VN1K was initiated, to 2026. Three things changed along the way — infrastructure, human resources, and government support — and together they determine whether a national 1000-genome project is feasible today.
VN1K in brief. Initiated in 2018 at the Bigdata Research Institute (now part of VinUniversity), VN1K is the first large-scale multi-omics and phenomics resource for the Vietnamese population: 1,011 individuals from 53 provinces and cities, more than 42 million genetic variants — about 8.5 million of them never recorded in major international databases — and the Vietnamese PanGenome Reference (VPR), the first population-specific reference genome for Vietnam. The work was published in Nature Communications in 2026 (VinUni announcement).
Eight years is a long time in genomics. When VN1K started in 2018, Vietnam had no population-scale genome reference of its own: genetic studies of Vietnamese people leaned on reference data built mostly from other ancestries, which is exactly the gap the project set out to close.
The environment around the project was just as challenging. Large-scale sequencing capacity was limited and often relied on partners abroad. Compute and storage for population-scale analysis were scarce — running joint genotyping across a thousand genomes, the way Part 1 describes, was out of reach for most teams. The number of people who could run a population-scale pipeline end to end could fit in a small room. And dedicated policy for genomic data and precision medicine was still at an early stage.
VN1K went ahead anyway. Between 2018 and publication, the team collected samples from 1,011 individuals across 53 provinces and cities and generated high-depth short-read whole-genome sequencing plus microarray, long-read whole-genome sequencing, RNA sequencing, and the first whole-genome methylation profile based on long reads. The VN1K data portal now exposes the resource to researchers.
In 2018, compute was the bottleneck. Most labs worked on small clusters or individual servers, data lived on local disks, and population-scale joint genotyping or cohort analytics required infrastructure that simply was not there.
The data center layer tells the same story in numbers:
| Data centers | 2018 | 2026 |
|---|---|---|
| Market | Grew ~12.7%/year from 2016, reaching ~US$858M in 2020 | Estimates vary by methodology: ~US$1.0–2.1B in 2025; capacity projected at ~950 MW by 2030 |
| Operators | Telco-grade facilities from Viettel, VNPT, FPT, and CMC | Viettel IDC, VNPT, FPT Telecom, and CMC Telecom lead by capacity; Hanoi Telecom and VNG operate as well; foreign players entering |
| Flagship projects | Viettel Hòa Lạc campus: ~800 racks when it opened (mid-2010s) | Viettel Hòa Lạc: 21,000 m², 30 MW, 2,400 racks; Viettel's 140 MW hyperscale campus under construction in Ho Chi Minh City; FPT AI Factory (NVIDIA GPU cloud) |
| National infrastructure | None | National Data Center No. 1 in Hòa Lạc; the Data Law (2024); data-localization rules |
| Pipeline | — | ~5.15M sq ft of upcoming white floor vs ~0.95M sq ft operational |
Today the market is anchored by four carrier-class operators — Viettel IDC, VNPT, FPT Telecom, and CMC Telecom — with Hanoi Telecom and VNG also operating facilities, and 33 commercial data centers in total. Foreign players such as Gaw Capital and Worldwide DC Solution (Singapore) have entered, and the announced pipeline (~5.15 million sq ft of white floor against ~0.95 million sq ft operational) is several times the current installed base. FPT's AI Factory adds the GPU layer: a US$200 million cloud with thousands of NVIDIA H100 GPUs, in service since January 2025 — capacity that matters for genomics, where tools like DeepVariant can run on GPUs.
For the genomics analysis stack, the picture is different too. High-throughput sequencing is accessible in-country, and the compute side can be built or rented locally: SLURM HPC clusters, GPU nodes for tools like DeepVariant, and S3-compatible object storage from domestic providers — the same architecture described in Part 2 and Part 3. The open-source stack this series proposes (Nextflow, nf-core, GLnexus/DPGT, Hail) runs on that infrastructure today; the proof-of-concept repositories referenced throughout the series are public and reproducible.
The point is not that infrastructure is free — it still takes capital and operational discipline. The point is that in 2018 the architecture was aspirational, and in 2026 it is deployable with local resources.
In 2018, expertise was concentrated in a small number of people, many of them abroad, and every project effectively had to train its own team.
By 2026, the picture is different. University programs, dedicated training initiatives, and the return of experienced diaspora have widened the pool considerably, and multi-institution collaboration has become the norm rather than the exception. The gap that remains is not headcount but experience at population scale — the kind only long-running projects can build.
The bottleneck is shifting. Finding people is no longer the hardest part; keeping them, organizing them across institutions, and giving them long-term projects is. That is a much better problem for a country to have.
In 2018, genomics in Vietnam was driven mainly by private-sector and university initiatives, and policy around genomic data was still taking shape.
By 2026, policy points in the same direction as the technology:
Genomics is also serving national missions directly: VinGenChip — the first biochip designed from Vietnamese genomic data, developed by GeneStory based on VN1K findings — is being used in the program to identify martyrs' remains, where millions of DNA samples are analyzed to match relatives with unidentified remains.
Put the three together and the conclusion is straightforward. In 2018, VN1K had to build its infrastructure, its team, and its methods from scratch. In 2026, a national project can stand on what VN1K and others have already built: the reference exists, the talent pool is wider, the tools are open source, and policy is aligned.
What remains is execution and governance — the themes of this series: standardized variant calling, scalable joint genotyping, a research-ready analytics layer, sovereign storage, and clear data access rules. Those are engineering and organizational problems, not scientific unknowns.
| Dimension | 2018 | 2026 |
|---|---|---|
| Infrastructure | Small clusters and local disks; population-scale analysis out of reach for most teams | SLURM HPC, GPUs, and domestic S3-compatible object storage; the open-source stack runs in-country |
| Human resources | A handful of specialists, many abroad; every project trained its own | Wider pool from university programs, training initiatives, and returning diaspora; multi-institution collaboration |
| Government support | Early stage; genomics driven by private and university initiatives | National healthcare and population program 2026–2035; High Technology Law 2025; ~US$3.8B of the 2026 budget for science, innovation, and digital transformation |
| Data and reference | No Vietnamese reference genome; foreign references | VPR, the first Vietnamese reference genome; 42M+ variants (~8.5M novel); imputation panel |
Key sources for this part — market figures vary between research firms because of methodology and scope:
In Part 5, we recap how to build the program, the optional extensions and future roadmap, and the questions ahead.
This is Part 4 of the series. Continue to Part 5 for the recap, the future roadmap, and the questions ahead.