Study of 1,000 Vietnamese genomes reveals more than 42 million genetic variants

A landmark study decoding 1,000 Vietnamese genomes has uncovered not only millions of genetic variants unique to the Vietnamese population but also laid the groundwork for precision medicine, artificial intelligence in biomedicine, and new health technologies designed to improve public healthcare.

Dr Vo Sy Nam discusses genome sequencing technologies with colleagues.
Dr Vo Sy Nam discusses genome sequencing technologies with colleagues.

More than 42 million genetic variants and previously unrecorded data

The first findings from the Vietnamese genome mapping project titled "VN1K is a Pangenome-Informed Multi-Omics and Phenomics Resource for the Vietnamese Population" have been published in the international scientific journal Nature Communications, marking a significant milestone in Viet Nam's efforts to build and harness its own national genetic database.

The project was led by Dr Vo Sy Nam, Director of the Centre for Biomedical Informatics at Vingroup Big Data Institute, together with Professor Vu Ha Van, Scientific Director of the VinBigData Institute. The study introduces the first comprehensive large-scale Vietnamese genomic, multi-omics, and phenotypic dataset, while successfully establishing the Vietnamese PanGenome Reference (VPR) — the country's first population-specific reference genome. The resource is expected to provide a critical scientific foundation for genetic research as well as applications in precision medicine, preventive healthcare and genome-based medical services tailored to the Vietnamese population.

Speaking to Nhan Dan Newspaper, Professor Vu Ha Van, Scientific Director of the VinBigData Institute under Vingroup, said the VN1K project analysed genetic data from 1,011 Vietnamese individuals originating from 53 of the country's 63 provinces and centrally governed cities (based on the administrative boundaries in place at the time of sample collection). The study identified more than 42 million genetic variants, including approximately 8.5 million variants that have never before been recorded in major international genomic databases.

These newly identified variants are important not only from a scientific perspective but also as a foundation for future research into disease susceptibility, drug response, immunity, ageing, and metabolism among Vietnamese people.

Key researchers participating in the 1,000 Vietnamese Genomes Project.
Key researchers participating in the 1,000 Vietnamese Genomes Project.

Drawing on this dataset, the research team developed the Vietnamese PanGenome Reference (VPR) — Viet Nam's first large-scale population reference genome. Compared with reference genomes derived primarily from European populations, the VPR significantly improves the accuracy of analysing and interpreting genetic variants in Vietnamese individuals.

According to Dr Vo Sy Nam, Director of the Biomedical Informatics Centre at VinBigData and Co-founder and Chief Technology Officer of GeneStory, one of the study's most significant findings was the need to reassess the suitability of gene panels recommended by the American College of Medical Genetics and Genomics (ACMG) for hereditary disease screening when applied to the Vietnamese population.

The Vietnamese genome project was launched in 2018 by the VinBigData Institute (now part of VinUniversity) with the goal of creating a large-scale genomic database specifically for Vietnamese people while advancing the application of genetics in healthcare and broader public-interest initiatives.

The researchers concluded that genetic screening strategies should be tailored to the unique genetic characteristics of each population rather than directly adopting models developed for other ethnic groups. A dedicated Vietnamese genomic database will help to optimise screening panels, improve diagnostic accuracy, and reduce the risk of missed or inaccurate disease risk assessments.

Another major discovery relates to the field of epigenetics. Although Vietnamese people share many genetic similarities with other populations worldwide, the study found that their epigenetic markers exhibit distinctive characteristics.

This suggests that interactions between genes and environmental factors may regulate gene activity differently in the Vietnamese population. The finding opens an important avenue for research into the relationship between genetics, environmental influences, and disease risk, ultimately supporting the development of more accurate health prediction models tailored to specific populations.

The project also yielded important insights into the population genetics of Vietnamese people, showing considerable genetic similarity with several populations across East and Southeast Asia. These findings reflect a long history of migration and genetic exchange within the region. Rather than being genetically isolated, the Vietnamese population has continuously exchanged genetic material with neighbouring communities throughout its history.

Such data provide a clearer picture of the origins and evolution of the Vietnamese population while establishing an important foundation for research into contemporary health-related genetic characteristics.

In addition, the study identified several genetic variants associated with metabolism and disease risk that differ from those found in other populations. These include variants linked to lipid metabolism, blood glucose regulation, enzymes involved in drug metabolism, and susceptibility to certain diseases.

"For example, gene groups associated with low-density lipoprotein (LDL) cholesterol metabolism, HbA1c — a key biomarker used to assess diabetes risk and monitor glycaemic control — as well as several enzymes related to the risk of liver cancer all display characteristics that are distinct within the Vietnamese population," Dr Nam said.

Dr Vo Sy Nam.
Dr Vo Sy Nam.

According to Dr Vo Sy Nam, the study also identified variants associated with the PMS2P1 and GCNT2 genes that occur at significantly higher frequencies in the Vietnamese population, whereas they are exceedingly rare among European populations, with frequencies of less than 0.5%. This finding underscores the importance of developing genetic screening strategies tailored to the unique characteristics of individual populations.

However, Dr Nam emphasised that the newly identified genes and genetic variants require further in-depth investigation to determine their biological significance, their association with disease risk, and their potential clinical applications.

Bringing precision medicine closer to the community

After eight years of developing the VN1K database of 1,000 Vietnamese genomes, the project's ambition extends well beyond academic publications. Its ultimate goal is to translate scientific discoveries into practical healthcare solutions that directly benefit the public.

Using data generated through VN1K, researchers have identified numerous genetic variants associated with adverse reactions to commonly prescribed medicines, including carbamazepine, allopurinol, and clopidogrel. Based on these findings, they have developed pharmacogenomic prediction models covering around 150 classes of medicines, enabling clinicians to select more appropriate treatments and tailor dosages to each patient's genetic profile.

One of VN1K's most notable applications is VinGenChip, the first biochip designed specifically using genomic data from the Vietnamese population. VinGenChip is currently being deployed in the government's Martyrs' Gene Bank Project and the nationwide 500-Day Campaign to Locate the Remains of Fallen Soldiers — an initiative involving millions of DNA samples. By simultaneously analysing a large number of highly informative genetic variants, the technology improves the accuracy of matching DNA samples from relatives with unidentified human remains while significantly reducing costs compared with many imported alternatives.

According to Professor Vu Ha Van, one of the project's key priorities is to use the newly identified genetic variants to develop healthcare products specifically designed for Vietnamese people, particularly in preventive medicine and personalised treatment.

Professor Vu Ha Van.
Professor Vu Ha Van.

One of the most immediate applications is genetic testing designed to predict an individual's risk of developing diseases, allowing earlier adjustments to lifestyle, diet, and exercise.

"Many countries, including the US, the UK, and France, have already developed genetic tests that estimate disease risk. However, most of these products are based on genetic data from European or North American populations. When applied to Vietnamese individuals, their predictive accuracy may be limited because of genetic differences between populations," Professor Van explained.

Beyond disease-risk prediction, genomic data also offers significant opportunities in pharmacogenomics. Using genetic information to predict how patients respond to medicines enables clinicians to make more informed treatment decisions tailored to each individual.

Professor Van cited the 'Right Medicine for Children' initiative, previously implemented to support children with epilepsy in remote mountainous areas. Through genetic testing, doctors were able to identify which children were suitable candidates for specific medications and which should avoid them because of the risk of serious adverse reactions. The programme enabled thousands of disadvantaged children with epilepsy to undergo genetic testing before treatment, reducing the likelihood of severe drug reactions while improving treatment outcomes.

Genomic data also has applications beyond disease management. It can support the development of personalised healthcare products, including genetic tests for children. By understanding a child's unique genetic profile, parents can make more informed decisions regarding nutrition, healthcare, and long-term wellbeing.

Nevertheless, Dr Nam believes that the greatest challenge today no longer lies in generating genomic data or developing new technologies but in ensuring that research findings are translated into practical tools that reach those who need them.

He noted that preventive medicine has yet to become a widespread practice in Viet Nam. Most people still seek medical attention only after symptoms appear, while proactive genetic testing to assess disease risk remains a relatively unfamiliar concept.

Expanding genomic data, advancing AI, and strengthening precision medicine

According to VinBigData's researchers, the greatest value of the VN1K project lies not in any single scientific discovery but in establishing the first comprehensive reference genomic database for the Vietnamese population.

Dr Nam explained that the newly published database of 1,000 Vietnamese genomes represents only the beginning. Expanding the dataset to 10,000, 100,000, or even millions of genomes would dramatically increase its scientific value. Such a resource would enable researchers to gain much deeper insights into the genetic architecture of the Vietnamese population, supporting studies of disease, immunity, ageing, and numerous other biomedical fields.

Looking ahead, VinBigData scientists identify the development of next-generation pangenome reference datasets as a key priority.

Whereas conventional reference genomes have typically been built from only a handful of representative individuals, a pangenome integrates genomic data from many individuals, providing a far more comprehensive representation of the genetic diversity within a population.

"This will provide an essential foundation for improving the accuracy of biomedical research involving the Vietnamese population," Dr Nam said.

Researchers from the project review genome sequencing results in the laboratory.
Researchers from the project review genome sequencing results in the laboratory.

Another area expected to deliver significant advances is the application of artificial intelligence (AI) to genomic data analysis. AI is already being widely used in biomedical research, from predicting disease risk and assisting diagnosis to analysing treatment responses. However, the single most important factor determining the effectiveness of AI models is the quality of the data used to train them.

Until now, Viet Nam has largely relied on AI models developed using overseas datasets. However, such models do not necessarily reflect the genetic characteristics of the Vietnamese population with sufficient accuracy.

"If we want to build AI models that work well for Vietnamese people, we need data from Vietnamese people. It is much the same as developing large language models for Vietnamese — you need Vietnamese-language data," Dr Nam explained.

In addition to advancing AI, the Vietnamese genomic database will significantly strengthen population genetics research. Previously, many domestic studies focused on individual genes or relatively small regions of the genome. With access to whole-genome data, Vietnamese scientists can now analyse billions of genetic data points for the first time, enabling them to uncover previously unknown biological patterns.

Another major advantage is that the dataset can be utilised by multiple research teams simultaneously. Each group can pursue different scientific questions, ranging from infectious diseases and cancer to cardiovascular disorders and immunology.

According to Dr Nam, the project will continue to expand its international collaborations in the coming years. These partnerships, however, will be driven by complementary expertise and shared research objectives. Some international groups excel in fundamental research, while others possess cutting-edge technologies or advanced analytical methods. By combining those capabilities with Vietnamese genomic data, Viet Nam will be able to produce higher-value scientific research and develop more advanced biomedical applications.

In the longer term, one of the project's key objectives is to strengthen domestic research capacity, particularly by training a new generation of specialists in bioinformatics — an interdisciplinary field that integrates biology, medicine, mathematics, data science, and artificial intelligence. This talent pool will be essential not only for making full use of existing genomic resources but also for generating new scientific discoveries and driving the future development of biotechnology in Viet Nam.

After eight years of implementation, Professor Vu Ha Van said he hopes the Vietnamese genomic database will gradually be translated into personalised healthcare solutions, enabling people to take a more proactive approach to disease prevention, receive treatments better suited to their individual genetic profiles, and improve the overall efficiency of healthcare delivery.

"It is important for society to recognise the value of long-term projects like this. Many scientific achievements cannot be produced in a short period of time — they require sustained investment and perseverance," Professor Van emphasised.

As the Vietnamese genomic database continues to expand, Viet Nam will be better positioned to develop medical solutions tailored to the genetic characteristics of its own population, steadily building domestic expertise in genomic technologies while strengthening its role in the global biotechnology sector.

The study was conducted by a team of more than 40 Vietnamese scientists, led by Dr Vo Sy Nam and Professor Vu Ha Van, with contributions from Associate Professor Le Duc Hau (Ha Noi University of Science and Technology), Associate Professor Nguyen Thuy Duong (Viet Nam Academy of Science and Technology), Associate Professor Le Thi Ly (Viet Nam National University, Ho Chi Minh City), Professor Le Sy Vinh (Viet Nam National University, Ha Noi), Professor Nguyen Thanh Liem (VinUniversity), Professor Tran Huy Thinh (Ha Noi Medical University), Associate Professor Nguyen Hoang Quan (The University of Queensland, Australia), and Associate Professor Luu Nguyen Hung (Houston Methodist Research Institute and Cornell University, US), together with many other outstanding early-career Vietnamese researchers.

Back to top