Addressing these barriers is essential to unlocking data resources, aiding the realisation of the objectives of the National Strategy on Artificial Intelligence to 2030, with a vision to 2045.
The National Strategy on Artificial Intelligence, approved by the Prime Minister under Decision No. 1671/QD-TTg (dated August 28, 2026), marks a shift from merely researching, developing, and applying AI towards comprehensive national transformation driven by AI research and development. Under this new approach, human resources, infrastructure, and data are identified as the foundations. The data and AI model pillar specifically calls for the development of data, data infrastructure, and models to serve AI research, development, training, testing, evaluation, and deployment.
This requirement translates into a series of tasks, including developing, opening, and sharing national databases, specialised databases, shared databases, and large, high-quality datasets; developing infrastructure and platforms for sharing data, products, and services, as well as data markets; and perfecting mechanisms for the opening, sharing, circulation, secure access to, and interoperability of data. Notably, the strategy sets out the task of refining regulations regarding property rights over data, intellectual property, and the lawful use of content for AI.
These tasks are directly relevant to the goal of developing Vietnamese-language AI models. By 2030, the strategy aims to create at least eight Vietnamese-language AI models of different scales, ranging from small models for edge devices and specialised tasks to very large models designed to address national-level foundational challenges and compete at the regional and international levels.
Building such models requires not only computing capacity and a team of experts but also a very large volume of high-quality and diverse data. However, in practice, automated data collection from the Internet and model training often involve copying works protected by copyright, including books, newspapers, images, videos, and software source code. Therefore, alongside the requirement to open up and exploit data, it is necessary to determine the scope of lawful use of protected content for training AI models.
According to Dr Hoang Minh Hieu, a full-time member of the National Assembly’s Committee on Law and Justice, current legislation lacks an obligation to proactively disclose information about training data. The Law on Artificial Intelligence prohibits the collection, processing, or use of data to develop, train, test, or operate AI systems in violation of laws on data, personal data protection, intellectual property, and cybersecurity; it also requires records to be retained and information provided upon request. These provisions provide a basis for inspection and enforcement but have not established a mechanism allowing rights holders to know whether their works have been used in training datasets.
Drawing on practical experience in development, the Legal AI Platform is recommended by the Ministry of Science and Technology as a priority for reviewing legal documents. Master Tran Van Tri, Director of the Viet Nam Law Media Joint Stock Company, said that the scope of data use for AI development needs to be clarified, particularly the distinction between publicly accessible data and data that may legally be copied, stored, and used for training. If this boundary is not clearly defined, it will be difficult to enforce authors’ and rights holders’ rights to reserve their copyright and related rights, as well as users’ obligations to pay royalties for data used for training.
Tri cited the input data source of the Law AI Platform as an example, which consists of legal normative documents and administrative documents that fall outside the scope of copyright protection. However, this does not mean that all products created from legal documents can be freely copied, as intellectual property rights may exist in the software, classification system, links concerning legal validity, amendment history, provision annotations, and editorial content surrounding those legal documents.
This highlights that determining the origin, legal status, and permissible scope of data exploitation is an important prerequisite before data is incorporated into AI training. Therefore, it is necessary to codify the obligation to ensure transparency regarding training data; AI model developers must retain information on data sources, the basis, and scope of use, and establish mechanisms to receive and address legitimate requests from rights holders.
The absence of transparency mechanisms will hinder the development of a data market serving AI, as AI developers will have little incentive to negotiate for usage rights if they can exploit data free of charge under the guise of research and technological development. At the same time, another risk is that Viet Nam’s digital archives, press, literature, and music could be automatically collected to train models hosted overseas without generating any revenue for authors or the country.
Experts believe that regulations should be supplemented to link compliance obligations concerning training data to the provision of AI products and services in the Vietnamese market, regardless of where the training takes place. Perfecting regulations in this direction would help prevent domestic businesses from being placed at a competitive disadvantage while also providing a tool to protect national data resources.
Given the absence of specific regulations on AI training-data transparency, lawyer Le Quang Vinh, Director of Bross and Partners Intellectual Property Company, advised businesses to proactively manage risks from the input data source. Accordingly, they need to clearly identify where data is collected from and retain information about the exploitation process to provide evidence should disputes arise.
Alongside improving the legal framework, Dr. Hoang Minh Hieu said the National Assembly should continue strengthening oversight of the implementation of policies and laws in this matter. The Government should also report to the National Assembly on impact assessments after a period of implementation, using specific quantitative indicators such as the number of disputes arising in connection with the exploitation of documents and input data; the number of registered declarations reserving copyright and related rights; the number of licensing transactions for training data that have been completed; and the average compliance costs incurred by domestic AI businesses.
Removing barriers to data exploitation is essential to realising the objectives of the National Strategy on Artificial Intelligence. When the rights and obligations of all parties are clearly defined, data can be shared and exploited safely and lawfully, thereby providing resources for research, training, and the development of Vietnamese AI models.