V-Bench launched to measure AI performance in Vietnamese
VOV.VN - Vietnam has introduced V-Bench, an open and non-profit benchmark designed to evaluate how well large language models (LLMs) understand and use Vietnamese, in a move aimed at establishing a local standard for AI assessment.
Developed by the AI Research Centre at VinUniversity (VinUni), V-Bench is intended to provide an independent reference framework for evaluating AI systems in Vietnamese, reducing reliance on international benchmarks that may not fully reflect the country’s linguistic and cultural context.
The benchmark comprises more than 40,000 questions and evaluation tasks, reviewed by a panel of 18 Vietnamese AI experts from both Vietnam and overseas. The project is operated on a non-profit basis and is freely available to researchers, developers and businesses.
Unlike many existing benchmarks that primarily measure academic knowledge, V-Bench is designed to assess a broader range of capabilities required for real-world applications in Vietnam.
Its evaluation framework covers cultural understanding, regional language variation, safety and responsible AI, as well as the ability to perform practical tasks in sectors such as education, healthcare and public services.
One of its distinguishing features is the inclusion of agentic AI evaluation, which measures a model’s ability to plan, reason and execute multi-step tasks rather than simply generate responses.
According to the development team, this broader approach is intended to assess how effectively AI models can operate in real-world Vietnamese environments, where language use is closely intertwined with cultural context, regional dialects and local knowledge.
V-Bench has been developed in line with international AI evaluation practices to ensure compatibility with global benchmarking efforts while addressing Vietnam-specific requirements. Initial benchmark results have already been released for 15 large language models.
The developers plan to expand the platform to evaluate multimodal AI systems capable of processing images, audio and video, as well as models designed to handle long-form Vietnamese documents and complex legal texts.
By establishing an evaluation framework rooted in Vietnam’s language and social context, V-Bench is expected to support the development of more capable and trustworthy AI systems and at the same time to strengthen the country’s role in AI research and innovation.