Skip to content
Dataset By Humans
On requestEvaluation & tests· Evaluation

Vietnamese LLM Evaluation Set with Reference Answers & Rubrics

Private, never-published Vietnamese prompts covering reasoning, culture, writing, public administration and safety, each with a reference answer and a grading rubric by native experts.

On request: this dataset is produced to your specification after you order. Typically 2–3 weeks for 200 prompts and 5–6 weeks for 1,000.

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
test cases2001,000Any quantity
What's included
  • 200 prompts in JSONL
  • Reference answer and rubric per prompt
  • Datasheet, schema, license and checksums
  • 1,000 prompts in JSONL
  • Everything in Starter
  • Categories weighted to your use case
  • One round of human grading of your model
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$590$2.95 / test case$2,290$2.29 / test caseQuote
Commercial license$1,790$8.95 / test case$6,900$6.90 / test caseQuote
Exclusive license$5,900$29.50 / test case$21,900$21.90 / test caseQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 2–3 weeks for 200 prompts and 5–6 weeks for 1,000.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Human grading of your model outputs
  • Pairwise preference judgements
  • Categories weighted to your use case
  • English parallel version
Request a custom quote

Overview

Public benchmarks leak into training data within months, and most multilingual benchmarks are machine-translated from English. The result is scores that say little about how a model behaves for real Vietnamese users.

This evaluation set is written in Vietnamese from the start by native experts: teachers, editors, lawyers and civil-service specialists. Every prompt comes with a reference answer and a points-based rubric, so human graders or an LLM judge can score answers consistently. Sets are delivered privately and never published.

Example records

Illustrative records in the delivered format. Request the sample pack for real records written by our experts.

  • Record 1 of 1 · illustrative

    {
      "id": "vle-000058",
      "prompt": "Giải thích sự khác nhau giữa Tết Nguyên Đán và Tết Trung Thu cho một người nước ngoài, trong khoảng 100 từ.",
      "rubric": [
        {
          "points": 2,
          "criterion": "Nêu đúng thời điểm theo âm lịch của cả hai dịp"
        },
        {
          "points": 2,
          "criterion": "Nêu ý nghĩa: năm mới, sum họp gia đình vs. lễ hội trăng rằm cho trẻ em"
        },
        {
          "points": 2,
          "criterion": "Có ít nhất hai phong tục đúng cho mỗi dịp"
        },
        {
          "points": 1,
          "criterion": "Văn phong dễ hiểu cho người nước ngoài, khoảng 100 từ"
        }
      ],
      "category": "cultural-knowledge",
      "difficulty": "easy",
      "max_points": 7,
      "reference_answer": "Tết Nguyên Đán là Tết cổ truyền mừng năm mới âm lịch, thường rơi vào cuối tháng 1 hoặc tháng 2 dương lịch; mọi người về quê sum họp, cúng tổ tiên, chúc Tết và lì xì. Tết Trung Thu diễn ra ngày rằm tháng Tám âm lịch, chủ yếu dành cho trẻ em: rước đèn, phá cỗ, ăn bánh trung thu và ngắm trăng."
    }

Get a free sample pack

We email you real test cases from this dataset.

Used only to send your samples. Privacy policy.

Technical specifications

Authors
Native Vietnamese experts: teachers, editors, lawyers, public-service specialists
Categories
Reasoning, cultural knowledge, writing and style, administration and law, everyday advice, safety
Fields
Prompt, category, difficulty, reference answer, rubric with points, notes for graders
Contamination
Never published; Exclusive sets are written for one buyer only
Review
A second expert answers each prompt blind to check the rubric
Format
JSONL (one record per line), UTF-8; CSV on request
AI use
Written by people; screened for LLM-generated text. Any AI assistance is disclosed in the datasheet
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Model selection

    Compare vendors and open models on the tasks your Vietnamese users actually ask about.

  • Regression testing

    Re-run the same held-out set after every fine-tune to catch quality drops.

  • LLM-as-judge calibration

    Rubrics and reference answers let you check whether an automatic judge agrees with humans.

  • Safety review

    Culturally specific safety prompts reveal failures that translated English sets miss.

How this data is made

  1. 1Test plan agreed with you: skills, coverage and difficulty mix
  2. 2Cases written by experienced QA engineers or domain experts
  3. 3Expected results or grading rubrics written for every case
  4. 4Each case executed or answered once by a second person to confirm it is solvable
  5. 5Ambiguous or leaky cases rewritten or dropped
  6. 6Never published online, so they stay out of training data
  7. 7Exported as JSONL or CSV, ready for your eval harness
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Prompts are original: not translated from English benchmarks and not taken from websites.
  • Facts that change over time (laws, prices, officials) carry a valid-as-of date.
  • Human grading of your model's outputs against the rubric is available as a service.

Delivery format

One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.

Each delivery contains:

vietnamese-llm-evaluation-v1/
├── cases.jsonl
├── rubrics/
├── SCHEMA.md
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Why buy an eval set instead of using public benchmarks?

Public sets end up in training data, so scores inflate without real improvement. A private set written for you stays a fair test.

Can you grade our model's answers?

Yes. Our experts can score your outputs against the rubric and report results per category.

Is an Exclusive set really never reused?

Yes. Exclusive prompts are written for your order only, are never licensed to anyone else and are not added to our catalog.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the test cases produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets