Evaluation and test case datasets for AI
Held-out test sets, QA test cases and agent tasks with rubrics, written to measure what your model really does.
3 datasets shown
- On request
Software QA Test Cases
Manual test cases and bug reports written by professional QA engineers for real app flows: sign-up, search, checkout, payments, forms and settings.
- Feature
- Checkout: apply voucher
- Expected result· 3 steps
- Voucher is not applied
From $490 · 300+ test cases
- On request
Vietnamese LLM Evaluation
Private, never-published Vietnamese prompts covering reasoning, culture, writing, public administration and safety, each with a reference answer and a grading rubric by native experts.
- Prompt
- Giải thích sự khác nhau giữa Tết Nguyên Đán và Tết Trung Thu cho một người nước ngoài, trong khoảng 100 từ.
- Reference answer
- Tết Nguyên Đán là Tết cổ truyền mừng năm mới âm lịch, thường rơi vào cuối tháng 1 hoặc tháng 2 dương lịch; mọi người về quê sum họp, cúng tổ tiên, chúc Tết và lì xì. Tết Trung Thu diễn ra ngày rằm tháng Tám âm lịch, chủ yếu dành cho trẻ em: rước đèn, phá cỗ, ăn bánh trung thu và ngắm trăng.
From $590 · 200+ test cases
- On request
AI Agent Task Suite
Realistic multi-step tasks for browser, tool-using and coding agents, each with a start state, allowed tools, a human reference solution and an objective success check.
- Goal
- Find the cheapest 1 kg bag of Robusta beans that ships for free and add two bags to the cart.
- Success check· 5 steps
- cart has exactly 2 units of the lowest-priced 1 kg Robusta SKU with free shipping; no order created
From $790 · 100+ test cases