AI Agent Task Suite: Multi-Step Tasks with Success Checks
Realistic multi-step tasks for browser, tool-using and coding agents, each with a start state, allowed tools, a human reference solution and an objective success check.
On request: this dataset is produced to your specification after you order. Typically 3–4 weeks for 100 tasks and 6–8 weeks for 500.
{
"id": "ats-000031",
"goal": "Find the cheapest 1 kg bag of Robusta beans that ships for free and add two bags to the cart.",
"skills": [
"search",
"filtering",
"comparison"
],
"constraints": [
"Do not place the order"
],
"environment": "web-browser",
"start_state": "Demo store, empty cart, signed in as test user",
"allowed_tools": [
"browser"
],
"success_check": {
"type": "state",
"assert": "cart has exactly 2 units of the lowest-priced 1 kg Robusta SKU with free shipping; no order created"
},
"human_time_seconds": 95,
"human_solution_steps": [
"Search 'robusta 1kg'",
"Filter: free shipping",
"Sort by price, low to high",
"Open the first result",
"Set quantity to 2 and add to cart"
]
}Pricing
Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.
| Feature | Starter | Pro | Custom |
|---|---|---|---|
| test cases | 100 | 500 | Any quantity |
| What's included |
|
| Your quantity, annotations and deadline |
| Datasheet, license, checksums | Included | Included | Included |
| Research license | $790$7.90 / test case | $2,990$5.98 / test case | Quote |
| Commercial license | $2,490$24.90 / test case | $8,900$17.80 / test case | Quote |
| Exclusive license | $7,900$79.00 / test case | $26,900$53.80 / test case | Quote |
How ordering works
- 1Choose a package and license and place your order. No payment is taken online.
- 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
- 3We produce and deliver in parts so you can review early. Typically 3–4 weeks for 100 tasks and 6–8 weeks for 500.
- 4Pay the balance and download the final dataset.
Custom quote
Need another quantity, an Enterprise license or annotations for this dataset? Available here:
- Human screen recordings of the solution
- Skill and difficulty labels
- Tasks for your own product
- Grading of your agent runs
Overview
Agents are judged by whether they finish real tasks, not by how fluent they sound. Building good agent tasks is slow: each needs a reproducible start state, a clear goal, and a success check that cannot be gamed.
Our task writers design tasks the way a QA lead designs acceptance tests. Each one is completed by a person first, timed, and paired with a machine-checkable success condition, so you can run thousands of agent attempts and score them automatically.
Example records
Illustrative records in the delivered format. Request the sample pack for real records written by our experts.
Record 1 of 1 · illustrative
{ "id": "ats-000031", "goal": "Find the cheapest 1 kg bag of Robusta beans that ships for free and add two bags to the cart.", "skills": [ "search", "filtering", "comparison" ], "constraints": [ "Do not place the order" ], "environment": "web-browser", "start_state": "Demo store, empty cart, signed in as test user", "allowed_tools": [ "browser" ], "success_check": { "type": "state", "assert": "cart has exactly 2 units of the lowest-priced 1 kg Robusta SKU with free shipping; no order created" }, "human_time_seconds": 95, "human_solution_steps": [ "Search 'robusta 1kg'", "Filter: free shipping", "Sort by price, low to high", "Open the first result", "Set quantity to 2 and add to cart" ] }
Get a free sample pack
We email you real test cases from this dataset.
Technical specifications
- Authors
- Task designers with QA or product backgrounds; each task solved by a second person
- Environments
- Web browser (demo sites we host or yours), REST APIs, spreadsheets, code repositories
- Fields
- Goal, start state, allowed tools, constraints, human solution steps, human time, success check
- Success checks
- State assertions, expected outputs or rubric, written to be run automatically
- Difficulty
- From 2-step lookups to 20+ step multi-app workflows
- Format
- JSONL (one record per line), UTF-8; CSV on request
- License
- Research, Commercial, Enterprise or Exclusive
Use cases
Agent benchmarking
Measure success rate, steps and cost per task for browser, desktop and API agents.
Training trajectories
Human reference solutions provide demonstrations for imitation learning and fine-tuning.
Product acceptance tests
Commission tasks for your own app to check that an agent can operate it before launch.
Failure analysis
Tasks are labelled by skill (search, form filling, comparison, multi-app) to locate weaknesses.
How this data is made
- 1Test plan agreed with you: skills, coverage and difficulty mix
- 2Cases written by experienced QA engineers or domain experts
- 3Expected results or grading rubrics written for every case
- 4Each case executed or answered once by a second person to confirm it is solvable
- 5Ambiguous or leaky cases rewritten or dropped
- 6Never published online, so they stay out of training data
- 7Exported as JSONL or CSV, ready for your eval harness
- 8Datasheet and license packaged with the delivery
Provenance & legal
- 100% made by people: no scraping, no generative AI
- Datasheet documenting how the data was made, checked and its limitations
- Commercial license that lets you keep models trained on the data
- Supports training-data documentation under the EU AI Act
Dataset-specific notes
- Tasks run against demo environments or sandboxes, never against live accounts or real payments.
- Each task records how long a person needed, to compare agent and human efficiency.
- Tasks that turn out to be ambiguous during the second person's run are rewritten or dropped.
Delivery format
One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.
Each delivery contains:
ai-agent-task-suite-v1/ ├── cases.jsonl ├── rubrics/ ├── SCHEMA.md ├── DATASHEET.md ├── LICENSE.pdf └── checksums.sha256
Frequently asked questions
Do you provide the environment too?
For web tasks we can host demo sites with reset scripts, or write tasks against your staging environment. API and repository tasks ship with setup scripts.
How are tasks scored?
Each task has a success check: a state assertion, an expected output or a short rubric. Most can be run without a human.
Can tasks be written in Vietnamese?
Yes. Goals and content can be in Vietnamese, English or both.
How does ordering and payment work?
Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.
What is the difference between non-exclusive and Exclusive?
With Research or Commercial we may license the test cases produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.
Can I order a different quantity or extra annotations?
Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.
Need a variation?
Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.
Request a quoteRelated datasets
Browse all datasets- Software QA Test CasesEvaluation & testsManual test cases and bug reports written by professional QA engineers for real app flows: sign-up, search, checkout, payments, forms and settings.On requestFrom $490
- Professional Reasoning TracesReasoningHow experienced accountants, supply-chain planners, marketers and engineers think through real work decisions, written out step by step by the professionals themselves.On requestFrom $790
- Code Debugging ReasoningReasoningRealistic bugs in Python, TypeScript, Java and SQL with the step-by-step reasoning a senior engineer used to find, explain and fix them, plus a test that proves the fix.On requestFrom $690
- Vietnamese LLM EvaluationEvaluation & testsPrivate, never-published Vietnamese prompts covering reasoning, culture, writing, public administration and safety, each with a reference answer and a grading rubric by native experts.On requestFrom $590