Skip to content
Dataset By Humans
On requestEvaluation & tests· Evaluation

AI Agent Task Suite: Multi-Step Tasks with Success Checks

Realistic multi-step tasks for browser, tool-using and coding agents, each with a start state, allowed tools, a human reference solution and an objective success check.

On request: this dataset is produced to your specification after you order. Typically 3–4 weeks for 100 tasks and 6–8 weeks for 500.

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
test cases100500Any quantity
What's included
  • 100 tasks in JSONL
  • Success check and human solution per task
  • Datasheet, schema, license and checksums
  • 500 tasks in JSONL
  • Everything in Starter
  • Environments and skills chosen by you
  • Human screen recordings for 20% of tasks
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$790$7.90 / test case$2,990$5.98 / test caseQuote
Commercial license$2,490$24.90 / test case$8,900$17.80 / test caseQuote
Exclusive license$7,900$79.00 / test case$26,900$53.80 / test caseQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 3–4 weeks for 100 tasks and 6–8 weeks for 500.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Human screen recordings of the solution
  • Skill and difficulty labels
  • Tasks for your own product
  • Grading of your agent runs
Request a custom quote

Overview

Agents are judged by whether they finish real tasks, not by how fluent they sound. Building good agent tasks is slow: each needs a reproducible start state, a clear goal, and a success check that cannot be gamed.

Our task writers design tasks the way a QA lead designs acceptance tests. Each one is completed by a person first, timed, and paired with a machine-checkable success condition, so you can run thousands of agent attempts and score them automatically.

Example records

Illustrative records in the delivered format. Request the sample pack for real records written by our experts.

  • Record 1 of 1 · illustrative

    {
      "id": "ats-000031",
      "goal": "Find the cheapest 1 kg bag of Robusta beans that ships for free and add two bags to the cart.",
      "skills": [
        "search",
        "filtering",
        "comparison"
      ],
      "constraints": [
        "Do not place the order"
      ],
      "environment": "web-browser",
      "start_state": "Demo store, empty cart, signed in as test user",
      "allowed_tools": [
        "browser"
      ],
      "success_check": {
        "type": "state",
        "assert": "cart has exactly 2 units of the lowest-priced 1 kg Robusta SKU with free shipping; no order created"
      },
      "human_time_seconds": 95,
      "human_solution_steps": [
        "Search 'robusta 1kg'",
        "Filter: free shipping",
        "Sort by price, low to high",
        "Open the first result",
        "Set quantity to 2 and add to cart"
      ]
    }

Get a free sample pack

We email you real test cases from this dataset.

Used only to send your samples. Privacy policy.

Technical specifications

Authors
Task designers with QA or product backgrounds; each task solved by a second person
Environments
Web browser (demo sites we host or yours), REST APIs, spreadsheets, code repositories
Fields
Goal, start state, allowed tools, constraints, human solution steps, human time, success check
Success checks
State assertions, expected outputs or rubric, written to be run automatically
Difficulty
From 2-step lookups to 20+ step multi-app workflows
Format
JSONL (one record per line), UTF-8; CSV on request
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Agent benchmarking

    Measure success rate, steps and cost per task for browser, desktop and API agents.

  • Training trajectories

    Human reference solutions provide demonstrations for imitation learning and fine-tuning.

  • Product acceptance tests

    Commission tasks for your own app to check that an agent can operate it before launch.

  • Failure analysis

    Tasks are labelled by skill (search, form filling, comparison, multi-app) to locate weaknesses.

How this data is made

  1. 1Test plan agreed with you: skills, coverage and difficulty mix
  2. 2Cases written by experienced QA engineers or domain experts
  3. 3Expected results or grading rubrics written for every case
  4. 4Each case executed or answered once by a second person to confirm it is solvable
  5. 5Ambiguous or leaky cases rewritten or dropped
  6. 6Never published online, so they stay out of training data
  7. 7Exported as JSONL or CSV, ready for your eval harness
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Tasks run against demo environments or sandboxes, never against live accounts or real payments.
  • Each task records how long a person needed, to compare agent and human efficiency.
  • Tasks that turn out to be ambiguous during the second person's run are rewritten or dropped.

Delivery format

One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.

Each delivery contains:

ai-agent-task-suite-v1/
├── cases.jsonl
├── rubrics/
├── SCHEMA.md
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Do you provide the environment too?

For web tasks we can host demo sites with reset scripts, or write tasks against your staging environment. API and repository tasks ship with setup scripts.

How are tasks scored?

Each task has a success check: a state assertion, an expected output or a short rubric. Most can be run without a human.

Can tasks be written in Vietnamese?

Yes. Goals and content can be in Vietnamese, English or both.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the test cases produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets