Skip to content
Dataset By Humans
On requestEvaluation & tests· Evaluation

Software QA Test Case Dataset for Web & Mobile Apps

Manual test cases and bug reports written by professional QA engineers for real app flows: sign-up, search, checkout, payments, forms and settings.

On request: this dataset is produced to your specification after you order. Typically 2–3 weeks for 300 cases and 5–6 weeks for 1,500.

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
test cases3001,500Any quantity
What's included
  • 300 test cases in JSONL
  • Type and priority labels
  • Datasheet, schema, license and checksums
  • 1,500 test cases in JSONL
  • Everything in Starter
  • App types and flows chosen by you
  • Matching bug reports for 10% of cases
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$490$1.63 / test case$1,890$1.26 / test caseQuote
Commercial license$1,490$4.97 / test case$5,900$3.93 / test caseQuote
Exclusive license$4,900$16.33 / test case$18,900$12.60 / test caseQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 2–3 weeks for 300 cases and 5–6 weeks for 1,500.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Gherkin format
  • Matching bug reports
  • Traceability to user stories
  • Your own product specs (under NDA)
Request a custom quote

Overview

AI tools that write tests or run QA on apps learn from public repositories, where test cases are sparse, inconsistent and mostly unit tests. What real QA teams write every day, structured manual test cases with preconditions, steps and expected results, is almost absent from public data.

This dataset is written by professional QA engineers against common app flows in e-commerce, banking, delivery and SaaS apps. It covers positive, negative, boundary and accessibility cases, and can include matching bug reports with severity and reproduction steps.

Example records

Illustrative records in the delivered format. Request the sample pack for real records written by our experts.

  • Record 1 of 1 · illustrative

    {
      "id": "qtc-001203",
      "type": "negative",
      "steps": [
        "Open the cart and tap Checkout",
        "Enter voucher code SAVE50K",
        "Tap Apply"
      ],
      "feature": "Checkout: apply voucher",
      "app_type": "e-commerce mobile app",
      "priority": "high",
      "author_role": "QA engineer, 5 years",
      "preconditions": [
        "User is signed in",
        "Cart contains 1 item worth 150,000 VND",
        "Voucher SAVE50K requires a minimum order of 200,000 VND"
      ],
      "expected_result": [
        "Voucher is not applied",
        "Message explains the minimum order value (200,000 VND)",
        "Order total stays 150,000 VND plus shipping"
      ]
    }

Get a free sample pack

We email you real test cases from this dataset.

Used only to send your samples. Privacy policy.

Technical specifications

Authors
QA engineers with 3+ years of manual and automation testing
App types
E-commerce, banking and wallets, delivery, booking, SaaS dashboards (web, iOS, Android)
Case types
Functional, negative, boundary, UI, accessibility, localisation (Vietnamese/English)
Fields
Feature, preconditions, steps, test data, expected result, priority, type; bug reports with severity
Review
A second QA engineer executes or walks through every case
Format
JSONL (one record per line), UTF-8; CSV on request
AI use
Written by people; screened for LLM-generated text. Any AI assistance is disclosed in the datasheet
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Test generation models

    Fine-tune models that turn a feature description or user story into a complete set of test cases.

  • QA agent evaluation

    Measure whether an AI agent can execute test steps and judge pass/fail like a QA engineer.

  • Bug report triage

    Train classifiers for severity, component and duplicate detection on clean, consistent reports.

  • QA team tooling

    Bootstrap a test library for a new product with professionally written examples.

How this data is made

  1. 1Test plan agreed with you: skills, coverage and difficulty mix
  2. 2Cases written by experienced QA engineers or domain experts
  3. 3Expected results or grading rubrics written for every case
  4. 4Each case executed or answered once by a second person to confirm it is solvable
  5. 5Ambiguous or leaky cases rewritten or dropped
  6. 6Never published online, so they stay out of training data
  7. 7Exported as JSONL or CSV, ready for your eval harness
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Written against feature specifications we design or against public demo apps; never against a client's private product.
  • Gherkin (Given/When/Then) format is available on request.
  • Test data such as card numbers and phone numbers is fictional or uses official test values.

Delivery format

One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.

Each delivery contains:

software-qa-test-cases-v1/
├── cases.jsonl
├── rubrics/
├── SCHEMA.md
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Can you write test cases for our own app?

Yes, under NDA. The cases then belong to you under an Exclusive license and are never added to our catalog.

Do you provide automated test scripts?

The core dataset is manual test cases. Playwright or Appium scripts for a subset can be quoted on request.

What languages are the cases written in?

English by default; Vietnamese or bilingual on request, which is useful for localisation testing.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the test cases produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets