Skip to content
Dataset By Humans
On requestReasoning· AI reasoning

Code Debugging & Review Reasoning Dataset

Realistic bugs in Python, TypeScript, Java and SQL with the step-by-step reasoning a senior engineer used to find, explain and fix them, plus a test that proves the fix.

On request: this dataset is produced to your specification after you order. Typically 3–4 weeks for 200 records and 6–8 weeks for 1,000.

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
examples2001,000Any quantity
What's included
  • 200 records in JSONL
  • Bug type and difficulty labels
  • Executed tests
  • Datasheet, schema, license and checksums
  • 1,000 records in JSONL
  • Everything in Starter
  • Languages and bug types chosen by you
  • Code review comments
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$690$3.45 / example$2,490$2.49 / exampleQuote
Commercial license$1,990$9.95 / example$7,900$7.90 / exampleQuote
Exclusive license$6,900$34.50 / example$23,900$23.90 / exampleQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 3–4 weeks for 200 records and 6–8 weeks for 1,000.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Bug type and difficulty labels
  • Code review comments
  • Multi-file repository bugs
  • Your languages and frameworks
Request a custom quote

Overview

Coding assistants are good at writing new code and weaker at finding why existing code fails. Public bug datasets are mined from open-source commits: they show the diff, but not the reasoning that led to it, and many are already in model training data.

Each record in this dataset is written by a senior engineer: a realistic piece of buggy code, the symptom a user would report, the step-by-step reasoning to locate the cause, the fix and a test that fails before and passes after. Every record is executed to confirm it.

Example records

Illustrative records in the delivered format. Request the sample pack for real records written by our experts.

  • Record 1 of 1 · illustrative

    {
      "id": "cdr-000412",
      "test": "assert add_tag('a') == ['a']\nassert add_tag('b') == ['b']",
      "symptom": "Tags from earlier calls show up in new results.",
      "bug_type": "mutable-default-argument",
      "language": "python",
      "buggy_code": "def add_tag(tag, tags=[]):\n    tags.append(tag)\n    return tags",
      "difficulty": "easy",
      "fixed_code": "def add_tag(tag, tags=None):\n    if tags is None:\n        tags = []\n    tags.append(tag)\n    return tags",
      "root_cause": "Mutable default argument shared across calls.",
      "author_role": "senior backend engineer, 9 years",
      "reasoning_steps": [
        "The list only grows across calls, so state is shared between calls.",
        "The only shared object is the default value of `tags`.",
        "Python evaluates default arguments once, when the function is defined, so every call reuses the same list."
      ]
    }

Get a free sample pack

We email you real examples from this dataset.

Used only to send your samples. Privacy policy.

Technical specifications

Authors
Senior software engineers (5+ years), reviewed by a second engineer
Languages
Python, TypeScript/JavaScript, Java, Go, SQL; others on request
Bug types
Logic, off-by-one, concurrency, null handling, SQL joins, API misuse, performance, security
Verification
Every test fails on the buggy code and passes on the fix (executed in CI)
Fields
Buggy code, symptom, reasoning steps, root cause, fix, test, bug type, difficulty
Format
JSONL (one record per line), UTF-8; CSV on request
AI use
Written by people; screened for LLM-generated text. Any AI assistance is disclosed in the datasheet
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Coding model fine-tuning

    Teach models to diagnose before they patch, with explicit reasoning about the root cause.

  • Code review assistants

    Train reviewers that explain why a change is risky, not just flag it.

  • Uncontaminated evaluation

    Exclusive, never-published bugs give an honest measure of debugging ability.

  • Developer education

    Explanations written for humans make good material for teaching junior engineers.

How this data is made

  1. 1Problems sourced or written by vetted domain experts
  2. 2Each solution written step by step by a person, never generated
  3. 3A second expert checks every step and the final answer
  4. 4Disagreements resolved by a lead reviewer
  5. 5Screened for copied text and LLM-generated writing
  6. 6Difficulty, topic and skill labels added
  7. 7Exported as JSONL with a schema you can rely on
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Code is written for the dataset, not copied from employers or open-source projects, so licensing is clean.
  • Every record is executed: the test must fail before the fix and pass after.
  • Security bugs are limited to defensive examples (input validation, injection fixes), never working exploits.

Delivery format

One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.

Each delivery contains:

code-debugging-reasoning-v1/
├── data/train.jsonl
├── data/sample.jsonl
├── SCHEMA.md
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Is the code copied from GitHub?

No. Engineers write the code for the dataset, so it is not in public training data and carries no open-source license obligations.

How do you check the fixes?

Each record ships with a test. We run it in CI: it must fail on the buggy version and pass on the fixed one.

Can you cover our stack?

Yes. Tell us the languages, frameworks and bug categories you care about and we will quote the mix.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the examples produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets