Skip to content
Dataset By Humans

How we make human data

What you buy is trust in where the data came from. Here is exactly how each photo, clip and written record is made, checked and documented.

Photos and video

  1. 1

    Capture

    Original smartphone photos taken by people in real places. No scraping, no stock, no generative AI. We only shoot where photography is permitted and ask permission in private spaces.

  2. 2

    De-duplicate

    Exact and near-duplicate images are removed so they do not inflate the dataset or leak across your train/test split.

  3. 3

    Filter

    Broken, corrupted and accidental shots are removed, unless imperfection is the point of the dataset, in which case it is labeled.

  4. 4

    Blur

    Faces, license plates, documents, screens and phone numbers are detected and blurred. Children are excluded.

  5. 5

    Review

    Every image is checked by a person. Missed details are blurred by hand or the image is dropped.

  6. 6

    Clean metadata

    GPS is removed; device and exposure EXIF is kept because it helps training. Location is recorded at region level only, never as coordinates.

  7. 7

    Label

    Topic labels, human-written captions or bounding boxes, depending on your order. Any AI assistance in labeling is disclosed in the datasheet.

  8. 8

    Document & deliver

    Images, metadata.jsonl, datasheet, license and checksums, delivered privately. Each customer copy is uniquely fingerprinted to deter leaks.

Reasoning, tests, Q&A and conversations

  1. 1

    Recruit and vet

    Authors are recruited for each field and checked: qualifications, years of practice and a paid trial task reviewed by a senior expert.

  2. 2

    Brief and guidelines

    Every project starts with written guidelines, a schema and worked examples agreed with you, so authors produce consistent records.

  3. 3

    Write

    Authors write each record themselves: the problem and every reasoning step, the test case and expected result, or the question and answer.

  4. 4

    Second review

    A second expert checks every record independently: re-solving problems, executing tests, verifying answers and citations.

  5. 5

    Screen

    Records are checked for copied text, LLM-generated writing and personal data. Anything that fails is rewritten or dropped.

  6. 6

    Document & deliver

    JSONL with a schema, datasheet, license and checksums. Each record notes the author's role and experience, never their name.

Our commitments

  • 100% made by people: no scraping, no stock, no generative AI
  • Every record checked by a second person
  • Faces and license plates blurred; GPS removed from every photo and clip
  • Text datasets use invented personal details only
  • Authors paid fairly and credited by role in the datasheet
  • A datasheet with every dataset, including known limitations
  • A takedown process for anyone who recognises themselves

Honest about limitations

Data made by a small team shares its devices, habits, locations and writing styles. That can introduce bias in framing, geography, phrasing and the kinds of problems chosen. Every datasheet states these limitations so you can account for them, and custom collections can be designed to widen coverage.

Read the detailed write-up