How we make human data
Photos and video
- 1
Capture
Original smartphone photos taken by people in real places. No scraping, no stock, no generative AI. We only shoot where photography is permitted and ask permission in private spaces.
- 2
De-duplicate
Exact and near-duplicate images are removed so they do not inflate the dataset or leak across your train/test split.
- 3
Filter
Broken, corrupted and accidental shots are removed, unless imperfection is the point of the dataset, in which case it is labeled.
- 4
Blur
Faces, license plates, documents, screens and phone numbers are detected and blurred. Children are excluded.
- 5
Review
Every image is checked by a person. Missed details are blurred by hand or the image is dropped.
- 6
Clean metadata
GPS is removed; device and exposure EXIF is kept because it helps training. Location is recorded at region level only, never as coordinates.
- 7
Label
Topic labels, human-written captions or bounding boxes, depending on your order. Any AI assistance in labeling is disclosed in the datasheet.
- 8
Document & deliver
Images, metadata.jsonl, datasheet, license and checksums, delivered privately. Each customer copy is uniquely fingerprinted to deter leaks.
Reasoning, tests, Q&A and conversations
- 1
Recruit and vet
Authors are recruited for each field and checked: qualifications, years of practice and a paid trial task reviewed by a senior expert.
- 2
Brief and guidelines
Every project starts with written guidelines, a schema and worked examples agreed with you, so authors produce consistent records.
- 3
Write
Authors write each record themselves: the problem and every reasoning step, the test case and expected result, or the question and answer.
- 4
Second review
A second expert checks every record independently: re-solving problems, executing tests, verifying answers and citations.
- 5
Screen
Records are checked for copied text, LLM-generated writing and personal data. Anything that fails is rewritten or dropped.
- 6
Document & deliver
JSONL with a schema, datasheet, license and checksums. Each record notes the author's role and experience, never their name.
Our commitments
- 100% made by people: no scraping, no stock, no generative AI
- Every record checked by a second person
- Faces and license plates blurred; GPS removed from every photo and clip
- Text datasets use invented personal details only
- Authors paid fairly and credited by role in the datasheet
- A datasheet with every dataset, including known limitations
- A takedown process for anyone who recognises themselves
Honest about limitations
Data made by a small team shares its devices, habits, locations and writing styles. That can introduce bias in framing, geography, phrasing and the kinds of problems chosen. Every datasheet states these limitations so you can account for them, and custom collections can be designed to widen coverage.
Read the detailed write-up