Real data, made by people
Dataset By Humans is an independent data studio. We produce datasets the slow way: people photograph and film the real world, write out how they reason, design tests and answer questions in their own field, and every record is checked by a second person before it reaches your training or evaluation set.
Why we exist
AI systems increasingly learn from scraped or synthetic data of unknown origin. That creates blind spots, especially for languages, regions and professions that are underrepresented online, and legal uncertainty for the teams building on it. As models get better, the scarce ingredient is no longer volume: it is data with real human judgement in it and a clear, documented chain from its author to your model.
What we make
- Photos and video of real places and activities in Southeast Asia, privacy-processed by hand.
- Reasoning data: step-by-step solutions and decision traces written by teachers, engineers and professionals.
- Evaluation and test cases: private LLM evaluation sets with rubrics, software QA test cases and agent tasks.
- Expert Q&A from lawyers, accountants, agronomists, clinicians and other qualified people.
- Conversations role-played by trained people for assistants and support bots.
Where we are going
We started with photos of Southeast Asia, where public datasets barely reach. We are now building a network of vetted professionals so that any team can order human-made data for a new language, profession or task, and know exactly who made it and how.
How we use AI ourselves
Our images are never generated or altered by AI, and our written records are not generated by language models: we screen submissions for machine-written text. We do use software, including AI-based detection, to find faces and plates to blur, and we may use AI tools to help draft text such as translations of this site. When AI assists with any part of a dataset, the datasheet says so.