Skip to content
Dataset By Humans
Coming soonVideo· Video

First-Person Household Task Video Dataset

Head- and chest-mounted clips of real household tasks in Southeast Asian homes, cooking, cleaning, laundry and small repairs, with step-by-step descriptions.

Coming soon: this dataset is produced to your specification after you order. Typically 3–4 weeks for 100 clips and 6–8 weeks for 500.

Pricing

Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.

Packages compared
FeatureStarterProCustom
clips100500Any quantity
What's included
  • 100 clips in JSONL
  • Step segments and descriptions
  • Datasheet, schema, license and checksums
  • 500 clips in JSONL
  • Everything in Starter
  • Tasks and settings chosen by you
  • Hand-object interaction labels
Your quantity, annotations and deadline
Datasheet, license, checksumsIncludedIncludedIncluded
Research license$790$7.90 / clip$2,990$5.98 / clipQuote
Commercial license$2,490$24.90 / clip$8,900$17.80 / clipQuote
Exclusive license$7,900$79.00 / clip$26,900$53.80 / clipQuote

How ordering works

  1. 1Choose a package and license and place your order. No payment is taken online.
  2. 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
  3. 3We produce and deliver in parts so you can review early. Typically 3–4 weeks for 100 clips and 6–8 weeks for 500.
  4. 4Pay the balance and download the final dataset.

Custom quote

Need another quantity, an Enterprise license or annotations for this dataset? Available here:

  • Step segments and descriptions
  • Hand-object interaction labels
  • Specific tasks and kitchens
  • Multi-view (two cameras)
Request a custom quote

Overview

Home robots and embodied AI learn from people doing tasks. Existing first-person video datasets are filmed mostly in Western homes, with different kitchens, tools and habits from those of the billion people in Southeast Asia.

Participants film themselves doing everyday tasks at home with head- or chest-mounted cameras: cooking Vietnamese dishes, washing up, hanging laundry, sweeping, small repairs. Each clip is split into steps and described by a person, and the household's privacy is protected.

Sample clips

Sample clips are emailed on request. The record below shows the metadata delivered with every clip.

  • Record 1 of 1 · illustrative

    {
      "fps": 30,
      "task": "cooking: fried rice with egg",
      "mount": "head",
      "steps": [
        {
          "end_s": 41,
          "start_s": 0,
          "description": "Crack two eggs into a bowl and whisk with chopsticks."
        },
        {
          "end_s": 96,
          "start_s": 41,
          "description": "Heat oil in a wok and scramble the eggs."
        },
        {
          "end_s": 230,
          "start_s": 96,
          "description": "Add cold rice, break up lumps and stir-fry."
        },
        {
          "end_s": 312,
          "start_s": 230,
          "description": "Season with fish sauce, add spring onion and plate."
        }
      ],
      "clip_id": "ehv-000045",
      "privacy": {
        "consent_form": true,
        "audio_removed": true,
        "faces_blurred": 0
      },
      "duration_s": 312,
      "resolution": "1920x1080"
    }

Get a free sample pack

We email you real clips from this dataset.

Used only to send your samples. Privacy policy.

Technical specifications

Capture
Head- or chest-mounted action cameras, 1080p or 4K, 30–60 fps, wide angle
Clip length
1–10 minutes per task
Tasks
Cooking, washing up, laundry, cleaning, tidying, small repairs, plant care
Participants
Adults in their own homes, paid and with signed consent
Annotations
Step segments with start/end times and a description per step
Privacy processing
Faces of others, documents and screens blurred; audio removed by default
License
Research, Commercial, Enterprise or Exclusive

Use cases

  • Robot learning

    Demonstrations of manipulation in real, cluttered homes for imitation learning.

  • Action recognition

    Step-level labels for recognising and segmenting household activities.

  • Video-language models

    Step descriptions for instruction following and procedural video QA.

  • Assistive technology

    Models that understand daily-living tasks for elderly and disability care.

How this data is made

  1. 1Filmed on a phone or action camera by a person
  2. 2Clips trimmed to the agreed length; duplicates removed
  3. 3Shaky, corrupted or off-spec clips filtered out
  4. 4Faces and license plates tracked and blurred frame by frame
  5. 5Every clip watched in full by a reviewer
  6. 6GPS and audio removed unless you ask to keep audio
  7. 7Human-written descriptions and labels per clip
  8. 8Datasheet and license packaged with the delivery
Read our full process

Provenance & legal

  • 100% made by people: no scraping, no generative AI
  • Datasheet documenting how the data was made, checked and its limitations
  • Commercial license that lets you keep models trained on the data
  • Supports training-data documentation under the EU AI Act
Licensing options

Dataset-specific notes

  • Every participant signs a consent form covering AI training and can withdraw future use of their clips.
  • Children and other household members are kept out of frame or blurred.
  • Sample clips are being prepared; register interest to receive them first.

Example metadata record

One line of metadata.jsonl per clip. Fields vary with the annotation options you choose; values shown are illustrative.

{
  "fps": 30,
  "task": "cooking: fried rice with egg",
  "mount": "head",
  "steps": [
    {
      "end_s": 41,
      "start_s": 0,
      "description": "Crack two eggs into a bowl and whisk with chopsticks."
    },
    {
      "end_s": 96,
      "start_s": 41,
      "description": "Heat oil in a wok and scramble the eggs."
    },
    {
      "end_s": 230,
      "start_s": 96,
      "description": "Add cold rice, break up lumps and stir-fry."
    },
    {
      "end_s": 312,
      "start_s": 230,
      "description": "Season with fish sauce, add spring onion and plate."
    }
  ],
  "clip_id": "ehv-000045",
  "privacy": {
    "consent_form": true,
    "audio_removed": true,
    "faces_blurred": 0
  },
  "duration_s": 312,
  "resolution": "1920x1080"
}

Each delivery contains:

egocentric-household-video-v1/
├── clips/
├── metadata.jsonl
├── DATASHEET.md
├── LICENSE.pdf
└── checksums.sha256

Frequently asked questions

Who appears in the videos?

Adult participants filming their own hands and surroundings. They are paid, sign consent for AI training and can withdraw.

Can we specify the tasks?

Yes. Send a task list (for example dishes to cook or objects to manipulate) and we will plan participants and homes around it.

When will sample clips be available?

Sample clips are in production. Register interest with the sample form and we will send them as soon as they are ready.

How does ordering and payment work?

Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.

What is the difference between non-exclusive and Exclusive?

With Research or Commercial we may license the clips produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.

Can I order a different quantity or extra annotations?

Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.

Need a variation?

Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.

Request a quote
Browse all datasets