Physical AI training data

Commission the egocentric demonstrations your policy is missing.

First-person human work, captured to your brief and delivered as rights-evidenced training data.

Machines are learning to work with their hands.

A robot learns a task the way a person does — by watching one.

The tasks people do without thinking are the ones machines find hardest.

Fourthuman

Why this is hard

AI learned from the internet. The physical world was never on it.

A model can read every sentence ever published and watch every video ever uploaded, and still not know how much force a drawer needs. Physical intelligence is not written down anywhere; it is performed, by people, and then it is gone.

Language models

Billions of words

Already collected, already public. The corpus existed before the models did.

Vision models

Billions of frames

Images and video scraped from an internet that was already there. Again, found rather than made.

Physical AI

No corpus

Nothing equivalent exists. Egocentric capture of ordinary human work has to be recorded on purpose, by someone, with consent.

The first two corpora were found. The third has to be made, and that asymmetry is the point.

One recording

What happens to sixty-one seconds of footage.

Five layers, one real recording, and every figure on this page computed from the file itself. Each step links to the page that sets it out in full.

Raw frame 150 of the reference clip, from the wearer's viewpoint: the left hand aligns and organizes fabric items on a workstation surface beside a sewing machine, the action the published annotation records for this segment. The same frame with FH-Ego v1 annotation drawn on: the left hand's 21 tracked joints, the connecting skeleton, its wrist triad axes, and a label giving per-hand confidence and the wrist position vector. Only the left hand is annotated in this frame. A header burnt into the image reads: Egocentric Bimanual Workstation Interaction, frame 150 of 1843, action organize_items.

50% annotated · raw at left, FH-Ego v1 overlay at right

Reference clip, sewing workstation · 21 3D joints per hand, left/right classified · the overlay is the engine's own output, untouched
What redaction covers

Faces — automatic, server-side at ingestion, before any human review, and it fails closed. measured

Licence plates — not implemented, and not claimed. none

chunks · drawn at true frame positions · derived

FRAME 0276 / 1843 · loading…

The annotation engine runs today on a clip we publish in full; the Capture Partner network is in pilot.

FRAME  / 1843 · · · speed · conf · loading…

sha256 of this file
file size
tracks
sampled every
frames
Sampled every tenth frame by the pipeline, so the path is drawn point-to-point between real measurements and nothing between them is invented. Positions and velocities are derived; they are the tracking model's own output, not an instrumented ground truth.

01 · Capture

One person, one camera, sixty-one seconds.

measured

A Capture Partner records the task on a head-mounted camera against a brief that fixed the task, the environment and the rig before anything was recorded. Drag the handle to move between the raw frame and the annotated one.

About this clip, before you go further

This clip predates our deployed consent flow, so it carries no consent receipt. That is the one thing wrong with the artifact this whole page is built on, and it is the next thing we fix.

How a recording is specified, consented and made

02 · Redact

Faces go before a person sees the footage. Number plates do not.

none — we hold no artifact for this step, and will not stage one

Redaction runs at ingestion, server-side, before any human review, and it fails closed — it refuses to emit a clip rather than emit an unredacted one.

What we blur, what we do not, and what happens when it fails

03 · Segment

Sixteen fixed windows, at their true frame positions.

derived

The recording is divided into sixteen chunks and each is given a verb-noun code. On this clip that division is a fixed window of 120 frames, not a boundary found in the footage, and the codes come from a fixed pool rather than a model. Content-aware segmentation is specified and not built. The bars below sit at their true frame positions.

How one recording becomes usable clips

04 · Annotate

Two hands, measured through sixty-one seconds of real work.

measured

Hand landmarks are computed every tenth frame — a three-hertz track, not thirty. Drag the timeline to move between sampled frames. The scrubber never interpolates: every position it shows is a row of the artifact, and there is nothing between two measurements to show.

FH-Ego v1, field by field

05 · Deliver

The file, and the hash that proves it is the file.

measured

The finished record is content-addressed: rehash the annotation file you download and you get the digest below. The trajectory below is drawn complete rather than animated, because a path easing into place invents positions between the ones we measured.

What arrives, and in what container

By the numbers

Every number here comes from a file you can open.

Clips commissionable across the catalogue

182,700

Target

Specified per brief before capture, across 6 capture domains. Every brief states its own rig, its annotation depth and where it stands.

Hours of egocentric capture, specified

1,505

Target

Volume is agreed in a brief and recorded against it. Simulated scenes are catalogued separately and are not counted here.

Delivery containers, each independently verified

5

Each has a writer and its own verifier, and every export is read back off disk before it ships.

Frames annotated end to end and published in full

1,843

One real recording at 30 fps, with its annotation and its checksum. Every engine figure on this page comes from that file.

Mean per-hand confidence

0.87

Arithmetic mean over the kinematic track. The worst value is published beside it rather than the mean alone.

Routes in the engine behind it

173

Across 24 durable stores — the consent ledger, the holdings register, the money book and the rest.

Catalogue

One annotation standard, across every domain we scope.

Each domain below is a capture specification you can commission, and each ships against FH-Ego v1, so a dataloader you write once works across every delivery.

Delivery formats Five containers ship. MCAP and WebDataset do not. What arrives, and in what container.

Tier 1 · Session

What the clip is

Scene, task label, frame count, frame rate, usability verdict, and the provenance block: clip id, content hash, cohort, QA verdict, and which fields were derived rather than measured.

Tier 2 · Action

What happens, and when

Frame bounds per action segment, hand laterality (left / right / both), scene semantics, camera motion, a verb-noun action code and a dense natural-language description.

Tier 3 · Kinematics

Where the hands are

21 3D joint landmarks per hand, camera-relative wrist position, frame-to-frame velocity, and a per-hand confidence; so you can filter on quality rather than trust it.

The catalogue also scopes agriculture and outdoor tasks, and pairwise preference ranking on policy rollouts.

Why egocentric

A head-mounted camera sees the task from where the hands are.

Teleoperation gives you a robot's own proprioception, and it costs one person driving one rig for every hour of demonstration it produces. Egocentric capture gives you human hands solving the task in the environment the task actually happens in, and it scales with people rather than with rigs. There is a real case against it (no joint torques, no action labels for free, a domain gap you have to close), and the published work arguing both sides is a category argument, not a Fourthuman result. We set it out on the method page.

Price

We publish what reaches the person holding the camera.

The rate that reaches the person holding the camera is published on the page addressed to them — so it is the one cost inside a Fourthuman quote you can check before you have spoken to us.

What moves the number · 1

The environment

An Indian household kitchen and an active distribution centre are not the same shoot. Access, scheduling and how many people have to consent all change with the room.

What moves the number · 2

The session

How long a session runs, how many of them, and how many Capture Partners the brief needs to reach the volume you are asking for.

What moves the number · 3

The annotation tier

Tier 1 is session metadata. Tier 3 is 21 3D joints per hand with per-hand confidence, and it costs more to produce and more to verify. You pick the tier in the brief.

A brief is scoped first and priced against that scope, so the number follows the specification rather than a list. The structural reason a bespoke brief is affordable here at all is the India–US wage differential.

FAQ

Frequently asked questions

How do you handle privacy and data protection?
Every clip passes automated face redaction server-side at ingestion, before any human review or delivery. YuNet runs where the host provides it — a DNN we chose because egocentric faces are off-axis, motion-blurred and partial, the angles a classical cascade misses. Where it does not, an older OpenCV cascade runs instead. The stage fails closed: if the detector cannot load, the clip fails the stage rather than passing through, so there is no unredacted path to delivery. The pass covers faces. What it covers, and how far, is set out in full on the redaction page. Every capture session produces a digital consent receipt persisted with a SHA-256 audit hash, revocable by the Capture Partner under the data-deletion rights in India's DPDP Act 2023. We hold no third-party attestation — no SOC 2 report, no ISO 27001 certificate — because that requires an external audit we have not commissioned. We will link the report here when we do.

Also answered in full: what we capture and in what form · what a Capture Partner is paid

What annotation schema do you use?
We use the Fourthuman Ego Schema v1, a 3-tier annotation schema: Tier 1 (Session): task_id, scene, task_label, usability flag. Tier 2 (Action): frame bounds, hand laterality (left/right/both), scene_semantic, camera_motion, verb-noun action code, dense NL description. Tier 3 (Kinematics): 21 3D joint landmarks per hand, camera-relative Cartesian position vector P=[X,Y,Z], and frame-to-frame velocity, each hand carrying a per-hand confidence for the frame. All exported as structured JSON alongside video streams.
What data formats do you deliver datasets in?
We deliver LeRobot v2.1 (pi0/openpi and GR00T pipelines; our output loads with the real lerobot package), HDF5 (EgoDex-style paired .hdf5+.mp4, ALOHA / RoboMimic-lineage labs), and the open FH-Ego v1 JSON schema. Zarr v2 (UMI / diffusion-policy layout) and TFRecord (RLDS-style step schema; not tfds.load-compatible) are produced on request. Every export is read back by an independent verifier before delivery. We do not offer MCAP or WebDataset: MCAP is a robotics logging container no training dataloader consumes, and WebDataset is an internal sharding optimization, not a distribution format.
What makes you different from a general annotation vendor?
We work only on Physical AI, not general-purpose annotation. We combine a capture network in India — paid at a rate we publish rather than quote privately — with automated 3D annotation pipelines. The India–US wage differential is the structural cost advantage; we publish the floor a Capture Partner is paid from rather than a comparison against a US teleoperation figure we have not independently verified. Unlike general-purpose BPO annotation vendors, we own the full stack: recruitment, capture hardware, QA and privacy pipeline, and AI annotation. Delivery is LeRobot v2.1, HDF5 (EgoDex-style) and the open FH-Ego v1 JSON schema, with Zarr v2 and TFRecord (RLDS-style) on request. We don't sell raw footage; we sell annotated, privacy-cleared, format-ready training data.
How does your pricing compare to US-based teleoperation data collection?
We publish the floor a Capture Partner is paid from rather than quoting a US teleoperation benchmark we have not independently verified — the range and each brief's own rate sit in the Capture Partner app, because they depend on the brief. The India–US wage differential is the structural cost advantage and it is real; what we won't do is put a number on someone else's cost base and present it as measured. Ask for a quote against your capture spec and compare it with what you pay today. We capture value by delivering annotated, format-ready data, not raw footage.

Get started

Tell us what you’re training.
We’ll scope the dataset.

Describe the policy and the embodiment you are training it for. We scope the capture specification, price it against your volume and format, and come back with what we can commit to, including the parts we cannot. One of the two founders reads every one of these.

What a brief costs is scoped before it is priced: how we price, and what moves the number.

Three fields: your name, your email, and what you are trying to build. Everything else helps us scope faster and none of it is required.

Volume, timeline, the constraint that matters most. Whatever tells us whether we can help.

What happens when you submit

When you press Send this brief, it is stored and you get a reference id on screen. Keep it — quote it if you write to us.

A person reads it, not a bot. You also get a confirmation email at the address you gave — if it does not arrive, check your spam folder, and keep the reference id on screen either way. If the store cannot be reached you get a button instead, which opens a pre-filled email to [email protected] with your text intact.

We aim to respond within 1–2 business days. If you have not heard from us by then, please feel free to follow up.