What we do

What happens between a brief and a delivery

Five stages, each with its own page. Every one says what it does not do as plainly as what it does. The annotation pipeline runs now; the Capture Partner network is still being built.

Delivery formats Five containers ship. MCAP and WebDataset do not. What arrives, and in what container.

The five stages, in the order a Recording travels them

Capture and redaction happen before anyone here opens the footage. Segmentation, annotation and delivery happen after.

Three more sit alongside them, and none of them is a step in the walk: how a clip is judged, which happens inside segmentation rather than after it — one analysis over one set of frames decides both where the task boundaries are and which spans are usable; the consent ledger, which is how a receipt is checked; and why egocentric, which is the category question.

What every layer shares

Three things hold across all five stages. Each is a field you can read in a delivered record or a check you can run against this site, not a promise.

What holdsWhere you can see it
The consent record travels with the Recording Consent attaches to the Recording, not to a Clip cut from it, so every Clip carries a consent_receipt_id back to one receipt — on the annotation record, the holdings row and the provenance block alike. Any receipt id is checkable at Check a receipt.
Every derived value carries the layer that produced it A record names its own model-generated fields in derived_fields, and caption_source names what wrote them — a model id, or heuristic_stub when it is development filler rather than a model. A derived value that does not say so is the one thing our own schema will not let us ship.
Nothing is interpolated Where this site shows a reading, it is an exact row from the published file — never a value blended between two. Two gates assert it and one of them forges a blended reading to prove the rejection fires.

Ready to scope your training dataset?

Send the task list, camera perspectives and format requirements. We come back with a scoped brief: what would be captured, at what tier, in what volume, and what it costs. There is no inventory to pick from; every engagement is specified before it is collected.

If you are working out whether first-person capture suits the model at all, physical AI training data takes that question on its own terms.