Prerequisites: Foundations of Embodied AI & Robotics Annotation for the labeling layers, and ideally Robot Failure Taxonomy & Edge-Case Diagnosis for the diagnostic mindset.
Egocentric footage is video shot from the perspective of the person doing the task, recorded by a head-mounted camera, smart glasses, or a headset worn during the activity. It is a growing source of manipulation data in embodied AI, and it is labeled under conventions that differ from anything in fixed-camera work.
Why first-person human footage is used to train robots
The obvious objection comes first: a human hand has five fingers and a robot gripper usually has two. Copying human grip mechanics directly does not work.
The data is collected anyway because of volume. Every hour of teleoperated robot demonstration needs a robot, a trained operator, and a controlled environment. Every hour of egocentric footage needs a person wearing a camera while doing something they were going to do regardless. The gap in cost and scale between those two is enormous, and large public research datasets built on this approach, including Ego4D and its multi-view successor Ego-Exo4D, established much of the labeling vocabulary the commercial work now uses.
What transfers across the embodiment gap is structure, not mechanics. How a task decomposes into steps, what order those steps occur in, which objects participate at which stage, where a person directs their attention before acting, and what a completed step looks like. That structural knowledge is what the annotation captures. Research on cross-embodiment transfer has shown that task structure learned from human footage can improve robot policy even when the physical bodies are different, which is the practical justification for collecting this data at scale.
Hand-object interaction
The first labeling layer identifies the hands and the objects they are engaging with.
Active and passive objects. An active object is one the hands are currently manipulating. A passive object is present in the scene without participating in the interaction. The distinction is about participation, and two common misreadings come from substituting something easier for it.
Motion is not the test. In head-mounted footage the camera moves constantly, so a stationary object sweeps across the frame and can look like the moving element in the shot. It is still passive.
Stillness is not the test either. A hand holding a jar steady while the other hand turns the lid is manipulating that jar. It is active despite being motionless, and a jar held in place is doing necessary work in the task.
Hand tracking and handedness. Projects commonly ask for hand keypoints, a set of tracked joint positions per hand, and for left and right to be identified consistently. Handedness matters more than it appears: which hand stabilizes and which hand acts is a meaningful pattern in manipulation data, and swapping the labels mid-sequence corrupts it.
Tools against targets. When a tool is involved, the tool and the object being worked on are usually labeled separately, and the boundary between them needs to stay clean at the point of contact. A screwdriver engaging a screw head is the moment those two boundaries are closest together and the moment the distinction matters most.
Critical frames and the point of no return
Fixed-camera action labeling often just needs a start time and an end time. Egocentric interaction labeling typically asks for three specific frames per interaction.
The pre-frame captures the state immediately before the action begins, generally the last frame before contact or before the motion starts.
The PNR frame, the point of no return, marks the moment the state change becomes irreversible. This is the one that requires thought. Approaching a jar and beginning to turn the lid are both reversible: stop, and nothing has changed. The instant the seal breaks is not reversible. That frame is the PNR. For a wire being cut, it is the moment the strands separate. For a drawer, it is the moment the drawer has moved enough that releasing it will not return it to closed.
The post-frame captures the completed state, usually where the hands disengage.
The reason the PNR is singled out is that it distinguishes an action attempted from an action that took effect. A model that learns only start and end times cannot tell a completed turn of a lid from a turn that slipped, because both occupy a similar span of time and look similar in motion. The irreversible moment is the evidence that the action landed.
Frame precision matters here in a way it does not for coarse segmentation, and this is where rubrics specify tolerances.
Try It: find the PNR
For each interaction, pick which step marks the point of no return before opening the reveal.
Interaction 1 — opening a jar.
- Hand reaches for the jar lid.
- Fingers grip the lid.
- Wrist rotates the lid counterclockwise.
- The seal breaks with an audible pop.
- The now-loose lid keeps spinning.
- The lid lifts off and is set aside.
Interaction 2 — cutting a wire with pliers.
- Hand positions the pliers around the wire.
- Jaws close around the wire without squeezing yet.
- Grip tightens and the pliers begin to compress the wire.
- The strands separate — the wire is now two pieces.
- The pliers open and release the cut ends.
Interaction 3 — opening a dresser drawer.
- Hand grips the drawer handle.
- The drawer slides out slightly — released now, it would glide back shut on its own.
- The drawer passes the point where it would self-close if released; from here it stays open without being held.
- The drawer is pulled the rest of the way open.
- The hand releases the handle.
See answer
Interaction 1 → step 4. Reaching, gripping, and turning the lid are all reversible — stop at any point and nothing has permanently changed. The seal breaking is the moment the jar’s state cannot be undone by reversing the motion.
Interaction 2 → step 4. Positioning and compressing the pliers are reversible right up until the strands actually separate. Once the wire is severed, releasing the pliers doesn’t restore it.
Interaction 3 → step 3. This is the least intuitive of the three because nothing dramatic happens visually at that instant — no snap, no pop. But step 2 is still fully reversible: let go and the drawer glides shut. Step 3 is the frame where that stops being true, which is exactly what a PNR marks, even when it isn’t the most visually obvious moment in the sequence.
Atomic actions and the keystep hierarchy
Continuous human activity gets structured into a hierarchy, usually three levels.
The goal is the overall activity. Repairing a bicycle. Assembling a piece of furniture. Preparing a specific dish.
Keysteps are meaningful stages within that goal. Removing the wheel. Attaching the backing panel. Reducing the sauce.
Atomic actions are the individual movements that make up a keystep, recorded as verb-noun pairs with timestamps. Loosen nut. Pull lever. Align bracket.
The verb-noun format is a constraint, not a stylistic preference. Free text descriptions of the same movement produce dozens of phrasings that cannot be aggregated across a dataset, so projects supply a fixed vocabulary for both the verb and the noun, and part of the skill is mapping an observed movement onto the nearest term in that vocabulary rather than inventing a more precise one.
A concrete example. The goal is “make a cup of coffee.” One keystep within that goal is “grind the beans,” which decomposes into atomic actions:
open bag→pour beans→close bag→close lid→press button
Another keystep, “pour the coffee,” might decompose into:
pick-up carafe→pour liquid→put-down carafe
The hierarchy gives a model three levels of understanding of the same footage: the overall activity, the procedural stages, and the individual hand movements.
Two recurring difficulties. Real activity does not proceed in clean sequence: people interleave steps, abandon one partway to do another, and return later. And the boundary between two consecutive atomic actions has the same property as the motion primitive boundaries covered in ROB1, which is that a defensible rule applied consistently matters more than locating an objectively correct frame that generally does not exist.
Gaze and intention
Some projects annotate where the camera wearer is looking. When the headset includes an eye tracker, that data gives a precise fixation point per frame. When eye tracking is not available, gaze is estimated from head orientation alone, which gives an approximate attention direction rather than a fixation. Projects using estimated gaze typically ask for a region or quadrant rather than a precise point, and state the expected granularity in the setup.
The value is in timing. People look at a target before reaching for it, so the gaze signal leads the physical action. A model trained on that relationship can begin predicting what someone is about to do rather than only classifying what they have already done, and anticipation is a capability that matters a great deal for a robot expected to work alongside a person.
Gaze labeling also captures attention that never turns into action, such as a glance at a tool that is considered and passed over. That is informative on its own, and it means gaze annotation cannot be shortcut by simply marking whatever the hands eventually touched.
What makes this footage difficult
Egocentric material has failure modes that fixed-camera work does not, and rubrics account for most of them explicitly.
Ego-motion. The camera moves with the wearer’s head. Object motion and camera motion have to be separated, and objects leave the frame entirely mid-interaction and come back, while remaining part of the same interaction throughout.
Motion blur. Head movement is fast and often produces several unusable frames exactly at the moment of interest, which is a real problem when the task is identifying a precise transition frame. Rubrics usually specify how to handle a PNR that falls inside a blurred stretch.
Self-occlusion. The wearer’s own hands and forearms block the view of what they are doing, most severely at the instant of contact. The hand is between the camera and the interaction by the nature of the setup.
Field of view. People do not keep their hands centered in view. Work happens at the edge of the frame and outside it, particularly during two-handed tasks.
None of these are annotation errors. They are properties of the medium, and knowing the rubric’s stated handling for each one is most of what separates efficient work from repeated escalation.
Egocentric video annotation flashcards
Camera motion makes stationary objects appear to move, and a hand holding something steady is manipulating it.
Not the start or the end. It's the instant the object's state changes in a way simple reversal can't undo, separating an action that took effect from one that was merely attempted.
The pre-frame captures the scene before the action, the PNR captures the irreversible change, and the post-frame captures the result. Together they give a model a before, a transition, and an after.
A constrained verb-noun structure keeps the same action labeled the same way across thousands of clips and many annotators. Free text produces fifty phrasings of one action that can't be aggregated.
The hierarchy runs broad to fine: the goal is what the person is ultimately doing (repairing a bicycle), a keystep is a stage within that (removing the wheel), and atomic actions are the individual movements that make it up (loosening a nut).
Visual attention shifts to a target ahead of the hand moving toward it, which is what makes gaze useful for predicting what someone is about to do, not just recognizing what they've already done.