The modules you’ve read so far deal with language: two responses, a rubric, a judgment about which one is better and why. Embodied AI data is a different kind of problem. The subject is a physical machine moving through real space, and a label has to describe what happened in three dimensions, across time, while the camera, the robot, and the objects in the scene are all capable of moving at once.
This module covers how that data is structured and what makes a label usable. It is background for Robot Failure Taxonomy & Edge-Case Diagnosis, which deals with the harder and better paid question of diagnosing why an attempt went wrong.
What this data is training
Robots that operate in unstructured environments are increasingly trained the way language models are: on large quantities of recorded demonstration instead of hand-written control code for each individual task.
Most of that demonstration comes from teleoperation. A human operator drives the robot through a task while the system records everything: video from one or several cameras, the angle of every joint, the position and rotation of the gripper, and often the force registered when the gripper contacts something. The result is a recording of a task performed through the robot’s own body.
On its own that recording is unlabeled. It becomes training signal once someone has marked what objects were involved, where the task divides into steps, which steps succeeded, and what instruction the whole sequence was carrying out. That work is the job.
Four layers of labeling sit on top of a single clip. Most projects use some combination of them, not all four at once.
The perception layer covers what is in the scene
The first question is what objects exist and where they are.
Two dimensional boxes are the simplest form and are drawn on the image plane. They tell a model where something appears in the picture with no information about how far away it is.
Three dimensional cuboids describe a volume in actual space. A cuboid carries a position, three dimensions, and an orientation. That orientation is the part annotators most often get wrong and it carries real consequence: a cuboid with the right position and size but the wrong rotation tells the model the object is turned a way it is not, and any grasp pose derived from it will approach the object at the wrong angle.
Segmentation masks trace an object at the pixel level instead of enclosing it in a box. They are slower to produce and used when shape matters, such as for deformable objects like cloth and cable where a box communicates almost nothing useful.
Two recurring complications:
Occlusion. When one object passes in front of another, you need to know which convention the project uses. Modal annotation labels only the portion you can see. Amodal annotation estimates the full extent of the object and labels it including the hidden part. Both are legitimate and projects specify one. The failure to avoid is drifting between them, because an object that appears to shrink whenever something passes in front of it teaches the model that objects change size for no physical reason.
Reflective and transparent surfaces. Depth sensors work by emitting light and measuring what comes back. Glass transmits that light to whatever sits behind it and a mirrored surface redirects it, so the sensor reports a distance to the wrong surface or reports nothing at all. A glass of water can look unambiguous in the color image while the depth channel in that region is empty or wrong. Knowing that the depth data can be missing for a clearly visible object prevents a great deal of wasted time.
The action layer divides continuous motion into steps
A demonstration arrives as unbroken motion. Training signal requires it divided into discrete, named, timestamped segments.
Those segments are motion primitives. A pick and place task typically decomposes into something close to: approach, align, grasp, lift, transport, place, release. The specific vocabulary comes from the project.
A sample vocabulary for a tabletop pick-and-place project might look like this:
approach·align·grasp·lift·transport·lower·place·release·retract
Each term has a project-supplied definition. “Approach” might mean “end effector moving toward the target with no orientation correction,” while “align” means “adjusting orientation without closing remaining distance.” The definitions matter more than the names.
The difficulty is not naming the segments. It is deciding exactly where one ends and the next begins. There is usually no single frame where approach objectively becomes align. The arm is moving continuously and the transition is gradual.
This is where new annotators look for a correct answer that does not exist, and the resolution is worth internalizing early. The rubric defines the boundary, your intuition does not. If the rubric says align begins when the gripper’s approach vector stops changing, that is the rule, and it applies on every clip whether or not it feels right on any particular one. Consistency is the property that makes the dataset usable. A defensible boundary applied the same way every time is worth considerably more than a perfect boundary applied unpredictably, because contradictory labels across similar clips teach a model that phase boundaries are arbitrary.
The sensor layer matters most when the video and the robot disagree
This is the idea that separates people who are effective at this work from people who are merely fast at it.
Alongside the video, the robot records its own state: joint angles, the pose of the end effector, whether the gripper is commanded open or closed, and frequently the force and torque registered at the wrist or the fingers.
Video and telemetry can tell different stories about the same moment, and the disagreement is often the most informative thing in the clip.
The clearest case is contact. From one camera angle, a gripper that has closed firmly around an object and a gripper that has closed just beside it can look identical. The force reading resolves it immediately. A real grasp presses on something and that pressure registers. Flat force readings through an apparent grasp mean the fingers closed on air.
The general principle: when the sensor channel is available, it is the stronger evidence about physical events. The camera shows you appearance. The force reading shows you what happened.
The language layer connects instructions to intent
The last layer connects physical behavior to words.
Most embodied datasets pair each demonstration with one or more natural language instructions describing what the robot was asked to do. Some tasks ask annotators to produce several different phrasings for the same recorded action.
That request is not redundancy for its own sake. A model trained on a single phrasing learns that phrasing. Real users ask for the same thing in wildly different ways, and the variance in the training data is what lets the model handle a request it has never seen worded that way before. Writing “pick up the red mug,” “grab the red cup from the table,” and “hand me that red mug” for one clip is teaching the model that all three point at the same behavior.
A related task type appears in footage where a person and a robot work together. There the labeling question shifts from what the robot did to what the human appeared to want, which is inference about intent instead of description of motion, and rubrics for it tend to be considerably more prescriptive as a result.
What makes an instruction fail review is usually vagueness. “Get the thing from over there” uses a pronoun with no referent and a spatial term with no anchor. “Pick up the red mug from the left side of the counter” gives the model a color, an object class, and a location. The instruction needs to be unambiguous to someone who cannot see the video.
A note on coordinate frames
Cuboid positions and orientations are always relative to a reference frame, and projects specify which one. Common choices are the world frame (a fixed origin in the room), the camera frame (relative to the camera’s viewpoint), or the robot base frame (relative to where the robot is mounted). Placing a cuboid in the wrong frame produces coordinates that are numerically valid and physically meaningless. The frame is usually stated once in the project setup and easy to miss on a second read.
What separates a usable label from a rejected one
Across all four layers, reviewers tend to reject work for the same handful of reasons.
Labels that are internally inconsistent across a batch cause more damage than labels that are uniformly slightly off. A model can absorb a consistent bias. It cannot absorb contradiction. Boundaries placed by feel rather than by the rubric’s definition produce exactly that contradiction. Cuboids with correct position and careless orientation pass a glance and fail their actual purpose. Ignoring the sensor channel when it is provided means treating appearance as truth on precisely the clips where appearance is misleading.
The common thread is that this work rewards disciplined application of someone else’s definitions over independent judgment about what looks right.
Embodied AI annotation flashcards
Orientation is the one that gets treated as optional, and the one that determines whether a learned grasp approaches the object correctly.
Modal annotation traces only the visible pixels when one object partially hides another. Amodal annotation estimates where the hidden portion continues and labels the whole object. Mixing the two within a dataset makes objects appear to change size.
Short, named, timestamped segments that break a continuous demonstration into steps a model can learn from, rather than an undifferentiated stream of motion.
There's seldom an objectively correct frame where one motion phase becomes the next. A defensible rule applied consistently beats an ideal boundary applied unpredictably.
A gripper closing beside an object looks like a successful grasp from most camera angles and registers as nothing on the sensor. When force data and video disagree, the force data wins.
Glass lets sensor light pass through and mirrors deflect it, so a clearly visible transparent object can have no usable depth behind it. That's a sensor limitation, not an annotation error.