Prerequisites: Foundations of Embodied AI & Robotics Annotation covers the labeling layers this module assumes, including motion primitives and sensor telemetry.
Standard annotation answers the question what happened. This module covers a different question, why it went wrong. Describing a sequence is observation. Classifying a failure is diagnosis, and it requires attributing an outcome to a cause in a system with several places where things can break.
The work is paid accordingly. Roles built around failure and edge-case diagnosis sit toward the Specialist and Subject Matter Expert end of the range, not in general annotation, and screening for them tests exactly the distinctions in this module.
Why failure footage is the valuable material
A robot that picks up a cup ten thousand times without incident produces a large quantity of data with limited information in it. The behavior already works, and repeating it mostly reinforces what the model has.
The episode where the grasp slips, or the arm misjudges depth against a reflective surface, or the system freezes because it cannot find a route, carries far more signal. It shows where the task stops working. A model that only sees the happy path has no representation of the boundary of its own competence and no example of what to do after something has gone wrong.
Failure footage is also scarcer. Most recorded demonstrations succeed, by design, which makes each failure worth more per clip and is why this category of work commands higher rates and tighter screening than descriptive labeling.
The diagnostic chain
Before reaching for a specific label, locate the failure in the pipeline. A robot acting on the world runs roughly three stages: it builds an understanding of the scene, it decides how to move through that scene, and it physically executes that decision.
Each stage has its own failure family, and the first question in any diagnosis is which stage broke.
The reason this ordering matters is that the stages are sequential. A wrong understanding of the scene produces a correct plan for a world that does not exist, which then executes into a collision. The visible event is at the end of the chain and the cause is often further back.
Perception failures
The robot’s model of the scene was wrong.
This family covers depth misjudged, so the gripper closes short of or past the object. Glare or lighting conditions that wash out a surface the system needed to see. Transparent and mirrored materials, which return unreliable depth for the reasons covered in ROB1. Objects that were present but never detected, and phantom detections of objects that were not there. Partial occlusion where the system lost track of an object once something passed in front of it.
The identifying characteristic is that the robot acted sensibly given what it believed, and what it believed was incorrect.
Planning and navigation failures
The scene was understood correctly and the route through it was wrong.
This covers collisions with obstacles the system had detected, paths that were valid but unnecessarily convoluted, approaches to an object from an angle that made the grasp impossible, and stalls where the system stopped making progress because it could not resolve a route.
That last case, where the robot simply stops, is worth separating from a robot that is briefly paused. Stalls and pauses look the same at a glance, which is why rubrics define idleness with a minimum duration rather than by eye. More on that below.
Manipulation and execution failures
Perception was right, the plan was right, and the physical execution failed.
This covers grasps that close off center, grips that slip during transport, objects gripped hard enough to deform or damage them, objects released early or in the wrong place, and losses of an object partway through a trajectory.
Within this family sits the distinction that most screening tests use to separate candidates who understand the taxonomy from those applying it by outcome.
The control test
A drop requires that the robot had the object.
If control over the object was never established, nothing was dropped. The event is a failed grasp, no matter how dramatic the object’s departure looks on video, and no matter that it ends up in the same place either way.
This is unintuitive on first encounter because outcome-based reasoning is the natural default. The mug is on the floor, so the robot dropped it. The taxonomy does not classify by where the object ended up. It classifies by whether the robot ever held it.
Two clips can look nearly identical and belong in different categories:
- The gripper closes, the mug shifts and squeezes free before the grip is ever stable, and it falls. Control was never established. This is a failed grasp.
- The gripper closes, the mug is held securely, the arm begins to lift and transport it, and partway through the motion the grip gives way. Control existed and was lost. This is a drop.
Where the threshold sits. Rubrics generally do not treat the instant of gripper closure as the moment control begins, because momentary contact during a botched grasp would then satisfy the definition and every slip would be logged as a drop. Instead they require the hold to persist for a defined continuous interval before control counts as established. The length of that interval is a project decision and it varies. Read it from the rubric rather than assuming a value, and apply the same number across every clip in the batch.
The useful habit when a case feels ambiguous is to stop asking what happened to the object and ask instead whether control was ever established. That answer resolves nearly every borderline case, and it resolves it the same way every time, which is the property the dataset needs.
Try It: drop or failed grasp?
For each clip description below, decide whether it is a drop or a failed grasp before opening the reveal.
Clip 1. The gripper closes around a phone. Before any lift begins, the phone slips free and falls flat on the table.
Clip 2. The gripper closes around a bottle, the arm lifts it thirty centimeters, and midway through transport the bottle slips out and falls.
Clip 3. The gripper closes around a ball, but the ball is slightly larger than the planned grip width. The fingers never achieve a stable squeeze, and the ball rolls free the instant the arm starts to move.
Clip 4. The robot lifts a box cleanly and is moving it toward the destination bin when a jolt in the arm’s motion causes the box to shift and fall off the end effector mid-trajectory.
Clip 5. The gripper closes around a pen, slightly off-center. The pen tips out of the grip immediately, with no interval where it was actually held.
See answer
Clip 1 is a failed grasp. No hold interval elapsed before the phone was gone. Control was never established.
Clip 2 is a drop. The object was lifted and transported, which means control existed. It was lost partway through.
Clip 3 is a failed grasp. The grip never stabilized. That the ball only moved once the arm started doesn’t matter. Nothing was ever under control.
Clip 4 is a drop. The box was successfully lifted and was in transit. Losing it mid-trajectory is a loss of established control, not a botched grasp.
Clip 5 is a failed grasp. No interval of stable holding occurred. An off-center grip that never holds is the same category as one that never closes correctly, regardless of how brief the visible “grip” looks on video.
Try It: annotate a clip
A one-minute clip from a fixed camera shows a robot arm moving a cube and a cylinder into a bin. Using the example rubric below, how many notable events does the clip contain, and what is each one?
Example rubric. The numbers apply to this exercise only. Every project sets its own.
- Failed grasp: control was never established. Control counts once the object has been held for 0.8 s after the gripper closes. Start: the gripper starts closing on the object. End: the object is at rest.
- Drop: control was established, then lost. Start: the object separates from the gripper. End: the object is at rest.
- Idle: no arm motion for 6 s or more. Start: the last productive motion ends. End: motion resumes.
At 00:05 the gripper closes on the cube, which rises with it for about half a second and falls back to the table. The arm grasps the cube again, carries it to the bin and releases it. From 00:22 it holds still over the bin until 00:34. It moves back, pauses for 2 seconds at 00:36, then grasps the cylinder at 00:40 and lifts it. At 00:45, partway to the bin, the cylinder slips out and falls. The arm carries on to the bin and opens over it with nothing in the fingers.
See answer
The clip contains three events.
- Failed grasp, 00:05 to 00:07. The cube rose with the gripper for about half a second, short of the 0.8 s hold the rubric needs, so control was never established and nothing was dropped.
- Idle, 00:22 to 00:34. Twelve seconds without motion, over the 6 s floor. The 2-second pause at 00:36 is under the floor, so it is normal execution and not an event.
- Drop, 00:45 to 00:46. The cylinder was lifted and carried for about four seconds before it slipped, so control existed and was lost.
A project with a recovery field would mark this clip as no recovery, because the arm never checks whether it is still holding the cylinder.
Instruction-action disconnect
The robot executed a motion competently and it was the wrong task.
Asked for the blue cup, it picks up the red one with a clean, stable, well oriented grasp. Every motion primitive in the sequence succeeded. A taxonomy that examines only physical execution records the episode as a success.
This family is easy to miss precisely because nothing looks wrong. Catching it requires reading the instruction that accompanied the clip and comparing it against what the robot did, which is a step annotators skip when the motion itself is clean. The failure lives in the mapping from language to action, not anywhere in the physical execution.
A subtler version: the instruction says “stack the red block on the blue block.” The robot stacks the blue block on the red one. Every grasp, lift, and place succeeds. The objects are in a stack. But the spatial relationship is inverted, and the episode is a failure that will not be caught by anyone who only watches the motion.
Recovery behavior
After something goes wrong, the informative question is whether the robot noticed.
A system that loses an object, detects the loss, and returns to attempt the grasp again is demonstrating error detection and correction. A system that loses an object and continues through the remainder of the sequence, carefully placing nothing at the destination, has no representation of its own state.
These two behaviors produce the same outcome and sit far apart in capability. Recovery labeling captures that difference, and it is frequently the field that carries the most weight in evaluation work, since error recovery is one of the harder capabilities to train and one of the most important to measure.
Worth distinguishing: a retry that happens because the robot detected a problem, and a retry that happens because the robot blindly repeats its routine. Only the first is recovery.
Why rubrics specify exact thresholds and counts
Several categories in this work depend on numbers the rubric supplies instead of on judgment.
Duration floors. Idleness is typically defined as a period of no productive action lasting at least some minimum, measured from the point the last productive action ended. Robots pause constantly while sensing, settling, or waiting on a state update, so without a floor every transition between phases would register as a stall.
Attempt counts. Repeated unsuccessful tries at the same goal are usually collapsed into one labeled event, not logged individually, with a minimum number of attempts before the category applies at all. Rubrics also specify whether attempts accumulate against the goal or reset when the robot changes its approach. That is a specification to read, not to infer.
The general point: these numbers exist so that two annotators watching the same clip produce the same label. They vary between projects, and carrying a threshold from one project into another is a reliable way to fail a calibration review.
What is not a failure
Two things that get tagged as defects and should not be.
Slowness. An action can be slow and productive at the same time. A careful, deliberate approach to a fragile object is correct behavior that happens to take longer. Speed is generally not part of the defect taxonomy, and tagging slow motion as a failure trains a model to associate caution with error.
Normal pauses. The gap between finishing one motion primitive and starting the next is execution, not idleness, which is the entire reason duration floors exist.
The distinction to hold onto is between unproductive time and slow productive time. Only the first belongs in the taxonomy.
Symptom or root cause
One structural question runs through everything above. When a perception error causes a manipulation failure, which do you record?
The robot misjudges depth against a glass container, closes its gripper short, and comes away with nothing. The symptom is in manipulation. The cause is in perception.
Rubrics differ. Some want the observed symptom, some want the root cause, some provide separate fields for each. What produces unusable data is annotators deciding individually which one is more interesting, because the batch then contains both conventions with no way to tell which was applied to any given clip.
Find the answer in the rubric before starting, and if the rubric is genuinely silent, raise it rather than resolving it privately.
When you cannot tell which stage failed
If you are stuck between two families, work backwards through the chain.
First: was the robot’s understanding of the scene correct? Check the perception data against reality. If the scene model was wrong, the failure is in perception regardless of what happened physically.
Second: given what the robot believed (even if that belief was wrong), was the planned route reasonable? If the scene was understood correctly and the path through it was bad, the failure is in planning.
Third: if perception was right and the plan was sensible, then whatever went wrong happened during physical execution.
This sequence eliminates one family at a time and produces a consistent answer. When the rubric asks for both symptom and root cause, the symptom is usually wherever the chain ended (execution) and the root cause is wherever it started (perception or planning).
Try It: classify the failure
For each scenario, classify it as perception, planning, manipulation, or instruction-action disconnect before opening the reveal.
Scenario 1. Glare on a countertop washes out a stack of boxes in the camera feed. The robot’s planned path runs directly into them because it never registered they were there.
Scenario 2. The robot correctly identifies a mug and a spoon on the table. It plots a route to the mug that swings unnecessarily close to the spoon, clipping it on the way.
Scenario 3. Perception and the planned path are both correct. The gripper closes on the target jar slightly off-center, and the jar slips free instead of being lifted.
Scenario 4. Instructed to pick up the green apple, the robot executes a clean, stable grasp, but on the red apple sitting next to it.
Scenario 5. A depth sensor misreads the distance to a glass beaker because the surface is transparent. The gripper closes several centimeters short of the beaker and comes away empty.
Scenario 6. Instructed to stack the red block on the blue block, the robot performs two flawless grasps and places, but stacks the blue block on the red one.
See answer
Scenario 1 is a perception failure. The scene was misread before any plan was made. The robot acted sensibly given what it believed; what it believed was wrong.
Scenario 2 is a planning failure. The scene was understood correctly. The chosen route through it was the problem.
Scenario 3 is a manipulation failure. Perception and planning were both right. The physical grasp itself failed to establish control.
Scenario 4 is an instruction-action disconnect. The execution was competent. The mapping from instruction to target object was wrong, which a taxonomy that only watches the motion would miss entirely.
Scenario 5 is a perception failure with a manipulation symptom. This is the ambiguous case the module covers directly: the visible event (gripper closes on nothing) is in manipulation, but the cause (bad depth estimate against glass) is in perception. Record whichever one the rubric asks for, and if it asks for both, perception is the root cause here.
Scenario 6 is an instruction-action disconnect. Every motion primitive succeeded and the objects ended up stacked. The spatial relationship specified in the instruction was inverted, which only shows up if you check the instruction against the result rather than only watching the motion.
Robot failure taxonomy flashcards
No established control means a failed grasp, not a drop, however the clip looks. A drop requires the robot had the object first.
Perception covers understanding the scene, planning covers the route through it, execution covers the physical attempt. The visible event is often downstream of the cause.
Every motion primitive can succeed and the episode still fail: asked for the blue cup, the robot flawlessly grasps the red one. The failure is in the mapping from language to intended action, not the physical execution.
A system that drops an object, notices, and retries is demonstrating error detection. One that completes a careful placing motion with an empty gripper has no idea the object is gone, identical outcome, very different capability.
A deliberate, careful approach is often preferable to a fast one. Tagging slowness as a defect trains a model to associate caution with error, the opposite of the intended signal.
Hold durations, idleness floors, and attempt counts vary by project. A duration floor is what separates a normal pause between task phases from a genuine stall.