Leela Yanamaddi
October 1, 2026

I’d choose egocentric video to teach task understanding, teleoperation logs to teach robot control, and both when you need the two to work together. First, define what your model should produce: task steps, motion paths, or gripper commands.
Here’s how I’d make the choice:
I’d also check labels, robot fit, sensors, data quality, safety, and cost before scaling. More footage isn’t the same as more usable training data. In the article’s illustrative $18,400 pilot, 500 recorded episodes at an 80% usable yield work out to $46 per usable episode.
Quick Comparison
| Criterion | Egocentric video | Teleoperation logs |
|---|---|---|
| Main use | Workflow and visual understanding | Robot execution and recovery |
| Action supervision | Inferred or annotated human actions | Logged robot commands |
| Contact feedback | Usually inferred from video | Measured if force or tactile sensors are installed |
| Robot fit | Needs action mapping and testing | Closer fit when collected on the target robot |
| Quality checks | Visibility, blur, labels, task boundaries | Those checks plus timing, calibration, state, and command logs |
| Cost planning | Include annotation and action mapping | Include robot access, operators, sensors, and review |
My next step: <u>test 30–50 episodes per major task variant</u>, then scale only if the data improves the behavior you need at a cost you can support.
Egocentric vs Teleoperation: Choose the Right Robot Training Data
Sensor data and labels are not the same thing. Camera frames and encoder readings are recorded measurements. Hand poses, task boundaries, and success labels may be estimated or annotated rather than recorded directly by sensors.
| Egocentric video | Teleoperation logs |
|---|---|
| Action labels: Task steps are inferred or manually annotated. | Action labels: Commands are logged directly; measuring motion requires separate sensors. |
| Sensor channels: RGB, plus optional audio, gaze, IMU, depth, or hand tracking. | Sensor channels: Cameras, joint states, gripper state, timestamps, end-effector pose, and optional force/tactile signals. |
| Contact information: Contact may be inferred visually; force and slip are generally unmeasured. | Contact information: Force or tactile measurements require sensors and synchronized logging. |
The distinction matters: observation, control, and feedback data support different learning targets.
First-person footage shows workflow order and object interactions across varied settings. But it does not provide robot commands or measured contact forces.
Occlusion, blur, and camera motion can hide the grasp point. Human reach and finger movements may also exceed what a robot can do.
Seeing an action is one thing. Turning that footage into robot action is another.
A useful log includes camera streams, operator commands, joint positions and velocities, gripper state, end-effector pose, timestamps, and outcomes.
Recording force, torque, tactile, and operator-side haptic signals requires instrumentation, calibration, and logging. Neither method guarantees complete, synchronized, training-ready data.
For teleoperation, check clock synchronization, camera calibration, operator latency, and controller limits.
Access to the target robot reduces embodiment mismatch - the gap between the body used to collect data and the robot learning from it. Still, changing the gripper, camera placement, or controller can limit transfer.
The next step is to match these signals to the task the robot needs to learn.
Choose data based on the behavior you need to supervise - not its file format. The table below matches each learning target to a data source. It gives you a starting point, not a complete training solution.
| What behavior do I need to supervise? | Best starting point | What it teaches | What it still misses |
|---|---|---|---|
| Workflow recognition | Egocentric human video | Activity labels and step boundaries | Robot actions, control timing, and state |
| Visual pretraining | Egocentric video, optionally mixed with robot video | Object, scene, and language-aligned representations | Reliable visibility through motion and occlusion |
| Robot control | Target-robot teleoperation logs | Commands aligned with observations, state, and outcomes | Coverage beyond demonstrated tasks |
| Grasping | Teleoperated robot demonstrations | Approach, closure, grasp outcomes, and corrections | Transfer across different grippers and reach limits |
| Force-sensitive manipulation | Teleoperation with force or tactile sensing | Contact, slip, and force-based corrections | Contact signals beyond the installed sensors |
| Failure recovery | Teleoperation with failures and corrections | Retries, recovery actions, and safe stops | Recovery scenarios absent from collection |
For warehouse workflow learning, label item identification, item handling, tote placement, and misplaced-inventory correction with timestamped boundaries, object identities, and outcomes. These labels support activity recognition and task decomposition. Test on new items and layouts.
These semantic labels describe the task, but they don't provide executable wrist trajectories or gripper forces. Control still requires action inference, retargeting, or robot demonstrations.
When the task shifts from understanding a scene to acting on it, switch from human video to robot state, commands, and feedback.
For grasp training, log observations, robot state, commands, and outcomes across approach, grasp, lift, and release, with synchronized timestamps. Label whether the object stayed secured, slipped, or was dropped, and retain failed attempts with corrections or safe stops so the policy can learn recovery. Insertion and fragile-object handling need force or tactile sensing. Record contact onset, load changes, and corrective motion alongside commands and tool pose; verify that the signals detect slip or excessive force when the policy must react. An outcome label alone does not explain when or how to adjust.
If one source isn't enough, pair human video with robot demonstrations. The robot demonstrations connect representations learned from human video to executable actions.
Shared representations need explicit embodiment mapping. Simply pooling recordings doesn't guarantee transfer between robots with different kinematics, sensors, or action spaces. Validate the representations from the target robot’s camera view, then test closed-loop execution on its hardware.
Create a collection brief that covers the task goal, start and end conditions, allowed objects and layouts, target robot and end effector, workspace, cycle time, action format, required sensors, coverage target, acceptance rules, and budget. Check that the data includes timestamped commands, robot states, gripper actions, success/failure labels, recovery actions, operator interventions, and termination labels.
Check for embodiment gaps in reach, joint layout, dexterity, camera placement, workspace limits, approach angle, and contact mechanics. Human video requires action inference or retargeting; robot demonstrations directly supervise execution. Choose the source that fits your needs, and budget for retargeting.
Use the deployment robot when possible. Otherwise, use one with documented equivalence in joints, end effector, sensors, control rate, and workspace. Treat retargeted human actions as hypotheses until you validate them on the target platform.
Once labels and embodiment fit are set, lock in the sensors and quality gates needed to support them.
Start with the sensors available at deployment. Add cameras to understand the workflow, and use proprioception, force-torque, or tactile sensing for grasp and contact. Add audio only if the task or debugging goal calls for it.
Save calibration versions, camera intrinsics and extrinsics, robot-to-camera transforms, frame rates, exposure settings, sensor serial numbers, and synchronization diagnostics. Set pass/fail gates before recording: complete start/end markers, visible task objects, synchronized streams, intact robot-state and action logs, and acceptable motion blur.
Set task-specific numeric thresholds, such as less than 1% missing state samples.
Reject trajectories with joint-limit violations, controller saturation, collisions, or implausible pose or velocity jumps.
Label episodes as successful, failed, interrupted, or recovery attempts only, and keep a rejection log. Record consent, data rights, retention, access controls, and required masking. Before collection, set rules for emergency stops, speed and force, exclusion zones, handling, and maintenance.
Test the collection spec in a small pilot before spending more.
Price the full workflow - not just operator time. Before committing the budget, check robot availability, operator capacity, facility access, annotation effort, and rare-case coverage. Start with 30–50 episodes per major task variant. The budget below is an illustrative pilot estimate, not a market quote.
Illustrative pilot budget: U.S. labor, existing robot access, 500 recorded 10-second episodes, two cameras, and an 80% usable-data yield.
| Cost category | Pilot assumption | Estimated cost |
|---|---|---|
| Hardware and mounting | Cameras, lighting, mounts, operator interface, basic safety equipment | $8,000 |
| Operators | 80 hours at $35/hour, including training and breaks | $2,800 |
| Annotation or retargeting | $6 per episode for labels, segmentation, or human-to-robot mapping | $3,000 |
| Calibration and synchronization | Initial calibration plus verification after setup changes | $1,500 |
| Maintenance and consumables | Gripper wear, fixtures, cables, repairs, and replacement objects | $1,200 |
| Storage and processing | Approximately 100 GB of compressed episode data plus backups and preprocessing | $300 |
| Quality review | 10% second-person review at $40/hour, including rejection analysis | $1,600 |
| Total | $18,400 |
Calculate unit costs separately for egocentric workflow clips, retargeted human labels, and teleoperation robot demonstrations. Counting only operator labor understates the cost per usable demonstration.
The checklist points to one rule: choose egocentric data for task understanding and teleoperation data for robot-specific execution and feedback. Use both to connect understanding with execution only when each source meets a separate need.
Judge the data by its usable supervision, embodiment fit, and required feedback - not its volume or collection cost.
Define the target output, then test the pipeline in a pilot. Scale only when the pilot shows that the data source improves the target metric enough to justify the cost per usable demo.
Human and robot bodies move differently, so copying joint angles won’t work. Motion retargeting translates human intent into the robot’s coordinate frame and joint space, often using inverse kinematics.
Filter out unstable trajectories and movements the robot can’t perform. Teleoperation works best for recording state-action pairs directly within the robot’s control loop. These pairs provide ground truth for imitation learning.
How much data your robot needs depends on the task and how mature your autonomy stack is. There’s no one-size-fits-all formula, but labs aim to bring teleoperation data down to 5% to 10% of total training data.
For new tasks, plan for 300 to 1,200 human demonstrations. Focus on quality, not just volume, and include robot edge cases and recovery actions. Intentional failure and recovery demonstrations should make up 5% to 30% of production training data.
Check your pipeline against your original quality requirements. Use your first successful run as the benchmark to see whether you can repeat the results. Look for missing data types, gaps in edge-case coverage, collection artifacts, and whether the pipeline works across different robots or environments.
Track human takeovers in intervention logs. Fewer takeovers for the same task variant are a primary sign of improvement. Use automated filters or manual review to remove weak episodes or give them less weight. Prioritize carefully selected, high-quality episodes over raw volume.