Leela Yanamaddi
October 3, 2026

I use human video to teach a VLA model what a task looks like - and teleoperation data to teach it how the robot should act. I start with a teleoperation-only baseline, then test human-to-robot sampling ratios of 1:4, 1:2, and 1:1. More footage earns its place only if robot trials show better results.
My checklist is simple:
My rule: <u>better video understanding is not proof of better robot control</u>. I approve the mix only after robot tests pass, then use failure logs to decide whether the next collection needs more visual context or more robot action data.
Human Video + Teleoperation: VLA Training Workflow
Once source roles are clear, bring both datasets into the same episode format. Start with the schema: alignment and masking work only when both datasets use the same fields.
Use one episode schema for every manipulation task. Keep the original instruction, normalize verb/object/source/destination/constraints, verify identities, and label matches as exact, parameterized, related, or unmatched. Set consistent episode boundaries, from approach through confirmed placement and retreat. Exclude reset footage.
Give every episode the same record format, including:
Leave missing fields blank. Do not replace them with zeros.
The shared schema makes later alignment and masking rules clear. Once the fields match, align time, viewpoint, and embodiment.
Correct each recording for clock offset, drift, frame rate, and command latency before matching episodes by task stage. Keep command timestamps separate from motion-onset estimates. Also keep observation-to-command pairs separate from checks that link commands to observed effects. Align independent episodes by approach, grasp, lift, transport, release, and retreat - not by frame number.
| Method | Inputs | Supervision | Main failure modes | Suitable use |
|---|---|---|---|---|
| Viewpoint canonicalization | Human video, camera calibration or estimated geometry, and a target camera convention | Images or clips transformed toward a common view | Warping artifacts, inpainted regions, distorted hands, and loss of depth cues | Visual representation learning and coarse human-robot visual alignment |
| Object-centric alignment | Object tracks, segmentation, keypoints, or estimated 3D object poses | Object-relative observations and stage correspondence | Identity swaps, tracking failures, and ambiguity during occlusion | Shared task understanding when cameras and bodies differ greatly |
| Hand-to-end-effector retargeting | Human wrist or hand pose, robot end-effector model, coordinate transforms, and workspace limits | Estimated robot end-effector trajectories | Infeasible reach, incorrect hand orientation, and mismatch in grasp geometry | Generating candidate robot actions from clear human manipulation |
| Inverse-kinematics conversion | Target end-effector poses, robot kinematic model, collision constraints, and joint limits | Robot joint targets or trajectories | No IK solution, singularities, collisions, and temporal jitter | Producing candidate robot trajectories after geometric alignment |
| Direct robot-state alignment | Robot camera frames, proprioception, actions, and synchronized timestamps | High-confidence observation-state-action tuples | Narrow visual coverage, sensor drift, and hardware-specific bias | Behavior cloning, fine-tuning, and policy evaluation |
Use the least aggressive alignment that works. Camera stabilization must preserve hand-object relationships. If it doesn't, favor object-centric alignment. Explicitly mark occluded segments and missing force signals.
Before accepting converted actions, check reachability, joint and velocity limits, acceleration, collisions, gripper compatibility, and trajectory continuity. An infeasible human path is not a robot target. Keep it as task context instead.
Then decide which aligned segments can supervise actions and which should remain context only.
Label segments as task-context, human-motion, or robot-action, and distinguish converted actions from measured robot actions. These labels determine which samples can support action loss, auxiliary loss, or evaluation only.
Use converted actions for auxiliary objectives only after validating the mapping and checking feasibility. Reserve robot-execution action loss for measured robot actions. Mask, exclude, or down-weight unclear segments. Keep correctly labeled robot failures for failure detection or recovery objectives, but never treat failed commands as successful demonstrations.
| Supervision tier | Representation learning | Action learning | Fine-tuning | Evaluation |
|---|---|---|---|---|
| Task context | Appropriate | No direct robot-action loss | Limited use for conditioning or auxiliary objectives | Human-video checks only |
| Human motion | Appropriate | Human-motion or embodiment-agnostic objectives only | Use cautiously with masking or source-specific heads | Motion prediction or alignment |
| Converted robot supervision | Appropriate | Validated, confidence-weighted auxiliary objectives only | Auxiliary objectives alongside measured robot data | Offline alignment metrics, with validation limits |
| Measured robot supervision | Appropriate | Required for robot-execution action loss | Primary source for robot-policy fine-tuning | Required for robot execution evaluation |
Only robot execution proves robot-policy success.
Use a teleoperation-only baseline as your reference run. Robot trajectories teach control, while human video provides context and supports auxiliary objectives. Track robot-data exposure separately from total compute.
With action labels separated from context, decide what each data source should teach the model.
Match demonstrations by instruction, object interaction, goal state, or motion primitive - not just task label. Then choose whether human footage shapes representations before or alongside robot-supervised control.
| Schedule | Implementation complexity | Possible gains | Main risks | Suitable dataset conditions |
|---|---|---|---|---|
| Human-video pretraining → robot fine-tuning | Low to medium | Builds visual, language, and task representations before robot data teaches control | Human-specific viewpoints or motion patterns may transfer poorly | Large human-video corpus, limited high-quality robot data, and reliable task or language annotations |
| Source-aware joint training | Medium to high | Both sources shape shared representations; measured robot actions teach control | Loss imbalance, negative transfer, and human-video dominance can weaken control | Well-matched demonstrations and separate source losses or masks |
| Interleaved training → robot-heavy adaptation | Medium | Alternates human-context learning with robot-control updates, then focuses on execution | Schedule tuning and loss of useful human-video features | Mixed datasets with useful human semantics and robot data for precise control |
Apply an action-validity mask and normalize action loss by valid targets - not total frames:
[ \mathcal{L}{\text{action}} = \frac{\sum{e,t} m_{e,t},\ell(\hat{a}{e,t}, a{e,t})} {\max\left(1,\sum_{e,t} m_{e,t}\right)} ]
For human-only frames, set (m_{e,t}=0). Don't substitute a fake zero action. Human clips can still teach language grounding, temporal-order prediction, or object-state prediction. Give each objective its own weight and track its contribution so dataset size doesn't quietly dictate what the model learns.
Next, set source ratios that keep human video from overwhelming robot supervision.
Test human-to-robot sampling ratios of 0:1, 1:4, 1:2, and 1:1 as ablations. State whether each ratio counts episodes, clips, frames, or optimizer updates. Start conservatively when annotations are weak or embodiments differ. Increase human exposure only when held-out robot success or transfer improves without unacceptable declines in completion time, action smoothness, or collision rate.
Record the exact training mix so you can trace robot gains to the data rather than extra compute.
Use explicit source quotas or a weighted sampler. Within each source, balance tasks, objects, scenes, operators, embodiments, and outcomes. Log source counts, valid action timesteps, robot-supervised updates, and compute. Also log train/validation overlap by task, object, scene, operator, embodiment, and outcome. Match robot exposure where possible. If human training adds compute, run a compute-matched robot-only control to separate the effect of human context from extra optimization.
After training on the data mix, test robot execution on unseen scenes and object instances. The goal is better robot control, not just better representations.
Lock the test set before training. Hold out entire sessions, scenes, and object instances, and keep each trajectory intact. Use a separate validation set to select models and tune thresholds. In a fixed evaluation manifest, record task IDs, instructions, camera configurations, and success rules. Include both in-domain and transfer settings.
Test teleoperation-only, mixed-data, and mixed-data checkpoints fine-tuned on robot data under identical deployment conditions. Keep hardware, camera placement, initial object arrangement, instructions, lighting, controller rate, timeout, reset procedure, and safety constraints the same. Report robot fine-tuning steps and duration separately.
Treat human-only checkpoints as robot policies only if they produce validated, executable actions. Otherwise, use them only as representations or initializations.
| Evaluation area | Teleoperation-only | Mixed-data | Mixed-data fine-tuned on robot data |
|---|---|---|---|
| Task success | Completion rate, pick success, placement accuracy | Same metrics | Same metrics after adaptation |
| Generalization | Exact-match and held-out transfer success | Related-task, new-object, scene, instruction, viewpoint, and embodiment success | Same metrics after adaptation |
| Safety and recovery | Interventions, recovery success, collisions, force violations, emergency stops | Same metrics and limits | Same metrics and limits |
| Efficiency | Execution time, path length, action smoothness | Same metrics and timeouts | Same metrics and timeouts |
| Robot data needed for a given success rate | Robot-demonstration count | Count at matched robot-data budgets | Count including adaptation data |
Define completion as grasping the target object, carrying it without dropping it, and releasing it within the target region. Report assisted completions separately from unassisted successes.
Log model version, task variant, scene and object IDs, timestamps, observations, actions, outcome, completion time, safety events, and failure reason. For repeated trials, report raw counts and 95% Wilson confidence intervals. When attempts are clustered, use resampling by scene or session. For skewed execution times and placement errors, report medians and IQRs.
Test how the mixed-data model performs when the task, object, scene, instruction, viewpoint, or embodiment changes. Change one factor at a time before testing combined changes. Disclose any controller calibration or retargeting used for cross-embodiment transfer.
Run one-change-at-a-time ablations for human sampling, language normalization, viewpoint alignment, retargeting, confidence filtering, and robot fine-tuning duration. Keep all other settings fixed. Report success, transfer, interventions, and safety - not just the aggregate score.
Set acceptance limits before inspecting results. Suggested starting gates - not universal standards - are:
Require no observed critical safety failures, but recognize that a small test cannot establish zero risk. Test bounded disturbances within predefined workspace, speed, and force limits. Stop immediately for critical events. Mark uncertain results inconclusive rather than accepting transfer gains that hide reliability or safety regressions.
After alignment, sampling, and evaluation, approve the data mix only when every check passes: tasks and instructions belong to the same task family, records are synchronized, labels can be traced to their sources, supervision tiers are kept separate, embodiment mappings are validated, missing labels are masked, and sampling ratios and source weights are documented.
Make the deployment decision separately, after controlled trials. Require measured transfer gains over teleoperation-only training, plus compliance with the preset in-domain and safety limits. Any hard safety failure is a no-go.
If the model fails, use its logs to shape the next collection brief. Identify missing objects, scenes, task stages, language variants, embodiment configurations, and recovery behaviors. Assign each gap a required supervision tier and data source. Collect visual context where coverage is weak, and robot-labeled trajectories where control or recovery needs correction.
For managed collection, define tasks, language prompts, viewpoints, metadata, quality checks, privacy and consent requirements, and pilot batch acceptance criteria before scaling. Use those criteria to accept the dataset or request targeted collection.
Watch for imprecise or jittery control, failures from viewpoint mismatches, and trouble recovering from mistakes. These signs may point to too much human video and too little robot-specific action supervision.
To check, review your data for synchronization errors, label noise, and calibration drift. If performance is still poor, prioritize high-quality teleoperation data. It provides the motor command ground truth that passive human video lacks.
Start with 30 to 50 episodes to set a VLA training baseline before scaling. Trust comes from data quality and repeatability - not a fixed number of trials. Use one successful run as your first quality benchmark, then compare later runs for consistency.
As your dataset grows to hundreds of episodes, keep it consistent and follow strict recording standards. Include full sensor logs, intervention markers, and recovery behaviors.
Use egocentric video to scale visual pretraining, teach task structure, or record a range of surroundings at a lower cost. Prioritize **robot action data - specifically teleoperation - **for high-fidelity control signals, grounding, closed-loop control, and precise recovery behaviors.
Let the model’s failures guide what you collect. If it struggles with perception or scene context, collect more video. If it struggles with execution, precision, or recovery, collect targeted teleoperation demonstrations that address those specific failures.