Leela Yanamaddi
September 28, 2026

Most teleoperation data fails for one simple reason: it looks complete, but it cannot teach, test, or explain anything with confidence.
If I had to boil this article down, I’d say this:
In other words: if your data is not time-aligned, complete, and labeled with enough detail, you do not have a training set. You have stored files.
A few numbers stand out:
The main takeaway is simple: keep fewer bad episodes, route more episodes to the right use case, and judge data by downstream model results - not by raw volume.
That is the frame I’d use for the rest of the piece.
The main problem usually isn't volume. Teams may collect hours of teleoperation data and still end up with training data that falls apart later. The damage often hides until model training or deployment. In practice, three issues show up again and again: poor demonstrations, missing state, and broken timing or logs.
A finished task isn't always a good training example.
If a warehouse mobile manipulator takes a jerky path to pick a tote, that run doesn't teach clean navigation. It teaches detours. If a factory arm opens and closes its gripper three times before it finally secures a part, that run doesn't teach firm grasping. It teaches hesitation. Behavior cloning doesn't clean this up for you. It copies what it sees. So once bad habits get into the dataset, the model treats them like normal task behavior.
Late-shift data is a good example. Operators are often slower, they make more corrections, and they miss grasps more often. That pattern can slip by unless the logs include operator ID and shift context.
Inconsistent operator style makes things worse. In a factory insertion task, one operator may line up the part by sight and move in slowly. Another may take a faster diagonal path and depend on contact force. If both runs get the same "success" label, the model sees two different action patterns but gets no clue about what matters. It can't tell whether the difference comes from intent, task demands, or operator skill.
That's how a dataset can grow large and still point the model in the wrong direction. It helps to separate:
When those cases get mixed together, behavior cloning gets contaminated. And even 5% erratic episodes can make a behavior-cloning model occasionally repeat erratic behavior.
Video-only data, or sessions with partial logs, often look usable at first glance. During training, they usually aren't.
Without joint state, gripper state, end-effector pose, force/torque, and task context, a model can't reliably link a robot action to the visual change that followed. That's the core issue: missing state breaks the connection between what the robot saw, what it did, and what happened next. Two camera frames can look the same while the robot is in two completely different physical conditions.
Synchronization errors add another problem. If camera, state, and action streams drift apart, the model starts pairing the wrong observation with the wrong action. Even small drift can hurt policy performance.
Then there are weak labels. Suppose a factory dataset marks both "failed insertion" and "successful insertion after three retries" as "success." Now the model, or even the evaluator, can't learn the difference. The same thing happens when task names are too vague. A label like "pick item" without object IDs, station IDs, or clear success rules makes filtering shaky later on.
If the logs don't support a confident label, it's better to mark the episode as unknown than force it into a false binary outcome. If these fields don't line up, the episode can't support dependable training or diagnosis.
Variable latency can wreck the feedback loop between the human and the robot.
An operator sends a correction, gets no prompt response, sends another one, and then both commands land at once. The trajectory now shows oscillation, overshoot, and sharp reversals. That doesn't always mean the operator performed poorly. It can mean the interface failed. Stale video makes this even worse because the operator may be reacting to a scene that has already changed.
Research ties human-to-robot latency to more teleoperation errors and longer mission completion times. Delays above about 200 milliseconds can push operators into slow, cautious behavior that leads to worse demonstrations. The practical move is simple: use task-specific latency budgets and drop sessions that go past them instead of mixing them into normal training data. If latency swings during a session, exclude that session or tag it as latency-affected for separate robustness work.
Intervention logs can fail in a different way. A record that only says "human took over" is too thin to help. Without the trigger, handoff, robot state, and outcome, that log can't support recovery-policy training or operations diagnosis.
"human took over"
Think about a warehouse robot that drifts off its planned path. That could mean obstacle avoidance. It could mean a localization failure. Or it could mean an operator takeover. If all three end up treated like standard demonstrations, the model learns the wrong lesson. Latency issues can turn otherwise usable sessions into misleading ones.
A teleoperation session is only usable when the data are complete, time-aligned, and labeled for a clear downstream job. So the next move isn't collecting more sessions. It's sorting the ones you already have based on what they can safely teach.
A simple way to do that is to use three gates: technical, behavioral, and semantic.
Technical gates are the easiest to check because machines can verify them. Every stream - camera, joint state, end-effector pose, gripper state, commanded actions, and force/torque when it matters - should run on a shared clock or follow a documented sync method. Aim for about 10 ms of alignment. More drift than that can damage policy performance. Each episode also needs a unique ID, clear start and end markers, reset status, and a plain terminal outcome. Keep records of latency, dropouts, packet loss, safety stops, controller faults, and operator takeovers too. Those details explain what happened in the session.
Once the streams line up, the next check is the motion itself. Does it make physical sense, or does something look off?
Behavioral gates catch the problems timestamps miss. Look for joint-limit violations, collisions, impossible velocities or accelerations, excessive jerk, command reversals, stalls, abrupt pauses, gripper chatter, and big gaps between command and motion. A clean successful trajectory with no corrections can feed behavior cloning. But a trajectory full of unexplained pauses and repeated joystick reversals - especially when they came from poor teleop mapping - should be rejected or marked low quality, even if the task still finished.
If the motion looks plausible, there's one more check: do the labels still say something useful?
Semantic gates make sure the episode can be understood later. A plain "success" or "failure" label usually isn't enough. Better outcome labels include nominal success, success with correction, partial completion, recoverable failure, safety intervention, and operator abort. Each episode should also include a task ID, object and scene context, software version, robot ID, site, and operator setup. Without that metadata, filtering or auditing the dataset later becomes a mess.
After a session passes through gating, send it to the job it actually fits.
A clean, synchronized success belongs in training. That same kind of episode, if it came from a held-out robot, operator, or site, belongs in evaluation. Failures with a known cause - perception, grasp, planning, control, hardware, latency, or operator error - fit failure analysis or robustness research. Sessions affected by latency, safety interventions, or incomplete logs are better kept for operations analytics, where they help expose system issues without polluting a supervised training set.
A good working rule is simple: don't delete an unusual episode until you know why it looks unusual and where it might still help. Keep the raw data, store the quality status, and record the reason it was excluded. A safety intervention, for example, may be wrong for nominal imitation learning but still useful for training a takeover detector or measuring when autonomous behavior starts to become unsafe.
| Data Category | What Was Captured | Why It May Be Unusable | Minimum Acceptance Condition | Appropriate Downstream Use |
|---|---|---|---|---|
| Video-only demonstration | One or more camera feeds | No action, state, or reliable causal link to motion | Recoverable action/state alignment and task labels | Visual review or perception research, not standard behavior cloning |
| Unsynchronized multimodal recording | Video, robot state, and commands | Observations may be paired with the wrong actions | Verified clock alignment, drift bounds, and gap report | Reprocess if recoverable; otherwise quarantine or failure analysis |
| Clean successful trajectory | Complete observation–action sequence | Can still mislead if labels are weak or the scene is overrepresented | Plausible dynamics, verified outcome, and traceable metadata | Imitation-learning training |
| Successful trajectory with corrections | Task success plus operator adjustments | Corrections are learned as intended nominal behavior, not recovery | Marked correction events, control-mode log, and segment-level labels | Recovery learning, robustness research |
| Failure demonstration with a causal label | Failed attempt with diagnosed cause | Failure actions are harmful as positive imitation targets | Reliable failure phase, cause category, and outcome label | Failure prediction, recovery learning, evaluation, or robustness research |
| Latency-corrupted session | Commands, video, and robot response under delay | Action–state pairing is corrupted or delay is variable | Measured latency, packet/frame logs, and stable delay model | Latency analysis, control improvement, or latency-aware training |
| Human safety intervention | Operator takeover or emergency action | Human action reflects safety policy, not task policy | Explicit intervention event, reason, and pre/post robot state | Safety review, intervention modeling, operations analytics |
| Incomplete episode | Partial trajectory or interrupted streams | Outcome and terminal state are unknown | Documented stop point and valid partial-progress label | Partial-task analysis or quarantine |
Teleoperation Data Quality Gates: From Raw Sessions to Usable Training Data
Stop bad data while you're collecting it, not after training. The goal is simple: keep bad demonstrations, missing state, bad sync, weak labels, and latency out of training in the first place. That means putting checks in place before, during, and after collection.
Every collection session should begin with a documented preflight check. Before an operator records even one demonstration, confirm that each required stream matches the expected schema. Then run a known-event test: trigger a visible mechanical movement or LED flash and compare when it shows up across all streams. That quick test can catch camera-to-action offset and cross-camera misalignment before anyone wastes operator time. In plain terms, this is your first line of defense against missing state, bad synchronization, and weak logs.
If the preflight fails, block collection. Don't count on cleanup later to fix a bad camera setup or a missing sync signal.
During the session, track frame rate, drops, drift, latency, and occlusion. Then respond based on severity: warn, pause, or terminate the episode. Flag any episode where frame loss goes above 2%. For tighter tasks, teams have reported targets as strict as a frame-drop rate below 0.1% and camera-to-action synchronization below 5 ms.
Preflight catches the obvious problems. Review catches the ones that slip through.
Start with automated completeness checks on every episode. Check for expected files, readable media, monotonic timestamps, frame counts, synchronization bounds, valid termination markers, and truncated streams. After that, apply automatic quality scores for visual sharpness, exposure, occlusion, trajectory smoothness, excessive pauses, collision events, and control-mode changes. Then sample 10%–20% of episodes for human review of task success, protocol compliance, operator strategy, and whether the action matches the observed state.
Human reviewers should label exact intervals with standard reason codes such as "camera occluded", "late action label", or "successful but off-protocol."
At the dataset level, review task and outcome balance, operator representation, duplicate or near-duplicate trajectories, and distribution shifts across collection batches or sites. A sudden shift in action distributions can point to a changed camera, a new operator cohort, or an updated task protocol. It doesn't automatically mean the data got better. That's why filters should be tested on your own robots instead of copied from another setup.
After filtering, check whether the rules improve robot performance.
Keep a trusted holdout set that stays separate from all filtering and training decisions. That set should cover normal tasks, hard conditions, multiple operators and sites, and the failure modes that matter. Then test candidate data slices against it: train or evaluate comparable models on different subsets and compare task success rate, recovery rate, intervention frequency, and collision rate.
Score data slices by how much they help task success, not just by how clean they look. A strict visual filter might cut blur but still do nothing for task success or recovery. If that happens, it may be removing useful variation instead of protecting model quality. Adjust thresholds based on measured outcomes, and update acceptance rules through version control rather than informal judgment. If rejection rates climb above 15%, treat that as a protocol or operator-training problem, not a cleanup problem.
A lot of teleoperation data goes bad for a simple reason: the collection pipeline misses the signals that matter. The issue usually isn’t that operators did a poor job. So the practical question is this: which sessions should stay, and which should be quarantined?
Interventions and failures help only when the record shows when they happened, why they happened, and how the episode ended. A failed run might be a bad fit for training, but still useful for recovery analysis or day-to-day ops. The answer is to put quality gates in place, route sessions by use case, and keep full episode records with aligned commands, motion, observations, phases, and outcomes.
Use the checklist below to make that call batch by batch.
Apply these gates before you accept any batch. If a required check fails, quarantine the affected episodes instead of pushing bad data downstream.
| Gate | What to Confirm |
|---|---|
| Streams | All required sensor, camera, robot-state, command, and safety streams are present |
| Timing | Timestamps are monotonic; drift, dropped frames, and sync error stay within the task budget |
| Latency | End-to-end latency is measured; sessions outside limits are flagged |
| Boundaries | Episode start, task phases, resets, aborts, interventions, and termination are clearly marked |
| Outcomes | Outcome can be verified from task-specific criteria |
| Interventions | Takeover reason, failure state, operator correction, and outcome are recorded |
| Metadata | Operator, robot, environment, task variant, software version, and collection conditions are complete |
| Separation | Clean demonstrations, failures, recoveries, safety stops, and defective sessions are routed separately |
| Review | A representative sample from every batch is inspected with a consistent rubric |
| Proof | Retain only slices that improve a held-out metric |
Keep a gate only if it improves held-out success, recovery rate, or intervention frequency.
Audit it before it enters your pipeline. Trainable data should meet five criteria: viewpoint, action, completeness, synchronization, and privacy.
Make sure sensor streams and action labels line up within 10 ms, camera calibration is correct, and runs stay consistent from one episode to the next. Filter out high-jerk trajectories, missing steps, and label noise.
Each episode should include:
Exclude data with label noise, skipped steps, camera calibration errors, timing gaps, or sensor streams that drift out of sync. Action and observation timestamps need to line up within 10 ms. If they don't, the data can get corrupted.
You should also remove demonstrations with incomplete state-action records. That includes cases with missing joint states, proprioception, or force data. Low-quality trajectories should be filtered out too, especially when the outcome isn't clear or intervention logs are missing.
Prioritize timestamp alignment so your action and observation streams stay in sync within 10 ms. If you can, use hardware-based Precision Time Protocol (IEEE 1588). That tight timing matters more than it may seem at first glance. If timestamps drift, your data can tell the wrong story about what the robot saw and what it did.
Check camera and sensor extrinsic calibration daily. Even a small drift of 3 to 5 mm can spoil precision data. In a setup that depends on exact motion and exact perception, a few millimeters is enough to throw things off.
It also helps to run automated filters that catch bad data before it spreads through the pipeline. Focus on flagging:
Keep labeling consistent across the full dataset. And for every intervention, make sure the audit record is complete. That record should include the robot state, intended action, operator correction, and final outcome.