Leela Yanamaddi
October 6, 2026

A dataset spec tells your data provider what to deliver - and tells you how to check it. I recommend agreeing on counts, labels, privacy rules, and pass/fail tests before collection starts.
For Physical AI projects, I use the spec to answer five questions:
For example, a project might require automated checks on 100% of files and independent labeling of 10% of clips. These are project requirements - not universal standards.
My rule: <u>approve a pilot before scaling collection</u>. A spec helps you reject unusable data early, but it does not guarantee model performance.
For Physical AI datasets, the spec must turn task goals into fields you can test and enforce. Classify every field as required, optional, conditional, or prohibited. Required fields are acceptance criteria - not extra documentation. That includes intended uses, usage rights, and acceptance tests.
Conditional fields apply only when relevant: audio metadata when audio is collected, object coordinates when detection is required, or consent records when people are identifiable. Prohibited content includes unredacted faces, readable badges, documents, screens, or other personal information without permission; unlicensed footage; and undocumented synthetic edits.
Define each field so teams can test it during collection, labeling, delivery, and acceptance.
Define the sample unit before setting counts. A complete picking episode isn't the same as a short clip. Specify minimum and target quantities, duration limits, and whether incomplete episodes count.
Set minimums or maximums by task, object category, contributor, facility, recording day, operating condition, and outcome. Record split proportions and deterministic assignment rules. Keep related episodes and near-duplicates in the same split, and hold out contributors or facilities when evaluation calls for it.
The thresholds below are project-specific requirements.
| Field | Status | Required value | Measurement method | Tolerance | Acceptance consequence |
|---|---|---|---|---|---|
| Complete episodes | Required | 10,000 minimum | Count valid episode IDs in the manifest | 0 below minimum | Reject or require replacement |
| Clip duration | Required | 10–60 seconds | Compare decoded timestamps | Up to 1% outside range | Quarantine outliers; reject if over 2% |
| Coverage quotas | Required | At least the approved minimum for every coverage group | Group manifest by quota fields | No group below minimum | Require supplemental collection |
| Required metadata | Required | All required fields populated | Validate against schema | No missing required fields | Reject affected records |
| Sensor alignment | Conditional: synchronized sensors collected | Maximum 50 ms offset | Compare synchronization markers | Maximum 50 ms offset | Reject or request reprocessing |
| Cross-split duplicates | Prohibited | No duplicate episodes across splits | Hash and perceptual-similarity checks | 0 cross-split duplicates | Rebuild splits and investigate leakage |
Use these rules to verify sample counts and split boundaries before annotation starts.
Specify the required container, codec, resolution, frame rate, color space, timestamp precision, and synchronization limits. The schema must define naming, stable IDs, data types, units, allowed values, and null rules. Require a manifest with paths, file sizes, checksums, split assignments, and schema versions.
Every file should trace back to its capture session, device, and environment. Provenance should include collection date and time, facility or environment ID, pseudonymous contributor ID, device and equipment versions, protocol and software versions, source record, redaction status, consent reference, license or permitted-use category, and retention or deletion status when applicable.
Audio and depth metadata are conditional on collecting those streams. Secondary scene descriptions are optional.
Display human-readable dates as October 6, 2026. Store machine timestamps in time-zone-aware ISO 8601 format, such as 2026-10-06T14:30:00-04:00. State units explicitly and format quantities as 10,000 or 2.5.
Require a versioned label dictionary for task actions, objects, hand or gripper identity, locations, and outcomes. Define boundaries that annotators can observe - for example, contact begins at the first frame of physical contact. State whether events can overlap and how to record occlusion, uncertainty, and unknown values. Unknown is not negative. Object tracks are conditional on localization needs; labels outside the approved ontology are prohibited.
Require annotator training, a double-labeling rate, an adjudication owner, and measurable agreement targets. Double-label 10% of clips, require outcome-label Cohen’s kappa of at least 0.85, and limit median boundary disagreement to five frames. Retain confidence, pseudonymous annotator IDs, review status, and correction history. Define confidence separately from measured accuracy.
This example is illustrative, not a universal schema. Its timestamps describe the labeled event interval; additional fields depend on the training objective.
{ "clip_id": "wh03_s014_c0021", "task": "single_item_pick", "start_time": "2026-10-06T14:30:12.400-04:00", "end_time": "2026-10-06T14:30:19.967-04:00", "action": "lift", "object_id": "obj_0042", "outcome": "success", "confidence": 0.94 }
Use these definitions to tie labels to the task rather than each annotator's interpretation.
These rules turn the earlier spec fields into pass/fail checks for dataset delivery.
For first-person warehouse video and similar physical AI datasets, measure quality by how reliably each clip supports the task. Check technical integrity, visibility, completeness, label agreement and boundary accuracy, coverage, and capture consistency.
For every requirement, specify the check, method, target, tolerance, reviewer, evidence, and consequence of failure. State whether each threshold applies to a clip, session, task, or full delivery. An overall pass can hide a serious failure in one task or location. Use stricter limits when errors could affect safe robot behavior. ISO/IEC 5259-2 defines a data-quality model and measurable characteristics for machine-learning data. ISO/IEC 5259-4 covers validation against specified targets.
Run automated checks on 100% of the delivery. For manual review, sample the larger of 200 clips or 2% of the delivery, then add targeted checks across tasks, cameras, warehouse zones, worker groups, and dates. Oversample low-light footage, occlusions, dropped frames, and clips flagged by automated tests.
The annotation lead should confirm that task-critical actions remain visible and episodes are complete. Record sample IDs and selection rules. A failed sample must trigger a batch-level investigation - not just fixes to individual clips.
Set privacy and safety rules before recording. Approve camera placement and recording areas, verify permission and workplace authorization, and specify how to handle faces, voices, badges, screens, barcodes, shipping labels, and customer information.
Require automatic redaction before delivery, followed by human verification. Set access controls, retention and deletion deadlines, and incident-reporting requirements.
A rights register must link each source and contributor to evidence covering ownership, training, evaluation, commercial use, modification, sublicensing, redistribution, geographic limits, and expiration. NIST’s AI Risk Management Framework connects these controls to privacy, intellectual-property rights, and documented data use.
Quarantine data with unknown rights or unresolved personal information. Do not allow conditional use.
Once quality, privacy, and rights requirements are set, use the same spec for acceptance testing. The supplier provides validation reports; the buyer independently reruns the tests. Apply the spec’s approved thresholds - not exceptions introduced after delivery. “Spec limits” means the exact target and tolerance set before collection.
| Requirement | Validation test and reviewer | Pass threshold | Evidence | Remediation |
|---|---|---|---|---|
| Schema and manifest | Data engineer validates rows and resolves paths | 100% valid; each row maps to exactly one media file | Validator output | Repair and resubmit |
| File opening and checksums | Data engineer opens every file and recomputes hashes | Zero unreadable files; 100% hash match | Media and checksum reports | Replace corrupted files |
| Media properties and timing | Data engineer checks encoding, duration, and timestamp relationships | All spec limits met | Automated inspection report | Reprocess or recollect |
| Duplicates | Data engineer checks hashes and perceptual fingerprints | Within-delivery duplicates below 0.5%; no duplicates without permission | Duplicate report | Remove duplicates and rebalance splits |
| Coverage quotas | Dataset owner compares counts by approved coverage group | Every mandatory minimum met | Coverage report | Recollect missing strata |
| Split leakage | Data engineer checks identities, sessions, and similar sequences | Zero prohibited overlap | Leakage report | Rebuild and retest |
| Annotations | Annotation lead audits labels and independent reviews | Approved agreement and boundary limits met | Adjudication records | Relabel affected batches |
| Human spot checks | Annotation lead reviews random and risk-based samples | Approved defect limit met; no missing task-critical labels | Signed review records | Expand audit and correct affected batch |
| Privacy | Privacy reviewer checks redaction and exposure risks | Zero unresolved critical exposures | Redaction logs and audit | Quarantine, redact, or delete |
| Rights and provenance | Rights reviewer matches files to permissions | 100% documented rights for intended use | Rights register and supporting agreements | Quarantine or remove unsupported files |
Define critical, major, and minor defects, along with contractual deadlines, to determine whether a batch is accepted, quarantined, or resubmitted.
Allow partial acceptance only for partitions that pass independently. Require new versions, checksums, and regression tests for corrected deliveries. Name the buyer’s final sign-off owner and waiver authority, and specify who pays for rework. Retain scripts, review decisions, and exception approvals.
Warehouse Video Dataset Coverage Plan
This example turns the earlier spec fields into a warehouse-picking request.
Hypothetical spec: Specify first-person warehouse video for action recognition and temporal segmentation; treat next-action prediction and workflow evaluation as secondary uses. Do not classify it as robot-control data unless robot commands, state, and calibration streams are collected separately.
Hypothetical coverage plan: Request 12,000 clips across navigation (2,000), item identification (2,000), reach-and-grasp (2,000), lift-and-place (2,000), scanning (1,500), and recovery/failure (2,500), with at least 25% cluttered scenes, 20% difficult objects, 20% low-light clips, and 30% combined failed, interrupted, or abandoned attempts. Count each clip under one primary task while retaining overlapping event labels; coverage groups may overlap. Set separate minimums for successful, incorrect, failed, interrupted, and abandoned attempts. Distribute coverage across shelves, bins, totes, pallets, carts, narrow aisles, and staging areas, including fragile, deformable, reflective, transparent, and oversized objects. Report both clip counts and recording hours.
Meeting the quotas isn't enough. The camera must also show the actions needed for the task.
Hypothetical recording plan: Require head-mounted RGB video at 1,920 × 1,080 and 30 frames per second. Document field of view, stabilization, exposure, motion, and audio; keep hands, target objects, and destinations visible. Define each session by one contributor, site, camera setup, and operating context. Exclude unsafe recordings and clips missing task-critical views. State which extra streams are required and synchronize timestamps across them.
Define events using the following observable boundaries. Store timestamps in milliseconds from the beginning of the source video. Link each
video_idto metadata, annotations, contributor ID, site ID, and review records. Allow simultaneous events and mark visibility asvisible,partially_occluded, oroff-camera; useuncertain_boundaryrather than inventing precision.
Event Start End Pick Worker begins orienting or reaching toward the intended item Attempt’s outcome becomes observable Grasp Hand or tool first contacts the item Item is stably held, or grasp clearly fails Lift Item leaves its support surface Lifting motion ends or transitions to carrying Scan Barcode is presented to the scanner Success, failure, or abandonment is observable Place Item contacts its destination Item is released, or placement fails Recovery Corrective behavior begins after an error Task resumes, succeeds, or is abandoned
With coverage and recording requirements set, the buyer can check whether the delivered data supports the intended task.
Hypothetical buyer acceptance checklist: Verify every warehouse coverage quota and completed privacy and usage-rights reviews for each accepted session. Set media tolerances and annotation thresholds from pilot results and deployment needs. Before acceptance, verify grasp visibility, scan evidence, and prediction context without future-label leakage; correct and retest failed batches.
Use this checklist to turn the spec into a buyer brief and acceptance gate.
Put this checklist in the provider brief. Give each requirement an owner, a measurable target, a test method, and evidence needed to pass.
Use the approved spec as your acceptance baseline. It defines training-ready data in testable terms rather than vague claims of “high quality.”
Map target capabilities to task categories, then set quotas for each category, setting, and time of day. Let evaluation needs guide data volume: start initial training with 50–200 hours per major task category, and track both hours and episodes. Use a pilot with 5–10 contributors to plan a 20–30% buffer for lost or invalid data.
For annotations, aim for 200–1,000 examples per schema. Allocate 25–35% to difficult or rare classes, with extra emphasis on edge cases.
Assess usable supervision, embodiment fit, and technical quality. First-person video helps models recognize workflows and understand what they see. But it lacks the robot commands and measured contact forces needed to execute tasks. Control, grasping, and failure recovery require teleoperation logs that synchronize robot states, commands, and outcomes.
Check that the data matches your robot’s kinematics, sensors, and workspace. Verify that streams are synchronized, hands and objects are visible, and trajectories are complete. Also check that checksums, timestamps, and label references are valid.
Revise your dataset spec when pilot runs, sample reviews, or quality assurance checks reveal unclear labels, weak category boundaries, or inconsistent annotations. Update collection briefs and specifications, too, when model failures during training point to missing objects, scenes, or task stages.
Never modify an existing release in place. Issue a new semantic version tag so releases stay unchanged and loaders work across dataset versions.