Leela Yanamaddi
October 8, 2026

Robots are doing paid work - but “deployed” doesn’t mean they work without human help. We compare seven programs using records available through October 8, 2026. Before you buy, ask for intervention rates, recovery times, staffing needs, and uptime, not just task counts.
GXO reports more than 100,000 totes moved by Digit. BMW reports more than 90,000 components moved during its 2025 Figure 02 pilot - but that does not confirm continued use in 2026. Across these cases, we found too little public cost and staffing data to calculate ROI.
Quick Comparison
| Program | What the records support | Task and human role |
|---|---|---|
| GXO - Digit | Commercial warehouse work | Tote transfer; on-site safety monitoring and help when tasks fail |
| BMW - Figure 02 | Completed 2025 production-line pilot | Sheet-metal placement; on-site supervision and recovery |
| Mercedes-Benz - Apollo | Factory pilot | Parts and tote handling; employee-led training through teleoperation |
| Jabil - Apollo | Announced factory validation pilot | Planned inspection, kitting, and assembly tasks; human support undisclosed |
| Amazon - warehouse fleets | Working fulfillment-center systems | Inventory and cart movement; fleet supervision, remote help, and local recovery |
| DHL - Locus AMRs | Production use at more than 40 sites reported in 2026 | Robots carry totes; people pick, verify, and handle problems |
| Tesla - Optimus | Internal testing and training | Reported parts handling; supervision details undisclosed |
Our focus is simple: <u>what work is documented, who helps when it stops, and what results you can check</u>. We separate production work from pilots, tests, and plans - and treat missing data as unknown, not proof of autonomy.
Human-in-the-Loop Robotics in 2026: Production vs. Pilots
GXO’s Digit deployment is the first verified warehouse case. Digit works in commercial service at GXO’s SPANX facility in Flowery Branch, Georgia. The site started with a proof-of-concept pilot in late 2023 and moved into regular operations on June 5, 2024. GXO announced a multi-year Robots-as-a-Service agreement on June 27, 2024.
Digit has a specific, limited job: AMRs deliver totes, and Digit unloads them onto a conveyor for downstream packing. It also temporarily stages overflow totes until capacity returns. This is bounded tote handling - not evidence that Digit picks, packs, or performs general warehouse labor. Public records describe the task more clearly than the staffing needed to support it.
Public descriptions say Digit works behind safety barriers, with an on-site human monitor handling exceptions and supervising safety. They don’t disclose the monitor-to-robot ratio, remote support, maintenance coverage, or how often people intervene.
Buyers should ask how the site handles failed grasps, delayed or misaligned AMR arrivals, and conveyor congestion. The public record doesn’t show the full exception workflow.
Agility reported more than 100,000 totes moved by November 2025 and about 98% task accuracy. Those figures show continued use, but task accuracy isn’t the same as availability. Public sources don’t disclose robot count, throughput, uptime, or the amount of human support required.
Request interventions per 1,000 totes, average recovery time, monitor workload, and throughput compared with the manual process. Public sources don’t report labor hours saved, injury reduction, payback, or RaaS pricing.
BMW moves the comparison from warehouse tote handling to line-side automotive assembly. Its August 2024 announcement described a multi-week test of Figure 02 at Plant Spartanburg in South Carolina. The robot inserted sheet-metal parts into fixtures while BMW assessed safe integration. BMW had no timetable for permanent deployment.
BMW and Figure later described a 2025 active-line pilot. Figure 02 retrieved sheet-metal parts from racks or bins and placed them into welding fixtures with millimeter-level precision. Industrial robots handled welding and downstream feeding.
Production-line activity is documented, but continued operation in 2026 is not established. BMW and Figure’s 2026 Figure 03 sequencing project is separate. It should not count toward this deployment or serve as evidence that Figure 02 operated in 2026.
Figure defines interventions as pauses or resets. The documented support is on-site supervision and exception handling; remote teleoperation is not documented. Staffing ratios, recovery times, and escalation procedures have not been publicly disclosed.
Buyers should ask who monitors the cell, handles physical recovery, and approves restarts. They should also check whether support staff are dedicated to the cell or shared across shifts.
BMW reports more than 30,000 BMW X3 vehicles supported, more than 90,000 components moved, about 1,250 operating hours, and about 1.2 million steps over roughly 10 months. The reported schedule ran Monday through Friday. A 10-hour shift does not mean 10 hours of uninterrupted robot runtime.
The 84-second cycle requirement, 37-second load time, 99%+ success target, and zero-intervention goal are targets - not achieved averages. Before treating the deployment as stable at scale, request cycle-time distributions, intervention logs, and line-stop minutes.
Mercedes-Benz and Apptronik announced an Apollo pilot in March 2024. Mercedes-Benz later named the Digital Factory Campus in Berlin-Marienfelde as its production-environment test site. This is a pilot, not a scaled rollout.
Apollo is intended to move components and kitted totes to the line and inspect parts, while employees handle skilled assembly. On-site data collection and employee-led training make this a training-heavy pilot - not a hands-off deployment.
The documented human role is training. Mercedes-Benz employees train Apollo through teleoperation and augmented reality, but supervision and recovery responsibilities remain undisclosed. Sources also don’t report staffing levels, operator-to-robot ratios, or routine remote-supervision policies.
Buyers should verify who clears safety stops, handles abnormal loads, and authorizes restarts. Early Apollo systems operated inside a protected zone with external sensors and light curtains. They paused when someone crossed the boundary. Mercedes-Benz has not disclosed the site’s full safety case or recovery procedures. Those support arrangements matter because the program’s operating scale is also undisclosed.
Public sources do not establish robot count, operating hours, throughput, intervention rate, labor savings, or ROI for Mercedes-Benz’s Apollo program.
Jabil moves Apollo from customer-site testing to factory validation on the manufacturing side. On February 25, 2025, Jabil announced that it would manufacture Apollo and put newly built units into a factory validation pilot. The evidence supports a validation pilot - not verified routine production in 2026.
Announced tasks include inspection, sorting, kitting, line-side delivery, fixture placement, and subassembly. Public materials do not identify a facility, city, line, or business unit.
Jabil and Apptronik have not disclosed how human intervention, teleoperation, exception handling, or safety supervision would work. Buyers therefore cannot verify who steps in, when they do so, or how much staffing and recovery time the workflow needs. This makes scale data more useful than the announcement’s wording.
The sources provide no robot count, operating hours, throughput, intervention rate, quality results, labor savings, or production outcome. Manufacturing Apollo at scale remains a stated goal, not a verified result.
Buyers should request cycle time, first-pass quality, and human-assisted actions before treating the pilot as production-ready.
Amazon moves the comparison from single-robot pilots to fleet-level supervision in working fulfillment centers. The verified cases covered here include Sequoia in Houston, Shreveport, and warehouse workflows with fleet supervision.
Sequoia was retrieving inventory for employees in Houston, Texas, in October 2023. Shreveport, Louisiana, opened in 2024 with robots helping move inventory, sort items, and package orders. Proteus moves carts toward outbound operations, while employees still pick items at robot-supported stations.
The question isn't whether robots run. It's how much human help they need when something goes wrong.
Historical exception-handling example: A 2019 North Haven report showed local exception handling: a worker recovered a fallen item while nearby robots slowed or rerouted under safety controls.
A November 2022 fleet-supervision study identifies Amazon as a commercial user of intermittent remote human assistance when robots face risk or cannot progress. The reviewed sources do not disclose Amazon's supervisor ratios, intervention volumes, response times, or restart procedures.
Shreveport reportedly uses about 10 times more robotics than earlier advanced Amazon facilities. Sequoia's capacity exceeds 30 million items. The Times reported a 25% drop in cost-to-serve after the facility opened in May 2024.
To compare operations, ask for staffing by function, intervention frequency, response time, and fallback behavior when communication is lost.
DHL moves the comparison beyond single-site pilots to a multi-site AMR network.
DHL has used Locus AMRs since 2017. Locus reports active fleets at more than 40 DHL-managed sites in 2026 - a production network, not a pilot. Associates pick and verify items, while robots carry totes and move orders through the warehouse.
The announced 5,000-robot target is planned expansion, not verified deployment.
People still handle picking, verification, and exceptions across the network.
Associates follow robot prompts. Supervisors use dashboards to monitor orders, picks, and performance. Public sources do not disclose intervention rates, escalation steps, or staffing ratios.
Confirmed site milestone: DHL’s Toledo, Spain, facility completed the partnership’s 500-millionth Locus-assisted pick on May 18, 2024, involving a consumer home-goods product. DHL’s June 2024 release reported deployments at more than 35 sites.
To estimate staffing needs, buyers should request pick rates from comparable sites, exception volumes, robot utilization, and training time.
Tesla’s Fremont case is an internal factory deployment, not a customer-site deployment. Its 2026 disclosures describe an initial Optimus production line in space formerly used for Model S/X manufacturing. Early units are used internally to collect training data and develop functionality. This supports internal testing, not verified production work.
Public reporting mentions battery-cell sorting and parts handling, but other tasks remain unverified. Tesla has not published a full task log or independently audited performance data. These tasks therefore remain reported or in training. Tesla says the initial factory skills are basic and will expand over time.
At Fremont, the main question is how much human supervision those tasks require.
Reports describe robots working in controlled, supervised areas, but Tesla has not disclosed its supervision model. It has not published an operator-to-robot ratio, shift coverage, intervention frequency, or recovery procedures.
Buyers should ask who handles failed grasps, misplaced parts, perception errors, blocked paths, e-stop recovery, and safety-zone violations. These are points to evaluate - not documented Tesla workflows. The missing detail is how people respond when the robots need help.
Tesla’s reported Fremont capacity of up to 1 million robots per year is rated capacity - not verified output or deployment scale. Public sources do not confirm robot count, throughput, uptime, labor savings, safety results, or intervention rates.
For staffing and intervention workflow, buyers still need answers to three questions: who steps in, when, and how often.
Human support starts with three questions: who monitors the robots, who can step in, and what public sources tell us about those roles.
The comparison below focuses on a cost buyers often underestimate: human intervention, not robot capability.
| Deployment | Human role | Intervention trigger | On-site or remote | Disclosed staffing data | Intervention authority |
|---|---|---|---|---|---|
| GXO - Digit | Safety monitoring and exception handling; direct teleoperation undisclosed | Full trigger list undisclosed | On-site monitor; remote support undisclosed | Monitor-to-robot ratio and shift coverage undisclosed | Restart and command permissions undisclosed |
| BMW - Figure, Spartanburg | Supervision and exception handling during the 2025 active-line pilot | Pauses or resets; detailed triggers undisclosed | On-site; remote teleoperation undocumented | Staffing ratios undisclosed | Recovery and restart permissions undisclosed |
| Mercedes-Benz - Apollo | Employee-led teleoperation and augmented-reality training; production support unverified | Training activities documented; live intervention triggers undisclosed | On-site training; remote support undisclosed | Undisclosed | Recovery and restart permissions undisclosed |
| Jabil - Apollo | Validation pilot announced; human-support model undisclosed | Undisclosed | Undisclosed | Undisclosed | Undisclosed |
| Amazon - warehouse robot fleet | Fleet supervision, intermittent remote assistance, and historical local exception handling | Risk or inability to progress; site-specific triggers undisclosed | Remote assistance and local recovery documented; current site assignments undisclosed | Supervisor ratios and shift coverage undisclosed | Site-specific restart permissions undisclosed |
| DHL - Locus AMRs | Associates pick, verify, and handle exceptions; supervisors monitor dashboards | Detailed exception triggers undisclosed | On-site associates; remote support undisclosed | Staffing ratios undisclosed | Control and restart permissions undisclosed |
| Tesla - Optimus | Training and data collection in internal development - not verified production support | Production intervention triggers unverified | Undisclosed | Undisclosed | Undisclosed |
Undisclosed means the reviewed public sources don’t provide that detail. Task results alone don’t establish who can step in or authorize a restart. Public sources usually confirm the task, but leave the intervention workflow unclear.
Buyers should check exception handling directly: who receives alerts, who assists remotely or responds on-site, who authorizes recovery, and who logs the event.
Request video and telemetry that let your team diagnose failures, along with explicit command permissions and logged control handoffs. Test control lag under load and robot behavior during network loss. Local interlocks, emergency stops, and protective responses must work independently of the remote connection.
Measure intervention frequency, recovery time, and peak concurrent alerts. Don’t infer staffing needs from robot count.
These measurements shift the focus from what robots can do to how often people still need to help.
Repeated output tells us more than an announcement, but it still leaves gaps. The table separates reported results from undisclosed metrics. The next comparison is straightforward: which deployments report enough scale to judge human support needs?
| Deployment | Reported scale | Measurement period | Operational result | Human-support metrics | Evidence strength | Operational advantage | Constraint |
|---|---|---|---|---|---|---|---|
| GXO - Digit, Flowery Branch, GA | Active units not disclosed | Commercial operation documented; exact measurement window not disclosed | More than 100,000 totes moved; 98% handling accuracy reported | Intervention rate, recovery time, and staffing not disclosed | Commercial operation documented; results vendor-reported and tracker-cited | Repeated output demonstrated | Hourly throughput and uptime not disclosed; results limited to specified tote tasks |
| BMW - Figure 02, Spartanburg, SC | Simultaneous active-unit count not disclosed | 10-month pilot; 10-hour weekday shifts; approximately 1,250 operating hours | More than 90,000 components moved; more than 30,000 BMW X3 vehicles supported; approximately 1.2 million steps | Intervention rate, recovery time, and staffing not disclosed | Strong completed-pilot evidence for the production-line pilot | Sustained output demonstrated | Uptime, cycle time, reject rate, and maintenance burden not disclosed; pilot does not establish current fleet scale |
| Amazon - warehouse fleet supervision | Not disclosed | Not disclosed | Site-level throughput not disclosed | Intervention frequency, recovery time, and staffing ratios not disclosed | Throughput, staffing, and recovery metrics not disclosed | Not quantified | Fleet totals cannot establish support workload or site-level performance |
| DHL - Locus AMR workflows | Not disclosed | Not disclosed | Site-level throughput and uptime not disclosed | Staffing ratios and recovery metrics not disclosed | Throughput, staffing, and recovery metrics not disclosed | Not quantified | Results depend on picking, replenishment, traffic, and exception handling |
GXO’s cumulative tote count and accuracy don't establish hourly throughput, intervention rate, or recovery time. Likewise, BMW’s scheduled shifts and operating hours don't establish uptime. Separately, a tracker reported 65,000 operating hours and nine committed Agility customer-facility deployments as of May 2026. That does not mean nine verified active sites, and those hours are not GXO-only.
Production status alone doesn't tell buyers how much support a deployment needs. Amazon and DHL lack disclosed throughput, staffing, and recovery metrics, which limits comparison with the quantified humanoid results. Output for a specific task also doesn't establish broader autonomy. Ask for records of failed grasps, misplaced parts, blocked routes, and stopped cycles. Digit’s reported November 2025 safety evaluation shows that an assessment took place - not that unrestricted shared-space operation was approved or that incidents fell by a measured amount.
The public record does not support an ROI calculation. Sources do not consistently disclose purchase or lease charges, integration engineering, tooling, facility changes, safety validation, connectivity, software, local and remote labor, training, maintenance, spare parts, or downtime costs. GXO’s multi-year commercial arrangement does not disclose contract value or payback. Compare automotive handling, tote transfer, and AMR picking separately, using matched task, shift, quality, and support logs. These missing details feed directly into the buyer checklist in the final section.
Across the cases above, documented human support matters more than robot branding. Look for recurring production work backed by records - not just the word “deployed.” Require details on the operator, site, robot, operating period, recurring task, and human-support model. Keep production work separate from pilots, validation runs, demos, and simulation studies.
Request robot- and shift-level logs that show commanded versus completed tasks, autonomous versus assisted completions, exception triggers, recovery time, downtime, and safety stops. Match those logs to staffing and training records, and define whether human corrections count as success. Use this evidence standard when comparing sites.
Calculate staffing needs from exception rate × average handling time. Then account for simultaneous incidents, response latency, travel, breaks, and safety coverage. Public evidence does not establish a validated robot-to-operator ratio that applies across the named deployments.
Require written procedures for safe shutdowns, remote diagnosis, physical recovery, lockout/tagout where applicable, and restart authorization. Test lost communications and overlapping failures. Specify who can stop the fleet, quarantine equipment, and authorize a restart. Remote support does not replace local safety and maintenance responsibilities. Without these procedures, reported autonomy does not mean the system is ready to operate.
For procurement, turn those records into acceptance criteria covering throughput, intervention limits, recovery targets, and staffing assumptions. Base approval on your own logs and recovery tests - not industry claims.
A successful demo isn’t enough. Verify that robots deliver repeatable results under normal operating conditions. Check connectivity across every zone and shift, integration with existing workflows, and trained teleoperator coverage whenever robots run. Safety procedures should be tested, uptime monitored in real time, and interventions logged and used to guide improvements.
Also check low-latency connectivity, end-to-end encryption, and role-based access control. Roll out in stages with clear gates for moving forward, and monitor field health signals before scaling.
Staffing depends on task complexity and autonomy maturity. Complex work often needs one operator per robot. With greater autonomy, one operator can supervise a small fleet - or dozens of robots - and step in for exceptions or recovery.
Set staffing levels around response needs and escalation paths, rather than a fixed ratio. Give each role a clear scope: local operators handle simple stops, technicians address repeated faults, and remote teleoperators manage complex edge cases.
Budget for human support as a core operating cost, not a temporary backup. Cover trained teleoperators and supervisors during operating hours, connectivity and monitoring tools, intervention routing and logging, plus safety training and safeguards.
Offset these costs with measured gains in uptime and faster recovery. Log interventions and use that data to improve autonomy. This can reduce future human assists and long-term labor needs.