TB-4162 · REV D · Technical newsheet
Industrial IoT & MonitoringDevice profile
Physical AI's Missing Verb: Why 'Verify' Belongs in the Definition
A BMW humanoid placed 90,000+ parts at 99% success, but success rates don't equal tolerance conformance. SkillReal argues physical AI needs a fourth verb: verify, with the gap in acceptance criteria.
By Amara Osei3 min read627 words
Features
- BMW Spartanburg humanoid: 10 months, 90,000+ components placed, >99% success per shift — a pose/execution metric, not a tolerance-conformance metric
- A robot repeatable to 0.1 mm typically holds absolute accuracy of only 0.5–1.5 mm; fixture wear, incoming variation, datum stack-up, and thermal drift widen the gap
- Proposed verification method: measure print-called critical characteristics on a full shift of parts and compare conformance rate against the system's reported success rate

Every published definition of physical AI uses the same three verbs: machines that perceive, reason, and act in the physical world. Shai Newman, CEO of SkillReal, argues that a production floor needs a fourth — verify — and that its absence from the definition has measurable consequences for anyone qualifying automation for release-critical work.
The gap shows up in a widely cited deployment. A humanoid robot at BMW's Spartanburg plant ran for ten months and placed more than 90,000 components at better than 99% success per shift. That figure describes execution — the robot completed its programmed motion and the part went in. It says nothing about whether the critical characteristics on those assemblies held tolerance, and nobody publishing the number claimed otherwise.
The two metrics diverge for unglamorous, well-understood reasons. A robot repeatable to 0.1 mm typically holds absolute accuracy between 0.5 and 1.5 mm. Fixtures wear. Incoming parts carry their own variation. Datums stack up. The cell drifts thermally across a shift. A machine can return to its programmed position every cycle and still leave an assembly out of specification.
This is the classic repeatability-versus-accuracy distinction that metrology labs have managed for decades, now relocated from CNC encoders and CMM programs to learning-based manipulation. Repeatability describes the spread around the robot's own mean position; accuracy describes the offset between that mean and the true coordinate frame established by the part's datums. Grippers, compliant wrists, and vision-guided pose estimation add error sources that a pose-based success metric never sees, because the metric is measured at the robot, not at the feature.
Newman's proposed test makes the distinction operational. Pull a shift of parts from the robot-fed station. Measure the critical characteristics the print calls out — not the robot's pose. Calculate the conformance rate and set it beside the success rate the system reported for the same shift. Where the two diverge is what the definition left out.
Most plants never run that comparison, and the reason is economic rather than conceptual: measuring a full shift of parts on a coordinate measuring machine costs more than the answer seems worth. Sampling plans trim the cost but leave the comparison statistically soft. In-line dimensional inspection, Newman argues, has changed that arithmetic — closed-loop, 100%-coverage measurement at production tact makes the shift-level conformance-versus-success comparison routine rather than a special study.
The argument lands hardest on procurement. Acceptance criteria for automation cells conventionally specify pose repeatability, mean time between failures, and cycle time. None of those predicts first-pass yield on the product. If the cell's incoming material varies, if the fixture wears at a known rate, if the thermal envelope shifts between the second and seventh hour of a shift, the conformance rate carries that variation forward into the plant's cost of quality. The robot's self-reported success rate does not.
There is also a traceability question. A system that logs what it did — commanded poses, gripper states, cycle completion — supports one class of audit. A system that can report what it produced — conformance against print callouts, per feature, per shift — supports release decisions and corrective action. Newman's closing formulation compresses the point: physical AI earns its place in production when a machine can report what it produced, not only what it did.
For engineers writing acceptance test procedures for robot-fed stations, the development raises a concrete question: will the fourth verb appear in the spec? Until conformance data sits beside success data in the acceptance criteria, the gap between 99% execution and actual tolerance hold remains an unmeasured liability — one that the standard three-verb definition of physical AI has no vocabulary to describe.
via Automation World (Source)
Filed under
- physical-ai
- robotics
- metrology
- dimensional-inspection
- conformance
More from Amara Osei
Show full bio
Senior reporter covering industry trends and analytics at Testbench Report.
22 articles
Application notes
- Robot Inspection System Aims at Zero-Downtime AI Vision QA
- Machine Vision's Post-IMTS 2026 Question: System, Not Sensor
- Industrial Machine Vision Cameras Headed for $5.61B by 2035
- Machine Vision Moves From Defect Detection to Autonomous Control
- Tech Briefs Convenes Executive Roundtable on AI in Machine Vision