Metric AI Lab
Metric is a Physical AI Research Lab. We are building the world's most accurate foundation model for robotic failure detection — a single evaluation layer that measures progress, flags success or failure, and explains what went wrong, for any robot.
Robotics today is fundamentally blind to its own performance. Teams rely on manual review, hand-crafted heuristics, and task-specific checks to know whether a robot succeeded, stalled, or failed. This patchwork doesn't scale, doesn't transfer across platforms, and leaves the most valuable signal in robotics — what actually happened — buried in raw episode data.
We're developing the first family of foundation models designed to understand robot behavior itself. Our approach treats every embodiment as equally important — robotic arms, mobile manipulators, and humanoids — enabling automatic progress measurement, success and failure classification, failure reason attribution, and precise failure timing, without building a custom evaluator for every task.
One Model, All Modalities
Instead of writing brittle, task-specific success checks, we deliver a single model that powers the full spectrum of robot evaluation workloads — from labeling thousands of teleoperation episodes overnight, to scoring policy rollouts during training, to detecting the exact second a grasp begins to slip on the production floor.
Our models are designed for real-world integration — fast enough to run alongside live robot operation, scalable to millions of episodes, and deployable in both cloud and edge environments. This means robotics teams can evaluate sensitive operational data in-house without sacrificing accuracy or speed.
Built for Real-World Impact
Knowing whether a robot succeeded — and why it failed — is mission-critical across the entire robotics stack. We're building for the full range of evaluation needs: data labeling and annotation for large-scale collection efforts, standardized benchmarking across platforms, reward modeling for policy training, and real-time success, progress, and failure monitoring in production.
The same capability extends further: anticipating dangerous states before they cause harm, and laying the actuarial foundation insurers will need to underwrite robotic deployments. Whether you're curating a training dataset, comparing two policies, or certifying a fleet of humanoids for deployment — this is the evaluation backbone Physical AI is missing.
Science Meets Engineering
Research excellence drives everything we do. We're building at the frontier of video understanding and embodied reasoning, combining large-scale multimodal learning with the robustness and latency required on real hardware.
Our approach is empirical and iterative. We measure success not only by benchmarks, but by how effectively our models improve real robot fleets, data pipelines, and training runs. The most important breakthroughs come from understanding what truly matters when judging behavior in the physical world.
We're building the ground truth layer for Physical AI — how the world measures, understands, and trusts robot behavior.If you're interested in joining us, get in touch:
hrant@metriclabs.ai
Hrant Davtyan, PhD
Founder & CEO