Simulation · Evaluation · Benchmarking · Post-training

Simulation & evaluation infrastructure for Physical AI.

Generate worlds, evaluate Vision-Language-Action policies on serverless GPUs, benchmark models head-to-head, and post-train, one reproducible loop from a task idea to a better-performing robot.

ROBIT live rollout, OpenVLA-7B evaluated on a LIBERO task, with success rate, control step, cost and a Pareto frontier

The bottleneck

Robotics models improve faster than teams can test, compare, and improve them.

Simulation, evaluation, data collection, and training still run as disconnected workflows, every iteration needs custom infrastructure and expensive hardware time.

01

Evaluation doesn’t scale

Custom scripts, hand-reviewed rollouts, and task-specific success criteria make it impossible to compare models consistently.

02

Simulation ≠ improvement

Policies run in sim, but failures rarely flow into better datasets or the next training run. It produces episodes, not learning.

03

Real-world testing is slow

Robots, operators, resets, and safety checks make it far too costly to validate every model checkpoint on hardware.

One platform, four connected stages

There is no universally best model, only the best one for your task.

01

Simulate

Generate simulation-ready tasks, worlds, and scenario variations across LIBERO, MuJoCo, Isaac Lab, and Genesis.

02

Evaluate

Run real policies in a closed loop on serverless GPUs, success, robustness, control rate, memory, and cost, with rollout video.

03

Benchmark

Compare models on a Pareto frontier, across embodiments and simulators, with transparent, priority-driven recommendations.

04

Post-train

Turn failures into curated data, fine-tune, and re-validate, one loop from a task idea to a better-performing robot.

New · Dataroom

Understand why policies fail, then turn that into better data.

Evaluation tells you a policy failed; the Dataroom tells you where and why. Import rollouts or any HuggingFace dataset, let an agent annotate phases, events, and failure modes, and curate the episodes worth retraining on, the missing bridge between a failed run and a better model.

New

Agent annotation

An agent watches each rollout, segments it into phases, marks grasp & release events from the state channels, and tags the failure type, with no hand-labeling.

New

Import any dataset

Pull LeRobot datasets straight from HuggingFace, including your own private & gated repos, and turn every episode into an inspectable clip.

New

Failure attribution at scale

Move past a single success number toward a quantified picture of where and why policies fail across many episodes.

New

Curate → retrain

Select the clips that matter, export a curated dataset, and feed the next training run, closing the failure → data → policy loop.

Bring your own data: connect HuggingFace New

Link your own HuggingFace account to import your private & gated datasets and push your recorded work to your own namespace. Your credentials stay yours, encrypted per account; shared model terms still run on the platform.

Decision-grade evidence

The numbers that decide deployability, not a single success rate.

Success + 95% CI

The bar, not just the mean, Wilson intervals at every sample size.

Control rate

Milliseconds per step, whether a policy can close a real robot’s loop.

Robustness Δ

Clean vs. seeded perturbations: action noise, occlusion, observation lag.

GPU memory

Peak VRAM per policy, the hardware tier it actually needs.

Cost / success

Dollars per successful episode, the number that decides deployability.

Reproducible

Frozen initial states plus a config hash make any run byte-identical.

How it works

Generate → Evaluate → Benchmark → Improve

01

Generate

A task, a world, and seeded variations.

02

Evaluate

Real rollouts on serverless GPUs.

03

Benchmark

Head-to-head on one frontier.

04

Improve

Post-train on what failed, re-test.

Built for the teams shipping Physical AI

A neutral platform, bring any model, any robot, any simulator.

VLA & foundation-model labs
Robot companies & integrators
Robotics researchers
Data & post-training teams

Evaluate your first policy on a real GPU.

Bring a model, pick a task, and get decision-grade evidence back, reproducible, and cheap enough to run continuously.