Axis Robotics has launched Axis Sim Dataset V1, one of many largest open-source simulation datasets for Franka arm manipulation, with the complete dataset, coaching code, and benchmarks publicly accessible. V1 is constructed from greater than 50,000 human-teleoperated simulation trajectories throughout 207 manipulation duties and 60,000+ scene variants on a simulated Franka Analysis Three arm.
This dataset drew over 160,000 downloads, making it probably the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with each end result open and verifiable.

Axis Robotics is constructing the last word compounding information engine for Bodily AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The corporate raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Community Ventures, 10Ok Ventures, and angel buyers.
A Guess In opposition to “Clear Knowledge Solely”
A standard assumption in robotics is that demonstrations have to be near-optimal to start with — filter all the way down to knowledgeable trajectories, standardize the setup, and discard something noisy earlier than it’s secure to mimic. Axis’s thesis runs the opposite approach: information high quality lives on the distribution degree, not the one trajectory. When a big and numerous sufficient crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives throughout coaching.
Axis Sim Dataset V1 places that thesis to a public check. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and power use, all collected by means of Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd slightly than a single knowledgeable workforce. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the College of Michigan, and different establishments.
Outcomes That Scale
On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Efficiency improves constantly as pretraining information scales from 25% to 100% of the dataset, with no saturation in sight, proof that the positive factors come from range and protection slightly than a one-off bump. The most important enhancements seem below digicam, sensor-noise, and structure perturbations, the precise axes Axis randomizes throughout era.

The workforce says V2 is already underway, scaling to 1.2 million trajectories throughout 1,200 duties, with cross-embodiment generalization and outcomes throughout a number of VLA fashions exhibiting that suboptimal simulation information trains strong insurance policies.
The Engine Behind the Dataset
The dataset is one output of a bigger, actively compounding information engine. The place a conventional information vendor collects to a hard and fast spec and stops, Axis makes use of mannequin efficiency and failure circumstances to find out what needs to be collected subsequent, so each coaching spherical informs the subsequent. That engine runs on a hybrid technique throughout 4 information strains, and all 4 now run at scale:
- Simulation: over 200,000 distributed contributors on Axis Hub, a top-Three dApp on Base, producing 4.7M+ trajectories throughout 13 embodiments.
- Selfish: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in actual properties and companies throughout 14 industries: 200,000+ hours already banked and rising by 4,000+ hours each day, with Vicon-verified hand pose.
- Loco-manipulation: 500+ hours combining mobility and dexterity on actual humanoids (Unitree G1, Booster T2) by means of hardware-agnostic teleoperation.
- Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge circumstances.
Each activity and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.
From Open Knowledge to Industrial Deployment
Past open-sourcing simulation information, Axis works straight with robotic embodiment corporations to construct personalized, embodiment-specific information pipelines and mannequin priors.
As Booster Robotics’ first sim-data companion, Axis rebuilt Booster’s actual workspace as a task-aligned digital twin, had distributed contributors gather 42,000+ simulation episodes on it, and distilled them right into a Booster-specific mannequin prior. With simply 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box π0.5, matching π0.5 utilizing half the real-world demonstrations.
Different companions span embodiment corporations (Feagine Robotics), mannequin corporations (Manycore Tech, Dexmal) and industrial automation (Lotus Automobiles, Geely Auto). Axis additionally provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Bodily AI’s Knowledge Basis
“The way forward for Bodily AI isn’t a static dataset you obtain as soon as,” mentioned Chris Feng, founding father of Axis Robotics. “It’s an engine that retains producing the information the mannequin wants subsequent. Scale will get you broad protection. Variety retains the noise unbiased. The closed loop turns each failure into progress. That’s what compounds.”
Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who’ve scaled shopper platforms to over 30 million customers. Its analysis is suggested by Jiachen Li, Assistant Professor at Georgia Tech.
Paper Hyperlink: https://arxiv.org/abs/2607.21588
Mission Web page: https://axisaiorg.github.io/AXIS-V1/
Dataset Hyperlink: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Training
BlockmanPR Read More








