CapraWorks · Robot learning
Policies come first. Morphologies later.
CapraMorph is a sim-first robotics research project. Control policies are developed and evaluated in MuJoCo against explicit exit criteria — hardware is the periodic materialisation of what survives the sim gates.
Phase 2 — picks it up, but cannot hold it: see log
01 · Method
Simulation is the primary instrument.
Development happens in simulation before anything is machined, printed or bought. The target platform is the SO-101, a low-cost desktop arm in the lineage of TheRobotStudio's SO-ARM100 — but the arm is the output of the process, not its starting point.
Train in sim
Policies are developed and evaluated in MuJoCo, where iteration is cheap and every run is measurable.
Gate on evidence
Each phase has explicit exit criteria, written down in advance. A phase either passes its gate or it does not.
Materialise periodically
Hardware builds happen only when the simulated evidence justifies them — the physical arm realises what already works.
02 · Sim gates
Most candidates don't make it. That is the point.
Candidate policies are trained in simulation and judged against predefined exit criteria. Most are discarded. The few that pass a gate define what the next phase — and eventually the hardware — has to support.
The gates are written before the runs, so a pass means something and a failure is information rather than embarrassment.
03 · Morphology
Hardware is downstream of policy.
The target morphology is the SO-101, a low-cost desktop arm descended from TheRobotStudio's SO-ARM100 design. It is deliberately modest hardware: cheap enough to rebuild, well-understood enough to model.
Physical builds are periodic, not continuous. When a policy survives the sim gates, the arm is assembled to meet it.
04 · Current status
The arm picks things up.
Phase 0 (simulation substrate) passed 2026-08-14. Phase 1 — reaching a point — passed 2026-08-16 at 86.6%. Phase 2, as of 2026-08-18: the arm reaches for a box from a random start, grips it and lifts it 92.5% of the time over 80 trials, by chaining a pose-reaching policy with a grasp routine proven in advance on 4,258 certified positions. We previously published 85% and then 88.8% here. Both were wrong: our scoring counted a single 10 ms instant of contact as a successful pick, where our own standard requires the box be held for 27 of any 30 consecutive control steps. On the correct criterion this phase does not pass its own gate. The log has the full retraction. The progress log has the video, the numbers, and the day spent certifying a grasp that turned out to be the box held by the fingertips — caught by watching the footage, not by reading the telemetry.
- Telemetry
- Machine-readable logs from every training and eval run.
- Video
- Rendered rollouts of the learned policy, straight from the simulator.
- Report
- Deterministic eval numbers against each phase's pre-written exit criteria.
Learned policies in simulation; no hardware yet — hardware is the reward for passing gates. There is no product, no mailing list and nothing to sign up for. The log says exactly where things stand.
05 · Contact
Correspondence
Questions about the project are welcome. The working repository is private for now.
hello@capramorph.com