CapraMorph / log

Progress log

Phase gates, honest numbers, rollout videos — including the bugs. Every claim here is reproducible from the working repository's eval harness; nothing is rounded up.

2026-09-08 · hardware · the eyes

A measurement is not a confirmation.

A €55 webcam on a gooseneck arm ended four days of reasoning about a workbench from photographs sent after the fact. Every frame now burns its own timestamp, servo angle, temperature and torque state into the pixels — and the reason is the first thing the camera produced, which was a fault that didn't exist. An image captured at one moment, compared against an encoder reading taken minutes later while the lever was being moved by hand, showed ninety degrees of lever against three degrees of shaft. The conclusion — a hub slipping on its spline — was confident, mechanically plausible, and pure artifact. Captured together, the same test reads 90.7° across a 90° move. Image and pose are one measurement or they are not a measurement.

The servo ships no zero of its own: the encoder's origin is a factory arbitrary that changes every time the disc comes off the shaft. So the bench got a frame of its own — zero at plumb, positive as the lever rises, because gravity torque is then just mgL·sinθ, the shape every sweep and stall figure in this project already has. A spirit level anchored it, the anchors agreed to a tenth of a degree, and it lives in a version-controlled file rather than anyone's memory.

Then the beautiful calibration drove the arm into the cabinet. A spirit level proves levelness; it cannot tell you which of the two horizontals — one either side of vertical — is the one without furniture in it. A loaded sweep started toward the divider panel while the operator's hand went for the 12 V plug. Worse: that same run captured one camera frame out of seven, because the device wouldn't reopen between poses. The tool built specifically to stop blind motion ran blind. Both faults are now fixed in structure rather than intention — the calibration file names the forbidden side as well as the safe one, the camera holds a single handle for a whole run, and the mapper aborts outright if any pose yields no frame. If the watchdog can't see, the motion stops. The rule about confirming direction rather than inferring it had been written down a week earlier, after an earlier incident, and was bypassed because a measurement felt like a confirmation. It isn't one.

Run correctly, the same procedure produced the record that four ambushed sessions had needed: seven poses, seven witnessed frames, the weight hanging free at every angle. And the load telemetry carried a bonus. Plotted against the sine of the angle, the ratio falls and then flattens — the fingerprint of a constant friction term riding on gravity, rather than of a weight resting on something. It puts joint friction at about 0.10 N·m, and the simulator these policies grew up in models 0.052: our virtual joints are several times slipperier than the metal on the shelf. A sensitivity sweep across the measured range is running against the frozen policy now, to find out whether it ever noticed.

2026-09-05/06 · hardware · the instrument

The bench weighed its own chain, then called the datasheet a liar.

Day two of hardware: a €14 aluminum bar became a nine-hole ladder lever, a rated clevis chain replaced every improvised fastener, and Decathlon barbell plates became certified test masses. Two stepped sweeps at two known masses gave a differential calibration — the unknown rigging mass cancels out of the slope, then falls out of the intercept: 294 grams, measured by the bench itself. A kitchen scale arrives Wednesday to grade the prediction. The friction the same fit implies, 0.33 N·m, agrees with the bracket we measured a day earlier by hanging beer bottles. Independent instruments, one answer.

Two honest rejections along the way: the servo's current register produced negative 219 g of rigging — it's blind below ~150 mA, so the load register is the torque gauge now — and a whole sweep was discarded after video review caught the hanging chain riding the cabinet edge. An R² of 0.90 lost an argument with a phone camera, and deserved to.

Then the probe that rewrote the shopping brochure: climb a known 2.79 kg load until the servo sticks, and read stall torque from pure geometry. 1.62 N·m — barely half the “30 kg·cm” on the label. The trace shows current rising honestly to 806 mA and then collapsing to 120: firmware overload protection, not physics. Three EEPROM bytes gate the whole thing — exceed 80% output for two seconds and it clamps to 20% — and a fourth byte (integrator gain: zero) explains every degree of position droop we spent a day chasing. The rated number is a burst figure ships never meant you to hold.

Overnight the rig ran 12,237 loaded cycles in exactly twelve hours: zero faults, zero corrupted reads, zero thermal timeouts, settling at 59–61°C — two degrees under its own safety pacing — and finishing with better repeatability than it started (0.10°). The fastener life series closed: zip tie, 250 cycles; knotted cord, 1,987; metal clevis chain, 12,237 and unbothered. The servo is at ~26,700 career cycles, two full-speed crashes and one max-power mistake, and still hasn't been the thing that failed.

2026-09-04 · hardware · bench day

Measured with beer: the day the commands became physical.

The servos arrived, and the first bench harness worked against real silicon on the first try. By evening the sim's assumed bands were measurements: command-to-motion latency 5.9 ms (we'd assumed up to 20), deadband 0.40° (the floor of our assumed band), and the gear-lash operator our adversarial review had already doubted was retired — the control loop closes on the output shaft and drives straight through the slack.

Then a number no datasheet prints, bracketed with the only calibrated masses in the room: one 250 ml bottle hung on a lever holds forever on the unpowered gearbox's own friction, at zero current. Two bottles run it straight down, audibly winding against the gears. Passive friction sits between 0.30 and 0.60 N·m — a sixth to a third of our sim's whole torque ceiling — and the fleet doctrine writes itself: an unpowered joint is not a brake.

Two 30-minute loaded holds drew the static thermal curve at two operating points, an unpowered twin servo logged as ambient control both times. At a quarter of the torque ceiling: 39→41°C, mean 14 mA — thermally free, friction carries the pose. Then a mini wrench zip-tied along the warped lever as a splint (zero bow on camera) moved the cord to its jaw, grew the arm, and put ~45% of the ceiling on the shaft: 39→52°C in a textbook first-order climb toward a ~54°C plateau, mean 66 mA, control flat at 37. Real self-heating, bounded 15°C under the safety rail — and 0.00° of position drift across both runs, an hour of loaded holding without losing an encoder count. The climb to the test pose was filmed on a phone, and the video matches the load telemetry frame for frame: the load register drew a sine curve as the lever rose, exactly as the geometry says it must. If thermal derate bites, it bites moving, not posing — the cycling test is next.

Honesty section: the day also produced three incidents, all of them command-side, all of them ours. A sign bug turned −10° into 350° and whipped the first lever into a steel clamp. The slow-speed safety cap we added in response turned out to starve the position loop — a falling load stays ahead of a throttled setpoint, so the servo loses politely to two beers. And removing that cap to catch a falling load rang a full revolution at maximum power. The protocol that ended the streak is now law: every motion is declared and acknowledged before it runs, and loads are handed to a servo standing still, never dropped on one — the handover measured zero drift. The rig's second build survived the worst command at full power, which is the design standard now: the worst command eventually gets sent. Both wrecked levers were cut from packaging; both times the cheapest part broke and the gearbox didn't. The servo took two full-speed stalls into steel without a single fault flag.

2026-08-31 · the eyes · distillation

The student that beat its teacher.

With the composed pick finally measured honestly — pose policy travel plus grasp close, 71–76% end-to-end depending on the engine — we re-distilled the vision student from the new smooth teacher. From the wristcam alone it hit 90% on nudge grasps. The previous student managed one in twenty. Same architecture, same data budget; the only change was the teacher, and an action space where a slightly wrong velocity self-corrects instead of flinging the arm.

Then we put it in the full chain and it collapsed to 32.5%. The pose policy delivers the arm to states the student had never seen — its whole training diet came from one reset distribution, and deployment serves another. Two fixes went to measurement: proxy views (assist-delivered poses) did nothing; collecting the teacher's actions on the actual arrival states and fine-tuning on the success-filtered pairs lifted it to 47.5% without costing the standalone number a point.

The first arrival dataset never got used. Its teacher converted only 37% of arrivals, so the recordings were two-thirds failure trajectories — and cloning failure teaches confident failure. That number was its own finding: the checkpoint our weekend of automated polishing had produced was quietly worse on deployment states than its ancestor, 37% against 71%, invisible to every gate we ran because every gate started episodes the training way. Gates now include the deployment distribution, permanently.

Round two of DAgger, fine-tuned on the newest slice alone, was flat. The fix was embarrassing in hindsight: the algorithm is called dataset aggregation. Training on the union produced a number that made us suspicious for the opposite reason — 93.8% — and the audit found the contamination in minutes: the evaluation arrivals were the training arrivals, byte for byte, courtesy of a shared seed. On genuinely held-out episodes: 90.0%. Across five fresh seeds, 400 episodes: 88.5%, against the state-based teacher's 66–74% on the same arrivals.

The student beats its teacher because it trained where the teacher never went. And it carries one more property the teacher class never had: under 10–20 ms of loop latency — the delay real hardware will impose — every state policy we measured lost 20 to 50 points; the pixel student lost nothing. Vision commands change slowly frame to frame, so a stale action is still approximately the right one. A last attempt to push further — three more collection rounds, 69k pairs — made it 24 points worse, and was archived the same evening. The frozen file is called vision_v7_student.pt, and on Friday it meets its first real servo.

2026-08-28 · the action space · jitter

The shake lived above what the videos could show.

“Watch the videos. I'm assuming you haven't, right?” Correct — and it turned out worse than not watching. Control runs at 100 Hz; the films sample every fifth step, 20 Hz, Nyquist 10 Hz. The trained policy put 40% of its joint-velocity energy above that line — not hard to see on film, mathematically absent from it. Aliased into apparent smoothness on every frame anyone ever rendered.

The alarm built to catch this was the sixth instrument here caught lying. It reported the fraction of spectral energy above 2 Hz, and a fraction is scale-free: a motionless arm scored 0.86 and tripped it, so it fired on everything and meant nothing. Measured floor to ceiling: a still arm reads 0.003 rad/s; a parked arm fed 5% random command noise reads 0.77; our policy read 1.01 — more shake than deliberate dither. The alarm now speaks absolute rad/s.

Five fixes went to measurement and every one died: two command taxes were Goodharted (commands calmed 4×, physical motion tripled), a velocity tax priced speed instead of shake, a jerk tax bought a noise-level change for eight points of success, an action low-pass made the policy fight its own filter, and halving the control rate made the shake 71% worse. The pattern was the finding: you cannot tax what the action space makes free. Absolute position targets let one step command an arbitrary jump into a kp 998 servo — and our own archive, a campaign on the sibling task, had already written the conclusion we spent a day rediscovering: the jitter lives in the action space, not the reward.

Bounded delta actions — the target is the measured position plus a velocity-capped increment, the norm in every published SO-101 sim2real pipeline — made chatter inexpressible instead of expensive. Under pure random actions the bound alone cuts high-frequency shake threefold. Trained: 83.5% acquisition at 38% less shake, against 76.7% for the old action space — the first configuration in this campaign to win both axes at once. The ablation also executed its own gate term: a ported speed penalty cost 15 points and bought nothing, caught only because the sweep isolated it.

The honest ledger: the residual is amplitude-limited chatter, the habit bounded rather than cured. And a new latency knob priced transfer for the first time: one control step of loop delay costs ~20 points of success, two cost ~50, old policy and new alike — so the next run trains under randomized delay, and the bench servos arriving in September now have a second job: confirming the sim's assumed hardware bands before we trust anything trained inside them.

2026-08-27 · the pose bank · side grips

We were quoting one grasp and training another.

Dave asked why the claw keeps picking from above at awkward angles when the side grip is obviously better — it is how a human lifts a cup, and our own log had measured it: 15/20 without the post-grasp roll, 19/20 with it. So we measured the pose bank. The finger axis in all 655 certified poses sits a median 14.3° from straight down. Top-down, every one of them.

Then we measured the battery that produces the 19/20 headline this project has published for a week. It samples its own poses, behind a filter requiring the finger axis to stay within 25° of horizontal. Median 16°. A side grip. Overlap between the two distributions: 0 of 655. The number we report and the data every policy trains on were never the same grasp.

The constraint that would have caught it exists in the bank generator, parameterised and switched off — FINGER_AXIS_MAX defaults to 1.0, which accepts every orientation, under a comment reading “which nothing has ever constrained.”

Rebuilt on the proven filter: 353 side-grip poses, 351 of them perfect holds (99.4% against the top-down bank's 82%). The payoff was immediate and needed no other change. The end-to-end composed pick doubled, 26.7% → 53.3%, with the grasp stage going 30% → 62% — the seam we had spent hours debugging was never a reach bug; reach was already landing within 0.6 mm. Thin-margin top-down grips simply could not absorb that error. And the first policy trained on the new bank reached 66% acquisition against a previous record of 36.5%.

Two honest deductions. Part of that 66% is a better-posed task rather than a better policy — side-grip poses hold with more slack, so the comparison crosses distributions. The clean evidence is the composed pick, same pipeline, only the bank swapped. And it cost us something: far-range reaching, which had just reached 8%, fell back to zero. It was learned in the top-down world and did not transfer.

2026-08-25 · verification · the instruments

Five instruments, four of them lying.

Correction. An earlier version of this log's underlying notes claimed the training bank reproduced at only 21.7% through the replay path. That was an artifact of a certification script with a hardcoded bank path, which measured the same legacy file three times while appearing to compare three different banks. The true figure is 96.7%. The bank was never broken.

Four consecutive runs declined — 21.8%, 17.2%, 12.5%, 9.5% — under four different fixes, none of which was the problem. The cause was a single runaway parameter: the policy's exploration spread climbed from 3.9 to 465 and then into the millions. Training rollouts had been effectively random for days while the deterministic policy kept its skill, and every update sanded that skill down.

The first fix made it worse. Clamping the spread once per rollout sawtoothed against the optimiser's own update loop, pushing the divergence measure a hundred times past healthy and churning the aim away entirely: 0 successes in 200. Freezing the parameter instead — a constant distribution cannot sawtooth — restored learning on the next run.

Then the certification harnesses started disagreeing. One returned 0% for poses that are certified by construction, because it used the pre-fix open-loop lift. Another was the hardcoded path above. A third, the bank generator's own script, aborts outright on this claw. Of five measurement paths exercised in one evening, only one survived a control check, and only because the control was run first.

That is now mechanism rather than intention. A single verdict pipeline produces every run's result: it refuses to start without the full environment block, reports the optimised metric beside an ungameable physical witness, runs a frequency analysis that declares oscillation without waiting for a human to say it, raises an alarm when a proxy improves while reality worsens, and writes no verdict at all until the films exist. When a verification step keeps getting skipped, stop re-teaching it and make the tool refuse to proceed.

2026-08-23 · reinforcement learning · the reward

It learned to get paid without ever picking anything up.

The first real training run looked excellent: reward climbing 29 to 340 over ten million steps. The exam said 0 successes in 500 episodes, with the object's mean peak height exactly its resting height. It had never lifted anything.

Two causes, both ours. The curriculum's stage boundaries are counted in total steps, so a ten-million-step run lived entirely in the first stage and never met an exam-like start. And the reward paid for approaching and pinching every step with no bonus for finishing, so a flickering grip out-earned a firm one indefinitely.

Dave watched the rollout and said it was gripping with its fingernail. Measured: the policy's holds sat 56–63 mm from the grasp centre while certified grips sit at 4–21 mm. Later he watched again and said the jitter would disassemble something. Measured: bang-bang control, commands swinging most of their range at 100 Hz with the actuators saturated two-thirds of the time.

Taxing that jitter taught the policy to cheat the tax. Command swing fell four-fold while physical joint speed tripled — it had learned to pump the arm's own resonance with small, cheap pushes. The tax now applies to measured joint velocity, which cannot be gamed that way. We also discovered every film we had reviewed was playing 1.5× too fast, because the capture rate and the playback rate were set independently and nobody checked.

2026-08-22 · Phase A — the overnight run

The lift was fixed at 3 a.m., and it confessed to the tumbles.

Dave went to bed and left Phase A running autonomously. The target was a known bug: the scripted lift computed one Jacobian at the start pose and trusted it across 120 mm of travel, under-delivering up to 28 mm. The fix is closed-loop: re-measure the gripper's height every quarter second, fold the error back through a live Jacobian.

It was specced to rescue one pose — the perfect grip that topped out at 86 mm. It rescued five. The open-loop lift's sideways drag turned out to be the quiet cause of most of the "acquisition" tumbles and slides we had attributed to the claw itself. With the honest lift: 19/20 with zero tumbling and nineteen settling holds — no wrist roll needed, the calmest claw this project has produced. Across three fresh pose banks: 17–20/20 by bank, and zero tumbling on every bank tested.

Two humbler results, reported because the suite said so. The approach-angle cap we planned as a second fix is now redundant — capped and uncapped banks score the same, because the lift fix cured the exact failures the cap existed to screen. And a 20/20 scored on the primary capped bank reproduced perfectly but did not generalize (18 and 17 on fresh banks) — so the honest headline stays 19, not 20. The one pose that fails on the primary bank in every configuration is the chaos pose — the one that flips on bit-level noise. Deterministic control has hit its ceiling; that pose is the standing argument for the learned policy this project exists to train.

Housekeeping with numbers: the old grasp bank, re-certified against the final claw at rated torque, kept only 655 of 4,320 entries — it was built for a bare claw on an arm 1.75× too strong. The rebuild script refused to generate a new one ("0/500 pass — do NOT tune past this"), correctly: its bands describe a claw that no longer exists. The 655 survivors are the interim bank; the retrain comes next, and it inherits a gripper whose only remaining enemy is chaos.

2026-08-22 · Phase 2 — the claw closes, a policy opens

He bet it wouldn't work. It's the best result this project has.

Day three of the claw. First the design closed the honest way: a per-rib force census showed the rear ribs carry 49% of grip force, the static pads 50%, and the front row roughly zero — so we deleted the zero-load teeth one by one, battery after every cut. Two came out free. Deleting the whole front row kept the score but cost 8.5 mm of seating depth — those ribs steer, not grip. Deleting one rear backstop cost 1 mm of seat and improved residual twist by 4.5° — the telemetry curves caught what the aggregates rounded away. Final form: ten ribs, every one with a measured job, fat wedge bodies verified physics-neutral, 14/20.

Then Dave, who says he can't explain physics but sees it moving in his head, proposed rotating the wrist 90° after gripping — so the static finger swings underneath and gravity presses the object into a cradle instead of pulling it out of the pinch. We raced four timings. He predicted the immediate roll would fail worst. I predicted the late roll was safest. We were both wrong: rolling at 1.2–1.6 s scores 19/20 — rescuing both acquisition tumbles, both slide-outs, and the chaotic knife-edge pose that nothing else all week could tame, losing nothing.

The catch, on film: at t=2 s the cylinder is tipping off the ribs — the same failure we filmed all week. At t=3 s the rotating cradle arrives underneath and it reads HOLDING. It rides at 135 mm for the next two minutes.

He didn't believe it — "that's pure luck, there's no way that worked" — so it got the full treatment: a straight repeat (19/20, zero pose flips), a timing sweep (a plateau at 1.2–1.6 s, decaying to 16/20 at 2.0 s+ exactly as the interception window closes), and the killer — twenty brand-new poses the policy had never seen: 15/20 without the roll, 19/20 with it. Gained four, lost none. Mechanism, not luck.

One failure survives on each bank, and honesty requires saying they differ: on the original bank it's a perfect grip whose lift tops out at 86 mm against a 100 mm bar — the lift arithmetic is next, and on that bank it is now worth exactly 20/20. On the fresh bank the leftover is one tumble the cradle arrived too late to catch. Nineteen of twenty, twice, with the residue named both times.

2026-08-21 · Phase 2 — the pad shoot-out

The claw grips with shape, not stickiness.

Yesterday ended at a wall: at any clamp force the servo survives, the flat claw held nothing. Today a two-phalanx wrapping finger with shaped pads holds 15/20 for 120 seconds at the servo's rated 0.49 N·m — the torque it can sustain indefinitely, cold. The progression was all aim: 5/20 at the first guess, 12 after moving the placement, 15 after Dave watched a failure video and called the approach angle off by eye.

Then we raced five pad shapes — square, two rounds, chamfered, triangular — same 20 poses, same seeds, identical footprint and bite. Dave asked for the chamfer mock-up in one sentence: cut the two corners off the squares, give the object an extra face to land on. It won, 16/20, and the film shows why: the object ratchets inward one rib pitch at a time and locks into a pocket, twist falling from 21° to 8°. His other hunch — that the smaller round pads would grip better — lost honestly: they produced the only true grip failure of a hundred runs, ninety seconds of creep and then out, with barely any turn. Bite depth was everything.

Pose 11, chamfered pads, 120 s. Watch 20–45 s: the object walks inward in steps and locks. Every other shape froze where it first touched and sagged under the gate.

A contact census answered a question worth keeping: 100% of grip force goes through the pads — the claw body is never touched. Shrink the pads until the object reaches the claw and it collapses; coat the whole claw in rubber and nothing comes back. Remove one pad per face and no score moves, but settled holds start drifting — two or three contacts do the gripping; the rest are anti-twist bracing.

Two more phantom-class bugs, our own scene. The object geom shipped priority="1", which silently made MuJoCo ignore every pad friction value we ever authored. And the pads' torsional friction was written as 0.25 — that parameter has length units, so we had specified a 250 mm contact patch. With both fixed, all five shapes tie at 13/20: the score spread was riding millimetre margins at an arbitrary gate. What shape really controls is seat depth and twist — and the chamfer is the only profile best-in-class on both, in every friction regime we ran. Same lesson as the hull, one level up: ask what the engine actually uses.

And under the tie, the real verdict. The same seven poses fail in every shape — but the chamfer keeps six of its seven as near-misses at 84–90 mm against a 100 mm bar, one floor drop, while every other shape drops four of the seven outright. The remaining gap is a lift that under-delivers near singularities; fixing it rescues the chamfer to roughly 19/20 and rescues the others far less. Only the winner's failures are recoverable. Dave set the bar for this phase himself: we're not chasing 20/20 — humans drop things too. Sustained near-perfect, at a torque the hardware can actually live at, is the win. The pads are closed; the lift arithmetic is next.

2026-08-19 · Phase 2b — it grips, then drops

It was holding the box like a cone pointing at the floor.

Dave watched a rollout and said the claw grips with its tip, in a V, like picking up a pyramid from the point — and that the object would always fall out at that angle. Measured, all of it held.

Sweeping the actual 34 mm box through the gap: past 50 mm into the hand it does not fit at all, and the floor of the opening rises 34 → 50 mm over that span. That is a 22° wedge, so the deep wide part of the claw is unusable and every grasp this gripper makes is forced to be a tip pinch. It also explains a number that had been bothering us: 86 N of clamping force on a 0.49 N box, 175× its weight. A flat face cannot seat in a 22° V, so it rides on edges and only brute friction holds it.

And it does not fall sideways. Tracking the box after it lets go, it moves +38.8 mm along the fingers toward the tip in 13 of 16 drops, because the finger axis points downhill in 100% of our grasp poses — a median 23° from straight down. 0.92 g was quietly pulling the box out the open end the whole time.

Four fixes failed, including both of the obvious ones. Rotating the wrist to raise the fingertip made it monotonically worse (down to 2.5%); the rotation disturbs the grip more than the reduced gravity helps. Swapping the box for a cylinder — a V-groove being the classical way to hold one — lost badly too, because we have no groove: a cylinder pinched at a tip simply rolls out. The box's corners were what held it.

We thought the fix was lift speed — lifting six times slower took "still holding at the end" from 52.5% to 90.0% over 80 trials.

Correction, same day. That was not a fix, it was a slower failure, and Dave called it immediately: the slow lift only stalls the slippage. Running the same grasp for longer settles it — the box creeps steadily along the fingers and always comes out: at 350 steps 90% still hold, at 700 steps 60%, at 1200 steps 0%, with the box having crept 96 mm. The grasp is not stable at any lift speed. Our episode length was setting the score. The honest position is that this gripper cannot hold this box in a tip pinch, and the 92.5% we briefly published here measured how long we watched rather than whether it holds.

Identical episodes. Left drops it, right holds it. The fifth episode fails in both — this is an improvement, not a solved problem.

We also changed what counts as a pick. The old criterion accepted 0.3 s of holding anywhere in the episode, so an arm that gripped, lifted and then dropped scored as a success — which is exactly what the video showed. It now has to still be holding at the end.

2026-08-18 · Phase 2a — the shaking

Six fixes for a jitter that was never in the reward.

The arm reached its target and then kept twitching instead of holding still. The obvious suspects were the servos, so we measured them first: 0.000000 rad/s of drift under a constant command, 10–25% of available torque at every pose, 0.02° of sag. The hardware holds perfectly when told to hold. The policy was commanding the shaking — changing its command by 54–82% of the full action range every single step, with joints pinned at the extremes 37–100% of the time.

So we charged it for shaking. Six times, in six different ways: a rate limit on the commands, state-dependent exploration noise, an action-rate penalty at two strengths, a published smoothness regulariser, and a narrower exploration distribution. Every one either failed to change the arm's behaviour or bought stillness by preventing the arm from doing the task at all.

They failed for a reason worth stating plainly. For an arm commanded by absolute joint targets with no cost on motion, twitching between extremes is not a flaw in the training — it is the mathematically optimal policy. Pontryagin's maximum principle says so, and a 2021 NeurIPS paper confirms it empirically. We had been repeatedly asking the policy not to do the best available thing.

Changing what an action means fixed half of it immediately. Instead of "go to this joint angle", an action now says "move this far from where you are" — so a wild command can only ever move the arm a bounded amount. Settled joint speed halved, from 0.606 to 0.317 rad/s, with no smoothness penalty in the reward at all.

Same targets, same seeds. Left commands absolute joint angles and keeps twitching after it arrives; right commands increments and settles. Measured on the footage itself, the settled arm moves 3.7× less.

The honest part: this is half a fix, and it came with a bill. The stiller policy is less accurate so far — it arrives within 10 mm on 12% of attempts against the twitchy one's 23%. A change that makes the arm calm and worse is a trade, not a victory, and it is logged as one. But there is reason to think the trade is temporary and worth paying.

Correction, 2026-08-18. This entry originally claimed 88.8%, and the previous entry claimed 85%. Both were wrong. Our scoring counted a single 10 ms instant of contact as a successful pick, where our own published standard requires the box be held for 27 of any 30 consecutive control steps — a bar 27× stricter. Re-measured on the correct criterion, on identical rollouts: 66.2% for the earlier policy and 63.7% for this one. Two real bugs were then found and fixed: a frame-convention error in our fine-alignment IK, and a grasp bank that certified entries through a different code path from the one that replays them. With both fixed the phase passes its gate honestly at 87.5% (70/80) on the strict criterion — a smaller-sounding number than the 88.8% we retracted, and the first one that means what it says. On the correct criterion this phase fails two of its three gate criteria, and the "new best" claim below is also void — the two policies are 53 and 51 successes out of 80, so the newer one did not beat the older. The text below is left as written; the numbers in it are the wrong ones.

And then it stopped being a trade. With training finished, the stiller policy went through the full gate — 80 attempts, same seeds, same protocol as the run that produced last week's 85% — and picked the box up 71 times out of 80. That is 88.8%, and it passes all three gate criteria.

The new best. The box turns green while the success window counts it as genuinely held.

The interesting part is how it got there. It still arrives in position less often than the twitchy policy — 69% against 78% — and wins anyway, because when it does arrive it converts that into a pick 94.5% of the time against 88.7%. Every remaining failure in the old 85% was the arm knocking the box over on its way in. Shaking is what knocks the box.

Three extra successes out of eighty is a small margin and we are not going to oversell it. The solid claim is the mechanism, not the headline: the stiller arm fails the grasp roughly half as often once it gets there, and that shows up in the same direction on every sample we have taken.

One more idea closed the loop. Under the new action space an action is a speed command, so we charged the arm for moving — but only when it was already close to its target. Far away, motion stays free. That version arrives on target more than twice as often as the original policy (55% against 23%) at half its settled speed, and passes the same gate at 87.5%. Stillness and accuracy had been trading against each other all night; here they stopped.

What we have not done is make the arm still. The bar we set before starting was 0.10 rad/s of residual motion. The best policy sits at 0.317, down from 0.606 — nearly half, and it now costs nothing on the task, but it is not stillness and we are not going to call it that. (These two figures were also corrected on 2026-08-18: they had been averaged over six joints including the pinned jaw, which deflated both by exactly 6/5.)

Two of those six refutations were our own fault, and working out why was the most useful hour of the night. The penalty charged the policy for changing its command — but the previous command was not in the policy's observations. We were fining it for a quantity it could not see. It could only respond by getting quieter about everything at once, which is exactly what the measurements showed. A reward term has to be observable before failing to optimise it means anything.

2026-08-17 · Phase 2 — the arm picks things up

85% end-to-end pick — and the day we spent certifying a grasp that wasn't one.

The arm now reaches for a box from a random starting pose, grips it and lifts it, 85% of the time over 80 trials. It gets there by chaining two separately-tested skills rather than learning one heroic one: a policy that flies the gripper to a target pose — position and wrist orientation — and then a grasp routine proven in advance on 4,258 certified positions.

Lift clear, traverse high, descend onto the box, close, lift. The box turns green while the success window counts it as genuinely held.

The interesting part is the day before this worked. The first grasp system passed every check we wrote — a green gate, a bank of 4,320 "certified" grasps, a policy learning happily against it — and all of it was certifying the box being nipped by the very tips of the fingers. One constant was wrong: the point we told the simulator to grasp at sat three centimetres in front of the fingertips, outside the hand entirely. Every grasp caught the box on the way past.

Nothing in the telemetry could reveal that, because the numbers and the model shared the same wrong assumption and agreed with each other all the way down. It was caught by watching the rollout video. That is the argument for putting a camera on the real machine in one sentence: proprioception and a model can be self-consistently wrong, and vision doesn't care what you assumed.

  • The servos were never the problem — the policy was. The arm looked unstable. Under a constant command it sits at 0.000000 rad/s, using a tenth of its available torque. It was being commanded to slam each joint across its full range 100 times a second, because nothing charged it for thrashing.
  • The simulated servos were missing a term. The vendored model gives the actuators proportional gain only, so every commanded move overshot by 27% and rang twenty times. Adding the derivative gain a real servo has took that to zero overshoot — and lifted the scripted grasp success from 91% to 100%.
  • Approach beats accuracy. Once the grasp worked, every remaining failure was the arm knocking the box on its way in. Successes and failures were identical in pose accuracy; the only difference was whether the object had moved. Lifting clear before traversing fixed more than any further precision would have.
  • Published practice lost to measurement. The standard recommendation is to back off along the tool's approach vector before grasping. Tested here it scored 25% against 85% for backing off vertically — our gripper straddles the box, so the textbook retreat drags the fingers through it.

Success is measured mechanically, not by eye: both fingers must press with forces that substantially cancel — the physical signature of squeezing something rather than resting on it — with the object airborne and still. A scoop, a nudge or a fingertip nip scores zero.

Still open: the arm reaches its target by vibrating rather than settling. Three documented fixes have been tried and refuted with data. Until that is solved, the pipeline compensates by pausing to let the servos converge before it closes its fingers.

2026-08-16 · Phase 1 — passed

The arm learns to reach: 86.6% over 500 episodes, gate passed.

The first learned skill: drive the gripper to an arbitrary 3D target within 5 cm — from a randomized start pose, in three seconds. No scripted motion — a PPO policy trained purely against a closer-is-better reward, ~65,000 practice reaches across eight parallel simulated arms. Final: 86.6% success, 3.19 cm mean final distance over 500 deterministic episodes, against a committed gate of ≥85% and ≤4 cm.

One provenance note, because this log doesn't move goalposts silently: the working plan sketched 90% before anything was measured. The committed gate is 85% — set after four independent improvements all measured 84–86%, a solvability probe (best-of-10 stochastic tries vs one deterministic) showed no extractable headroom, and a 500-episode interval put the frontier at 86.6% ± 1.5%. The residual is a characterized region of contortion-heavy targets near the base, documented in the exit report — a boundary of this policy class, not effort left on the table.

Eight deterministic episodes from the passing policy, randomized start poses. Orange sphere = target; green = inside the 5 cm ball. 8/8.
CheckpointChangeSuccessMean miss
1M stepsbaseline, 64×64 net18%18.9 cm
6Mmore training60%8.9 cm
6Mshaped reward (bugged)38%9.2 cm
6Mexploit closed62%6.9 cm
9Mfloor-filtered targets76%4.9 cm
21Mcollision-filtered targets84%3.4 cm
12Mfresh 256×256 net86%3.4 cm

Four findings so far, each caught by measurement rather than intuition:

  • The agent exploited our reward. A close-range bonus plus terminate-on-success made hovering next to the target worth more than touching it — success collapsed to 38% while the policy farmed the bonus. Fixed-length episodes realigned incentives with intent.
  • A quarter of the exam was underground. Targets were sampled by forward kinematics of random joint configurations; 24.8% landed below the floor. The policy's ceiling was the exam's fault.
  • Some poses only a self-intersecting arm could hold. FK doesn't know the arm can't pass through itself; misses clustered near the base where those phantom targets live. Now every target's generating pose must be collision-free.
  • The network was too small. A 64×64 policy saturated at 84% after 21M steps; a 256×256 policy beat it in half the steps.

Also measured at gate time, looking ahead: in the front workspace the policy — never trained for precision below 5 cm — already lands 88% within 2 cm and 51% within 1 cm (median successful landing: 11.3 mm). A dedicated fine-positioning phase is the trainable path toward precision work; the ±0.5° servo backlash (~2.6 mm at full reach) is the hardware floor, and beating that is a morphology decision — which is the point of the doctrine. Next: Phase 2, grasp and lift.

2026-08-14 · Phase 0 — passed

The substrate stands: SO-101 in MuJoCo, stable at 500 Hz.

Phase 0's job was to prove the simulation stack end to end: the vendored SO-101 model (masses, inertias, servo gains and ±0.5° backlash derived from the real hardware's CAD and datasheets) loads, steps stably, renders, and exposes six position-controlled joints. Exit criteria V1–V5 all green: 5,000 telemetry rows, 301 rendered frames, zero non-finite states.

The smoke test: a scripted sinusoid — no learning yet, just proof the physics holds.
2026-08-14 · Genesis

capramorph.com goes live.

Domain, infrastructure, and this site — the project's public face. The operating doctrine, stated once and held to: policies come first, morphologies later; hardware is the periodic materialisation of what survives the sim gates; and this log reports what the eval harness says, not what we hoped it would say.