Phase gates, honest numbers, rollout videos — including the bugs.
Every claim here is reproducible from the working repository's eval harness;
nothing is rounded up.
2026-09-08 · hardware · the eyes
A measurement is not a confirmation.
A €55 webcam on a gooseneck arm ended four days of reasoning
about a workbench from photographs sent after the fact. Every frame
now burns its own timestamp, servo angle, temperature and torque
state into the pixels — and the reason is the first thing the
camera produced, which was a fault that didn't exist. An image
captured at one moment, compared against an encoder reading taken
minutes later while the lever was being moved by hand, showed
ninety degrees of lever against three degrees of shaft. The
conclusion — a hub slipping on its spline — was
confident, mechanically plausible, and pure artifact. Captured
together, the same test reads 90.7° across a 90° move.
Image and pose are one measurement or they are not a
measurement.
The servo ships no zero of its own: the encoder's origin is a
factory arbitrary that changes every time the disc comes off the
shaft. So the bench got a frame of its own —
zero at plumb, positive as the lever rises, because
gravity torque is then just mgL·sinθ, the shape every
sweep and stall figure in this project already has. A spirit level
anchored it, the anchors agreed to a tenth of a degree, and it
lives in a version-controlled file rather than anyone's memory.
Then the beautiful calibration drove the arm into the cabinet.
A spirit level proves levelness; it cannot tell you which
of the two horizontals — one either side of vertical —
is the one without furniture in it. A loaded sweep started toward
the divider panel while the operator's hand went for the 12 V
plug. Worse: that same run captured one camera frame out of seven,
because the device wouldn't reopen between poses. The tool built
specifically to stop blind motion ran blind. Both faults are now
fixed in structure rather than intention — the calibration
file names the forbidden side as well as the safe one, the camera
holds a single handle for a whole run, and the mapper aborts
outright if any pose yields no frame. If the watchdog
can't see, the motion stops. The rule about confirming
direction rather than inferring it had been written down a week
earlier, after an earlier incident, and was bypassed because a
measurement felt like a confirmation. It isn't one.
Run correctly, the same procedure produced the record that four
ambushed sessions had needed: seven poses, seven witnessed frames,
the weight hanging free at every angle. And the load telemetry
carried a bonus. Plotted against the sine of the angle, the ratio
falls and then flattens — the fingerprint of a constant
friction term riding on gravity, rather than of a weight resting on
something. It puts joint friction at about
0.10 N·m, and the simulator these policies grew
up in models 0.052: our virtual joints are several times slipperier
than the metal on the shelf. A sensitivity sweep across the
measured range is running against the frozen policy now, to find
out whether it ever noticed.
2026-09-05/06 · hardware · the instrument
The bench weighed its own chain, then called the datasheet a liar.
Day two of hardware: a €14 aluminum bar became a nine-hole
ladder lever, a rated clevis chain replaced every improvised
fastener, and Decathlon barbell plates became certified test
masses. Two stepped sweeps at two known masses gave a
differential calibration — the unknown
rigging mass cancels out of the slope, then falls out of the
intercept: 294 grams, measured by the bench
itself. A kitchen scale arrives Wednesday to grade the
prediction. The friction the same fit implies, 0.33 N·m,
agrees with the bracket we measured a day earlier by hanging beer
bottles. Independent instruments, one answer.
Two honest rejections along the way: the servo's current
register produced negative 219 g of rigging —
it's blind below ~150 mA, so the load register is the torque
gauge now — and a whole sweep was discarded after video
review caught the hanging chain riding the cabinet edge. An
R² of 0.90 lost an argument with a phone camera, and
deserved to.
Then the probe that rewrote the shopping brochure: climb a
known 2.79 kg load until the servo sticks, and read stall
torque from pure geometry. 1.62 N·m —
barely half the “30 kg·cm” on the
label. The trace shows current rising honestly to
806 mA and then collapsing to 120: firmware overload
protection, not physics. Three EEPROM bytes gate the whole thing
— exceed 80% output for two seconds and it clamps to 20% —
and a fourth byte (integrator gain: zero) explains every degree of
position droop we spent a day chasing. The rated number is a
burst figure ships never meant you to hold.
Overnight the rig ran 12,237 loaded cycles in exactly
twelve hours: zero faults, zero corrupted reads, zero thermal
timeouts, settling at 59–61°C — two
degrees under its own safety pacing — and finishing with
better repeatability than it started (0.10°). The
fastener life series closed: zip tie, 250 cycles; knotted cord,
1,987; metal clevis chain, 12,237 and unbothered. The servo is at
~26,700 career cycles, two full-speed crashes and one max-power
mistake, and still hasn't been the thing that failed.
2026-09-04 · hardware · bench day
Measured with beer: the day the commands became physical.
The servos arrived, and the first bench harness worked against real
silicon on the first try. By evening the sim's assumed bands were
measurements: command-to-motion latency 5.9 ms
(we'd assumed up to 20), deadband 0.40° (the
floor of our assumed band), and the gear-lash operator our adversarial
review had already doubted was retired — the control loop closes
on the output shaft and drives straight through the slack.
Then a number no datasheet prints, bracketed with the only
calibrated masses in the room: one 250 ml bottle hung on a lever
holds forever on the unpowered gearbox's own friction, at zero
current. Two bottles run it straight down, audibly winding against the
gears. Passive friction sits between 0.30 and
0.60 N·m — a sixth to a third of our sim's
whole torque ceiling — and the fleet doctrine writes itself: an
unpowered joint is not a brake.
Two 30-minute loaded holds drew the static thermal curve at two
operating points, an unpowered twin servo logged as ambient control
both times. At a quarter of the torque ceiling:
39→41°C, mean 14 mA — thermally
free, friction carries the pose. Then a mini wrench zip-tied along
the warped lever as a splint (zero bow on camera) moved the cord to
its jaw, grew the arm, and put ~45% of the ceiling on the shaft:
39→52°C in a textbook first-order climb toward a
~54°C plateau, mean 66 mA, control flat at 37.
Real self-heating, bounded 15°C under the safety rail —
and 0.00° of position drift across both runs,
an hour of loaded holding without losing an encoder count. The
climb to the test pose was filmed on a phone, and the video matches
the load telemetry frame for frame: the load register drew a sine
curve as the lever rose, exactly as the geometry says it must. If
thermal derate bites, it bites moving, not posing — the
cycling test is next.
Honesty section: the day also produced three incidents, all of
them command-side, all of them ours. A sign bug turned −10°
into 350° and whipped the first lever into a steel clamp. The
slow-speed safety cap we added in response turned out to starve the
position loop — a falling load stays ahead of a throttled
setpoint, so the servo loses politely to two beers. And removing that
cap to catch a falling load rang a full revolution at maximum power.
The protocol that ended the streak is now law: every motion is
declared and acknowledged before it runs, and loads are
handed to a servo standing still, never dropped on one
— the handover measured zero drift. The rig's second build
survived the worst command at full power, which is the design
standard now: the worst command eventually gets sent. Both wrecked
levers were cut from packaging; both times the cheapest part broke
and the gearbox didn't. The servo took two full-speed stalls into
steel without a single fault flag.
2026-08-31 · the eyes · distillation
The student that beat its teacher.
With the composed pick finally measured honestly — pose policy
travel plus grasp close, 71–76% end-to-end depending on the engine
— we re-distilled the vision student from the new smooth teacher.
From the wristcam alone it hit 90% on nudge grasps. The
previous student managed one in twenty. Same architecture, same data
budget; the only change was the teacher, and an action space where a
slightly wrong velocity self-corrects instead of flinging the arm.
Then we put it in the full chain and it collapsed to
32.5%. The pose policy delivers the arm to states the
student had never seen — its whole training diet came from one
reset distribution, and deployment serves another. Two fixes went to
measurement: proxy views (assist-delivered poses) did nothing; collecting
the teacher's actions on the actual arrival states and
fine-tuning on the success-filtered pairs lifted it to 47.5% without
costing the standalone number a point.
The first arrival dataset never got used. Its teacher converted only
37% of arrivals, so the recordings were two-thirds failure trajectories
— and cloning failure teaches confident failure. That number was
its own finding: the checkpoint our weekend of automated polishing had
produced was quietly worse on deployment states than its
ancestor, 37% against 71%, invisible to every gate we ran because every
gate started episodes the training way. Gates now include the
deployment distribution, permanently.
Round two of DAgger, fine-tuned on the newest slice alone, was flat.
The fix was embarrassing in hindsight: the algorithm is called
dataset aggregation. Training on the union produced a number
that made us suspicious for the opposite reason — 93.8% —
and the audit found the contamination in minutes: the evaluation
arrivals were the training arrivals, byte for byte, courtesy of a
shared seed. On genuinely held-out episodes:
90.0%. Across five fresh seeds, 400 episodes:
88.5%, against the state-based teacher's 66–74%
on the same arrivals.
The student beats its teacher because it trained where the teacher
never went. And it carries one more property the teacher class never
had: under 10–20 ms of loop latency — the delay real
hardware will impose — every state policy we measured lost 20 to
50 points; the pixel student lost nothing. Vision
commands change slowly frame to frame, so a stale action is still
approximately the right one. A last attempt to push further —
three more collection rounds, 69k pairs — made it 24 points
worse, and was archived the same evening. The frozen file is
called vision_v7_student.pt, and on Friday it meets its
first real servo.
2026-08-28 · the action space · jitter
The shake lived above what the videos could show.
“Watch the videos. I'm assuming you haven't, right?” Correct
— and it turned out worse than not watching. Control runs at
100 Hz; the films sample every fifth step, 20 Hz, Nyquist
10 Hz. The trained policy put 40% of its joint-velocity energy
above that line — not hard to see on film, mathematically
absent from it. Aliased into apparent smoothness on every frame anyone
ever rendered.
The alarm built to catch this was the sixth instrument here caught lying.
It reported the fraction of spectral energy above 2 Hz, and a
fraction is scale-free: a motionless arm scored 0.86 and tripped it, so it
fired on everything and meant nothing. Measured floor to ceiling: a still arm
reads 0.003 rad/s; a parked arm fed 5% random command noise reads 0.77;
our policy read 1.01 — more shake than deliberate dither.
The alarm now speaks absolute rad/s.
Five fixes went to measurement and every one died: two command taxes were
Goodharted (commands calmed 4×, physical motion tripled), a velocity tax
priced speed instead of shake, a jerk tax bought a noise-level change for
eight points of success, an action low-pass made the policy fight its own
filter, and halving the control rate made the shake 71%
worse. The pattern was the finding: you cannot tax what the action
space makes free. Absolute position targets let one step command an arbitrary
jump into a kp 998 servo — and our own archive, a campaign on the
sibling task, had already written the conclusion we spent a day rediscovering:
the jitter lives in the action space, not the reward.
Bounded delta actions — the target is the measured position plus a
velocity-capped increment, the norm in every published SO-101 sim2real
pipeline — made chatter inexpressible instead of expensive. Under pure
random actions the bound alone cuts high-frequency shake threefold. Trained:
83.5% acquisition at 38% less shake, against 76.7% for the
old action space — the first configuration in this campaign to win both
axes at once. The ablation also executed its own gate term: a ported speed
penalty cost 15 points and bought nothing, caught only because the sweep
isolated it.
The honest ledger: the residual is amplitude-limited chatter, the habit
bounded rather than cured. And a new latency knob priced transfer for the
first time: one control step of loop delay costs ~20 points of
success, two cost ~50, old policy and new alike — so the next
run trains under randomized delay, and the bench servos arriving in September
now have a second job: confirming the sim's assumed hardware bands before we
trust anything trained inside them.
2026-08-27 · the pose bank · side grips
We were quoting one grasp and training another.
Dave asked why the claw keeps picking from above at awkward angles when
the side grip is obviously better — it is how a human lifts a cup, and
our own log had measured it: 15/20 without the post-grasp roll,
19/20 with it. So we measured the pose bank. The finger axis in all
655 certified poses sits a median 14.3° from straight down.
Top-down, every one of them.
Then we measured the battery that produces the 19/20 headline this project
has published for a week. It samples its own poses, behind a filter
requiring the finger axis to stay within 25° of horizontal.
Median 16°. A side grip. Overlap between the two distributions:
0 of 655. The number we report and the data every policy
trains on were never the same grasp.
The constraint that would have caught it exists in the bank generator,
parameterised and switched off — FINGER_AXIS_MAX defaults to
1.0, which accepts every orientation, under a comment reading
“which nothing has ever constrained.”
Rebuilt on the proven filter: 353 side-grip poses, 351 of them
perfect holds (99.4% against the top-down bank's 82%). The payoff was
immediate and needed no other change. The end-to-end composed pick
doubled, 26.7% → 53.3%, with the grasp stage going
30% → 62% — the seam we had spent hours debugging was never a reach
bug; reach was already landing within 0.6 mm. Thin-margin top-down grips
simply could not absorb that error. And the first policy trained on the new
bank reached 66% acquisition against a previous record of
36.5%.
Two honest deductions. Part of that 66% is a better-posed task rather than a
better policy — side-grip poses hold with more slack, so the comparison
crosses distributions. The clean evidence is the composed pick, same pipeline,
only the bank swapped. And it cost us something: far-range reaching, which had
just reached 8%, fell back to zero. It was learned in the top-down world and
did not transfer.
2026-08-25 · verification · the instruments
Five instruments, four of them lying.
Correction. An earlier version of this
log's underlying notes claimed the training bank reproduced at only 21.7%
through the replay path. That was an artifact of a certification script with a
hardcoded bank path, which measured the same legacy file three times while
appearing to compare three different banks. The true figure is
96.7%. The bank was never broken.
Four consecutive runs declined — 21.8%, 17.2%, 12.5%, 9.5% — under
four different fixes, none of which was the problem. The cause was a single
runaway parameter: the policy's exploration spread climbed from 3.9 to 465 and
then into the millions. Training rollouts had been effectively random for days
while the deterministic policy kept its skill, and every update sanded that
skill down.
The first fix made it worse. Clamping the spread once per rollout sawtoothed
against the optimiser's own update loop, pushing the divergence measure a
hundred times past healthy and churning the aim away entirely: 0 successes in
200. Freezing the parameter instead — a constant distribution cannot
sawtooth — restored learning on the next run.
Then the certification harnesses started disagreeing. One returned 0% for
poses that are certified by construction, because it used the pre-fix open-loop
lift. Another was the hardcoded path above. A third, the bank generator's own
script, aborts outright on this claw. Of five measurement paths exercised in
one evening, only one survived a control check, and only
because the control was run first.
That is now mechanism rather than intention. A single verdict pipeline
produces every run's result: it refuses to start without the full environment
block, reports the optimised metric beside an ungameable physical witness,
runs a frequency analysis that declares oscillation without waiting for a human
to say it, raises an alarm when a proxy improves while reality worsens, and
writes no verdict at all until the films exist. When a verification step keeps
getting skipped, stop re-teaching it and make the tool refuse to proceed.
2026-08-23 · reinforcement learning · the reward
It learned to get paid without ever picking anything up.
The first real training run looked excellent: reward climbing 29 to 340 over
ten million steps. The exam said 0 successes in 500 episodes,
with the object's mean peak height exactly its resting height. It had never
lifted anything.
Two causes, both ours. The curriculum's stage boundaries are counted in total
steps, so a ten-million-step run lived entirely in the first stage and never met
an exam-like start. And the reward paid for approaching and pinching every step
with no bonus for finishing, so a flickering grip out-earned a firm one
indefinitely.
Dave watched the rollout and said it was gripping with its fingernail.
Measured: the policy's holds sat 56–63 mm from the grasp centre while
certified grips sit at 4–21 mm. Later he watched again and said the
jitter would disassemble something. Measured: bang-bang control, commands
swinging most of their range at 100 Hz with the actuators saturated
two-thirds of the time.
Taxing that jitter taught the policy to cheat the tax. Command swing fell
four-fold while physical joint speed tripled — it had learned to
pump the arm's own resonance with small, cheap pushes. The tax now applies to
measured joint velocity, which cannot be gamed that way. We also discovered
every film we had reviewed was playing 1.5× too fast, because the capture
rate and the playback rate were set independently and nobody checked.
2026-08-22 · Phase A — the overnight run
The lift was fixed at 3 a.m., and it confessed to the tumbles.
Dave went to bed and left Phase A running autonomously. The target was a
known bug: the scripted lift computed one Jacobian at the start pose and
trusted it across 120 mm of travel, under-delivering up to 28 mm.
The fix is closed-loop: re-measure the gripper's height every quarter
second, fold the error back through a live Jacobian.
It was specced to rescue one pose — the perfect grip that topped out at
86 mm. It rescued five. The open-loop lift's sideways
drag turned out to be the quiet cause of most of the "acquisition"
tumbles and slides we had attributed to the claw itself. With the honest
lift: 19/20 with zero tumbling and nineteen settling holds —
no wrist roll needed, the calmest claw this project has produced. Across
three fresh pose banks: 17–20/20 by bank, and zero tumbling on
every bank tested.
Two humbler results, reported because the suite said so. The
approach-angle cap we planned as a second fix is now redundant — capped
and uncapped banks score the same, because the lift fix cured the exact
failures the cap existed to screen. And a 20/20 scored on the primary
capped bank reproduced perfectly but did not generalize (18 and 17 on
fresh banks) — so the honest headline stays 19, not 20. The one pose that
fails on the primary bank in every configuration is the chaos
pose — the one that flips on bit-level noise. Deterministic control has
hit its ceiling; that pose is the standing argument for the learned
policy this project exists to train.
Housekeeping with numbers: the old grasp bank, re-certified against
the final claw at rated torque, kept only 655 of 4,320 entries — it was
built for a bare claw on an arm 1.75× too strong. The rebuild script
refused to generate a new one ("0/500 pass — do NOT tune past this"),
correctly: its bands describe a claw that no longer exists. The 655
survivors are the interim bank; the retrain comes next, and it inherits
a gripper whose only remaining enemy is chaos.
2026-08-22 · Phase 2 — the claw closes, a policy opens
He bet it wouldn't work. It's the best result this project has.
Day three of the claw. First the design closed the honest way: a
per-rib force census showed the rear ribs carry 49% of grip force, the
static pads 50%, and the front row roughly zero — so we deleted the
zero-load teeth one by one, battery after every cut.
Two came out free. Deleting the whole front row kept the score but cost
8.5 mm of seating depth — those ribs steer, not grip. Deleting one
rear backstop cost 1 mm of seat and improved residual
twist by 4.5° — the telemetry curves caught what the aggregates
rounded away. Final form: ten ribs, every one with a measured
job, fat wedge bodies verified physics-neutral, 14/20.
Then Dave, who says he can't explain physics but sees it moving in
his head, proposed rotating the wrist 90° after gripping — so the
static finger swings underneath and gravity presses the object
into a cradle instead of pulling it out of the pinch. We raced four
timings. He predicted the immediate roll would fail worst. I predicted
the late roll was safest. We were both wrong: rolling at
1.2–1.6 s scores 19/20 — rescuing both acquisition
tumbles, both slide-outs, and the chaotic knife-edge pose that nothing
else all week could tame, losing nothing.
The catch, on film: at t=2 s the cylinder is tipping
off the ribs — the same failure we filmed all week. At t=3 s the
rotating cradle arrives underneath and it reads HOLDING. It rides at
135 mm for the next two minutes.
He didn't believe it — "that's pure luck, there's no way that
worked" — so it got the full treatment: a straight repeat
(19/20, zero pose flips), a timing sweep (a plateau at 1.2–1.6 s,
decaying to 16/20 at 2.0 s+ exactly as the interception window
closes), and the killer — twenty brand-new poses the policy
had never seen: 15/20 without the roll, 19/20 with it. Gained
four, lost none. Mechanism, not luck.
One failure survives on each bank, and honesty requires saying they
differ: on the original bank it's a perfect grip whose lift tops out at
86 mm against a 100 mm bar — the lift arithmetic is next, and
on that bank it is now worth exactly 20/20. On the fresh bank the
leftover is one tumble the cradle arrived too late to catch. Nineteen
of twenty, twice, with the residue named both times.
2026-08-21 · Phase 2 — the pad shoot-out
The claw grips with shape, not stickiness.
Yesterday ended at a wall: at any clamp force the servo survives, the flat
claw held nothing. Today a two-phalanx wrapping finger with shaped pads holds
15/20 for 120 seconds at the servo's rated 0.49 N·m —
the torque it can sustain indefinitely, cold. The progression was all aim:
5/20 at the first guess, 12 after moving the placement, 15 after Dave watched a
failure video and called the approach angle off by eye.
Then we raced five pad shapes — square, two rounds, chamfered, triangular —
same 20 poses, same seeds, identical footprint and bite. Dave asked for the
chamfer mock-up in one sentence: cut the two corners off the squares, give the
object an extra face to land on. It won, 16/20, and the film
shows why: the object ratchets inward one rib pitch at a time and locks
into a pocket, twist falling from 21° to 8°. His other hunch — that the smaller
round pads would grip better — lost honestly: they produced the only true grip
failure of a hundred runs, ninety seconds of creep and then out, with barely any
turn. Bite depth was everything.
Pose 11, chamfered pads, 120 s. Watch 20–45 s: the object
walks inward in steps and locks. Every other shape froze where it first touched
and sagged under the gate.
A contact census answered a question worth keeping: 100% of grip force
goes through the pads — the claw body is never touched. Shrink the pads
until the object reaches the claw and it collapses; coat the whole claw in rubber
and nothing comes back. Remove one pad per face and no score moves, but settled
holds start drifting — two or three contacts do the gripping; the rest
are anti-twist bracing.
Two more phantom-class bugs, our own scene.
The object geom shipped priority="1", which silently made MuJoCo
ignore every pad friction value we ever authored. And the pads' torsional
friction was written as 0.25 — that parameter has length units, so we
had specified a 250 mm contact patch. With both fixed, all five
shapes tie at 13/20: the score spread was riding millimetre margins at
an arbitrary gate. What shape really controls is seat depth and twist — and the
chamfer is the only profile best-in-class on both, in every friction regime we
ran. Same lesson as the hull, one level up: ask what the engine actually uses.
And under the tie, the real verdict. The same seven poses fail in every
shape — but the chamfer keeps six of its seven as near-misses at
84–90 mm against a 100 mm bar, one floor drop, while every
other shape drops four of the seven outright. The remaining gap is a lift that
under-delivers near singularities; fixing it rescues the chamfer to roughly
19/20 and rescues the others far less. Only the winner's failures are
recoverable. Dave set the bar for this phase himself: we're not chasing 20/20 —
humans drop things too. Sustained near-perfect, at a torque the hardware can
actually live at, is the win. The pads are closed; the lift arithmetic is next.
2026-08-19 · Phase 2b — it grips, then drops
It was holding the box like a cone pointing at the floor.
Dave watched a rollout and said the claw grips with its tip, in a V, like
picking up a pyramid from the point — and that the object would always fall out
at that angle. Measured, all of it held.
Sweeping the actual 34 mm box through the gap: past 50 mm
into the hand it does not fit at all, and the floor of the opening rises
34 → 50 mm over that span. That is a 22° wedge, so
the deep wide part of the claw is unusable and every grasp this gripper
makes is forced to be a tip pinch. It also explains a number that had been
bothering us: 86 N of clamping force on a 0.49 N box,
175× its weight. A flat face cannot seat in a 22° V, so it rides on edges and only
brute friction holds it.
And it does not fall sideways. Tracking the box after it lets go, it moves
+38.8 mm along the fingers toward the tip in 13 of 16 drops,
because the finger axis points downhill in 100% of our grasp poses
— a median 23° from straight down. 0.92 g was quietly
pulling the box out the open end the whole time.
Four fixes failed, including both of the obvious ones. Rotating the wrist to
raise the fingertip made it monotonically worse (down to 2.5%); the
rotation disturbs the grip more than the reduced gravity helps. Swapping the box
for a cylinder — a V-groove being the classical way to hold one — lost badly too,
because we have no groove: a cylinder pinched at a tip simply rolls out. The box's
corners were what held it.
We thought the fix was lift speed — lifting six times slower
took "still holding at the end" from 52.5% to 90.0% over 80 trials.
Correction, same day. That was not a fix, it
was a slower failure, and Dave called it immediately: the slow lift only stalls the
slippage. Running the same grasp for longer settles it — the box creeps steadily
along the fingers and always comes out:
at 350 steps 90% still hold, at 700 steps 60%, at 1200 steps 0%,
with the box having crept 96 mm. The grasp is not stable at any lift
speed. Our episode length was setting the score. The honest position is
that this gripper cannot hold this box in a tip pinch, and the 92.5% we briefly
published here measured how long we watched rather than whether it holds.
Identical episodes. Left drops it, right holds it. The fifth episode
fails in both — this is an improvement, not a solved problem.
We also changed what counts as a pick. The old
criterion accepted 0.3 s of holding anywhere in the episode, so an
arm that gripped, lifted and then dropped scored as a success — which is exactly
what the video showed. It now has to still be holding at the end.
2026-08-18 · Phase 2a — the shaking
Six fixes for a jitter that was never in the reward.
The arm reached its target and then kept twitching instead of holding still.
The obvious suspects were the servos, so we measured them first: 0.000000
rad/s of drift under a constant command, 10–25% of available torque at
every pose, 0.02° of sag. The hardware holds perfectly when told to hold. The
policy was commanding the shaking — changing its command by 54–82% of the full
action range every single step, with joints pinned at the extremes
37–100% of the time.
So we charged it for shaking. Six times, in six different ways: a rate limit on
the commands, state-dependent exploration noise, an action-rate penalty at two
strengths, a published smoothness regulariser, and a narrower exploration
distribution. Every one either failed to change the arm's behaviour or bought
stillness by preventing the arm from doing the task at all.
They failed for a reason worth stating plainly. For an arm commanded by
absolute joint targets with no cost on motion, twitching between extremes
is not a flaw in the training — it is the mathematically optimal
policy. Pontryagin's maximum principle says so, and a 2021 NeurIPS paper
confirms it empirically. We had been repeatedly asking the policy not to do the
best available thing.
Changing what an action means fixed half of it immediately. Instead of
"go to this joint angle", an action now says "move this far from where you are" —
so a wild command can only ever move the arm a bounded amount. Settled joint speed
halved, from 0.606 to 0.317 rad/s, with no smoothness
penalty in the reward at all.
Same targets, same seeds. Left commands absolute joint angles and
keeps twitching after it arrives; right commands increments and settles. Measured
on the footage itself, the settled arm moves 3.7× less.
The honest part: this is half a fix, and it came with a bill.
The stiller policy is less accurate so far — it arrives within 10 mm on 12% of
attempts against the twitchy one's 23%. A change that makes the arm calm and worse
is a trade, not a victory, and it is logged as one. But there is reason to think
the trade is temporary and worth paying.
Correction, 2026-08-18. This entry
originally claimed 88.8%, and the previous entry claimed 85%. Both were
wrong. Our scoring counted a single 10 ms instant of contact as a
successful pick, where our own published standard requires the box be held for 27
of any 30 consecutive control steps — a bar 27× stricter. Re-measured on the
correct criterion, on identical rollouts: 66.2% for the earlier
policy and 63.7% for this one. Two real bugs were then found and fixed: a frame-convention error in our
fine-alignment IK, and a grasp bank that certified entries through a different
code path from the one that replays them. With both fixed the phase passes its
gate honestly at 87.5% (70/80) on the strict criterion —
a smaller-sounding number than the 88.8% we retracted, and the first one that
means what it says. On the correct criterion this phase
fails two of its three gate criteria, and the "new best" claim
below is also void — the two policies are 53 and 51 successes out of 80, so the
newer one did not beat the older. The text below is left as written; the numbers in
it are the wrong ones.
And then it stopped being a trade. With training finished, the stiller policy
went through the full gate — 80 attempts, same seeds, same protocol as the run
that produced last week's 85% — and picked the box up 71 times out of
80. That is 88.8%, and it passes all three gate criteria.
The new best. The box turns green while the success window counts it
as genuinely held.
The interesting part is how it got there. It still arrives in position
less often than the twitchy policy — 69% against 78% — and wins anyway, because
when it does arrive it converts that into a pick 94.5% of the time against
88.7%. Every remaining failure in the old 85% was the arm knocking the box
over on its way in. Shaking is what knocks the box.
Three extra successes out of eighty is a small margin
and we are not going to oversell it. The solid claim is the mechanism, not the
headline: the stiller arm fails the grasp roughly half as often once it gets
there, and that shows up in the same direction on every sample we have taken.
One more idea closed the loop. Under the new action space an action is a
speed command, so we charged the arm for moving — but only when it was
already close to its target. Far away, motion stays free. That version arrives on
target more than twice as often as the original policy (55% against 23%) at
half its settled speed, and passes the same gate at 87.5%.
Stillness and accuracy had been trading against each other all night; here they
stopped.
What we have not done is make the arm still. The bar we set before
starting was 0.10 rad/s of residual motion. The best policy sits at
0.317, down from 0.606 — nearly half, and it now costs nothing
on the task, but it is not stillness and we are not going to call it that. (These two figures were also corrected on 2026-08-18: they had been averaged
over six joints including the pinned jaw, which deflated both by exactly 6/5.)
Two of those six refutations were our own fault, and
working out why was the most useful hour of the night. The penalty charged the
policy for changing its command — but the previous command was not in the policy's
observations. We were fining it for a quantity it could not see. It could only
respond by getting quieter about everything at once, which is exactly what the
measurements showed. A reward term has to be observable before failing to optimise
it means anything.
2026-08-17 · Phase 2 — the arm picks things up
85% end-to-end pick — and the day we spent certifying a grasp that wasn't one.
The arm now reaches for a box from a random starting pose, grips it and lifts
it, 85% of the time over 80 trials. It gets there by chaining two
separately-tested skills rather than learning one heroic one: a policy that flies
the gripper to a target pose — position and wrist orientation —
and then a grasp routine proven in advance on 4,258 certified positions.
Lift clear, traverse high, descend onto the box, close, lift. The
box turns green while the success window counts it as genuinely held.
The interesting part is the day before this worked. The first grasp system
passed every check we wrote — a green gate, a bank of 4,320 "certified" grasps, a
policy learning happily against it — and all of it was certifying the box
being nipped by the very tips of the fingers. One constant was wrong: the
point we told the simulator to grasp at sat three centimetres in front of the
fingertips, outside the hand entirely. Every grasp caught the box on the way
past.
Nothing in the telemetry could reveal that, because the numbers and the model
shared the same wrong assumption and agreed with each other all the way down. It
was caught by watching the rollout video. That is the argument for putting a camera
on the real machine in one sentence: proprioception and a model can be
self-consistently wrong, and vision doesn't care what you assumed.
The servos were never the problem — the policy was. The arm
looked unstable. Under a constant command it sits at 0.000000 rad/s, using
a tenth of its available torque. It was being commanded to slam each joint
across its full range 100 times a second, because nothing charged it for
thrashing.
The simulated servos were missing a term. The vendored model
gives the actuators proportional gain only, so every commanded move overshot by
27% and rang twenty times. Adding the derivative gain a real servo has took that
to zero overshoot — and lifted the scripted grasp success from 91% to 100%.
Approach beats accuracy. Once the grasp worked, every
remaining failure was the arm knocking the box on its way in. Successes and
failures were identical in pose accuracy; the only difference was whether the
object had moved. Lifting clear before traversing fixed more than any further
precision would have.
Published practice lost to measurement. The standard
recommendation is to back off along the tool's approach vector before grasping.
Tested here it scored 25% against 85% for backing off vertically — our gripper
straddles the box, so the textbook retreat drags the fingers through it.
Success is measured mechanically, not by eye: both fingers must press with
forces that substantially cancel — the physical signature of squeezing something
rather than resting on it — with the object airborne and still. A scoop, a nudge or
a fingertip nip scores zero.
Still open: the arm reaches its target by vibrating rather than settling. Three
documented fixes have been tried and refuted with data. Until that is solved, the
pipeline compensates by pausing to let the servos converge before it closes its
fingers.
2026-08-16 · Phase 1 — passed
The arm learns to reach: 86.6% over 500 episodes, gate passed.
The first learned skill: drive the gripper to an arbitrary 3D target within
5 cm — from a randomized start pose, in three seconds. No scripted motion —
a PPO policy trained purely against a closer-is-better reward, ~65,000 practice
reaches across eight parallel simulated arms. Final:
86.6% success, 3.19 cm mean final distance over 500 deterministic
episodes, against a committed gate of ≥85% and ≤4 cm.
One provenance note, because this log doesn't move goalposts silently: the
working plan sketched 90% before anything was measured. The committed gate is
85% — set after four independent improvements all measured 84–86%, a
solvability probe (best-of-10 stochastic tries vs one deterministic) showed no
extractable headroom, and a 500-episode interval put the frontier at
86.6% ± 1.5%. The residual is a characterized region of
contortion-heavy targets near the base, documented in the exit report — a
boundary of this policy class, not effort left on the table.
Eight deterministic episodes from the passing policy, randomized
start poses. Orange sphere = target; green = inside the 5 cm ball. 8/8.
Checkpoint
Change
Success
Mean miss
1M steps
baseline, 64×64 net
18%
18.9 cm
6M
more training
60%
8.9 cm
6M
shaped reward (bugged)
38%
9.2 cm
6M
exploit closed
62%
6.9 cm
9M
floor-filtered targets
76%
4.9 cm
21M
collision-filtered targets
84%
3.4 cm
12M
fresh 256×256 net
86%
3.4 cm
Four findings so far, each caught by measurement rather than intuition:
The agent exploited our reward. A close-range bonus plus
terminate-on-success made hovering next to the target worth more than
touching it — success collapsed to 38% while the policy farmed the bonus.
Fixed-length episodes realigned incentives with intent.
A quarter of the exam was underground. Targets were sampled
by forward kinematics of random joint configurations; 24.8% landed below the
floor. The policy's ceiling was the exam's fault.
Some poses only a self-intersecting arm could hold. FK
doesn't know the arm can't pass through itself; misses clustered near the base
where those phantom targets live. Now every target's generating pose must be
collision-free.
The network was too small. A 64×64 policy saturated at 84%
after 21M steps; a 256×256 policy beat it in half the steps.
Also measured at gate time, looking ahead: in the front workspace the policy —
never trained for precision below 5 cm — already lands 88% within
2 cm and 51% within 1 cm (median successful landing:
11.3 mm). A dedicated fine-positioning phase is the trainable path toward
precision work; the ±0.5° servo backlash (~2.6 mm at full reach) is the
hardware floor, and beating that is a morphology decision — which is
the point of the doctrine. Next: Phase 2, grasp and lift.
2026-08-14 · Phase 0 — passed
The substrate stands: SO-101 in MuJoCo, stable at 500 Hz.
Phase 0's job was to prove the simulation stack end to end: the vendored
SO-101 model (masses, inertias, servo gains and ±0.5° backlash derived from the
real hardware's CAD and datasheets) loads, steps stably, renders, and exposes
six position-controlled joints. Exit criteria V1–V5 all green: 5,000 telemetry
rows, 301 rendered frames, zero non-finite states.
The smoke test: a scripted sinusoid — no learning yet, just proof
the physics holds.2026-08-14 · Genesis
capramorph.com goes live.
Domain, infrastructure, and this site — the project's public face. The
operating doctrine, stated once and held to: policies come first, morphologies
later; hardware is the periodic materialisation of what survives the sim gates;
and this log reports what the eval harness says, not what we hoped it would say.