Ship agentsthat getbetter.
Self-improvement stopped being a training-time event and became an operating property — a loop distributed across four substrates: harness, skills, memory, and evaluators. Twelve illustrated essays map the loop, its failure modes, and the verification that bounds it.
The Essays
The Self-Improving Stack
Self-improvement stopped being a training-time event and became an operating property — a loop distributed across four substrates: harness, skills, memory, and evaluators.
Automation Ends Where Agency Begins
The agency/automation line is a definable engineering boundary — independence of decision, not tool-calling — and where you draw it fixes architecture, evaluation, and what to fear.
Harnesses That Build Themselves
Harness design is model-specific and iterable, so it is becoming something agents do to themselves — weakness mining in, human authorship out.
Skills Are Compound Interest
Skills — versioned, composed, evolved procedural knowledge — are the capability capital that survives model swaps and compounds across tasks.
Memory You Can Train
Agent memory is crossing from retrieval plumbing to trained behavior — metamemory, memory-as-action, learned context representations, and governed persistent state.
The World Model Turn
Next-state prediction is consolidating perception, prediction, and control — world models are becoming agent infrastructure rather than a research aesthetic.
Training on Your Own Exhaust
Learning from self-generated experience has named, measurable failure modes — collapse, addiction, hacking — and the guardrails are becoming standard equipment.
The Verification Ceiling
Verification is the binding constraint on agent autonomy — it no longer comes free with generation, and building it is the actual moat.
Judges on Trial
The measurement layer of the agent economy fails basic measurement science — reliability is not validity, saturation hides regressions, and eval-portfolio design is now a discipline.
The Red Queen Loop
Under self-improvement, static evaluators get consumed — evaluation must co-evolve with the agent or the loop optimizes the test.
The Serving Loop
Inference stopped being a cost center — speculators train on live traffic and search primitives are trained in; the serving stack is inside the learning loop.
After Autoregression
Parallel and any-order generation is the first credible bid against left-to-right — simultaneously a training story, a transparency story, and a serving story.
“A team that cannot name its loop's signal, update surface, guardrail, and check does not have an improvement loop; it has an anecdote.”
This series reads the July 2026 literature as one system: the substrates that improve, the pathologies that consume naive loops, and the evaluators that must run alongside the agents they judge. Every essay closes with a Loop Card a team can act on this quarter.
The Improvement Loop · the deployment companion to Harness Engineering and The Long Bet