Contents

The series, in order.

12 essays in 4 parts — the frame first, then the substrates that improve, the loops that fail and get guarded, and the stack beneath. Read it front to back, or pull whatever thread your production system touches first.

IThe FrameProve the loop exists; draw the agency line.
  1. 1
    The Self-Improving Stack

    Self-improvement stopped being a training-time event and became an operating property — a loop distributed across four substrates: harness, skills, memory, and evaluators.

    SEAG· 14 min· Jiang et al. (2026) — Era-of-Experience survey · Yu et al. (2026) — HORIZON · Kulikov et al. (2026) — Autodata
  2. 2
    Automation Ends Where Agency Begins

    The agency/automation line is a definable engineering boundary — independence of decision, not tool-calling — and where you draw it fixes architecture, evaluation, and what to fear.

    AGFN· 9 min· Xing, Deng & Hou (2026) — Critique of Agent Model · Chen (2026) — the co-failure ceiling
IIThe SubstratesWhat improves: harness, skills, memory, world models.
  1. 3
    Harnesses That Build Themselves

    Harness design is model-specific and iterable, so it is becoming something agents do to themselves — weakness mining in, human authorship out.

    SEAG· 10 min· Zhang et al. (2026) — Self-Harness · Kassianik & Saglam et al. (2026) — FAPO · Ivison & Yin et al. (2026) — TMax
  2. 4
    Skills Are Compound Interest

    Skills — versioned, composed, evolved procedural knowledge — are the capability capital that survives model swaps and compounds across tasks.

    SKSE· 9 min· Wang & Yan et al. (2026) — MetaSkill-Evolve · Berthon et al. (2026) — Skill Neologisms · Kim et al. (2026) — HASTE
  3. 5
    Memory You Can Train

    Agent memory is crossing from retrieval plumbing to trained behavior — metamemory, memory-as-action, learned context representations, and governed persistent state.

    MESK· 12 min· Wu et al. (2026) — AutoMem · Eyuboglu, Ehrlich & Arora et al. (2025) — Cartridges
  4. 6
    The World Model Turn

    Next-state prediction is consolidating perception, prediction, and control — world models are becoming agent infrastructure rather than a research aesthetic.

    WM· 9 min· Orca Team (2026) — world foundation model
IIIThe LoopsSelf-training pathologies, the verification ceiling, judges, co-evolution.
  1. 7
    Training on Your Own Exhaust

    Learning from self-generated experience has named, measurable failure modes — collapse, addiction, hacking — and the guardrails are becoming standard equipment.

    RLSE· 11 min· Che & Wu (2026) — Greed Is Learned · Yuan et al. (2026) — diversity collapse · Zhong et al. (2026) — HSIR
  2. 8
    The Verification Ceiling

    Verification is the binding constraint on agent autonomy — it no longer comes free with generation, and building it is the actual moat.

    EVAG· 10 min· Qwen Team (2026) — Verification Horizon · Kwok et al. (2026) — LLM-as-a-Verifier · Adamczewski et al. — MirrorCode · Hans & Bilionis (2026) — Paper-replication
  3. 9
    Judges on Trial

    The measurement layer of the agent economy fails basic measurement science — reliability is not validity, saturation hides regressions, and eval-portfolio design is now a discipline.

    EV· 12 min· Norman, Rivera & Hughes (2026) — the judge audit · Zeng & Papailiopoulos (2026) — BENCHPRESS · Nadgir et al. (2026) — life after saturation
  4. 10
    The Red Queen Loop

    Under self-improvement, static evaluators get consumed — evaluation must co-evolve with the agent or the loop optimizes the test.

    SEEV· 8 min· Iacob, Jovanović & Shen et al. (2026) — Red Queen Gödel Machine
IVBeneathThe serving loop, and the paradigm bet under everything.
  1. 11
    The Serving Loop

    Inference stopped being a cost center — speculators train on live traffic and search primitives are trained in; the serving stack is inside the learning loop.

    IERL· 9 min· Hamid & Orney et al. (2026) — SPIRAL · Wang & Bie et al. (2026) — Aurora
  2. 12
    After Autoregression

    Parallel and any-order generation is the first credible bid against left-to-right — simultaneously a training story, a transparency story, and a serving story.

    DF· 8 min· Nie et al. (2026) — iLLaDA · Engels, McDougall & Chughtai et al. (2026) — DiffusionGemma transparency