Part III · The Loops · 9 of 12EV

Judges on Trial

The measurement layer of the agent economy fails basic measurement science — reliability is not validity, saturation hides regressions, and eval-portfolio design is now a discipline.

Figure 1: A gavel balanced on a ruler, on a bell curve — judgment resting on measurement resting on statistics. In 2026 the middle layer is the one that fails quietly.
Anchor papers
Norman, Rivera & Hughes (2026) — the judge auditZeng & Papailiopoulos (2026) — BENCHPRESSNadgir et al. (2026) — life after saturation
12 min read2,720 words↳ Reading order: ← 8 · 10 →

§1 · The judge economy

Somewhere in your stack right now, a language model is grading another language model, and a number from that grading will move a decision — which checkpoint ships, which prompt survives, which vendor wins the bake-off. LLM-as-a-Judge is no longer an evaluation technique; it is the measurement layer of the agent economy. And the standard way that layer gets validated is a statistic that a first-year methods course would flag: exact-match agreement with human labels, uncorrected for chance.

Norman, Rivera and Hughes (2026), at Berkeley's School of Information, put the whole layer on trial — the largest systematic evaluation of LLM-as-a-Judge to date: 21 judges from nine providers, across MT-Bench, JudgeBench, and RewardBench, under three protocols — agreement, consistency, and a bias audit — totalling 118 runs and approximately 541,000 individual judgments. The cohort includes the April 2026 frontier. This is not an adversarial stress test or a cherry-picked failure reel; it is the field's own instruments, measured the way psychometrics has measured instruments for a century.

The verdict gives this series its cleanest statement of a theme that will recur for three articles: the things we use to measure improvement are themselves unmeasured. An improvement loop is only as trustworthy as its evaluator V — and this article is about what happens when you audit V itself.

[← B2] The Agent Evaluation Crisis established the June state of this problem for agent benchmarks. This article is the July sequel at the instrument level: not "are agent evals hard," but "do our judges measure anything at all." Verification of individual outputs — a different question — is treated in [← 8] The Verification Ceiling.

Key Takeaway 1

LLM judges are the de facto measurement layer of the agent economy, and they are validated in practice with chance-uncorrected agreement — a statistic that systematically overstates discriminative ability. The largest audit to date covers 21 judges, nine providers, and ~541,000 judgments; its findings are the subject of this article.

§2 · Reliability without validity

Four findings emerge from the Berkeley audit, consistent across the full cohort (see Figure 2).

First: kappa deflation is universal. Correct agreement for chance — replace raw exact-match with Cohen's κ — and every judge in the cohort drops, by 33 to 41 percentage points on MT-Bench. A judge with an impressive raw agreement score may owe a large share of it to the coin-flip floor. Nothing about this is exotic; chance correction is the first move in any serious inter-rater methodology. The finding is not that the correction is possible but that the field skips it, and that the skipped correction is large enough to reorder conclusions.

Second: rankings do not transfer. Judge rankings shift by up to 14 positions across the three benchmarks. The "best judge" is a property of the benchmark you happened to validate on, not of the judge. Any team that picked its production judge from a single leaderboard inherited this instability silently.

Third: the consistency–bias paradox. Two production-deployed judges in the cohort combine high test–retest reliability (above 0.95) with severe position bias (above 0.10). That combination is worse than noise: a judge that flips a coin at least errs symmetrically, while a judge that reliably prefers the first-listed answer will reliably reward whatever pipeline puts its candidate first. Reliability without validity is exactly the failure mode the paper's title names — the instrument repeats itself beautifully and measures the wrong thing.

Fourth: verbosity bias, the failure everyone expected, is small — under 0.011 across the cohort under a single pairwise rubric. This matters as method, not just as relief: the audit can tell the difference between the biases the community assumes and the ones the data shows. The authors distill the whole protocol into a Minimum Viable Validation Protocol — the checklist any team can run before trusting a judge with a shipping decision.

FindingMeasured effectOperational consequence
Kappa deflation (universal)❌ −33 to −41 pp, exact-match → Cohen's κ (MT-Bench)raw agreement overstates every judge; correct for chance before believing a number
Ranking instability⚠️ up to 14 positions across benchmarks"best judge" is benchmark-relative; re-rank on your own distribution
Consistency–bias paradox❌ test–retest > 0.95 with position bias > 0.10 (two production judges)repeatability is not correctness; audit bias separately from reliability
Verbosity bias✅ < 0.011 under a single pairwise rubricthe feared bias is minor in this cohort; spend audit budget on position, not length
Figure 2: The four findings of the Berkeley judge audit, redrawn from Norman, Rivera & Hughes (2026). The instrument layer is reliable, unstable, and biased in exactly the combination that raw agreement scores cannot see.
Key Takeaway 2

Chance-corrected, the judge layer deflates by 33–41 percentage points; rankings move by up to 14 positions between benchmarks; and two production judges pair >0.95 repeatability with >0.10 position bias. Reliability and validity are different properties — and only one of them is being checked in practice.

§3 · Saturation hides regressions

The judge audit covers the instruments; the second failure is in the lifecycle of the tests themselves. The reflex when a benchmark saturates is retirement: accuracy hits the ceiling, the benchmark is declared dead, a harder one replaces it. Nadgir et al. (2026) — a Princeton-centered team including Kapoor and Narayanan — argue that this reflex confuses one dimension with all of them, and they demonstrate it on CORE-Bench Hard, a benchmark for computational reproducibility of scientific code.

Their case study surfaces six dimensions that remain measurable after accuracy saturates: construct-validity threats such as shortcuts; out-of-distribution generalizability; efficiency; reliability; the relative contribution of the model versus the scaffold; and uplift from human–agent collaboration. Each is invisible on a leaderboard that reports one number. Two of their findings deserve emphasis. The shortcuts they surface in CORE-Bench Hard were difficult to anticipate with less capable agents — construct-validity problems appear as capability rises, which means a saturated benchmark is precisely where you can finally see them. And the saturated benchmark remained useful: they publish a repaired CORE-Bench v1.1 and an out-of-distribution suite, CORE-Bench OOD, plus a small randomized experiment measuring human–agent collaboration — an afterlife richer than the benchmark's headline era.

For an improvement loop, the operational lesson is sharp: a loop whose CHECK is a saturated-accuracy benchmark is blind on six axes at once (see Figure 3). The system can regress in efficiency, reliability, or scaffold-dependence — the dimensions production actually feels — while the accuracy needle sits pinned at the ceiling reading "fine."

LAUNCH SATURATION AFTERLIFE accuracy separates models headline era accuracy at ceiling usual reflex: retire it construct validity / shortcuts OOD generalizability efficiency · reliability model vs scaffold human–agent uplift six dimensions stay measurable after the ceiling — the stage the field throws away
Figure 3: The benchmark lifecycle with the stage the field discards. The six afterlife dimensions are those documented on CORE-Bench Hard by Nadgir et al. (2026); saturation ends accuracy's usefulness, not the benchmark's.
Key Takeaway 3

Saturation is the end of one measurement, not of measurement. Six dimensions — shortcuts, OOD transfer, efficiency, reliability, model-vs-scaffold, human uplift — stay informative past the accuracy ceiling, and the construct-validity problems only become visible there. Retiring a saturated benchmark discards the part of its life that audits your loop.

§4 · The portfolio view

If judges are unstable and benchmarks saturate, the next question is an economist's: how much measurement do you actually need to buy? Zeng and Papailiopoulos (2026) at Microsoft Research answer it with data. They compile a public score matrix of 84 frontier models on 133 benchmarks — 2,604 filled cells, 23.3% of the matrix — and find the matrix is approximately rank-2: a model's scores across all 133 benchmarks are largely determined by just two numbers, and two factors already explain over 90% of the variation among models on shared benchmarks.

Their system, BENCHPRESS, turns that structure into a tool: a logit-space rank-2 matrix-completion method that recovers held-out scores to within 4.6 points, with a confidence layer that flags which predictions can be trusted. The practical payoff is the subset result: five benchmarks — GPQA-D, HLE, Codeforces, MMLU-Pro, ARC-AGI-1 — recover the rest of a model's public scorecard to within 3.93 points; a cheaper set (GPQA-D, MMLU-Pro, Aider Polyglot, MATH-500, AIME 2026) predicts within 4.55 (see Figure 4). A release that runs 40-plus benchmarks is, on this evidence, mostly re-measuring two latent factors it already measured.

The rank-2 result cuts two ways, and the second is the one this series cares about. If most benchmarks are redundant projections of two factors, then a small portfolio buys almost all the signal — that is the efficiency reading. But it also means most of the eval surface cannot see anything the two factors do not see. Zhang et al. (2026) make the hidden dimension concrete with their Generalization Spectrum: for each training example, a controlled suite of test variants at increasing transfer distance — exact recall, implementation transfer across languages, context transfer under complete narrative re-framing, category-matched in-domain problems, and an unpaired baseline. Tracking performance across that spectrum reveals how far learning extends, per sample — a dimension invisible to every aggregate score in the 133-benchmark matrix. A measurement layer can be efficient and still be flat.

84 models × 133 benchmarks 23.3% filled (2,604 cells) rank-2 two factors ≈ >90% of shared-benchmark variation the 5-benchmark portfolio GPQA-D · HLE Codeforces · MMLU-Pro ARC-AGI-1 recovers the scorecard to within 3.93 points
Figure 4: The portfolio result of Zeng & Papailiopoulos (2026): an 84×133 score matrix (23.3% filled) is approximately rank-2, and a five-benchmark subset recovers full scorecards to within 3.93 points. Efficiency for the buyer — and a warning that most of the eval surface is redundant.
Key Takeaway 4

Eval portfolios have low-rank structure: two factors explain over 90% of shared-benchmark variation, and five well-chosen benchmarks recover a scorecard within ~4 points. Buy measurement like a portfolio — and remember that whatever the two factors miss, 133 redundant benchmarks miss too; per-sample transfer distance is one documented blind spot.

§5 · Fixing the instruments

The constructive wing of this literature is already building better instruments, and the designs share one move: decompose the opaque judgment into pieces that can be audited.

Cho et al. (2026) do it at the score level. BINEVAL replaces the holistic judge score with atomic binary questions: a meta-prompt generates fine-grained evaluation questions for the task, an LLM answers each independently per output, and the verdicts aggregate into interpretable, multi-dimensional scores. Across SummEval, Topical-Chat, and QAGS it matches or outperforms strong baselines including UniEval and G-Eval, with its strongest results on factual consistency — and, tellingly for §2's themes, it better matches human score distributions and avoids the ceiling effects that make prior judges unable to separate borderline from clearly flawed outputs. The same question-level feedback then feeds prompt improvement: the instrument doubles as the improvement signal, which is precisely the dual use this series' loops need.

Seddik and Fard (2026) do it at the representation level. Their axiomatic framework scores latent thought representations against four functional axioms — Causality, Minimality, Separability, Stability — each with a quantitative measure computed independently of downstream benchmark accuracy. Auditing open-weight LLMs across 23 reasoning tasks, they find no candidate satisfies all four axioms; representations distinguish task type reliably but cannot distinguish between two questions within the same task; and the failures are consistent across dense, reasoning-distilled, and RL-trained families — structural, not a size effect. Benchmark accuracy masks all of this. It is the same lesson as kappa deflation, one level down: the score can look fine while the thing the score is supposed to certify is absent.

Even the humble similarity metric earns a place in this audit. Li and Mehta (2026) — a practitioner survey out of BlackRock — organize the full lineage of text-similarity metrics, from lexical overlap through embedding-based to transformer-driven methods, under a classification that separates untrained from trained approaches. The reminder matters because similarity metrics sit inside judge pipelines and reward models as silent components; a measurement layer is only as sound as its least-audited stage.

Key Takeaway 5

The fixes share a shape: decompose, then audit the pieces. Binary-question evaluation restores interpretability and distribution-match; axiomatic representation metrics expose failures accuracy masks; even similarity metrics need lineage-aware selection. Instruments earn trust component by component, not holistically.

§6 · A shared failure vocabulary

Measurement science needs one more thing: a common language for what is being measured. Albayaydh, Zhao and Flechais (2026) at Oxford supply it, synthesizing 27 benchmark, taxonomy, and audit papers from 2023–2026 — spanning 19 distinct benchmarks — into a single cross-cutting taxonomy of agent limitations. Their six failure clusters: tool invocation and parameter-level errors; planning and constraint-satisfaction failures; long-horizon degradation from context accumulation; multi-agent coordination failures; safety and security failures under adversarial or underspecified conditions; and — closing the circle of this article — measurement validity problems, as a failure cluster of the field itself.

That last cluster is the right note to end on. When a synthesis of the evaluation literature must reserve one of its six categories for the invalidity of evaluations, the meta-problem has become a first-class problem. The vocabulary matters operationally: a team that logs its loop's failures against a shared taxonomy can compare across projects, notice when a "new" failure is cluster three wearing a costume, and — most relevant here — separate the agent failed from the measurement failed, which §2 showed are routinely conflated.

The loop card for this article's substrate follows directly. The evaluator V is not a fixture; it is a component with a maintenance schedule.

Loop Card · the eval-validity loop

SIGNAL — judge verdicts on a frozen, stratified anchor set with trusted human labels, logged every evaluation cycle.
UPDATE — the evaluator configuration V: judge choice, rubric, and benchmark portfolio (portfolio per the rank-2 result of Zeng & Papailiopoulos, 2026).
GUARDRAIL — the Minimum Viable Validation Protocol of Norman, Rivera & Hughes (2026): chance-corrected κ, cross-benchmark rank check, and a position-bias audit; a judge that passes raw agreement but fails the bias audit does not ship.
CHECK — quarterly: Cohen's κ (not raw agreement) on the anchor set, plus portfolio-predicted vs actually-measured scores within the completion tolerance; investigate any drift before trusting the quarter's improvement claims.

Key Takeaway 6

The field now has a shared taxonomy of agent failures — and its authors had to make "measurement validity" one of the six clusters. Treat your evaluator as a maintained component: re-validate it on a cadence, with chance-corrected statistics, or your improvement loop optimizes an instrument, not a capability.

What comes next

This article audited static instruments — judges and benchmarks that hold still while we measure them. But the systems from D1 do not hold still: a self-improving agent applies optimization pressure to whatever evaluates it, and a static V is not a maintained instrument but a consumable one. What happens when the test must learn back — when evaluation has to co-evolve with the agent it judges — is the next article.

References

  1. Norman, J. D., Rivera, M. U., & Hughes, D. A. (2026). Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias. arXiv preprint. arXiv:2606.19544.
  2. Zeng, Y., & Papailiopoulos, D. (2026). You Don't Need to Run Every Eval. arXiv preprint. arXiv:2606.24020.
  3. Nadgir, N., Kapoor, S., Liu, K., Kirgis, P., Orona, M., Rabanser, S., Bayer, T., Shetty, A., Ling, Y., Chan-Sew, D., Nakagawa, R., Utpala, S., Siegel, Z. S., & Narayanan, A. (2026). Life After Benchmark Saturation: A Case Study of CORE-Bench. arXiv preprint. arXiv:2606.26158.
  4. Zhang, J., Cheng, Z., Chen, S., Zhang, G., Huang, W., Liu, J., He, J., & Cai, T. (2026). The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms. arXiv preprint. arXiv:2606.25450.
  5. Cho, S., Chawla, K., Cai, P., Liu, Z., Zhu, C., Zhang, S.-X., & Sahu, S. (2026). Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement. arXiv preprint. arXiv:2606.27226.
  6. Seddik, F., & Fard, F. (2026). Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs. arXiv preprint. arXiv:2606.27378.
  7. Albayaydh, W., Zhao, R., & Flechais, I. (2026). Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents. arXiv preprint. arXiv:2607.05775.
  8. Li, M., & Mehta, D. (2026). A Review of Evaluation Metrics for Text Similarity. SSRN preprint, BlackRock, Inc.