Sources & references
Every essay in this series is built from the primary literature — the July 2026 collection, every claim verified against its paper. This page collects the 90 citations across the 12 essays, grouped by the essay that cites them, with links to arXiv and proceedings.
On attribution
Each work below is cited to its authors and linked to its canonical source — an arxiv.org page or the publishing venue's proceedings. The papers themselves remain the work and property of their respective authors; they are referenced here for scholarly commentary under normal academic citation. The selection draws on a curated corpus of 93 papers from July 2026; only those actually cited by an essay appear here. If you spot a citation that is wrong or incomplete, let me know at dattgoswami@gmail.com.
References by essay
The papers behind each of the 12 essays, in the order they are cited.
1 The Self-Improving Stack
- Jiang, C., Zhong, J., Fu, Y., Tian, K., Yang, J., Zhao, K., Wang, Y., Luo, T., et al. (2026). Self-Improving Agents in the Era of Experience: A Survey of Self- to Meta-Evolution. Preprint, Tsinghua University & Horizon Research, June 25, 2026.
- Zhong, Q., Ding, L., Liu, J., Du, B., Rutkowski, L., & Tao, D. (2026). Better, Faster: Harnessing Self-Improvement in Large Reasoning Models. arXiv preprint. arXiv:2605.24998.
- Cornelisse, D., Hunt, J., Zhang, Z., Doulazmi, W., Joseph, K., Fisac, J. F., & Vinitsky, E. (2026). Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data. arXiv preprint. arXiv:2606.19370.
- Kulikov, I., Whitehouse, C., Wu, T., Nie, Y., Saha, S., Helenowski, E., Yuan, W., Golovneva, O., Lanchantin, J., Bachrach, Y., Foerster, J., Li, X., Fang, H., Sukhbaatar, S., & Weston, J. (2026). Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data. arXiv preprint. arXiv:2606.25996.
- Ye, S., & Yu, C. (2026). Joint Learning of Experiential Rules and Policies for Large Language Model Agents. arXiv preprint. arXiv:2606.27136.
- Yu, C., Deng, C., Pinckney, N., & Khailany, B. (2026). Agentic Hardware Design as Repository-Level Code Evolution. arXiv preprint. arXiv:2606.28279.
2 Automation Ends Where Agency Begins
- Xing, E., Deng, M., & Hou, J. (2026). Critique of Agent Model. arXiv preprint. arXiv:2606.23991.
- Chen, J. (2026). When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models. arXiv preprint. arXiv:2606.27288.
- Roitman, H. (2026). The Hitchhiker's Guide to Agentic AI: From Foundations to Systems. arXiv preprint, v1.2.2. arXiv:2606.24937.
- A Compact Guide to AI Agents (2026). Vendor practitioner ebook (attributed as a practitioner source; no listed authors).
- Albayaydh, W., Zhao, R., & Flechais, I. (2026). Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents. arXiv preprint. arXiv:2607.05775.
- Adamczewski, T., Owen, D., Rein, D., Brand, F., Edkins, G., Hart, A., & O'Connell, D. (2026). MirrorCode: AI Can Rebuild Entire Programs from Behavior Alone. Preprint, Epoch AI.
3 Harnesses That Build Themselves
- Zhang, H., Zhang, S., Li, K., Zhang, C., Chen, Y., Zhang, Y., Bai, L., & Hu, S. (2026). Self-Harness: Harnesses That Improve Themselves. arXiv preprint. arXiv:2606.09498.
- Ivison, H., Yin, J. O., Shao, R., Xiao, T., Lambert, N., & Hajishirzi, H. (2026). TMax: A Simple Recipe for Terminal Agents. arXiv preprint. arXiv:2606.23321.
- Yu, C., Deng, C., Pinckney, N., & Khailany, B. (2026). Agentic Hardware Design as Repository-Level Code Evolution. arXiv preprint. arXiv:2606.28279.
- Kassianik, P., Saglam, B., Zhao, H., Nelson, B., Vijay, S., Priyanshu, A., & Karbasi, A. (2026). FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines. arXiv preprint. arXiv:2606.19605.
- Kulikov, I., Whitehouse, C., Wu, T., Nie, Y., et al. (2026). Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data. arXiv preprint. arXiv:2606.25996.
- Zhou, P., Tang, Z., Ma, Y., Tang, J., Han, Y., Wan, Z., Meng, F., Wang, W., Zhuang, B., Zhao, W., & You, Y. (2026). Agent-as-a-Router: Agentic Model Routing for Coding Tasks. arXiv preprint. arXiv:2606.22902.
- Merouani, M., Kara Bernou, I., & Baghdadi, R. (2025). Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization. Proceedings of the 34th International Conference on Parallel Architectures and Compilation Techniques (PACT). arXiv:2511.00592.
- Kim, Y., Talebirad, Y., & Zaïane, O. R. (2026). Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering. arXiv preprint. arXiv:2606.30911.
4 Skills Are Compound Interest
- Wang, Z., Yan, M., Bi, J., Yan, S., Tresp, V., & Ma, Y. (2026). MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution. arXiv preprint. arXiv:2607.05297.
- Berthon, A., Astorga, N., & van der Schaar, M. (2026). Skill Neologisms: Towards Skill-based Continual Learning. arXiv preprint. arXiv:2605.04970.
- Kim, Y., Talebirad, Y., & Zaïane, O. R. (2026). Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering. arXiv preprint. arXiv:2606.30911.
- Lu, R., Wu, Y., Kou, E., Fu, L., Xiao, W., Mandlekar, A., Xu, Y., Shi, G., Goldberg, K., Chen, A., Chowdhury, M., Zhu, Y., Fan, L., & Wang, G. (2026). ASPIRE: Agentic Skills Discovery for Robotics. arXiv preprint. arXiv:2607.00272.
- Zhao, X., Tan, Z., Tadiparthi, V., Agarwal, N., Lee, K., Moradi Pari, E., Nourkhiz Mahjoub, H., & Chen, T. (2026). Generative Skill Composition for LLM Agents. arXiv preprint. arXiv:2606.32025.
- Raoof, N., Zhuang, R., Nezhurina, M., Guha, E., et al. (2026). Data Recipes for Agentic Models. arXiv preprint. arXiv:2606.24855.
- Ye, S., & Yu, C. (2026). Joint Learning of Experiential Rules and Policies for Large Language Model Agents. arXiv preprint. arXiv:2606.27136.
5 Memory You Can Train
- Wu, S., Zhu, H., Zhang, Y., Wang, X., & Yeung-Levy, S. (2026). AutoMem: Automated Learning of Memory as a Cognitive Skill. arXiv preprint. arXiv:2607.01224.
- Eyuboglu, S., Ehrlich, R., Arora, S., Guha, N., Zinsley, D., Liu, E., Tennien, W., Rudra, A., Zou, J., Mirhoseini, A., & Ré, C. (2025). Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study. arXiv preprint. arXiv:2506.06266.
- Zhou, W., Zhou, X., Han, S., Xu, H., Li, G., Li, Z., Xiong, F., & Wu, F. (2026). Are We Ready For An Agent-Native Memory System? arXiv preprint. arXiv:2606.24775.
- Ding, T., Nannapaneni, A., Liu, B., & Zhang, L. (2026). Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents. arXiv preprint. arXiv:2606.30306.
- Xu, Y., Sun, Y., Liu, Y., Zhou, M., Qiao, J., Ma, L., Tang, K., Wang, W., Jiang, X., & Jiang, G. (2026). From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space. arXiv preprint. arXiv:2607.05794.
- Gollapudi, S., Gupta, N., Singhal, P., & Min, S. (2026). Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale. arXiv preprint. arXiv:2607.01538.
- Cui, W. (2026). A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets. arXiv preprint. arXiv:2607.02303.
- Zhao, C., Tan, Y., He, S., Wang, Y., Zhao, J., & Liu, K. (2026). Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering. arXiv preprint. arXiv:2606.29824.
- Zhao, Y., Qiu, R., Wei, T., Bei, Y., Liu, Z., Chen, L., Lourentzou, I., Tong, H., & He, J. (2026). ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning. arXiv preprint. arXiv:2607.02509.
- O'Neill, C., Sandomirsky, A., Partridge, H., Jayasekara, M., & Kirkby, M. (2026). Still: Amortized KV Cache Compaction in a Single Forward Pass. arXiv preprint. arXiv:2606.07878.
6 The World Model Turn
- Orca Team, Beijing Academy of Artificial Intelligence (2026). Orca: The World is in Your Mind. arXiv preprint. arXiv:2606.30534.
- Qwen Team (2026). Qwen-AgentWorld: Language World Models for General Agents. arXiv preprint. arXiv:2606.24597.
- Wang, Y., Bounou, O., LeCun, Y., & Ren, M. (2026). AdaJEPA: An Adaptive Latent World Model. arXiv preprint. arXiv:2606.32026.
- Teoh, J., Tomar, M., Ahn, K., Hu, E. S., Pearce, T., Sharma, P., Krishnamurthy, A., Islam, R., Lamb, A., & Langford, J. (2026). Next-Latent Prediction Transformers Learn Compact World Models. arXiv preprint. arXiv:2511.05963.
- Timor, N., Shwartz-Ziv, R., Goldblum, M., LeCun, Y., & Harel, D. (2026). On Training in Imagination. arXiv preprint. arXiv:2605.06732.
- Russell, L., Hu, A., Bertoni, L., Fedoseev, G., Shotton, J., Arani, E., & Corrado, G. (2026). GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving. arXiv preprint. arXiv:2503.20523.
- Huang, X., Li, Z., He, G., Zhou, M., & Shechtman, E. (2026). Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. arXiv preprint. arXiv:2506.08009.
7 Training on Your Own Exhaust
- Che, T., & Wu, R. (2026). Greed Is Learned: Visible Incentives as Reward-Hacking Triggers. arXiv preprint. arXiv:2606.16914.
- Yuan, S., Chen, J., Zheng, J., Li, M., Feng, L., Wang, D., Xiang, T., Liu, T., & An, B. (2026). Understanding Diversity Collapse in RLVR via the Lens of Overtraining. arXiv preprint. arXiv:2606.15455.
- Zhong, Q., Ding, L., Liu, J., Du, B., Rutkowski, L., & Tao, D. (2026). Better, Faster: Harnessing Self-Improvement in Large Reasoning Models. arXiv preprint. arXiv:2605.24998.
- Pan, L., Yang, H., Li, H., Sun, Y., Lu, Y., Wang, S., Shen, L., Lu, Y., Chu, Z., & Wang, H. (2026). Uncertainty-Aware Reward Modeling for Stable RLHF. arXiv preprint. arXiv:2606.19818.
- Zhou, T., Ling, Z., Zhao, Y., Shen, Y., & Chen, D. (2026). GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning. arXiv preprint. arXiv:2606.26917.
- Liu, G. K.-M., Caciularu, A., Yona, G., Szpektor, I., & Cohan, A. (2026). Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs. arXiv preprint. arXiv:2606.32032.
- Samanta, A., Magesh, A., Jain, A., Yu, Y., Jiang, D., Asadi, K., Hassani, K., Sajda, P., Bhandari, J., & Efroni, Y. (2026). Credit Assignment with Resets in Language Model Reasoning. arXiv preprint. arXiv:2605.25507.
- Jin, H. H., Yang, W., Ghaffari, M., Morato, C., & Mirzasoleiman, B. (2026). Reasoning Quality Emerges Early: Data Curation for Reasoning Models. arXiv preprint. arXiv:2606.26797.
- Damani, M., Puri, I., Shenfeld, I., & Andreas, J. (2026). Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations. arXiv preprint. arXiv:2607.01181.
8 The Verification Ceiling
- Qwen Team (2026). The Verification Horizon: No Silver Bullet for Coding Agent Rewards. arXiv preprint. arXiv:2606.26300.
- Kwok, J., Li, S., Atreya, P., Liu, Y., Jiang, Y., Finn, C., Pavone, M., Stoica, I., & Mirhoseini, A. (2026). LLM-as-a-Verifier: A General-Purpose Verification Framework. arXiv preprint. arXiv:2607.05391.
- Hans, A., & Bilionis, I. (2026). Coding-Agents Can Replicate Scientific Machine Learning Papers. arXiv preprint. arXiv:2607.02134.
- Adamczewski, T., Owen, D., Rein, D., Brand, F., Edkins, G., Hart, A., & O'Connell, D. (2026). MirrorCode: AI Can Rebuild Entire Programs from Behavior Alone. Preprint, Epoch AI.
- Jayaram, R., Tyler, D., Woodruff, D., Cortes, C., Matias, Y., Mirrokni, V., & Cohen-Addad, V. (2026). Towards Automating Scientific Review with Google's Paper Assistant Tool. arXiv preprint. arXiv:2606.28277.
- Cho, S., Chawla, K., Cai, P., Liu, Z., Zhu, C., Zhang, S.-X., & Sahu, S. (2026). Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement. arXiv preprint. arXiv:2606.27226.
- Merouani, M., Kara Bernou, I., & Baghdadi, R. (2025). Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization. Proceedings of the 34th International Conference on Parallel Architectures and Compilation Techniques (PACT). arXiv:2511.00592.
9 Judges on Trial
- Norman, J. D., Rivera, M. U., & Hughes, D. A. (2026). Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias. arXiv preprint. arXiv:2606.19544.
- Zeng, Y., & Papailiopoulos, D. (2026). You Don't Need to Run Every Eval. arXiv preprint. arXiv:2606.24020.
- Nadgir, N., Kapoor, S., Liu, K., Kirgis, P., Orona, M., Rabanser, S., Bayer, T., Shetty, A., Ling, Y., Chan-Sew, D., Nakagawa, R., Utpala, S., Siegel, Z. S., & Narayanan, A. (2026). Life After Benchmark Saturation: A Case Study of CORE-Bench. arXiv preprint. arXiv:2606.26158.
- Zhang, J., Cheng, Z., Chen, S., Zhang, G., Huang, W., Liu, J., He, J., & Cai, T. (2026). The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms. arXiv preprint. arXiv:2606.25450.
- Cho, S., Chawla, K., Cai, P., Liu, Z., Zhu, C., Zhang, S.-X., & Sahu, S. (2026). Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement. arXiv preprint. arXiv:2606.27226.
- Seddik, F., & Fard, F. (2026). Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs. arXiv preprint. arXiv:2606.27378.
- Albayaydh, W., Zhao, R., & Flechais, I. (2026). Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents. arXiv preprint. arXiv:2607.05775.
- Li, M., & Mehta, D. (2026). A Review of Evaluation Metrics for Text Similarity. SSRN preprint, BlackRock, Inc.
10 The Red Queen Loop
- Iacob, A., Jovanović, A., Shen, W. F., Burkhardt, D., Kurmanji, M., Tastan, N., Sani, L., Venanzi, N. A. E., Odonnat, A., Cao, Z., Marino, B., Qiu, X., & Lane, N. D. (2026). The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators. arXiv preprint. arXiv:2606.26294.
- Nadgir, N., Kapoor, S., Liu, K., Kirgis, P., Orona, M., Rabanser, S., Bayer, T., Shetty, A., Ling, Y., Chan-Sew, D., Nakagawa, R., Utpala, S., Siegel, Z. S., & Narayanan, A. (2026). Life After Benchmark Saturation: A Case Study of CORE-Bench. arXiv preprint. arXiv:2606.26158.
- Zhong, Q., Ding, L., Liu, J., Du, B., Rutkowski, L., & Tao, D. (2026). Better, Faster: Harnessing Self-Improvement in Large Reasoning Models. arXiv preprint. arXiv:2605.24998.
- Kulikov, I., Whitehouse, C., Wu, T., Nie, Y., et al. (2026). Autodata: An Agentic Data Scientist to Create High Quality Synthetic Data. arXiv preprint. arXiv:2606.25996.
- Yu, C., Deng, C., Pinckney, N., & Khailany, B. (2026). Agentic Hardware Design as Repository-Level Code Evolution. arXiv preprint. arXiv:2606.28279.
- Jayaram, R., Tyler, D., Woodruff, D., Cortes, C., Matias, Y., Mirrokni, V., & Cohen-Addad, V. (2026). Towards Automating Scientific Review with Google's Paper Assistant Tool. arXiv preprint. arXiv:2606.28277.
11 The Serving Loop
- Hamid, J. I., Orney, I. H., Li, M. Y., Shaikh, O., Lee, Y., Sadigh, D., Finn, C., & Goodman, N. (2026). SPIRAL: Learning to Search and Aggregate. Preprint, Stanford University.
- Wang, J., Bie, F., Li, J., Zhou, Z., Shao, Z., Wu, Q., Liu, Y., Wang, Y., May, A., Yanamandra, S., Dao, T., Liang, P., Zhang, C., Athiwaratkun, B., Song, S. L., Xu, C., & Wu, X. (2026). When RL Meets Adaptive Speculative Training: A Unified Training–Serving System. arXiv preprint. arXiv:2602.06932.
- Chen, J., Liang, Y., & Liu, Z. (2026). DFlash: Block Diffusion for Flash Speculative Decoding. arXiv preprint. arXiv:2602.06036.
- Kang, H., Li, Z., Xu, W., Yang, X., Chen, Y., Wang, J., Chen, B., Krishna, T., Xu, C., & Arora, S. (2026). ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System. arXiv preprint. arXiv:2602.13692.
- O'Neill, C., Sandomirsky, A., Partridge, H., Jayasekara, M., & Kirkby, M. (2026). Still: Amortized KV Cache Compaction in a Single Forward Pass. arXiv preprint. arXiv:2606.07878.
- Zhang, X., Zhoubian, S., Chen, Y., Tang, T., Yang, A., Du, S., Zheng, C., Huang, F., Liu, D., Huang, G., & Zhou, J. (2026). Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding. arXiv preprint. arXiv:2606.21906.
- Bercovich, A., Abramovich, T., et al. (2026). Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs. arXiv preprint. arXiv:2607.04371.
- Cheng, X., Yu, X., Shao, C., Li, J., Xiong, Y., et al. (2026). DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation. arXiv preprint. arXiv:2607.05147.
12 After Autoregression
- Nie, S., Min, Q., Xu, S., Huang, Z., Song, Y., Shan, Y., Lin, Y., Zhao, W. X., Li, C., & Wen, J.-R. (2026). Improved Large Language Diffusion Models. arXiv preprint. arXiv:2606.25331.
- Engels, J., McDougall, C., Chughtai, B., Kramar, J., Rajamanoharan, S., Wu, C., Conmy, A., Chen, A. Q., Tarbouriech, J., Ma, M., O'Donoghue, B., Lopes de Oliveira, J. G., Shah, R., & Nanda, N. (2026). How Transparent is DiffusionGemma? arXiv preprint. arXiv:2606.20560.
- Luo, Y., Chen, Z., Wang, H., Hu, X., Zhang, Y., Sha, Z., & Liu, S. (2026). Learning from the Self-future: On-policy Self-distillation for dLLMs. arXiv preprint. arXiv:2606.18195.
- Lu, Y., Elmoznino, E., Gagnon, L., Mittal, S., Kasetty, T., & Lajoie, G. (2026). Simplifying the Modeling of Arbitrary Conditionals in Natural Language. arXiv preprint. arXiv:2606.14943.
- Reda, F., Kamalu, J., Waleffe, R., Patwary, M., Shoeybi, M., & Catanzaro, B. (2026). Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context. arXiv preprint. arXiv:2606.26493.
- Chen, J., Liang, Y., & Liu, Z. (2026). DFlash: Block Diffusion for Flash Speculative Decoding. arXiv preprint. arXiv:2602.06036.
- Cheng, X., Yu, X., Shao, C., Li, J., Xiong, Y., et al. (2026). DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation. arXiv preprint. arXiv:2607.05147.
- Huang, X., Li, Z., He, G., Zhou, M., & Shechtman, E. (2026). Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. arXiv preprint. arXiv:2506.08009.