LLM Watch
Weekly newsletter about the most important AI research with a focus on Large Language Models (LLMs). Get insight on the cutting edge of AI from a human perspective.
- Indexed issues, last 90 days
- 6
- Latest publication
- Sep 30, 2026
- Audience
- Checking…
- Earliest in this view
- Jul 20, 2026
Latest issues
Just-in-Time Memory (opens the original)
Read excerpt
Most agent memory systems curate at write time. When a task finishes, the system reads the trajectory and distills it into a reflection, a workflow, a skill or a reasoning strategy, then stores that artifact for similarity retrieval later. The authors of Just-in-Time Memory name two costs of this design. Anything the distiller discards is gone for every future task that needed it. A single fixed summary also has to serve every future query, although one trajectory can teach several lessons. Thei
Papers You Should Know About (opens the original)
Read excerpt
Welcome to this issue of LLM Watch. The six pieces this week cover how to secure and check agents in production. They look at where agent code runs, which inputs judges and planners should trust, and how cheaply you can re-test an agent that changes every week.Showing a video judge the agent’s execution log makes open-weight judges accept most failed clips.A registry check placed before any gate blocks calls to tools that do not exist.One edited worker description can cut a multi-agent system’s
The Harness-Maxxing Trap (opens the original)
Read excerpt
In early August, a 27-billion-parameter open-weight model went head-to-head with a frontier model on Terminal-Bench 2.0. People were sceptical: Yes, the model was good. But not that good. Something wasn’t right. Take this GitHub issue: a developer trying to reproduce Qwen’s 53.5% SWE-bench Pro score got roughly 28% with a bash-only agent. Adding a single str_replace file-edit tool took it to 50.7%. One tool, one line in the harness, nearly doubled the score. Sounds reliab
LLM Watch Weekly: The Measurement Problem (opens the original)
Read excerpt
Welcome, Watcher! This week in LLM Watch:Enabling web search on ChatGPT reduced benchmark accuracy by up to 8 percentage points, and repeated runs of the same prompt produced inconsistent answers on up to 21% of prompts - raising hard questions about how we evaluate deployed AI systems.A neuro-symbolic RAG framework that compiles retrieved text into executable Prolog modules achieves 61.1% accuracy on ShARC, outperforming a standard RAG baseline’s 42.8% - without any domain-specific training.A n
Goal Engineering, or: Are We There Yet? (opens the original)
Read excerpt
The vocabulary around agents has added a floor roughly every quarter. Prompt engineering, then context engineering, then harness engineering, then loop engineering, and now, since about June, goal engineering. It would be easy to file this under branding churn, another layer of nouns invented to sell the same thing.Except that in the space of about eight weeks this spring, three different coding agents shipped a feature called /goal, and they shipped it to solve a problem none of the earlier lay
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.