Skip to content
HeyJared

LLM Watch

Weekly newsletter about the most important AI research with a focus on Large Language Models (LLMs). Get insight on the cutting edge of AI from a human perspective.

Newsletter · By Pascal Biese · English · Paid tier available · Official site

Indexed issues, last 90 days
6
Latest publication
Sep 30, 2026
Audience
Checking…
Earliest in this view
Jul 20, 2026

Latest issues

  1. Issue · Sep 30, 2026

    Just-in-Time Memory (opens the original)

    Excerpt

    Read excerpt

    Most agent memory systems curate at write time. When a task finishes, the system reads the trajectory and distills it into a reflection, a workflow, a skill or a reasoning strategy, then stores that artifact for similarity retrieval later. The authors of Just-in-Time Memory name two costs of this design. Anything the distiller discards is gone for every future task that needed it. A single fixed summary also has to serve every future query, although one trajectory can teach several lessons. Thei

  2. Issue · Sep 25, 2026

    Papers You Should Know About (opens the original)

    Excerpt

    Read excerpt

    Welcome to this issue of LLM Watch. The six pieces this week cover how to secure and check agents in production. They look at where agent code runs, which inputs judges and planners should trust, and how cheaply you can re-test an agent that changes every week.Showing a video judge the agent’s execution log makes open-weight judges accept most failed clips.A registry check placed before any gate blocks calls to tools that do not exist.One edited worker description can cut a multi-agent system’s

  3. Issue · Aug 25, 2026

    The Harness-Maxxing Trap (opens the original)

    Excerpt · Critical tone

    Read excerpt

    In early August, a 27-billion-parameter open-weight model went head-to-head with a frontier model on Terminal-Bench 2.0. People were sceptical: Yes, the model was good. But not that good. Something wasn’t right. Take this GitHub issue: a developer trying to reproduce Qwen’s 53.5% SWE-bench Pro score got roughly 28% with a bash-only agent. Adding a single str_replace file-edit tool took it to 50.7%. One tool, one line in the harness, nearly doubled the score. Sounds reliab

  4. Issue · Aug 7, 2026

    LLM Watch Weekly: The Measurement Problem (opens the original)

    Excerpt · Critical tone

    Read excerpt

    Welcome, Watcher! This week in LLM Watch:Enabling web search on ChatGPT reduced benchmark accuracy by up to 8 percentage points, and repeated runs of the same prompt produced inconsistent answers on up to 21% of prompts - raising hard questions about how we evaluate deployed AI systems.A neuro-symbolic RAG framework that compiles retrieved text into executable Prolog modules achieves 61.1% accuracy on ShARC, outperforming a standard RAG baseline’s 42.8% - without any domain-specific training.A n

  5. Issue · Jul 29, 2026

    Goal Engineering, or: Are We There Yet? (opens the original)

    Excerpt · Critical tone

    Read excerpt

    The vocabulary around agents has added a floor roughly every quarter. Prompt engineering, then context engineering, then harness engineering, then loop engineering, and now, since about June, goal engineering. It would be easy to file this under branding churn, another layer of nouns invented to sell the same thing.Except that in the space of about eight weeks this spring, three different coding agents shipped a feature called /goal, and they shipped it to solve a problem none of the earlier lay

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about LLM Watch