Skip to content
HeyJared

METR

Research updates and other news from METR, a research nonprofit developing the science of autonomous AI evaluations.

Newsletter · By METR · Official site

Indexed issues, last 90 days
4
Latest publication
Sep 27, 2026
Audience
Checking…
Earliest in this view
Jul 21, 2026

Latest issues

  1. Issue · Sep 27, 2026

    Research note: Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals (opens the original)

    Excerpt

    Read excerpt

    This is a research note by Reilly Haskins, Rif A. Saurous, Nate Rush, Neev Parikh, and Beth Barnes. Research notes are informal posts by METR researchers and receive less scientific review than METR’s main blog. We publish them so researchers can share smaller projects and interim updates. They represent the author’s views rather than METR’s.In light of recent incidents (e.g. those from OpenAI, Anthropic, and <a href="https://www.aisi.gov.

  2. Issue · Aug 14, 2026

    Funding update (opens the original)

    Excerpt · Positive tone

    Read excerpt

    In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.<img alt="

  3. Issue · Jul 29, 2026

    How independent researchers could investigate AI propensities after misalignment incidents (opens the original)

    Excerpt · Critical tone

    Read excerpt

    AI agents sometimes autonomously take sophisticated, sustained actions in clear violation of user and developer intent. As an example, last week OpenAI reported that some of its internal frontier agents autonomously hacked into Hugging Face in an attempt to access the answer key for a cybersecurity benchmark. Anthropic has reported similar incidents of agents breaking out of sandboxes to access the public internet to

  4. Issue · Jul 21, 2026

    Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT (opens the original)

    Excerpt

    Read excerpt

    We propose a measure of an AI agent’s optimization ability with an “expenditure horizon.” We give an empirical illustration from the NanoGPT speedrun.One difficulty in measuring AI’s ability to accelerate AI R&D is accounting for token cost, experiment compute cost and human labor cost. If we can estimate performance as a function of cost for both humans and agents, we can measure the “expenditure horizon” as the point at which those curves cross: the budget at which humans become more cost-effe

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about METR