Skip to content
HeyJared

Redwood Research Blog

Narrations of Redwood Research blog posts.Redwood Research is a research nonprofit based in Berkeley. We investigate risks posed by the development of powerful artificial intelligence and techniques for mitigating those risks.

Podcast · By Redwood Research · British English · US · Official site

Indexed episodes, last 90 days
12
Latest publication
Sep 23, 2026
Audience
Checking…
Earliest in this view
Jul 23, 2026

Latest episodes

  1. Episode · Sep 23, 2026

    “Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, Nathan Sheffield, Ryan Greenblatt (opens the original)

    Episode notes · Critical tone

    Read excerpt

    Subtitle: We should have a strong presumption that latent reasoning architectures would make oversight far more difficult. Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight mu

  2. Episode · Sep 23, 2026

    “Astra is much better at reasoning with filler tokens than previous models” by Dylan Xu, Sebastian Prasanna, Alek Westover (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    We measure GPT-6-Astra's capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra's performance (e.g., improving from ~60% to ~90% o

  3. Episode · Sep 11, 2026

    “CoT controllability evals seem very under-elicited” by Arun Jose (opens the original)

    Episode notes · Critical tone

    Read excerpt

    Subtitle: Simple prompt optimizations can improve model capability to control their reasoning. The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their curren

  4. Episode · Sep 10, 2026

    “Proposal for tracking the effects of architecture on monitorability” by Ryan Greenblatt, Alek Westover, Lukas Finnveden (opens the original)

    Episode notes

    Read excerpt

    Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward).[1] As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs

  5. Episode · Aug 27, 2026

    “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” by Ryan Greenblatt (opens the original)

    Episode notes · Critical tone

    Read excerpt

    We recently published the report from our brief independent investigation into this incident. You can read the full report here. Here is our tweet thread summarizing what we found: METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7 to 13 (the period OpenAI defin

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about Redwood Research Blog