Skip to content
HeyJared

Mechanical Dreams

An automatically generated podcast about machine learning and natural language processing. The two fictional hosts talk about papers that I want to learn more about on my way to work. It's not good, but it's useful.

Podcast · By Mechanical Dirk · American English · Official site

Indexed episodes, last 90 days
14
Latest publication
Aug 23, 2026
Audience
Checking…
Earliest in this view
Jul 12, 2026
The latest indexed work is over 30 days old. There may be a gap in what we hold.

Latest episodes

  1. Episode · Aug 23, 2026

    The State-Prediction Separation Hypothesis (opens the original)

    Episode notes

    Read excerpt

    In this episode: • Introduction to State-Prediction Separation: Professor Norris and Linda introduce the paper and the core problem of standard Transformers conflating next-token prediction and future state representation. • The SPS Mechanism: Linda explains how the authors introduce a dummy predict token to separate the input state stream from the prediction stream. • Experimental Results and Baselines: The hosts discuss the massive data efficiency gains and how clever ablations like Delayed St

  2. Episode · Aug 22, 2026

    Pretraining Mixture of Experts for Emergent Modularity (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    In this episode: • Introduction to Monolithic MoEs: Linda introduces EMO and Professor Norris discusses the memory bottlenecks of standard monolithic Large Language Models. • The Document-Level Routing Constraint: Linda explains EMOs core mechanism of restricting tokens in a document to a shared pool of experts. • Overcoming Training Instability: The hosts discuss the conflict with local load balancing and how global load balancing and dynamic pool sizing solved it. • Semantic Specialization and

  3. Episode · Aug 21, 2026

    Muon-Optimized Distillation and Quantization for LLM Deployment (opens the original)

    Episode notes · Positive tone

    Read excerpt

    In this episode: • Introduction & The Edge Deployment Problem: Professor Norris and Linda introduce the episode's paper, focusing on the difficulty of running large language models on edge devices with limited memory and compute. • The Multi-Teacher Distillation Pipeline: Linda explains the dual-teacher approach, using Llama 4 Scout for synthetic data generation and Llama 3.3 70B for logit-based knowledge distillation to a 3B parameter student. • Hyperparameter Surprises with Optuna: The hosts d

  4. Episode · Aug 20, 2026

    Continuous-Query Limited Memory Language Models (opens the original)

    Episode notes · Positive tone

    Read excerpt

    In this episode: • Welcome & The Parametric Memory Bottleneck: Professor Norris and Linda introduce the episode's topic, discussing the fundamental flaws of storing facts in model weights. • From Relational to Continuous Queries: Linda explains the limitations of prior Limited Memory Language Models (LMLMs) and introduces Co-LMLM's continuous-query approach. • Under the Hood: Training and The Token: The hosts dive into the technical details of how the model uses hidden states as retrieval querie

  5. Episode · Aug 8, 2026

    QuaRot (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    In this episode: • Introduction to Quantization and Outliers: Norris and Linda introduce the episode's paper, QuaRot, and discuss the main hurdle in LLM inference: the memory bottleneck and the pesky outlier features in activations. • The Magic of Hadamard Rotations: Linda explains the core mechanism of QuaRot, using randomized Hadamard transformations to eliminate outliers through computational invariance. • Taming the Attention Mechanism and KV Cache: The hosts dive into the complexities of qu

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about Mechanical Dreams