Mechanical Dreams
An automatically generated podcast about machine learning and natural language processing. The two fictional hosts talk about papers that I want to learn more about on my way to work. It's not good, but it's useful.
- Indexed episodes, last 90 days
- 14
- Latest publication
- Aug 23, 2026
- Audience
- Checking…
- Earliest in this view
- Jul 12, 2026
Latest episodes
The State-Prediction Separation Hypothesis (opens the original)
Read excerpt
In this episode: • Introduction to State-Prediction Separation: Professor Norris and Linda introduce the paper and the core problem of standard Transformers conflating next-token prediction and future state representation. • The SPS Mechanism: Linda explains how the authors introduce a dummy predict token to separate the input state stream from the prediction stream. • Experimental Results and Baselines: The hosts discuss the massive data efficiency gains and how clever ablations like Delayed St
Pretraining Mixture of Experts for Emergent Modularity (opens the original)
Read excerpt
In this episode: • Introduction to Monolithic MoEs: Linda introduces EMO and Professor Norris discusses the memory bottlenecks of standard monolithic Large Language Models. • The Document-Level Routing Constraint: Linda explains EMOs core mechanism of restricting tokens in a document to a shared pool of experts. • Overcoming Training Instability: The hosts discuss the conflict with local load balancing and how global load balancing and dynamic pool sizing solved it. • Semantic Specialization and
Muon-Optimized Distillation and Quantization for LLM Deployment (opens the original)
Read excerpt
In this episode: • Introduction & The Edge Deployment Problem: Professor Norris and Linda introduce the episode's paper, focusing on the difficulty of running large language models on edge devices with limited memory and compute. • The Multi-Teacher Distillation Pipeline: Linda explains the dual-teacher approach, using Llama 4 Scout for synthetic data generation and Llama 3.3 70B for logit-based knowledge distillation to a 3B parameter student. • Hyperparameter Surprises with Optuna: The hosts d
Continuous-Query Limited Memory Language Models (opens the original)
Read excerpt
In this episode: • Welcome & The Parametric Memory Bottleneck: Professor Norris and Linda introduce the episode's topic, discussing the fundamental flaws of storing facts in model weights. • From Relational to Continuous Queries: Linda explains the limitations of prior Limited Memory Language Models (LMLMs) and introduces Co-LMLM's continuous-query approach. • Under the Hood: Training and The Token: The hosts dive into the technical details of how the model uses hidden states as retrieval querie
QuaRot (opens the original)
Read excerpt
In this episode: • Introduction to Quantization and Outliers: Norris and Linda introduce the episode's paper, QuaRot, and discuss the main hurdle in LLM inference: the memory bottleneck and the pesky outlier features in activations. • The Magic of Hadamard Rotations: Linda explains the core mechanism of QuaRot, using randomized Hadamard transformations to eliminate outliers through computational invariance. • Taming the Attention Mechanism and KV Cache: The hosts dive into the complexities of qu
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.