Skip to content
HeyJared

Snacks Weekly on Data Science

This podcast is about making data science and machine learning knowledge accessible and less intimidating. Every week, I will handpick one selected industrial tech blog to break it down.

Podcast · By Pan Wu · English · Official site

Indexed episodes, last 90 days
13
Latest publication
Sep 28, 2026
Audience
Checking…
Earliest in this view
Jul 6, 2026

Latest episodes

  1. Episode · Sep 28, 2026

    Measuring AI Infrastructure Impact with Causal Inference [Meta] (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    In this episode, we discuss how Meta measures the business value of shared ML infrastructure when that value appears indirectly through downstream model launches. Instead of relying on adoption, usage, or traditional infrastructure metrics, the team built an observational causal framework around individual model launches. By comparing feasible adopters with comparable non-adopters and controlling for major confounding factors, the team estimated the incremental revenue associated with infrastruc

  2. Episode · Sep 21, 2026

    Building Taxonomies from Unstructured Text Using LLMs [Microsoft] (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    In this episode, we discuss two LLM-assisted pipelines for turning unstructured text into a structured taxonomy with minimal manual annotation. The bottom-up approach lets semantic representations and clustering discover categories from the data, making it useful for open-ended exploration and relatively efficient at scale. The top-down approach starts with LLM-generated categories and recursively classifies and subdivides the corpus, making it more controllable and easier to align with business

  3. Episode · Sep 14, 2026

    Building Semantic IDs for Product Understanding at Scale [Instacart] (opens the original)

    Episode notes

    Read excerpt

    In this episode, we talk about how Instacart identified a limitation of a standard taxonomy: it could classify products, but it couldn't connect them the way real shoppers do. The team developed a solution around Semantic IDs, which compress product embeddings into hierarchical discrete codes, while contrastive learning and taxonomy-based supervision make those codes more semantically meaningful. The result is a representation that can support better product understanding, catalog quality, and r

  4. Episode · Sep 7, 2026

    Agent-driven ML Exploration [Faire] (opens the original)

    Episode notes · Positive tone

    Read excerpt

    In this episode, we discuss how Faire addresses a fundamental bottleneck in machine learning development: there are more potential experiments than human scientists can realistically run. They developed Autoscience, giving an AI agent a structured environment where it can plan, execute, evaluate, remember, and continuously refine experiments. What makes the approach effective is not the agent alone, but the combination of a reliable experimentation harness, integration with the existing ML stack

  5. Episode · Aug 31, 2026

    Building food metadata with LLM juries [DoorDash] (opens the original)

    Episode notes · Neutral tone

    Read excerpt

    In this episode, we discuss how DoorDash built a reliable metadata generation platform that could operate across millions of constantly changing menu items. The solution combined multimodal generation, automated evaluation through LLM juries, optimization loops for improving model behavior, and distributed infrastructure for cost-efficient scaling. By transforming unstructured menu information into precise, structured attributes, DoorDash unlocked many new product capabilities that ultimately he

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Not enough subject data for this period yet.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about Snacks Weekly on Data Science