Skip to content
HeyJared

Machine Learning with PyTorch

I’m Abdulsalam Bande, a computer scientist. This channel explains machine learning concepts in PyTorch, with a focus on covering PyTorch’s built-in functions and practical implementations. For any collaborations, you can reach me at: annual03froze [at] icloud.com.

YouTube · Official site

Indexed videos, last 90 days
5
Latest publication
Sep 1, 2026
Audience
~3.4K subscribers
Earliest in this view
Jul 16, 2026

Latest videos

  1. Video · Sep 1, 2026

    Smarter MoE Routing for Your PyTorch Models (opens the original)

    Excerpt · Neutral tone · 14 min · 136 views by Sep 30, 2026

    Read excerpt

    In this video, we explore Path-Constrained Mixture-of-Experts (PathMoE), a paper from Apple and Google that studies the full sequence of expert choices a token makes across Transformer layers. We cover the intuition behind MoE routing, expert paths, routing entropy, and path concentration, then look at how PathMoE shares router parameters across blocks of layers to create more coordinated routing. We also build a small PyTorch implementation from scratch and visualize how expert paths become mor

  2. Video · Aug 19, 2026

    nn.Sequential vs nn.ModuleList in PyTorch — What’s the Difference? (opens the original)

    Excerpt · 2 min · 70 views by Sep 30, 2026

    Read excerpt

    nn.Sequential or nn.ModuleList — which should you use? In this quick tutorial, we see how nn.Sequential automatically passes data through modules in order, while nn.ModuleList registers modules but gives you control over how they are executed.

  3. Video · Aug 16, 2026

    FlashAttention Finally Makes Sense: A Step-by-Step Numerical Example (opens the original)

    Excerpt · Positive tone · 12 min · 183 views by Sep 30, 2026

    Read excerpt

    FlashAttention is often described as “attention, but faster”—but how does it actually work? In this video, I explain FlashAttention step by step using a complete numerical example. We begin with standard attention and then compute the same output block by block using tiling and online softmax. You will learn: - Why FlashAttention tracks the running maximum `m` - What the denominator `l` and accumulator `acc` represent - Why old values must be rescaled when the running maximum changes - How Flash

  4. Video · Jul 28, 2026

    How EpiCache Makes LLM Memory Efficient (opens the original)

    Excerpt · Neutral tone · 21 min · 122 views by Sep 30, 2026

    Read excerpt

    How can Large Language Models remember long conversations without storing massive KV caches? In this video, we break down EpiCache, a memory-efficient approach that introduces episodic KV caches for long-context conversations. Instead of rebuilding or storing the full KV cache, EpiCache organizes conversation history into semantic episodes, retrieves only the most relevant memory, and keeps memory usage bounded. In this video you'll learn: What problem EpiCache solves Why KV caches become a bott

  5. Video · Jul 16, 2026

    Why vLLM Is So Fast (Explained Simply) (opens the original)

    Title only · 14 min · 297 views by Sep 30, 2026

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

~3.4K subscribers

Measured Sep 19, 2026

Source's subscribers, not the number who saw an individual piece.

How this was measured

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about Machine Learning with PyTorch