Machine Learning with PyTorch
I’m Abdulsalam Bande, a computer scientist. This channel explains machine learning concepts in PyTorch, with a focus on covering PyTorch’s built-in functions and practical implementations. For any collaborations, you can reach me at: annual03froze [at] icloud.com.
- Indexed videos, last 90 days
- 5
- Latest publication
- Sep 1, 2026
- Audience
- ~3.4K subscribers
- Earliest in this view
- Jul 16, 2026
Latest videos
Smarter MoE Routing for Your PyTorch Models (opens the original)
Read excerpt
In this video, we explore Path-Constrained Mixture-of-Experts (PathMoE), a paper from Apple and Google that studies the full sequence of expert choices a token makes across Transformer layers. We cover the intuition behind MoE routing, expert paths, routing entropy, and path concentration, then look at how PathMoE shares router parameters across blocks of layers to create more coordinated routing. We also build a small PyTorch implementation from scratch and visualize how expert paths become mor
nn.Sequential vs nn.ModuleList in PyTorch — What’s the Difference? (opens the original)
Read excerpt
nn.Sequential or nn.ModuleList — which should you use? In this quick tutorial, we see how nn.Sequential automatically passes data through modules in order, while nn.ModuleList registers modules but gives you control over how they are executed.
FlashAttention Finally Makes Sense: A Step-by-Step Numerical Example (opens the original)
Read excerpt
FlashAttention is often described as “attention, but faster”—but how does it actually work? In this video, I explain FlashAttention step by step using a complete numerical example. We begin with standard attention and then compute the same output block by block using tiling and online softmax. You will learn: - Why FlashAttention tracks the running maximum `m` - What the denominator `l` and accumulator `acc` represent - Why old values must be rescaled when the running maximum changes - How Flash
How EpiCache Makes LLM Memory Efficient (opens the original)
Read excerpt
How can Large Language Models remember long conversations without storing massive KV caches? In this video, we break down EpiCache, a memory-efficient approach that introduces episodic KV caches for long-context conversations. Instead of rebuilding or storing the full KV cache, EpiCache organizes conversation history into semantic episodes, retrieves only the most relevant memory, and keeps memory usage bounded. In this video you'll learn: What problem EpiCache solves Why KV caches become a bott
Why vLLM Is So Fast (Explained Simply) (opens the original)
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
~3.4K subscribers
Measured Sep 19, 2026
Source's subscribers, not the number who saw an individual piece.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.