SemiAnalysis
Bridging the gap between the world's most important industry, semiconductors, and business.
- Indexed issues, last 90 days
- 3
- Latest publication
- Sep 28, 2026
- Audience
- 253K subscribers
- Earliest in this view
- Sep 18, 2026
Latest issues
How GLM5.3 Sparse Attention Affects HBM Memory Usage (opens the original)
Read excerpt
How Sparse Attention Affects DRAM/NAND Memory How does sparse attention affect the TAM of memory, including HBM and NAND? Sparse attention selects top-k most relevant tokens to attend to, reducing the memory consumption and bandwidth requirements during the core Scaled Dot-Production Attention (SDPA) operation. However, the efficiency improvement doesn’t directly translate to overall memory savings in practice. Concretely, the top-k selection operation typically requires the full context to be i
ClusterMAX 3.0: The Industry Standard GPU Cloud Rating System Returns (opens the original)
Read excerpt
This post has bonus content for paid subscribers. Upgrade to get full access.In 8 months since our last major release of ClusterMAX, slavering investors have just about run out of pockets to stuff che
Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading (opens the original)
Read excerpt
Engram extends standard token embeddings with learned multi-token lookups. Recurring local patterns retrieve vectors directly, reducing the need to reconstruct them through attention and feed-forward layers. With Engram model architecture optimization, it allows for lower HBM capacity to be needed for models at the same quality. This does not mean there won’t be an insane demand for HBM but it just means that model architecture will continue to innovate around constraints.This model architecture
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
253K subscribers
Measured Sep 19, 2026
Source's subscribers, not the number who saw an individual piece.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.