The Kaitchup – AI on a Budget
Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated.
- Indexed issues, last 90 days
- 16
- Latest publication
- Sep 30, 2026
- Audience
- Checking…
- Earliest in this view
- Aug 20, 2026
Latest issues
Qwen3.8 Flash Next GGUF Benchmark: Q4 to Q1 Accuracy and Token Efficiency (opens the original)
Read excerpt
Quantizing Qwen3.8 Flash Next is almost unavoidable if you want to run it locally. The BF16 original model occupies about 354 GB.But after quantization, how much of the model survives?A small file is useful only if the model can still do the work. And accuracy is only half of the question. If the compressed model needs twice as many output tokens to solve the same tasks, some of the apparent efficiency gain disappears. As we will see here, this “efficiency” evaluation is especially relevant for
ThinkingCap-Qwen3.8-27B: Less Thinking, Similar Accuracy (opens the original)
Read excerpt
<img alt="" class="sizing-normal" height="815" src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png"
Boosting Agentic Coding with LLM Retries: Lessons from Qwen3.8 27B on DeepSWE (opens the original)
Read excerpt
<img alt="" class="sizing-large" height="699.7252747252747" src="https://substackcdn.com/image/fetch/$s_!rTz9!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec9fdf39-a926-4ea6-b24a-720ff491cd45_9
Fastest Qwen3.8 27B Quantization? Benchmarking NVFP4, INT4, GSQ, and MTP (opens the original)
Read excerpt
Last week, I published my analysis of several quantized versions of Qwen3.8 27B.Unsloth NVFP4, NVIDIA NVFP4, AWQ INT4, GSQ 3-bit, Escha-W2, and my mixed INT3/INT4 models use very different quantization recipes. Some quantize activations, some keep them at higher precision, some preserve more layers in FP8 or BF16, and some include MTP weights while others do not.They also run at very different speeds.MTP can increase decoding throughput substantially by predicting several tokens at once, but the
Bonsai 2 27B: Qwen3.8 in 5.9 GB, but What About Agentic Coding? (opens the original)
Read excerpt
<img alt="" class="sizing-normal" height="815" src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png"
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.