Skip to content
HeyJared

The Kaitchup – AI on a Budget

Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated.

Newsletter · By Benjamin Marie · English · Paid tier available · Official site

Indexed issues, last 90 days
16
Latest publication
Sep 30, 2026
Audience
Checking…
Earliest in this view
Aug 20, 2026

Latest issues

  1. Issue · Sep 30, 2026

    Qwen3.8 Flash Next GGUF Benchmark: Q4 to Q1 Accuracy and Token Efficiency (opens the original)

    Excerpt · Neutral tone

    Read excerpt

    Quantizing Qwen3.8 Flash Next is almost unavoidable if you want to run it locally. The BF16 original model occupies about 354 GB.But after quantization, how much of the model survives?A small file is useful only if the model can still do the work. And accuracy is only half of the question. If the compressed model needs twice as many output tokens to solve the same tasks, some of the apparent efficiency gain disappears. As we will see here, this “efficiency” evaluation is especially relevant for

  2. Issue · Sep 26, 2026

    ThinkingCap-Qwen3.8-27B: Less Thinking, Similar Accuracy (opens the original)

    Excerpt · Positive tone

    Read excerpt

    <img alt="" class="sizing-normal" height="815" src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png"

  3. Issue · Sep 24, 2026

    Boosting Agentic Coding with LLM Retries: Lessons from Qwen3.8 27B on DeepSWE (opens the original)

    Excerpt

    Read excerpt

    <img alt="" class="sizing-large" height="699.7252747252747" src="https://substackcdn.com/image/fetch/$s_!rTz9!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec9fdf39-a926-4ea6-b24a-720ff491cd45_9

  4. Issue · Sep 23, 2026

    Fastest Qwen3.8 27B Quantization? Benchmarking NVFP4, INT4, GSQ, and MTP (opens the original)

    Excerpt

    Read excerpt

    Last week, I published my analysis of several quantized versions of Qwen3.8 27B.Unsloth NVFP4, NVIDIA NVFP4, AWQ INT4, GSQ 3-bit, Escha-W2, and my mixed INT3/INT4 models use very different quantization recipes. Some quantize activations, some keep them at higher precision, some preserve more layers in FP8 or BF16, and some include MTP weights while others do not.They also run at very different speeds.MTP can increase decoding throughput substantially by predicting several tokens at once, but the

  5. Issue · Sep 19, 2026

    Bonsai 2 27B: Qwen3.8 in 5.9 GB, but What About Agentic Coding? (opens the original)

    Excerpt

    Read excerpt

    <img alt="" class="sizing-normal" height="815" src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png"

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about The Kaitchup – AI on a Budget