Hands On "AI Engineering"
Hands On "AI Engineering Course' With "AI Powered Quiz" Implementation, Learn How to Build from Scratch.
- Indexed issues, last 90 days
- 16
- Latest publication
- Sep 28, 2026
- Audience
- Checking…
- Earliest in this view
- Aug 10, 2026
Latest issues
Lesson 17: Deterministic Checks — The Fast, Cheap First Layer (opens the original)
Read excerpt
Free Resources: Build AI Agents That Remember Users Across SessionsYour agent can follow a rule p
Lesson 16: The Eval Case Intake Process (opens the original)
Read excerpt
Today we build:An IntakeTool that runs a reported query and suggests — never finalizes — a dispositionA four-way outcome: true negative, known failure, content gap, or needs reviewOne real human override of the tool’s own suggestion, proving judgment isn’t automated awayWhy This MattersDays 14 and 15 turned discovered failures into golden cases by hand — notice something odd, run it manually, inspect the trace, decide what “correct” means, write a GoldenC
Lesson 15: Writing Your First Golden Eval Cases (opens the original)
Read excerpt
Today we build:7 new golden cases, closing Day 14’s coverage gaps in cancellation, pricing, and out-of-scopeThe first real eval run — every case actually scored against the pipeline, not just browsedThree genuinely new, previously undiscovered failures, found in the act of writing the casesWhy This MattersDay 14’s dataset existed but had never been run — 8 cases sitting in a browsable table, never checked against a single real answer. Writing new cases to close coverage gaps forces exactly t
Lesson 14: Designing the Golden Dataset Schema (opens the original)
Read excerpt
Today we build:A GoldenCase schema with real validation — invalid categories, missing fields, and duplicate IDs are actually rejectedA persisted dataset that saves to JSON and reloads with a verified round tripEight real cases, each one traced back to a specific finding from Days 7, 11, and 12 — nothing invented for this lessonWhy This MattersDay 13’s single EvalCase lived inline in a Python list — fine for one case, unworkable past a handful. A real golden dataset needs to sur
Lesson 13: Why AI Outputs Need Eval Cases, Not Unit Tests (opens the original)
Read excerpt
Today we build:An EvalCase format that checks contracts — required facts, prohibited phrases, a citation — instead of exact textA hypothetical rephrased synthesis agent, standing in for what real generation will eventually do to Day 7’s wordingDirect, run proof that a unit test breaks on rewording while an eval case doesn’t, on the exact same outputWhy This MattersEvery test written since Day 6 has checked exact equality: does intent_label equal "refund_policy
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.