Decoding AI Magazine

Decoding AI Magazine

Home
Notes
Chat
LLM Engineer's Handbook
Agent Engineering Course
Roadmaps
Perks
Contact Me
Archive
About

AI Evals Foundations

Why RAG Has Exactly 6 Failure Modes. No More, No Less.
A complete guide for evaluating your retrieval-augmented generation systems.
Mar 17 • Paul Iusztin
Our LLM Judge Passed Everything. It Was Wrong.
Align your evaluator with human judgment, or don't trust it at all.
Mar 10 • Paul Iusztin
How to Design Evaluators That Catch What Actually Breaks
The practical guide to code-based checks, LLM judges, and rubrics for real-world AI apps
Mar 3 • Paolo Perrone
Generate Synthetic Datasets for AI Evals
5 strategies from cold start to 450 diverse inputs in minutes
Feb 24 • Paul Iusztin
No Evals Dataset? Here's How to Build One from Scratch
Build evaluators to signal problems that users actually care about. Step-by-step guide.
Feb 17 • Paul Iusztin
Integrating AI Evals Into Your AI App
The holistic guide: From optimization to production monitoring
Feb 10 • Paul Iusztin
Behind the Scenes of AI Observability in Production
What actually works after 6 months of trial and error
Feb 3 • Alejandro Aboy
© 2026 Paul-Emil Iusztin · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture