Vikram Mandyam

I write about software, cloud, books, emacs, productivity, photography and fitness

What I learned from the book "AI Engineering" by Chip Huyen

[Vikram Mandyam] / 2026-08-10


Photo from the book’s O’Reilly page

Introduction

I read “AI Engineering: Building Applications with Foundation Models” by Chip Huyen last year, and it made it to my books of the year 2025 list. A paragraph in a year-end list did not do it justice, so here is a longer write-up.

Between conceptual reading, side projects, and building agents at work, I have spent a good part of the last couple of years with LLMs. This book is the one that made me grok(no pun intended) the concepts!

Chip opens the book with a line that sums up the shift really well:

The availability and accessibility of powerful foundation models lead to three factors that, together, create ideal conditions for the rapid growth of AI engineering as a discipline:

∙ Factor 1: General-purpose AI capabilities

∙ Factor 2: Increased AI investments

∙ Factor 3: Low entrance barrier to building AI applications

‐ Chip Huyen, AI Engineering

That is exactly the question this book answers. In many ways, it does for AI applications what Martin Kleppmann’s “Designing data-intensive applications ” did for data systems - it takes a field that is moving at a breakneck speed and distills it into principles that will outlive the tools.

TLDR; Key topics covered in the book

The book has ten chapters, and they roughly follow the journey of building an AI application:

The two ideas that stayed with me the most come from the evaluation chapters and the final chapter.

Evaluation is not an afterthought

In traditional software, we write tests. In traditional ML, we have accuracy, precision and recall. But how do you evaluate an essay, a summary or an agent that books your travel? There can be many correct answers, and they can all look completely different.

Not having a reliable evaluation pipeline is one of the biggest blockers to AI adoption.

‐ Chip Huyen, AI Engineering

Chip makes a compelling case that evaluation should come first, not last. She calls this “evaluation-driven development” - much like test-driven development, you decide how your application will be judged before you build it.

Before investing time, money, and resources into building an application, it’s important to understand how this application will be evaluated. I call this approach evaluation-driven development.

‐ Chip Huyen, AI Engineering

Some things I found really useful here:

Here’s the thing - with LLMs, it is very tempting to change a prompt, look at three outputs, and declare victory. It works great for a demo. It falls apart the moment real users show up.

It’s easy to build a cool demo with foundation models. It’s hard to create a profitable product.

‐ Chip Huyen, AI Engineering

There is a reason why coding assistants were among the first AI applications to take off - generated code can be checked by running it. Many valuable applications are not being built simply because nobody has figured out how to evaluate them.

Building the architecture step by step

The final chapter is where everything comes together. Rather than presenting a giant architecture diagram upfront, Chip builds it up one step at a time, adding a component only when there is a problem that needs it:

  1. Start simple: A user query goes to the model, and the response comes back.
  2. Enhance context: Give the model the information it needs through retrieval and tools. Context construction is “like feature engineering for foundation models.”
  3. Put in guardrails: Protect both the inputs (leaking private data, prompt attacks) and the outputs (bad formatting, toxic or made-up responses). Handle the reliability vs latency tradeoff.
  4. Add a model router and gateway: Send different queries to different models, and have one place to manage access, cost and fallbacks. A router typically consists of an intent classifier
  5. Reduce latency with caches: Exact and semantic caching, so you don’t pay for the same answer twice.
  6. Add agent patterns: Loops, parallel steps and write actions - which unlock a lot of power.

On top of all this sit monitoring, observability. The whole 6 steps is usually built on an orchestration framework e.g. LangGraph.

As a builder, this resonated with me. It is the same lesson we learned with distributed systems - don’t start with the most complicated architecture, evolve it as the problems show up. A recurring theme in the book is:

Many AI challenges are, at their core, system problems.

‐ Chip Huyen, AI Engineering

Conclusion

What makes this book valuable isn’t how much it covers, it’s the judgment. The tools and models mentioned in the book will change, but the principles - evaluate first, evolve the architecture as needed, and others - will hold for a long time.

And, what about you? How are you using LLMs and AI agents?

Happy building!