The AI-engineering job market, measured
Learn to build the production AI systems employers actually ask for.
There are 22 modules live right now, free to read, with no account, and more to come. Each one takes a real piece of an AI system, like the ground-truth dataset, the evals, the tracing, the harness or the guardrails, and walks you through building it the way it gets built at work, including the failure cases and the trade-offs.
Free · no account · over 12 hours of reading · save any section as you go
Inside the modules
M1 · Datasets & Evaluation
M13 · Classification
M3 · How LLMs Actually Run Things
Models · Vision-language models (VLMs)
M9 · Harness & Reliability
M6 · Data & State
M12 · LLMOps & Cost
Models · Whisper
M4 · The Model as a Component
M15 · Verification
M2 · Observability & Traceability
Models · Large language models (LLMs)
M10 · AI Security & Guardrails
M16 · Memory
M5 · Prompt & Context Engineering
Models · CLIP
M8 · Experimentation
M1 · Datasets & Evaluation
M11-1 · Production, Deployment & Scale, Part 1
M3 · How LLMs Actually Run Things
Models · Transformer and attention
M14 · Routing
M7 · AI Product Framing
M2 · Observability & Traceability
M4 · The Model as a Component25 of the 162 figures drawn for the lessons and the model articles. Each one opens the page it belongs to.
From M1 · Datasets & Evaluation
I was building an AI native Computer-Use Agent, which in this case was used to QA chatbots directly from the UI. It reads the screen and decides what to click or type. One small part of it had to answer a yes/no question. Is there more conversation below what I can see, or is this the whole thing?
In early testing it looked done. Then it started getting things wrong. When a user asked the LLM to output something in markdown, the background the text sat on turned the same color as the scroll button, and the detector missed it. Seeing one failure told me the detector could be wrong, and it told me nothing about how often, or for which pattern of cases.

What the sweep found
The areas AI-engineering postings ask for.
- Datasets & Evals
- Observability & Traceability
- LLM Application Engineering
- RAG & Retrieval Engineering
- Agentic Systems & Tool Use
- Prompt & Context Engineering
- Production AI Systems
- AI Security, Safety & Guardrails
- AI Infrastructure & Inference
- LLMOps / AI Dev Lifecycle
- Working as a Production AI Engineer
The whole curriculum
Every module the course covers, with the written ones filled in. The foundation comes first, the rails run through everything under them, and each capability repeats the same loop on its own subject.
Foundation
Rails
Capabilities
- M13 Classification, written
- M14 Routing, written
- M15 Verification, written
- M16 Memory, written
- M17 Planning, written
- M18 Tool Use, written
- M19-1 RAG & Retrieval, Part 1, written
- M19-2 RAG & Retrieval, Part 2, written
- M20 Visual Grounding & VLMs, written
- M21 Fine-Tuning, being written
- M22 Generation, being written
- M23 Extraction, being written
- M24 State Understanding, being written
- M30 Agents, being written
Compose
- M25 AI System Design & Trade-offs, being written
- M26 AI Infrastructure & Inference, being written
- M31 Orchestration & Multi-Agent Systems, being written
- M27 Capstone · Build One End-to-End, being written
Enterprise
- M28 Enterprise, Compliance & Governance, being written
Professional
- M29 Working as a Production AI Engineer, being written
start here22 written11 still being written
What you walk away with
You walk away with production-ready AI knowledge.
Production-ready goes beyond what an AI hobbyist builds in their projects. It means you learn what makes good AI systems employers hire for. Evals, observability, security, harnesses and much much more.
Each module ends with something you can put in front of a hiring manager. Three of them, from modules that are live now:
M1 · Datasets & Evaluation
A labeled test set of 20 to 40 cases, with your recall and precision before and after a change, and the list of misses that shows the pattern you found.
M2 · Observability & Traceability
The trace of one failing run with every step's input and reasoning, the step that went wrong, the fix, and that failing input added to your eval set.
M13 · Classification
A classifier written down as decisions: the label set, what counts as one case, the threshold that auto-acts, the band a person reviews, and the cost per decision.
Who this is for
- Software engineers who want to move into AI engineering.
- CS students and new grads aiming at AI roles out of school.
- Anyone who has shipped a demo and been asked how they know it works.
This covers building systems on top of models. Training foundation models from scratch is a different field and I don't cover it.
Who is writing this
I started taking AI seriously in 2022, did an R&D internship at an AI lab as the first undergrad they had taken, and became an Applied AI Engineer straight out of undergrad. Today I work in big tech on AI systems, including computer-use agents and AI tools for developer productivity, and I'm doing a master's in computer science and data science alongside it. The modules are the material I wanted when I was trying to get in.
Questions
Is this another bootcamp?
No cohort, no schedule, no 12-week promise. Modules you enter by need.
What does it cost?
Nothing right now. Every written module is readable in full, with no account and no email required. If that changes I'll say so on this page before it does.
I already work as a software engineer.
Start at M1 for the eval discipline, then jump to whatever you're building. The foundation modules assume you can already write and ship code, so they spend their time on the parts that are new: evals, tracing, harnesses, guardrails.
Will this get me a job?
It gets you the artifacts and the vocabulary. The market runs on two clocks. Postings ask for 2–8 years of software engineering but only 1–3 years of LLM-specific work, most commonly 2, because the field is barely two years old (five role families, n=493, measured 27 July 2026). What this makes you is the candidate with a working eval harness when everyone else brought a demo.
Do I need a degree?
31.2% of postings mention one. 68.8% don't.
One dot = one posting. Hover for the receipt.
Start with the module everything else builds on.
M1 is 8 minutes. It ends with a labeled test set and a number for how often your system is wrong.
Follow along
I post the work as it ships: new modules, what the data says, and what I got wrong on the way. The Discord is open if you want to follow along or ask something.
Join the Discord →Watch on TikTok →
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.



