MLGuerrillaStart with M1 →
Applied AI engineer·4 min read·Updated 17 September 2026

What an applied AI engineer actually does.

An applied AI engineer builds features and products on top of foundation models like Claude and GPT. The work is integration, evaluation, and reliability. You take a capable general model, give it the right context and tools, prove it works, and keep it working in production. Most of the day is software engineering. You call a model you did not train, and you build everything around it so it behaves.

The role

The role in one paragraph.

Applied AI engineering is the practice of building software that uses large language models to do useful work. The model is a component you call through an API. Your job is everything around it, from choosing the model to shaping its input, giving it tools and data, checking its output, and running it reliably at a reasonable cost. Some teams call this an AI engineer, a GenAI engineer, or an LLM engineer. The title varies. The work is the same.

Day to day

What the job looks like day to day.

The work spans a handful of repeatable activities.

  • Ship model-backed features that handle messy, real-world input.
  • Design the prompt and context a model sees, so its answers stay on task.
  • Build evaluations that measure whether a change made the system better or worse.
  • Add retrieval so the model can answer from your own data.
  • Give the model tools and let it take actions, then keep those actions safe.
  • Add tracing and monitoring so you can see what the system did and why.
  • Watch latency and cost, and tune both without hurting quality.
  • Work across product, design, and data to agree on what good looks like.

Core skills

The skills the role is built on.

A few skills carry most of the work. They are learnable, and the list is shorter than most job descriptions suggest.

  • Working with LLMs and their APIs, including streaming, tool calls, and structured output.
  • Prompt and context engineering, so the model gets what it needs and nothing it does not.
  • Evaluation, meaning the datasets and metrics that tell you whether the system works.
  • Retrieval and RAG, so answers stay grounded in real, current data.
  • Agents and tool use, so the model can plan and take actions on its own.
  • Production reliability: retries, timeouts, guardrails, and graceful failure.
  • Observability, so you can trace a bad output back to its cause.
  • Python and general software engineering, the foundation the rest sits on.
  • Cross-functional collaboration, because the hard calls are rarely purely technical.

This lines up with what hiring asks for. In 102 AI Engineer postings read in full, 4 skills appeared in at least 70% of them: LLMs (general), Cross-functional collaboration, Production deployment of AI/ML, AI agents / agentic systems. See the full breakdown on the data page, or the method behind it in the provenance notes.

Demand figure from data/demand-sweep2.json (AI Engineer role, n=102, measured 2026-07-27).

Adjacent roles

How it differs from adjacent roles.

The applied AI engineer sits between several roles. Here is the short version of where the lines fall.

  • Applied AI engineer. Builds products on top of existing models. Owns integration, prompting, evaluation, and reliability.
  • ML engineer. Trains, optimizes, and serves models. Owns pipelines, features, and model performance.
  • Data scientist. Analyzes data and runs experiments to answer questions and inform decisions.
  • Research engineer. Develops new models and methods, often to push the state of the art.
  • Software engineer. Builds general software. May add AI features, though the model is not the center of the work.
  • AI product manager. Decides what to build and why. Owns the roadmap and the requirements.

Projects

Projects that prove you can do the work.

The fastest way to show the skills is to build small systems that use them. A few that map well to real job requirements:

  • An evaluation harness for one LLM feature, with a labeled dataset and a metric you trust. Start with Datasets & Evaluation.
  • A retrieval assistant that answers from your own documents, with citations. Data & State covers the storage and state behind it.
  • A small agent that uses a tool or two to finish a task, with guardrails on what it can do. The Model as a Component and Prompt & Context Engineering set this up.
  • Tracing added to any of the above, so you can replay what happened on a bad run. See Observability & Traceability.
  • An experiment that compares two prompts or two models on the same eval, with a clear winner. Experimentation walks through the method.

Where to learn it

Where to learn each skill.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.