The role
What an AI engineer actually does
Most AI engineer jobs are software engineering jobs where the hard part is the model. You take a capable model you did not train and wire it into a product through an API. Your job is to make the result reliable enough to ship.
Day to day, the work is:
- writing the application code that calls the model
- designing the prompts and the context the model sees
- measuring whether the output is good enough, and fixing it when the quality drifts
This role sits apart from two others people confuse it with. A research scientist trains new models. A traditional ML engineer owns the training and data pipelines that produce a model. An AI engineer mostly consumes a model that already exists and is judged on the product built around it.
What employers ask for
The skills that actually show up in job postings
We read 102 live job postings titled "AI Engineer" in full and coded every skill each one named (measured 2026-07-27, full method at /provenance, full analysis at /data). A short list of skills carried most of the market:
- Working with LLMs: 77%
- Cross-functional collaboration: 76%
- Shipping AI to production: 72%
- Agents and agentic systems: 71%
- Ownership and autonomy: 60%
- Python: 59%
- APIs and microservices: 51%
In total, 7 skills appeared in at least half the postings. Of the 91 skills we tracked, 52 appeared in fewer than 15% of postings. For a learner, that is the whole strategy. The core list is short. Most of the exotic tools that fill "become an AI engineer" listicles barely register with employers. Learn the short list well before you spend time on the long tail.
Starting from zero
A learning path for beginners
If you are starting without a software job, work through these in order. Each step should end in something that runs, code you can point to.
- 01Learn Python. It is the default language for this work and shows up in most postings. Get comfortable with functions and data structures, and with calling libraries from your own code.
- 02Learn how the web fits together. HTTP requests and JSON, and how one service calls another through an API. You will use this every day.
- 03Call a model through its API. Sign up for a hosted model API, then call it from your own code and read the structured output it returns. This is the moment the role starts to feel real.
- 04Learn prompt and context engineering. How you word the instruction and what information you put in front of the model changes the output more than any other single thing you control.
- 05Learn to evaluate output. Build a small labeled dataset of inputs and correct answers. Pick a metric and score the model against it. This is the skill that separates people who can improve a system from people who only tweak it and hope.
- 06Ship one small project where people can use it. Deploy it somewhere public. Shipping to production was named in 72% of the postings we read, and it is the step most self-taught learners skip.
- 07Add observability. Log every run so that when something breaks in front of a user, you can find out why. At a minimum, capture:
- the input and the output
- the exact prompt that was sent, plus the model and its settings
- the cost and latency of the run
Expect this to take months of steady work. Depth on the short list beats a shallow tour of everything.
Already shipping software
A faster path for working developers
If you already ship software, steps 1 and 2 are mostly behind you. Spend your time on the parts that are specific to models, in this order:
- 01Build one real LLM feature inside a product you already understand. A search box that answers in plain language, or a support reply drafter. Something with real inputs.
- 02Learn evaluation early. It is the skill experienced developers skip most often, because grading fuzzy model output works differently from a normal unit test. You first define what a good answer looks like and collect examples. Then you score the model against them.
- 03Learn prompt and context engineering against your own feature, so you can see the output change as you change the input.
- 04Add tracing and monitoring to that feature while it runs. Watch a few real runs end to end. You will find failures you would never have guessed.
- 05Learn where models fail. The common failure modes are:
- wrong answers stated with confidence
- latency spikes under load
- cost that climbs with usage
Knowing these failure modes is most of what makes a feature production-ready.
Your portfolio
Build proof before you apply
One finished project that runs in front of users will do more for you than several half-built demos. In your writeup, include the eval you used and a short account of a real failure you found and fixed. A hiring manager can tell the difference between someone who has run a system and someone who has only followed a tutorial, and the failure story is where that difference shows.
Applying
How to run the job search
- Target roles built around the short skill list. The long-tail requirements in a posting are usually nice-to-haves you can pick up on the job.
- Read the full job description. Titles for this role vary widely, so the title alone tells you little. The skills underneath move very little, which is why a short, well-built skill set travels across many postings.
- Speak to the skills that were actually named. In interviews and on your resume, lead with shipping to production, which 72% of the postings named, and with evaluation, a skill self-taught candidates often lack.
Next steps
Where to go next
- The full measured breakdown of what these postings require
- Datasets & Evaluation for step 5, the evaluation skill
- Prompt & Context Engineering for step 4
- Observability & Traceability for step 7
- The full curriculum
Coming soon
New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.
