Learn inference engineering, from zero.
A course for developers who can code but have never run an AI model. Run models, measure them, break them on purpose, and learn how they are served at scale.
- 11parts
- 49chapters
- 2free chapters
- 0GPUs to start
01How you learn
One loop, every chapter.
01Learn
One idea at a time, drawn out.
02Check
Answer questions on a real server log, and see why when you miss one.
03Build it
Write the real parts yourself, with tests that tell you when they work.
04Break it
Change one setting, watch what fails, and say why.
05Fix it
Work through real outages, from the first alert to the fix.
02Demos
See what makes a model fast or slow.
Both run right here in your browser.
MeasuredSource
LLaMA 2 7B, 4-bit (Q4_0), 128 tokens, flash attention off, in llama.cpp. Tesla T4, 16 GB GDDR6 on Linux, build d32e03f: speed, memory, checked 2026-10-09. RTX 4060 Laptop, 8 GB GDDR6 on Windows, build 2cce9fd: speed, memory, checked 2026-10-09. Apple M2 Ultra, 76-core GPU on macOS, build 8e672ef: speed, checked 2026-10-02. The builds and backends differ, so the runs compare closely but not exactly.
One model writes the same 128 tokens on a Mac, a Windows laptop and a Linux cloud GPU. Watch which one finishes first.
Tesla T4Linux · 320 GB/s
0 of 128 tokens writtenFinishes at 2.76 sRTX 4060 LaptopWindows · 256 GB/s
0 of 128 tokens writtenFinishes at 2.20 sApple M2 UltramacOS · 800 GB/s
0 of 128 tokens writtenFinishes at 1.36 s
2.5x the memory bandwidth, 2.0x the speed.In chapter 1 you measure the speed of your own laptop, on Windows, Mac or Linux.
Illustration, not measuredSource
Ollama runs one request at a time by default: OLLAMA_NUM_PARALLEL is 1 at v0.35.0 (envconfig/config.go). Each request takes 4 s here, about one 230-token answer at the Windows laptop's measured 58.22 tokens/s. Shown at 4x speed. Real servers also slow each request down when slots share one GPU. In part 4 you load test your own server.
Eight people ask at once.
Your laptop serves one request at a time by default. The rest wait in line. Give the server more slots and watch the line move.
- Mean wait
- 14.0 s
- Last one done
- 32 s
More slots are not free. Each one needs its own memory.
03The course
From your laptop to production.
11 parts and 49 chapters, from beginner to expert. Chapters 1 and 2 are free.
01Start here: the whole gameBeginner3 chapters
- 1
Your first model, measured
Run a small model on your laptop, call it over HTTP, and measure its speed.
Run itFree - 2
What just happened: tokens, the loop, the two phases, and "will it fit?"
Explain how a model writes one token at a time, why the first token takes longest, and whether a model file fits in GPU memory.
Run itFree - 3
Python for JavaScript developers (optional bridge)
Read and write the Python the labs use, starting from the JavaScript you already know.
Build it
02Just enough ML for inferenceBeginner4 chapters
- 4
PyTorch and Transformers for inference (no training)
Load a model with PyTorch and write your own decode loop on a laptop, with no training.
Build it - 5
A model is a file of numbers
Open a model's files and explain what its weights, config and tokenizer are and how big they are.
Work it out - 6
Tokens in depth: tokenizers, templates, sampling, thinking and tool-call text
Predict token counts, format prompts correctly, control sampling, and find the thinking and tool-call text in a raw output.
Run it - 7
Inside a transformer, for engineers
Follow one token through a transformer and point to where the weights and the KV cache come from.
Work it out
03Memory, bandwidth and speedBeginner to intermediate5 chapters
- 8
Your first GPU, free: open it, check it, time it right
Open a free GPU, check its driver and versions, time its work correctly, and read its memory.
Run it - 9
Will it fit? Weights and number formats
Work out how much GPU memory a model's weights need in any number format, and pick a GPU that fits.
Work it out - 10
The KV cache: memory that grows with every token
Size the KV cache for any model, context length and user count, and predict when memory runs out.
Work it out - 11
What a GPU is, for software developers
Read a GPU spec sheet and explain why memory bandwidth, not compute, limits most inference.
Work it out - 12
Why decode is slow: the roofline in plain words
Predict a model's fastest possible tokens per second on a GPU, and explain why batching raises throughput.
Work it out
04Serve it and measure itBeginner to intermediate5 chapters
- 13
Linux, Docker and Kubernetes in one weekend
Run, inspect, break and fix an inference-shaped service in Docker and on a local Kubernetes cluster, with no GPU.
Run it - 14
vLLM on a free GPU without pain or a surprise bill
Install vLLM on a free GPU within an hour, with pinned versions and a shutdown rule, and fix the setup errors you meet.
Run it - 15
Your first vLLM server
Serve a model with vLLM's OpenAI-style API and read its startup log and main settings.
Run it - 16
Load testing that tells the truth
Run a load test that is repeatable, measure its noise, and read a real latency and throughput curve.
Run it - 17
Queues, overload and SLOs
Predict when a server falls over, read the signs of an overload incident, and protect the server with limits and SLOs.
Fix it
05Inside the engineIntermediate5 chapters
- 18
Batching: from one request to hundreds
Build a continuous batcher with tests, and measure what batching gives one user and what it gives everyone.
Build it - 19
Paged KV memory and preemption
Build a KV block allocator with preemption, and tune an engine's memory so requests stop getting preempted.
Build it - 20
Scheduling: chunked prefill, CPU overhead, CUDA graphs, the TTFT-ITL trade
Stop long prompts from freezing other users, and explain where CPU overhead, CUDA graphs and compilation fit.
Run it - 21
Prefix caching: reuse instead of recompute
Make repeated prompt prefixes nearly free, design prompts that hit the cache, and measure the hit rate.
Run it - 22
Read a real engine, then choose one
Trace a decode step through a small engine's code, then choose a production engine with evidence.
Run it
06Faster and cheaperIntermediate to advanced4 chapters
- 23
Quantization: smaller models, same answers?
Serve a quantized model, measure its gains, and prove that its answers hold up.
Run it - 24
Speculative decoding: faster when it works, slower when not
Turn on speculative decoding, measure how often the draft is accepted, and find where it stops paying.
Run it - 25
Structured output and multi-LoRA serving
Return JSON and tool calls that always parse, and serve many adapters on one base model.
Run it - 26
Serving reasoning and tool-calling models
Serve reasoning and tool-calling models so the right fields come back, cap thinking cost, and prove tool-call accuracy with a fixed test set.
Run it
07Production: ship it, keep it up, keep it correctIntermediate to advanced9 chapters
- 27
Go for platform engineers (optional bridge)
Read and write enough Go to build a gateway or a Kubernetes controller in it.
Build it - 28
The API layer: streaming, cancellation, limits, auth, TLS
Build a production API in front of an engine, with streaming, cancellation, rate limits, keys and TLS.
Build it - 29
Containers, Kubernetes and operators for GPU inference
Run an engine on Kubernetes with GPUs, honest health checks and safe rollouts, and write a small operator.
Build it - 30
Routing and load balancing for LLMs
Route requests so caches stay warm and no replica melts, behind a gateway that can fail over between providers.
Build it - 31
Autoscaling, cold starts and the OS under the server
Scale on the right signal, cut cold starts from minutes to seconds, and explain the operating-system work inside a cold start.
Run it - 32
Observability: see performance, quality and hardware
Build dashboards and alerts that catch latency, capacity, quality and GPU problems early.
Run it - 33
Correctness, reliability and incidents
Keep answers right and the service up with deploy gates and load shedding, and work real incidents from the first alert to the postmortem.
Fix it - 34
The serving plan: requirements, capacity and cost
Turn a vague request into a serving plan with SLOs, a model choice, a fleet size and a cost per million tokens.
Work it out - 35
Fleet and capacity: regions, quotas and failing GPUs
Keep enough healthy GPUs in the right regions, and pull a failing node out of service safely.
Fix it
08Scale outAdvanced to expert3 chapters
- 36
Multi-GPU basics: tensor parallelism
Split a model across GPUs and predict when more GPUs stop helping.
Run it - 37
Big models: pipeline and expert parallelism, MoE
Explain how the largest mixture-of-experts models run across many GPUs, and read a real production design.
Work it out - 38
Disaggregation, KV tiers and long context
Decide when to split prefill from decode, and when to move the KV cache to CPU memory or disk.
Run it
09GPU performance and kernelsAdvanced to expert4 chapters
- 39
Profiling: find where the time goes
Profile a request from Python down to GPU kernels and name the bottleneck with evidence.
Run it - 40
Numerics and determinism
Explain why one prompt can give different outputs, and test numerics across batch sizes and versions.
Run it - 41
GPU programming for inference engineers
Write and time your first GPU kernels, and explain how threads, warps and memory set their speed.
Build it - 42
Attention kernels and backends
Explain FlashAttention and paged-attention kernels, and fix unsupported-backend errors.
Build it
10Beyond chat: speech, vision and customer work (electives)Intermediate3 chapters
- 43
Speech: speech-to-text and streaming text-to-speech
Serve speech-to-text and text-to-speech models to real-time targets, and measure the time to first audio.
Run it - 44
Vision, image generation and embeddings
Serve vision-language, image generation and embedding models, and explain how their bottlenecks differ from chat.
Run it - 45
The proof of concept: from a customer brief to a memo
Turn a customer's brief into a working proof of concept, a benchmark and a one-page memo.
Build it
11Show your work: proof, interviews and the capstoneBeginner to advanced4 chapters
- 46
The job market and your track
Read a job post's real level and skills, and name the track and the portfolio you will build.
Show it - 47
Proof of work: open source and write-ups
Publish work a hiring manager can check in minutes, including a pull request a maintainer reviewed.
Show it - 48
Interview practice
Practice system-design, coding and debugging interviews under time limits, and finish one take-home.
Show it - 49
Capstone: ship a production inference service
Ship a production inference service for one real use case, proven by numbers, a live drill and a design review.
Show it
04Pricing
Start free today.
Chapters 1 and 2
Free
- Every lesson and check
- Runs on your laptop
- No GPU needed
Full course
Price coming soon
- All 11 parts and 49 chapters
- A graded check in every chapter
- A free GPU from part 3
05Questions
Before you start.
Do I need a GPU?
Not to start. Chapter 1 runs on your laptop. Part 3 opens with a free GPU.
Do I need to know machine learning?
No. Part 2 teaches just enough ML for inference.
Which languages do I need?
You should be able to code. Chapter 1 uses curl, JavaScript and Python. Chapter 3 is an optional Python bridge for JavaScript developers.
Is it free?
Chapters 1 and 2 are free.
Do I have to pass a check to move on?
No. The next chapter is always open. If you miss a check, it shows you which parts to go back to.
Start with chapter 1.
Run your first model on your laptop and measure its speed.