Learn inference engineering, from zero.

A course for developers who can code but have never run an AI model. Run models, measure them, break them on purpose, and learn how they are served at scale.

  • 11parts
  • 49chapters
  • 2free chapters
  • 0GPUs to start

01How you learn

One loop, every chapter.

01Learn

One idea at a time, drawn out.

How one prompt becomes an answer
Your promptwhat you ask the modelRead the whole promptprefill: all at onceWrite the next tokendecode: one at a timeAnswer streams backtoken by token1 send2 first token3 streamuntil doneWhat is a GPU?AGPUisachipbuiltformathfirst tokenthen the rest, one at a time

02Check

Answer questions on a real server log, and see why when you miss one.

Work out tokens per second from a server log
1 read the last lineserver log, times in nanosecondsillustration, not measured"done"truethe last line"total_duration"5000000000the whole request"load_duration"1600000000loading the model"prompt_eval_duration"400000000reading the prompt"eval_count"150tokens written"eval_duration"3000000000time spent writing2 answer the questionHow many tokens per second did the model write?3 a wrong answer tells you why30 tokens/sYou divided by total_duration. It alsocounts loading and reading the prompt.50 tokens/s150 tokens in 3.0 s of writing

03Build it

Write the real parts yourself, with tests that tell you when they work.

You write the scheduler
waitingR5R6R7Your scheduleryour code, with testsslots on the GPUR1R5R3R4R2 finished, R5 took its slot1 pick who runs2 fill a free slot3 run one step, again

04Break it

Change one setting, watch what fails, and say why.

Make it fail, then explain why
1 guess what happensbigger batchon the GPUwaitingfirst token: fastsmaller batchon the GPUwaitingfirst token: slow2 make the batch smaller3 explain why: too few slots for the traffic

05Fix it

Work through real outages, from the first alert to the fix.

A real outage, from alert to fix
time to first tokentimealert line1 alert fires2 dig in: GPU memory is full3 cap the prompt lengththen write it up

02Demos

See what makes a model fast or slow.

Both run right here in your browser.

MeasuredSource

LLaMA 2 7B, 4-bit (Q4_0), 128 tokens, flash attention off, in llama.cpp. Tesla T4, 16 GB GDDR6 on Linux, build d32e03f: speed, memory, checked 2026-10-09. RTX 4060 Laptop, 8 GB GDDR6 on Windows, build 2cce9fd: speed, memory, checked 2026-10-09. Apple M2 Ultra, 76-core GPU on macOS, build 8e672ef: speed, checked 2026-10-02. The builds and backends differ, so the runs compare closely but not exactly.

One model writes the same 128 tokens on a Mac, a Windows laptop and a Linux cloud GPU. Watch which one finishes first.

  1. Tesla T4Linux · 320 GB/s

    0 of 128 tokens writtenFinishes at 2.76 s
  2. RTX 4060 LaptopWindows · 256 GB/s

    0 of 128 tokens writtenFinishes at 2.20 s
  3. Apple M2 UltramacOS · 800 GB/s

    0 of 128 tokens writtenFinishes at 1.36 s

2.5x the memory bandwidth, 2.0x the speed.In chapter 1 you measure the speed of your own laptop, on Windows, Mac or Linux.

03The course

From your laptop to production.

11 parts and 49 chapters, from beginner to expert. Chapters 1 and 2 are free.

  1. 01Start here: the whole gameBeginner3 chapters
    1. 1

      Your first model, measured

      Run a small model on your laptop, call it over HTTP, and measure its speed.

      Run itFree
    2. 2

      What just happened: tokens, the loop, the two phases, and "will it fit?"

      Explain how a model writes one token at a time, why the first token takes longest, and whether a model file fits in GPU memory.

      Run itFree
    3. 3

      Python for JavaScript developers (optional bridge)

      Read and write the Python the labs use, starting from the JavaScript you already know.

      Build it
  2. 02Just enough ML for inferenceBeginner4 chapters
    1. 4

      PyTorch and Transformers for inference (no training)

      Load a model with PyTorch and write your own decode loop on a laptop, with no training.

      Build it
    2. 5

      A model is a file of numbers

      Open a model's files and explain what its weights, config and tokenizer are and how big they are.

      Work it out
    3. 6

      Tokens in depth: tokenizers, templates, sampling, thinking and tool-call text

      Predict token counts, format prompts correctly, control sampling, and find the thinking and tool-call text in a raw output.

      Run it
    4. 7

      Inside a transformer, for engineers

      Follow one token through a transformer and point to where the weights and the KV cache come from.

      Work it out
  3. 03Memory, bandwidth and speedBeginner to intermediate5 chapters
    1. 8

      Your first GPU, free: open it, check it, time it right

      Open a free GPU, check its driver and versions, time its work correctly, and read its memory.

      Run it
    2. 9

      Will it fit? Weights and number formats

      Work out how much GPU memory a model's weights need in any number format, and pick a GPU that fits.

      Work it out
    3. 10

      The KV cache: memory that grows with every token

      Size the KV cache for any model, context length and user count, and predict when memory runs out.

      Work it out
    4. 11

      What a GPU is, for software developers

      Read a GPU spec sheet and explain why memory bandwidth, not compute, limits most inference.

      Work it out
    5. 12

      Why decode is slow: the roofline in plain words

      Predict a model's fastest possible tokens per second on a GPU, and explain why batching raises throughput.

      Work it out
  4. 04Serve it and measure itBeginner to intermediate5 chapters
    1. 13

      Linux, Docker and Kubernetes in one weekend

      Run, inspect, break and fix an inference-shaped service in Docker and on a local Kubernetes cluster, with no GPU.

      Run it
    2. 14

      vLLM on a free GPU without pain or a surprise bill

      Install vLLM on a free GPU within an hour, with pinned versions and a shutdown rule, and fix the setup errors you meet.

      Run it
    3. 15

      Your first vLLM server

      Serve a model with vLLM's OpenAI-style API and read its startup log and main settings.

      Run it
    4. 16

      Load testing that tells the truth

      Run a load test that is repeatable, measure its noise, and read a real latency and throughput curve.

      Run it
    5. 17

      Queues, overload and SLOs

      Predict when a server falls over, read the signs of an overload incident, and protect the server with limits and SLOs.

      Fix it
  5. 05Inside the engineIntermediate5 chapters
    1. 18

      Batching: from one request to hundreds

      Build a continuous batcher with tests, and measure what batching gives one user and what it gives everyone.

      Build it
    2. 19

      Paged KV memory and preemption

      Build a KV block allocator with preemption, and tune an engine's memory so requests stop getting preempted.

      Build it
    3. 20

      Scheduling: chunked prefill, CPU overhead, CUDA graphs, the TTFT-ITL trade

      Stop long prompts from freezing other users, and explain where CPU overhead, CUDA graphs and compilation fit.

      Run it
    4. 21

      Prefix caching: reuse instead of recompute

      Make repeated prompt prefixes nearly free, design prompts that hit the cache, and measure the hit rate.

      Run it
    5. 22

      Read a real engine, then choose one

      Trace a decode step through a small engine's code, then choose a production engine with evidence.

      Run it
  6. 06Faster and cheaperIntermediate to advanced4 chapters
    1. 23

      Quantization: smaller models, same answers?

      Serve a quantized model, measure its gains, and prove that its answers hold up.

      Run it
    2. 24

      Speculative decoding: faster when it works, slower when not

      Turn on speculative decoding, measure how often the draft is accepted, and find where it stops paying.

      Run it
    3. 25

      Structured output and multi-LoRA serving

      Return JSON and tool calls that always parse, and serve many adapters on one base model.

      Run it
    4. 26

      Serving reasoning and tool-calling models

      Serve reasoning and tool-calling models so the right fields come back, cap thinking cost, and prove tool-call accuracy with a fixed test set.

      Run it
  7. 07Production: ship it, keep it up, keep it correctIntermediate to advanced9 chapters
    1. 27

      Go for platform engineers (optional bridge)

      Read and write enough Go to build a gateway or a Kubernetes controller in it.

      Build it
    2. 28

      The API layer: streaming, cancellation, limits, auth, TLS

      Build a production API in front of an engine, with streaming, cancellation, rate limits, keys and TLS.

      Build it
    3. 29

      Containers, Kubernetes and operators for GPU inference

      Run an engine on Kubernetes with GPUs, honest health checks and safe rollouts, and write a small operator.

      Build it
    4. 30

      Routing and load balancing for LLMs

      Route requests so caches stay warm and no replica melts, behind a gateway that can fail over between providers.

      Build it
    5. 31

      Autoscaling, cold starts and the OS under the server

      Scale on the right signal, cut cold starts from minutes to seconds, and explain the operating-system work inside a cold start.

      Run it
    6. 32

      Observability: see performance, quality and hardware

      Build dashboards and alerts that catch latency, capacity, quality and GPU problems early.

      Run it
    7. 33

      Correctness, reliability and incidents

      Keep answers right and the service up with deploy gates and load shedding, and work real incidents from the first alert to the postmortem.

      Fix it
    8. 34

      The serving plan: requirements, capacity and cost

      Turn a vague request into a serving plan with SLOs, a model choice, a fleet size and a cost per million tokens.

      Work it out
    9. 35

      Fleet and capacity: regions, quotas and failing GPUs

      Keep enough healthy GPUs in the right regions, and pull a failing node out of service safely.

      Fix it
  8. 08Scale outAdvanced to expert3 chapters
    1. 36

      Multi-GPU basics: tensor parallelism

      Split a model across GPUs and predict when more GPUs stop helping.

      Run it
    2. 37

      Big models: pipeline and expert parallelism, MoE

      Explain how the largest mixture-of-experts models run across many GPUs, and read a real production design.

      Work it out
    3. 38

      Disaggregation, KV tiers and long context

      Decide when to split prefill from decode, and when to move the KV cache to CPU memory or disk.

      Run it
  9. 09GPU performance and kernelsAdvanced to expert4 chapters
    1. 39

      Profiling: find where the time goes

      Profile a request from Python down to GPU kernels and name the bottleneck with evidence.

      Run it
    2. 40

      Numerics and determinism

      Explain why one prompt can give different outputs, and test numerics across batch sizes and versions.

      Run it
    3. 41

      GPU programming for inference engineers

      Write and time your first GPU kernels, and explain how threads, warps and memory set their speed.

      Build it
    4. 42

      Attention kernels and backends

      Explain FlashAttention and paged-attention kernels, and fix unsupported-backend errors.

      Build it
  10. 10Beyond chat: speech, vision and customer work (electives)Intermediate3 chapters
    1. 43

      Speech: speech-to-text and streaming text-to-speech

      Serve speech-to-text and text-to-speech models to real-time targets, and measure the time to first audio.

      Run it
    2. 44

      Vision, image generation and embeddings

      Serve vision-language, image generation and embedding models, and explain how their bottlenecks differ from chat.

      Run it
    3. 45

      The proof of concept: from a customer brief to a memo

      Turn a customer's brief into a working proof of concept, a benchmark and a one-page memo.

      Build it
  11. 11Show your work: proof, interviews and the capstoneBeginner to advanced4 chapters
    1. 46

      The job market and your track

      Read a job post's real level and skills, and name the track and the portfolio you will build.

      Show it
    2. 47

      Proof of work: open source and write-ups

      Publish work a hiring manager can check in minutes, including a pull request a maintainer reviewed.

      Show it
    3. 48

      Interview practice

      Practice system-design, coding and debugging interviews under time limits, and finish one take-home.

      Show it
    4. 49

      Capstone: ship a production inference service

      Ship a production inference service for one real use case, proven by numbers, a live drill and a design review.

      Show it

04Pricing

Start free today.

  • Chapters 1 and 2

    Free

    • Every lesson and check
    • Runs on your laptop
    • No GPU needed
  • Full course

    Price coming soon

    • All 11 parts and 49 chapters
    • A graded check in every chapter
    • A free GPU from part 3

05Questions

Before you start.

Do I need a GPU?

Not to start. Chapter 1 runs on your laptop. Part 3 opens with a free GPU.

Do I need to know machine learning?

No. Part 2 teaches just enough ML for inference.

Which languages do I need?

You should be able to code. Chapter 1 uses curl, JavaScript and Python. Chapter 3 is an optional Python bridge for JavaScript developers.

Is it free?

Chapters 1 and 2 are free.

Do I have to pass a check to move on?

No. The next chapter is always open. If you miss a check, it shows you which parts to go back to.

Start with chapter 1.

Run your first model on your laptop and measure its speed.