Compilers for Machine Learning
A hands-on one semester course where students build their own compiler from scratch, starting from elementwise programs and ending with training SOTA LLMs on GPUs. This course aggressively builds on the previous week, and is an exercise in slop management. If you let any slop in early, it will compound and you will not finish the class.
Weeks 1–2: UOps & Rewrites
uops · elementwise ops · symbolic · rewrites · renderer
Students become familiar with the UOp and write a compiler capable of compiling and simplifying elementwise programs. We introduce the graph structure and rewriting, and they can compile simple CPU programs to C.
UOp(op, src:tuple[UOp, ...], arg)
recursive properties:
shape:tuple[int, ...]
dtype(bool,int,float)
addrspace(MEM,ALU)
Ops introduced:
CALL(CallArg)/PARAM(ParamArg)
BUFFER/INDEX/LOAD/STORE
RECIP...unary/DIV...binary/WHERE...ternary
This can render (loopless) C functions that act on memory and immediates.
Weeks 3–4: Loops & Movement Ops
loops · movement ops · reduction · convs · rangeify
Now we introduce movement ops and loops. Note how you can derive gemms and convs from the basic movement ops. Without memory hierarchies, things are slow — but this compiler is now capable of compiling any model, splitting it into kernels, and compiling it to C.
Ops introduced:
RESHAPE/EXPAND/PERMUTE/FLIP/PAD/SHRINK (spec on arg)
REDUCE (axis count / op)
RANGE/END
STAGE/ALLOC
Weeks 5–6: Memory Hierarchies & Fast Kernels
memory hierarchies · upcasting · fast GEMMs · flash attention
Now things get fast. Still only on CPU, but we can now produce SOTA-competitive C code for GEMMs, reduces, and convs.
No new ops these weeks. Many students will realize they have to rewrite their week 3-4.
Weeks 7–8: GPUs & Hardware Accelerators
call · GPUs · hardware accelerators · tensor cores
Here we add kernels, GPUs, axis mappings, and tensor cores. This can produce decent torch-competitive CUDA/HIP code now.
Weeks 9–10: Real Models & Autodiff
real models · autodiff
Here we implement a real LLM + autodiff to train models.
Weeks 11–End: Student Projects
Use your compiler to implement a paper, port it to strange hardware, mostly anything
At a Glance
| Weeks | Topics | Milestone |
|---|---|---|
| 1–2 | UOps, elementwise ops, symbolic, rewrites | Compile & simplify elementwise programs to C |
| 3–4 | Loops, movement ops, reduction, rangeify | Compile any model to C (slow, but correct) |
| 5–6 | Memory hierarchies, upcasting, GEMMs, convs | SOTA-competitive CPU code |
| 7–8 | Call, GPUs, accelerators, tensor cores | Torch-competitive CUDA/HIP code |
| 9–10 | Real models, autodiff | Train a real LLM |
| 11+ | Student projects | Choose a project and extend your compiler |