Compilers for Machine Learning

A hands-on one semester course where students build their own compiler from scratch, starting from elementwise programs and ending with training SOTA LLMs on GPUs. This course aggressively builds on the previous week, and is an exercise in slop management. If you let any slop in early, it will compound and you will not finish the class.


Weeks 1–2: UOps & Rewrites

uops · elementwise ops · symbolic · rewrites · renderer

Students become familiar with the UOp and write a compiler capable of compiling and simplifying elementwise programs. We introduce the graph structure and rewriting, and they can compile simple CPU programs to C.

UOp(op, src:tuple[UOp, ...], arg)

recursive properties:

shape:tuple[int, ...]
dtype(bool,int,float)
addrspace(MEM,ALU)

Ops introduced:

CALL(CallArg)/PARAM(ParamArg)
BUFFER/INDEX/LOAD/STORE
RECIP...unary/DIV...binary/WHERE...ternary

This can render (loopless) C functions that act on memory and immediates.

Weeks 3–4: Loops & Movement Ops

loops · movement ops · reduction · convs · rangeify

Now we introduce movement ops and loops. Note how you can derive gemms and convs from the basic movement ops. Without memory hierarchies, things are slow — but this compiler is now capable of compiling any model, splitting it into kernels, and compiling it to C.

Ops introduced:

RESHAPE/EXPAND/PERMUTE/FLIP/PAD/SHRINK (spec on arg)
REDUCE (axis count / op)
RANGE/END
STAGE/ALLOC

Weeks 5–6: Memory Hierarchies & Fast Kernels

memory hierarchies · upcasting · fast GEMMs · flash attention

Now things get fast. Still only on CPU, but we can now produce SOTA-competitive C code for GEMMs, reduces, and convs.

No new ops these weeks. Many students will realize they have to rewrite their week 3-4.

Weeks 7–8: GPUs & Hardware Accelerators

call · GPUs · hardware accelerators · tensor cores

Here we add kernels, GPUs, axis mappings, and tensor cores. This can produce decent torch-competitive CUDA/HIP code now.

Weeks 9–10: Real Models & Autodiff

real models · autodiff

Here we implement a real LLM + autodiff to train models.

Weeks 11–End: Student Projects

Use your compiler to implement a paper, port it to strange hardware, mostly anything


At a Glance

Weeks Topics Milestone
1–2 UOps, elementwise ops, symbolic, rewrites Compile & simplify elementwise programs to C
3–4 Loops, movement ops, reduction, rangeify Compile any model to C (slow, but correct)
5–6 Memory hierarchies, upcasting, GEMMs, convs SOTA-competitive CPU code
7–8 Call, GPUs, accelerators, tensor cores Torch-competitive CUDA/HIP code
9–10 Real models, autodiff Train a real LLM
11+ Student projects Choose a project and extend your compiler
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论