Interpretable
a crash course in ~100 hours

Understand what's happening inside the model.

Transformer fundamentals → how LLMs are actually trained → mechanistic interpretability → safety, steering, and editing. Visual-first lessons, interactive toys, problem sets, and the real literature — from Attention Is All You Need to the 2026 Transformer Circuits papers on emotions and the global workspace.

Part 0Foundations

The minimum viable math: geometry, probability, optimization.

Part 1Transformer Fundamentals

Tokens, attention, the residual stream, and how training shapes it all.

Part 2How LLMs Are Actually Made

Base models → SFT → RLHF → RLVR, plus inference and reliability.

Part 3Mechanistic Interpretability: The Core

Features, circuits, superposition, SAEs, and causal methods.

Part 4Frontier Interpretability

Circuit tracing, functional emotions, and the global workspace (2025–2026).

Part 5Safety, Steering & Editing

Applying interpretability: behavior control, weight editing, and the safety landscape.

Built in the spirit of applying all of this to AI safety first — with side quests into performance & reliability, steering & character, and models that learn on the fly. Full plan in CURRICULUM.md.