| 8/20/26 |
Introduction |
Heckman & Zhao |
A0 out |
|
|
|
| 8/25/26 |
Scaling Laws & Model Codesign |
Heckman |
|
Training Compute-Optimal Large Language Models (Chinchilla); Are Emergent Abilities of Large Language Models a Mirage?. Optional: Scaling Laws for Neural Language Models; Scaling Data-Constrained Language Models; The Llama 3 Herd of Models; Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws |
|
Slides |
| 8/27/26 |
Model Architectures & Trade-offs |
Heckman |
A1 out |
DeepSeek-V3 Technical Report (architecture sections only); GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. Optional: Fast Transformer Decoding: One Write-Head is All You Need; Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity; Mixtral of Experts; Mamba: Linear-Time Sequence Modeling with Selective State Spaces; RoFormer: Enhanced Transformer with Rotary Position Embedding |
|
Slides |
| 9/1/26 |
ML Systems and MLOps |
Zhao |
|
We Have No Idea How Models will Behave in Production until Production; Roofline: An Insightful Visual Performance Model for Multicore Architectures |
|
Slides, Annotated |
| 9/3/26 |
Model Evaluation |
Heckman |
A1 due; A1 quiz in class; A2 out |
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena; Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. Optional: A Careful Examination of Large Language Model Performance on Grade School Arithmetic (GSM1k); Establishing Task Scaling Laws via Compute-Efficient Model Ladders; Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference; Paloma: A Benchmark for Evaluating Language Model Fit |
|
Slides |
| 9/8/26 |
Data Synthesis & Curation |
Heckman |
|
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale. Optional: DataComp-LM: In search of the next generation of training sets for language models; DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining; Self-Instruct: Aligning Language Models with Self-Generated Instructions; The Curse of Recursion: Training on Generated Data Makes Models Forget; Textbooks Are All You Need |
|
Slides |
| 9/10/26 |
Data Infrastructure I |
Zhao |
|
Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics |
|
Slides, Annotated |
| 9/15/26 |
Data Infrastructure II |
Zhao |
A2 due; A2 quiz in class; A3 out |
Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training |
|
Slides, Annotated |
| 9/17/26 |
Training Systems I |
Zhao |
|
The Ultra-Scale Playbook: Training LLMs on GPU Clusters - Section on Data Parallelism |
|
Slides, Annotated |
| 9/22/26 |
Training Systems II |
Zhao |
|
The Ultra-Scale Playbook: Training LLMs on GPU Clusters - Section on Tensor and Pipeline Parallelism |
|
Slides, Annotated |
| 9/24/26 |
Training Systems III |
Zhao |
A3 due; A3 quiz in class; A4 out |
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models |
|
Slides, Annotated |
| 9/29/26 |
Project Day |
Heckman |
|
|
|
Slides |
| 10/1/26 |
Fine-tuning & PEFT |
Heckman |
|
LoRA: Low-Rank Adaptation of Large Language Models. Optional: QLoRA: Efficient Finetuning of Quantized LLMs; S-LoRA: Serving Thousands of Concurrent LoRA Adapters; LIMA: Less Is More for Alignment; Reducing Text Bias in Synthetically Generated MCQAs for VLMs in Autonomous Driving |
Sa., Bl. C., Mi. |
Slides |
| 10/6/26 |
Inference-Time Compute |
Heckman |
A4 due; A4 quiz in class; A5 out |
FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. Optional: s1: Simple test-time scaling; RouteLLM: Learning to Route LLMs with Preference Data; Fast Inference from Transformers via Speculative Decoding; Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads; Chain-of-Thought Prompting Elicits Reasoning in Large Language Models |
Dav., Ja., Di. |
|
| 10/8/26 |
Reading Day (campus) — no class |
|
|
|
|
|
| 10/13/26 |
Model Quantization |
Heckman |
Mega assignment concept document due |
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. Optional: GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers; LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale; SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models; DeepSeek-V3 Technical Report (FP8 sections) |
O., Dan., L. |
|
| 10/15/26 |
Inference Systems I |
Zhao |
|
Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM). Optional: Orca: A Distributed Serving System for Transformer-Based Generative Models; FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness |
Ti., Ni., Ab. |
|
| 10/20/26 |
Inference Systems II |
Zhao |
A5 due; A5 quiz in class; A6 out |
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. Optional: Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve; Splitwise: Efficient Generative LLM Inference Using Phase Splitting |
An. D., Te., Bl. D. |
|
| 10/22/26 |
Inference Systems III |
Zhao |
|
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving. Optional: SGLang: Efficient Execution of Structured Language Model Programs |
Su., E., At. |
|
| 10/27/26 |
Mega Assignment Pitch Day |
Heckman |
Mega assignment pitches in class |
|
|
|
| 10/29/26 |
Post-training: RLHF, GRPO and RLVR |
Heckman |
|
Direct Preference Optimization: Your Language Model is Secretly a Reward Model; DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. Optional: Training language models to follow instructions with human feedback (InstructGPT); DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models; KTO: Model Alignment as Prospect Theoretic Optimization; Constitutional AI: Harmlessness from AI Feedback; Tulu 3: Pushing Frontiers in Open Language Model Post-Training |
No., Ne., Se. |
|
| 11/3/26 |
Agentic AI |
Zhao |
A6 due; A6 quiz in class; A7 out |
ReAct: Synergizing Reasoning and Acting in Language Models. Optional: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering; Agentix: An Efficient Serving Engine for LLM Agents as General Programs |
Z., Ad., Bi. |
|
| 11/5/26 |
Prompt Engineering & Tools |
Heckman |
|
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. Optional: Self-Consistency Improves Chain of Thought Reasoning in Language Models |
Y., Jo., Ma. |
|
| 11/10/26 |
Sustainability and ML |
Zhao |
|
Sustainable AI: Environmental Implications, Challenges and Opportunities. Optional: Carbon Emissions and Large Neural Network Training; Chasing Carbon: The Elusive Environmental Footprint of Computing |
Mo., H., Be. |
|
| 11/12/26 |
Applications: Robotics |
Heckman |
|
π0: A Vision-Language-Action Flow Model for General Robot Control. Optional: OpenVLA: An Open-Source Vision-Language-Action Model; RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control; Diffusion Policy: Visuomotor Policy Learning via Action Diffusion; Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT) |
Al., An. W., Sa. |
|
| 11/17/26 |
Applications: Vision at Scale |
Heckman |
|
Sigmoid Loss for Language Image Pre-Training (SigLIP). Optional: DINOv2: Learning Robust Visual Features without Supervision; Learning Transferable Visual Models From Natural Language Supervision (CLIP); Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution; GAIA-1: A Generative World Model for Autonomous Driving; Cosmos World Foundation Model Platform for Physical AI |
Bl. C., Mi., Dav. |
|
| 11/19/26 |
Applications: GeoAI |
Heckman |
|
GraphCast: Learning skillful medium-range global weather forecasting. Optional: A Foundation Model for the Earth System (Aurora); SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery; Foundation Models for Generalist Geospatial Artificial Intelligence (Prithvi) |
Ja., Di., O. |
|
| 11/24/26 |
Fall break — no class |
|
|
|
|
|
| 11/26/26 |
Fall break — no class |
|
|
|
|
|
| 12/1/26 |
Last Lecture |
Heckman & Zhao |
|
|
|
|
| 12/3/26 |
Poster Session |
Heckman & Zhao |
|
|
|
|
| 12/9/26 |
Final Exam (Wednesday) |
|
|
|
|
|