Two Chinese outlets ran nearly the same headline on September 18: the first session of Tang Jie’s “Advanced Machine Learning” course at Tsinghua was standing-room-only. Tencent News called it 挤爆了 — “packed to bursting.” That’s not usually how you’d describe a syllabus reading.
Then people actually read the syllabus, and the reaction made more sense. Students in this course write a tokenizer and a Transformer from scratch and train a 0.1-billion-parameter model with it. They write their own Triton attention kernel and benchmark training and inference across multiple GPUs. They clean raw data by hand, chase down scaling laws, and try to predict what a bigger model would do before they ever train it. On one shared base model, they run SFT, DPO, and RLVR side by side to see how each actually changes behavior. Then they build an environment and a harness, train a long-context agent inside it, and get that agent to grade its own output.
That’s not a semester project. That’s most of a small research lab’s roadmap, assigned as coursework.
Who assigns homework like this
Tang Jie isn’t a random professor who got ambitious. He’s faculty in Tsinghua’s Computer Science department, runs the Knowledge Engineering Group’s research on the AI side, holds IEEE, ACM, and AAAI fellowships, and built AMiner back in 2006 — the kind of citation graph tool most of us have used without knowing whose lab made it. He’s also co-founder and chief scientist of Zhipu AI (Z.AI), the company behind the GLM model family, which he and Li Juanzi spun out of that same Tsinghua lab in 2019. So when he tells a room of grad students to build a Triton kernel by hand, he’s not describing an academic exercise he read about. He’s describing what his own team does before breakfast.
Two details in the version of this story going around — teams of two or three, and a NeurIPS-format English report with a live demo — didn’t show up in either of the Tsinghua news pieces I could find. I’m not throwing them out; they fit the rest of the syllabus too well to be invented, and course pages like this rarely make it into press coverage in full. Just don’t treat them as confirmed the way the technical milestones are.
Tang also used the moment to say something more specific about where he thinks the field is heading, and this part is on record. At Beijing’s AGI-Next Summit back in January, he told the room the chat paradigm is “基本已经探索完了” — basically fully explored. Not “solved,” not “boring,” just: this particular curve has flattened, and Zhipu is steering budget toward coding, agents, and reasoning instead. His picture of what comes next is multimodality, memory, and self-improvement — models that keep state across sessions and get better at a task without a human retraining them in between. The “knowledge → coding → digital world → physical world” framing you’ll see floating around, and the claim that AI could run your whole PC autonomously in a year or two — I couldn’t pin either of those to an actual quote. They read like a reasonable extrapolation of what he has said, dressed up as something he said. Worth knowing the difference before you repeat it.
None of that changes the interesting part, though: the syllabus. If a professor with a foundation-model lab thinks this is the right way to teach people about LLMs in 2026, it’s worth asking what happens if you try it without the lab.
So what would it actually take to do this yourself
You don’t have Tsinghua’s GPU cluster, TAs, or a grade forcing you to finish. You do have evenings, weekends, and a credit card that can rent GPUs by the hour. Here’s the version of this course I’d actually run for myself, milestone by milestone, with what I’d cut if I were doing it solo instead of in a team.
1. Tokenizer and Transformer from scratch. Don’t start from a blank file — start from Karpathy’s minbpe for the tokenizer and nanoGPT (or the newer build-nanogpt) for the model, read every line, then rewrite them yourself without the tab open. Train your own byte-pair tokenizer on a few hundred megabytes of text first; a from-scratch BPE merge loop is short enough to actually understand start to finish:
def get_pair_counts(ids):
counts = {}
for pair in zip(ids, ids[1:]):
counts[pair] = counts.get(pair, 0) + 1
return counts
def merge(ids, pair, new_id):
out, i = [], 0
while i < len(ids):
if i < len(ids) - 1 and (ids[i], ids[i+1]) == pair:
out.append(new_id); i += 2
else:
out.append(ids[i]); i += 1
return out
Get that working before you touch attention. Then scale the model up toward 0.1B params — that’s within reach of a single rented A100 or H100 for a weekend, not a cluster.
2. Triton kernel plus multi-GPU benchmarking. Skip writing attention from raw CUDA — start from the official fused-attention tutorial in the Triton repo, get it numerically matching PyTorch’s SDPA, then benchmark it against SDPA and FlashAttention-2 at a few sequence lengths. For the multi-GPU piece, you don’t need eight GPUs; two rented GPUs and torchrun --nproc_per_node=2 running data-parallel training is enough to produce a real scaling curve and force you to understand gradient sync, not just imagine it.
3. Data cleaning and scaling laws. Pull a slice of FineWeb-Edu instead of scraping your own — the cleaning problem is the same, the crawling problem isn’t worth your time. Do the boring parts anyway: dedup with MinHash, filter with a small quality classifier, look at what got removed. Then train three or four model sizes (10M, 30M, 100M, maybe 300M params) on the same data recipe, plot loss against compute, and fit a power law the way the Chinchilla paper does. The point isn’t the exact exponent you get — it’s noticing that a curve fit on tiny models actually predicts your next size up.
4. SFT, DPO, and RLVR on one base model. Use TRL for all three so the comparison is about the method, not the tooling. SFT on a small instruction set like a filtered slice of UltraChat. DPO on a preference-pairs dataset. RLVR on something with a checkable answer — grade-school math or a code-execution task — so “reward” means “the test passed,” not “a judge model liked it.” Run all three from the identical checkpoint and compare on the same eval set. This is the milestone that actually teaches you something the other three don’t: SFT, DPO, and RL move a model’s behavior in genuinely different ways, and you won’t believe that until you’ve watched it happen to your own model.
5. Agent harness with self-evaluation. Build the simplest possible loop — plan, call a tool, observe, repeat — around your own model or a small open one, extend its usable context with a scratchpad or retrieval step, then add a self-critique pass where the model scores its own trajectory against a rubric before it stops. Measure whether self-eval actually catches bad runs or just rubber-stamps them. My bet, and the reason this milestone is on the syllabus at all, is that you’ll find it does both, inconsistently — which is the real state of the art right now, not a bug in your implementation.
6. Write it up like it’s real. Even solo, force yourself into a short structured report — problem, method, plots, what didn’t work — and record a two-minute demo. Nobody’s grading it, but a report is what turns six weekends of GPU rental into something you can point at later.
Budget and timeline, honestly: a solo engineer with a day job can get through a trimmed version of all six milestones in ten to twelve weeks of evenings, spending somewhere around $300–600 total on rented GPU time if you’re disciplined about turning instances off. A pair can split it — one person owns pretraining and the kernel, the other owns data, alignment, and the agent — and go faster without either person understanding it shallower. What I wouldn’t cut, even under time pressure: milestone 4. Reading about the difference between SFT and RL is not the same as watching your own model behave differently after each one.
I’m half-tempted to actually run this on myself. If I do, you’ll get the version with worse GPUs and no TAs — but the same six milestones.