All drills

Build GPT from Scratch

What you'll be able to do

Build the engine inside ChatGPT, Gemini, and Llama — a GPT you write yourself, from the tokenizer up to text generation, so the transformer is never a black box again.

The core architecture behind every modern LLM — OpenAI's GPT, Google's Gemini, Anthropic's Claude, and Meta's Llama all run decoder-only transformers like the one you build here.
Start this internship
Create an account to unlock the 10 sections, the workbench, and AskThili.
Begin

Sections

1. Building GPT from Scratch
🔒 locked
2. Lesson 1 - What a language model is
🔒 locked
3. Lesson 2 - Text to tokens
🔒 locked
4. Lesson 3 - Tokens to vectors
🔒 locked
5. Lesson 4 - Self-attention
🔒 locked
6. Lesson 5 - Look back, never forward
🔒 locked
7. Lesson 6 - The transformer block
🔒 locked
8. Lesson 7 - Assemble & train the GPT
🔒 locked
9. Lesson 8 - Make it talk (decoding strategies)
🔒 locked
10. Lesson 9 - Ship your mini-GPT
🔒 locked

Dig deeper

📄Attention Is All You Need (Vaswani et al., 2017) — the transformer
paper
📄Language Models are Unsupervised Multitask Learners (Radford et al., 2019) — GPT-2
paper
🔗The Illustrated Transformer (Jay Alammar) — a visual walkthrough of attention
article

Part of these learning paths

I'm a developer and I want to fine-tune small LLMs for my own use case
View path →