{"schemaVersion":"1.0","type":"TechArticle","types":["Article","TechArticle"],"slug":"build-a-small-llm-from-scratch-a-tested-gpt-in-pytorch-62ccl","url":"https://zyvop.com/build-a-small-llm-from-scratch-a-tested-gpt-in-pytorch-62ccl","title":"Build a Small LLM From Scratch: A Tested GPT in PyTorch","subtitle":"Build a compact GPT-style transformer in PyTorch, train it on Tiny Shakespeare, test the critical pieces, and see exactly how an 813K-parameter language model learns.","tldr":"A practical, fully tested GPT implementation in PyTorch with causal attention, training, sampling, baselines, checkpoints, and 25 tests. Trained on Tiny Shakespeare with 813K parameters and a final validation loss near 1.78.","keywords":["Tutorial"],"entities":["Pradeep Kumar","Tutorial","ZyVOP"],"keyTakeaways":["You can train a real language model on a laptop CPU during a coffee break.","It will have every part a big one has.","It's a decoder-only transformer in PyTorch, the same family as GPT-2, with 813,440 parameters."],"headings":["What a language model actually does","The architecture","Setup","Step 1: Data","Step 2: The model","Causal self-attention","MLP, block, and the full model","Step 3: Test it before you train it","Step 4: Train","Step 5: Is 1.78 any good?","Step 6: Generate text","What this model can and can't do","How to scale it up","Tests","Gotchas I hit while building it","Download and run","FAQ","References"],"outboundLinks":["https://github.com/pradeep200892/tinygpt","https://arxiv.org/abs/1706.03762","https://github.com/karpathy/char-rnn","https://github.com/karpathy/nanoGPT","https://arxiv.org/abs/2203.15556"],"contentText":"You can train a real language model on a laptop CPU during a coffee break. It won't be smart. It will have every part a big one has. This post builds one. It's a decoder-only transformer in PyTorch, the same family as GPT-2, with 813,440 parameters. It trains on Tiny Shakespeare and ends at a validation loss of 1.78. All the code ran before it went into this post. The full project is in tinygpt.zip: model, training script, sampler, baselines, 25 tests, and the trained checkpoint (out/best.pt). Every number below comes from one run on a single CPU core, with Python 3.12.3 and PyTorch 2.14.0. What a language model actually does A language model predicts the next token. That's the whole job. Feed it To be or not to b and it returns a probability for every possible next character. Training pushes those probabilities toward what the text really says. The loss is cross-entropy: the negative log of the probability the model gave the correct character. Loss also gives you a floor to compare against. A model that guesses uniformly across our 65 characters scores ln(65) = 4.174. An untrained network should land right there, and ours does. Chat models work the same way, with subword tokens, billions of parameters, and more training stages on top. The core loop doesn't change. We'll use characters instead of subword tokens. The vocabulary is 65 symbols, so there's no tokenizer to train. The price is that the model spends capacity learning to spell. GPT-2 used byte pair encoding with a 50,257-token vocabulary instead. The architecture flowchart TD A[\"Character ids, shape (B, T)\"] --&gt; B[\"Token embedding + position embedding\"] B --&gt; C[\"Block 1\"] C --&gt; D[\"Block 2\"] D --&gt; E[\"Block 3\"] E --&gt; F[\"Block 4\"] F --&gt; G[\"Final LayerNorm\"] G --&gt; H[\"Linear head, tied to the token embedding\"] H --&gt; I[\"Logits, shape (B, T, 65)\"]Each block has two sublayers. Both are wrapped in a residual connection, and both normalize their input first. flowchart LR x[\"x\"] --&gt; ln1[\"LayerNorm\"] ln1 --&gt; att[\"Causal self-attention\"] att --&gt; add1((\"+\")) x --&gt; add1 add1 --&gt; ln2[\"LayerNorm\"] ln2 --&gt; mlp[\"MLP: 128 to 512 to 128\"] mlp --&gt; add2((\"+\")) add1 --&gt; add2 add2 --&gt; out[\"block output\"]Moving LayerNorm to the input of each sublayer is one of the changes the GPT-2 paper made to the original transformer. It's why this design is called pre-LN. The original Transformer paper, Attention Is All You Need, applied it after. Setup pip install torch numpy pytest unzip tinygpt.zip cd tinygpt python get_data.pyget_data.py downloads Tiny Shakespeare from Andrej Karpathy's char-rnn repo. The file is 1,115,394 bytes. On macOS with a Python from python.org, that download can fail with a CERTIFICATE_VERIFY_FAILED error. The script catches this and falls back to curl. The FAQ at the end explains the real fix. Files in the project: data.py: character tokenizer, train/validation split, batching model.py: the GPT itself train.py: training loop, learning-rate schedule, evaluation sample.py: text generation from a checkpoint checkpoint.py: save and load baselines.py: unigram and bigram losses for comparison tests/test_tinygpt.py: 25 tests Step 1: Data data.py \"\"\"Character tokenizer and batching for tinygpt.\"\"\" from pathlib import Path import torch class CharTokenizer: \"\"\"Maps every distinct character in the corpus to an integer id.\"\"\" def __init__(self, chars): self.chars = list(chars) self.stoi = {c: i for i, c in enumerate(self.chars)} self.itos = {i: c for i, c in enumerate(self.chars)} @classmethod def from_text(cls, text): return cls(sorted(set(text))) def start_id(self): \"\"\"Id that starts generation: newline if the corpus has one, else id 0.\"\"\" return self.stoi.get(\"\\n\", 0) @property def vocab_size(self): return len(self.chars) def encode(self, s): return [self.stoi[c] for c in s] def decode(self, ids): return \"\".join(self.itos[int(i)] for i in ids) def load_corpus(path, val_fraction=0.1): \"\"\"Read a text file and return (tokenizer, train_ids, val_ids). The split is contiguous: the last `val_fraction` of the file is validation data. Shuffling characters would leak train text into val. \"\"\" text = Path(path).read_text(encoding=\"utf-8\") tok = CharTokenizer.from_text(text) ids = torch.tensor(tok.encode(text), dtype=torch.long) n_val = int(len(ids) * val_fraction) return tok, ids[:-n_val], ids[-n_val:] def get_batch(data, batch_size, block_size, device=\"cpu\", generator=None): \"\"\"Sample `batch_size` random windows. x is a window of block_size tokens. y is the same window shifted one position to the right, so y[t] is the token that follows x[t]. \"\"\" hi = len(data) - block_size - 1 if hi &lt;= 0: raise ValueError(f\"need at least {block_size + 2} tokens, got {len(data)}\") starts = torch.randint(0, hi, (batch_size,), generator=generator) x = torch.stack([data[s : s + block_size] for s in starts]) y = torch.stack([data[s + 1 : s + 1 + block_size] for s in starts]) return x.to(device), y.to(device)The tokenizer sorts the distinct characters in the file and numbers them. Sorting makes the mapping deterministic. start_id returns the newline's id if the corpus has one, and generation starts from it when you give no prompt. from data import load_corpus tok, train_ids, val_ids = load_corpus(\"data/input.txt\") print(tok.vocab_size, len(train_ids), len(val_ids)) print(repr(\"\".join(tok.chars))) print(tok.encode(\"First\")) print(tok.decode(tok.encode(\"First\")))65 1003855 111539 \"\\n !$&amp;',-.3:;?ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz\" [18, 47, 56, 57, 58] FirstThe vocabulary holds a newline, a space, a handful of punctuation marks, the digit 3, and both alphabets. That's 65 symbols. The last 10% of the text becomes validation data. The split is contiguous on purpose. A random split would scatter overlapping windows of the same passages across both sets, and validation loss would stop measuring anything. get_batch samples random windows. The target y is the input x shifted one position right. It refuses data shorter than block_size + 2 tokens. import torch from data import load_corpus, get_batch tok, train_ids, val_ids = load_corpus(\"data/input.txt\") torch.manual_seed(1337) x, y = get_batch(train_ids, batch_size=1, block_size=8) print(repr(tok.decode(x[0]))) print(repr(tok.decode(y[0])))\"Let's he\" \"et's hea\"Look at the two strings. From L the target is e, from Le it's t, from Let it's '. One window of 8 characters gives you 8 training examples at once. Step 2: The model model.py, imports and config \"\"\"A small GPT: decoder-only transformer with pre-LayerNorm blocks.\"\"\" import math from dataclasses import dataclass, asdict import torch import torch.nn as nn import torch.nn.functional as F@dataclass class GPTConfig: vocab_size: int = 65 block_size: int = 128 # maximum context length n_layer: int = 4 n_head: int = 4 n_embd: int = 128 dropout: float = 0.1 def to_dict(self): return asdict(self)The config is small on purpose: 4 layers, 4 heads, 128 dimensions, and a context of 128 characters. Each head works on 128 / 4 = 32 dimensions. Causal self-attention model.py, the attention layer class CausalSelfAttention(nn.Module): def __init__(self, cfg): super().__init__() assert cfg.n_embd % cfg.n_head == 0, \"n_embd must divide by n_head\" self.n_head = cfg.n_head self.qkv = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=False) self.proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=False) self.attn_drop = nn.Dropout(cfg.dropout) self.resid_drop = nn.Dropout(cfg.dropout) # Lower-triangular matrix: position t may look at positions 0..t only. mask = torch.tril(torch.ones(cfg.block_size, cfg.block_size)) self.register_buffer( \"mask\", mask.view(1, 1, cfg.block_size, cfg.block_size), persistent=False ) def forward(self, x): B, T, C = x.shape hs = C // self.n_head # head size q, k, v = self.qkv(x).split(C, dim=2) # (B, T, C) -&gt; (B, n_head, T, hs) q = q.view(B, T, self.n_head, hs).transpose(1, 2) k = k.view(B, T, self.n_head, hs).transpose(1, 2) v = v.view(B, T, self.n_head, hs).transpose(1, 2) att = (q @ k.transpose(-2, -1)) * (hs ** -0.5) # (B, nh, T, T) att = att.masked_fill(self.mask[:, :, :T, :T] == 0, float(\"-inf\")) att = self.attn_drop(F.softmax(att, dim=-1)) y = att @ v # (B, nh, T, hs) y = y.transpose(1, 2).contiguous().view(B, T, C) # merge heads return self.resid_drop(self.proj(y))One linear layer produces queries, keys and values together, and split cuts them apart. The view and transpose calls reshape each from (B, T, 128) to (B, 4, T, 32), so every head gets its own 32 dimensions. The scores are q @ k.T, scaled by hs ** -0.5. The scaling comes from the scaled dot-product attention in the Transformer paper. The authors suspect that for large head sizes the dot products grow in magnitude, which pushes softmax into regions with extremely small gradients. The mask is the part that makes this a language model. Position t may look at positions 0..t and nothing later. Here's the mask for 4 positions, and what softmax does with it: import torch mask = torch.tril(torch.ones(4, 4)) print(mask) torch.manual_seed(0) scores = torch.randn(4, 4).masked_fill(mask == 0, float(\"-inf\")) print(torch.softmax(scores, dim=-1).round(decimals=2))tensor([[1., 0., 0., 0.], [1., 1., 0., 0.], [1., 1., 1., 0.], [1., 1., 1., 1.]]) tensor([[1.0000, 0.0000, 0.0000, 0.0000], [0.5400, 0.4600, 0.0000, 0.0000], [0.4500, 0.0900, 0.4600, 0.0000], [0.1300, 0.4100, 0.3600, 0.0900]])Masked scores become -inf before softmax, so they turn into exact zeros. The second row spreads its weight across positions 0 and 1 only. Before rounding, every row sums to 1. PyTorch ships a fused version of this in scaled_dot_product_attention, with an is_causal flag. We write it by hand so you can see every step. A test later in this post checks that our layer matches PyTorch's output. MLP, block, and the full model model.py, MLP and block class MLP(nn.Module): def __init__(self, cfg): super().__init__() self.fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=False) self.proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=False) self.drop = nn.Dropout(cfg.dropout) def forward(self, x): return self.drop(self.proj(F.gelu(self.fc(x))))class Block(nn.Module): def __init__(self, cfg): super().__init__() self.ln1 = nn.LayerNorm(cfg.n_embd) self.attn = CausalSelfAttention(cfg) self.ln2 = nn.LayerNorm(cfg.n_embd) self.mlp = MLP(cfg) def forward(self, x): x = x + self.attn(self.ln1(x)) x = x + self.mlp(self.ln2(x)) return xThe MLP expands each position from 128 to 512 dimensions, applies GELU, and projects back. Attention mixes information across positions. The MLP transforms each position on its own. The block adds each sublayer's output to its input. Those residual paths let gradients flow straight through a deep stack. model.py, the GPT class class GPT(nn.Module): def __init__(self, cfg): super().__init__() self.cfg = cfg self.wte = nn.Embedding(cfg.vocab_size, cfg.n_embd) # token embeddings self.wpe = nn.Embedding(cfg.block_size, cfg.n_embd) # position embeddings self.drop = nn.Dropout(cfg.dropout) self.blocks = nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]) self.ln_f = nn.LayerNorm(cfg.n_embd) self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False) self.lm_head.weight = self.wte.weight # weight tying self.apply(self._init_weights) # GPT-2 trick: shrink the layers that write into the residual stream, # so the stream's variance doesn't grow with depth. for name, p in self.named_parameters(): if name.endswith(\"attn.proj.weight\") or name.endswith(\"mlp.proj.weight\"): nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer)) @staticmethod def _init_weights(module): if isinstance(module, (nn.Linear, nn.Embedding)): nn.init.normal_(module.weight, mean=0.0, std=0.02) def forward(self, idx, targets=None): B, T = idx.shape assert T &lt;= self.cfg.block_size, f\"sequence length {T} &gt; block_size {self.cfg.block_size}\" pos = torch.arange(T, device=idx.device) x = self.drop(self.wte(idx) + self.wpe(pos)) for block in self.blocks: x = block(x) logits = self.lm_head(self.ln_f(x)) # (B, T, vocab) loss = None if targets is not None: loss = F.cross_entropy(logits.view(-1, logits.size(-1)), targets.view(-1)) return logits, loss @torch.no_grad() def generate(self, idx, max_new_tokens, temperature=1.0, top_k=None, generator=None): assert temperature &gt; 0, \"use top_k=1 for greedy decoding\" was_training = self.training self.eval() for _ in range(max_new_tokens): ctx = idx[:, -self.cfg.block_size :] # crop to the window logits, _ = self(ctx) logits = logits[:, -1, :] / temperature # last position only if top_k is not None: kth = torch.topk(logits, min(top_k, logits.size(-1))).values[:, [-1]] logits = logits.masked_fill(logits &lt; kth, float(\"-inf\")) probs = F.softmax(logits, dim=-1) nxt = torch.multinomial(probs, num_samples=1, generator=generator) idx = torch.cat([idx, nxt], dim=1) self.train(was_training) return idx def num_params(self): # parameters() counts the tied embedding/head matrix once return sum(p.numel() for p in self.parameters())Three details in this class matter more than they look. First, the token embedding and the output head share one weight matrix. The input side turns an id into a vector, and the output side turns a vector back into scores over ids. Tying them saves 65 × 128 = 8,320 parameters, and the original Transformer paper shares its embedding and pre-softmax matrices the same way. Second, initialization: every weight starts from a normal distribution with standard deviation 0.02. The two projections that write into the residual stream, attn.proj and mlp.proj, get that value divided by sqrt(2 * n_layer). The GPT-2 paper describes scaling residual layers by 1/√N, where N is the number of residual layers. Each block contributes two, hence the 2 *. nanoGPT implements it the same way. Third, generate crops the context to the last block_size tokens. The position embedding table only has 128 rows. Let's count parameters and check the untrained loss. import math import torch from data import load_corpus, get_batch from model import GPT, GPTConfig tok, train_ids, val_ids = load_corpus(\"data/input.txt\") torch.manual_seed(1337) model = GPT(GPTConfig(vocab_size=tok.vocab_size)) print(f\"{model.num_params():,} parameters\") print(\"token embedding :\", model.wte.weight.numel()) print(\"position embed :\", model.wpe.weight.numel()) print(\"one block :\", sum(p.numel() for p in model.blocks[0].parameters())) print(\"final LayerNorm :\", sum(p.numel() for p in model.ln_f.parameters())) x, y = get_batch(train_ids, 32, 128) logits, loss = model(x, y) print(logits.shape) print(f\"loss {loss.item():.4f} ln(65) = {math.log(65):.4f}\")813,440 parameters token embedding : 8320 position embed : 16384 one block : 197120 final LayerNorm : 256 torch.Size([32, 128, 65]) loss 4.1827 ln(65) = 4.1744The numbers line up with the architecture: Part Parameters Token embedding (65 × 128) 8,320 Position embedding (128 × 128) 16,384 Each block (4 of them) 197,120 Final LayerNorm 256 Total 813,440 A block holds 65,536 attention weights (4 × 128²), 131,072 MLP weights (8 × 128²), and 512 LayerNorm parameters. The untrained loss of 4.1827 sits right on ln(65) = 4.1744. Step 3: Test it before you train it Training a broken transformer wastes an afternoon and doesn't announce itself. The loss still goes down. Three checks catch most bugs first. The first is the one you just saw: initial loss near ln(vocab size). The second is a leak test. Change a token at position 10 and confirm that nothing before position 10 moves. tests/test_tinygpt.py def test_model_cannot_see_the_future(): torch.manual_seed(0) model = GPT(small_cfg()).eval() a = torch.randint(0, 20, (1, 16)) b = a.clone() b[0, 10] = (b[0, 10] + 1) % 20 # change the token at position 10 la, _ = model(a) lb, _ = model(b) assert torch.allclose(la[:, :10], lb[:, :10], atol=1e-6) # earlier positions unchanged assert not torch.allclose(la[:, 10:], lb[:, 10:]) # position 10 onward changesA model that can see the answer gets near-zero training loss and generates junk. This test is what tells you the mask works. The third check is overfitting. A healthy model can memorize one small batch. def test_can_overfit_one_batch(): torch.manual_seed(0) model = GPT(small_cfg()) opt = torch.optim.AdamW(model.parameters(), lr=3e-3) x = torch.randint(0, 20, (4, 16)) y = torch.randint(0, 20, (4, 16)) first = model(x, y)[1].item() for _ in range(300): _, loss = model(x, y) opt.zero_grad() loss.backward() opt.step() assert first &gt; 2.5 assert loss.item() &lt; 0.1If the loss can't reach zero on 4 sequences, something is wrong in the forward pass or the optimizer. Step 4: Train train.py, schedule, optimizer and evaluation def get_lr(step, max_steps, lr, min_lr, warmup): \"\"\"Linear warmup, then cosine decay from lr down to min_lr.\"\"\" if step &lt; warmup: return lr * (step + 1) / warmup if step &gt;= max_steps: return min_lr progress = (step - warmup) / (max_steps - warmup) return min_lr + 0.5 * (1 + math.cos(math.pi * progress)) * (lr - min_lr)def make_optimizer(model, lr, weight_decay, betas=(0.9, 0.99)): \"\"\"AdamW that applies weight decay to matrices only, not to biases or norms.\"\"\" decay = [p for p in model.parameters() if p.dim() &gt;= 2] no_decay = [p for p in model.parameters() if p.dim() &lt; 2] groups = [ {\"params\": decay, \"weight_decay\": weight_decay}, {\"params\": no_decay, \"weight_decay\": 0.0}, ] return torch.optim.AdamW(groups, lr=lr, betas=betas)@torch.no_grad() def estimate_loss(model, splits, batch_size, block_size, eval_iters, device): model.eval() out = {} for name, data in splits.items(): losses = torch.zeros(eval_iters) for i in range(eval_iters): x, y = get_batch(data, batch_size, block_size, device) losses[i] = model(x, y)[1].item() out[name] = losses.mean().item() model.train() return outThe learning rate warms up over 100 steps to 1e-3, then follows a cosine curve down to 1e-4. Warmup keeps the first updates small while Adam's statistics settle. AdamW uses betas of (0.9, 0.99). The second value matches nanoGPT's Shakespeare config, which raises it because each step sees few tokens. Weight decay of 0.1 applies to weight matrices only. Biases and LayerNorm parameters are excluded, and a test checks the split. estimate_loss averages over several batches and switches to eval mode, which turns dropout off. It switches back afterward. train.py, the training loop (excerpt from main) torch.manual_seed(args.seed) device = pick_device(args.device) out = Path(args.out) out.mkdir(parents=True, exist_ok=True) tok, train_ids, val_ids = load_corpus(args.data) cfg = GPTConfig( vocab_size=tok.vocab_size, block_size=args.block_size, n_layer=args.n_layer, n_head=args.n_head, n_embd=args.n_embd, dropout=args.dropout, ) model = GPT(cfg).to(device) opt = make_optimizer(model, args.lr, args.weight_decay) splits = {\"train\": train_ids, \"val\": val_ids} print(f\"device={device} params={model.num_params():,} vocab={tok.vocab_size} \" f\"train_tokens={len(train_ids):,} val_tokens={len(val_ids):,}\") sample_gen = torch.Generator(device=device).manual_seed(0) prompt = torch.tensor([[tok.start_id()]], device=device) history, samples, best_val = [], [], float(\"inf\") t0 = time.time() for step in range(args.max_steps + 1): if step % args.eval_interval == 0 or step == args.max_steps: losses = estimate_loss(model, splits, args.batch_size, args.block_size, args.eval_iters, device) text = tok.decode(model.generate(prompt, 200, generator=sample_gen)[0].tolist()) history.append({\"step\": step, **losses, \"seconds\": round(time.time() - t0)}) samples.append({\"step\": step, \"text\": text}) print(f\"step {step:5d} | train {losses['train']:.4f} | val {losses['val']:.4f} \" f\"| {time.time() - t0:6.0f}s\", flush=True) if losses[\"val\"] &lt; best_val: best_val = losses[\"val\"] save_checkpoint(out / \"best.pt\", model, tok, step, best_val) (out / \"history.json\").write_text(json.dumps(history, indent=2)) (out / \"samples.json\").write_text(json.dumps(samples, indent=2)) if step == args.max_steps: break lr = get_lr(step, args.max_steps, args.lr, args.min_lr, args.warmup) for g in opt.param_groups: g[\"lr\"] = lr x, y = get_batch(train_ids, args.batch_size, args.block_size, device) _, loss = model(x, y) opt.zero_grad(set_to_none=True) loss.backward() torch.nn.utils.clip_grad_norm_(model.parameters(), args.grad_clip) opt.step()Each step draws a batch, computes the loss, backpropagates, clips the gradient norm to 1.0, and updates. Every 250 steps it evaluates, generates a 200-character sample, and saves best.pt if validation loss improved. A step processes 32 × 128 = 4,096 characters. Two thousand steps is 8,192,000 characters, or about 8.2 passes over the 1,003,855-character training split. Run it: python train.py --data data/input.txt --out outdevice=cpu params=813,440 vocab=65 train_tokens=1,003,855 val_tokens=111,539 step 0 | train 4.1835 | val 4.1799 | 13s step 250 | train 2.3962 | val 2.4186 | 160s step 500 | train 2.1546 | val 2.1950 | 307s step 750 | train 1.9662 | val 2.0432 | 450s step 1000 | train 1.8321 | val 1.9598 | 595s step 1250 | train 1.7527 | val 1.8790 | 740s step 1500 | train 1.6836 | val 1.8391 | 891s step 1750 | train 1.6436 | val 1.8080 | 1044s step 2000 | train 1.6215 | val 1.7838 | 1188s best val loss 1.7838 (2.573 bits/char)That took 1,188 seconds on one CPU core. The best validation loss was 1.7838, which is 2.573 bits per character. The samples that train.py writes to out/samples.json show the model learning in stages. These are the first 110 characters at four checkpoints: step 0 t3ZxjmhPPdkARxLsXltUWHoHca ?eghkyy&amp;&amp;NdkIgAlsvSBUMyccRD$BygXOfntw kVrM h&amp;?j SJR!sOEHelC!BTg3xUhPRCs3Ma--3KypMrW step 250 Inol3 f pifie-mily mpoennd shay tinte. The beome igle inkee he them. KES: S y tonoong sof themyowangh. HOn d step 1000 DORY: Not fieve my that that Romest thear not ut . LUCED: What the gelivl the. Shou jroses. Find With you lets step 2000 Eing Noices you, for villing a by like whose mork hard, Shall appy'd mired: thou almp the place! CORIOLANUS:At step 0 it's noise drawn from all 65 symbols. By step 250 it has word-length chunks, sentence-ending periods, and even a stray speaker tag (KES:). By step 1000 it knows that a name in capitals followed by a colon starts a speech. By step 2000 the format is solid, and many of the words are still not real ones. The training loss ends at 1.6215 and validation at 1.7838. That gap of 0.16 is small, and validation was still falling at the last checkpoint, so overfitting isn't what limits this run. The learning rate also decays to 1e-4 by step 2,000. That makes the flattening curve a poor sign that the model has run out of things to learn. Here's a tighter validation estimate on 200 batches instead of 40: import math import torch from checkpoint import load_checkpoint from data import load_corpus from train import estimate_loss model, tok, meta = load_checkpoint(\"out/best.pt\") _, train_ids, val_ids = load_corpus(\"data/input.txt\") torch.manual_seed(123) val = estimate_loss(model, {\"val\": val_ids}, 32, 128, 200, \"cpu\")[\"val\"] print(f\"val loss {val:.4f} perplexity {math.exp(val):.2f} bits/char {val / math.log(2):.3f}\")val loss 1.7899 perplexity 5.99 bits/char 2.582Perplexity of 5.99 means the model is about as uncertain as a fair choice between six characters at each step. The 1.7899 differs from the 1.7838 in the log only because it samples different batches. Step 5: Is 1.78 any good? A loss means nothing without a baseline. Two are easy to compute on the same validation split. The unigram model predicts from character frequencies alone. The bigram model predicts from the previous character alone. baselines.py def unigram_loss(train_ids, val_ids, vocab): counts = torch.bincount(train_ids, minlength=vocab).float() + 1 # add-one smoothing probs = counts / counts.sum() return -probs[val_ids].log().mean().item()def bigram_loss(train_ids, val_ids, vocab): pairs = train_ids[:-1] * vocab + train_ids[1:] counts = torch.bincount(pairs, minlength=vocab * vocab).view(vocab, vocab).float() + 1 probs = counts / counts.sum(dim=1, keepdim=True) # P(next | previous) return -probs[val_ids[:-1], val_ids[1:]].log().mean().item()python baselines.pyuniform guess : 4.1744 unigram : 3.3473 bigram : 2.4819The bigram table, just 65 × 65 counts, reaches 2.48. The transformer's 1.78 comes from using more than the previous character. That gap is what attention over 128 characters of context buys you. Step 6: Generate text model.py, the sampling loop inside generate ctx = idx[:, -self.cfg.block_size :] # crop to the window logits, _ = self(ctx) logits = logits[:, -1, :] / temperature # last position only if top_k is not None: kth = torch.topk(logits, min(top_k, logits.size(-1))).values[:, [-1]] logits = logits.masked_fill(logits &lt; kth, float(\"-inf\")) probs = F.softmax(logits, dim=-1) nxt = torch.multinomial(probs, num_samples=1, generator=generator) idx = torch.cat([idx, nxt], dim=1)The model outputs scores for every position, and we keep the last one. Dividing by temperature reshapes the distribution. Below 1 it sharpens, above 1 it flattens. Top-k keeps only the k highest scores and masks the rest to -inf. Then softmax turns scores into probabilities and multinomial draws one character. The code asserts that temperature is positive, so greedy decoding uses top_k=1 instead. sample.py \"\"\"Generate text from a trained checkpoint. python sample.py --ckpt out/best.pt --prompt \"ROMEO:\" --tokens 400 \"\"\" import argparse import torch from checkpoint import load_checkpoint def main(): ap = argparse.ArgumentParser() ap.add_argument(\"--ckpt\", default=\"out/best.pt\") ap.add_argument(\"--prompt\", default=None) ap.add_argument(\"--tokens\", type=int, default=400) ap.add_argument(\"--temperature\", type=float, default=0.8) ap.add_argument(\"--top-k\", type=int, default=40) ap.add_argument(\"--seed\", type=int, default=42) ap.add_argument(\"--device\", default=\"cpu\") args = ap.parse_args() model, tok, _ = load_checkpoint(args.ckpt, args.device) gen = torch.Generator(device=args.device).manual_seed(args.seed) if not args.prompt: ids = [tok.start_id()] else: unknown = sorted({c for c in args.prompt if c not in tok.stoi}) if unknown: raise SystemExit(f\"prompt has characters the model has never seen: {unknown}\") ids = tok.encode(args.prompt) idx = torch.tensor([ids], device=args.device) out = model.generate(idx, args.tokens, args.temperature, args.top_k, generator=gen) print(tok.decode(out[0].tolist())) if __name__ == \"__main__\": main()Temperature 0.8 with top-k 40 is the default: python sample.py --prompt \"ROMEO:\" --tokens 200ROMEO: As lord shall, you for the read, one and bed forther Bear head laist to dreath be that God, Why so, upon this up enemer fament. FLAUDE: I lord, the the fall for not presence forfe, The cuntracia, buTemperature 1.2 with no top-k filter (--top-k 65 keeps all 65 characters): python sample.py --prompt \"ROMEO:\" --tokens 200 --temperature 1.2 --top-k 65ROMEO: As loud shadford: whose chard's, Statenty no fe thy eread? First MAUXEqd: Sorrow that's planly so, Two to it'u Let it them could to hinde lilenct, a who and for neWrp Suanntifhed but nighnie his, buAnd greedy decoding, which always takes the most likely character: python sample.py --prompt \"ROMEO:\" --tokens 200 --top-k 1ROMEO: The shall the soul the soul the soul the soul the son, And the shall the some the so the soul the stand The shall the so the so the so the so the soul The shall the shall the shall the shall the stayThe default sample has the shape of a play: speaker names in capitals, colons, short lines, and a few real words like \"lord\", \"head\" and \"God\". It has no grammar and no meaning. At 1.2 it invents words, and greedy decoding falls into a loop. Greedy decoding always picks the single most likely character, so once a phrase repeats, the repeated context makes the next repeat even more likely. Sampling adds randomness that can break the cycle. What this model can and can't do It learned the surface of Shakespeare: layout, capitalization, common letter patterns. It didn't learn to write sentences. That's expected. The model has 813,440 parameters and read about a million characters, eight times. I ran one training run with one seed and didn't tune any hyperparameters. Everything ran on CPU only. How to scale it up The nanoGPT README gives useful reference points on the same dataset. Its CPU example (4 layers, 4 heads, 128 dimensions, context 64) takes about 3 minutes and reaches a loss of 1.88. Its GPU config (6 layers, 6 heads, 384 dimensions, context 256) reaches 1.4697 in about 3 minutes on one A100. The settings differ from ours, so treat those as rough comparisons. The same README lists GPT-2 at 124M parameters. That's about 150 times larger than tinygpt. It says train.py reproduces that model on OpenWebText in about 4 days on one node with 8 A100 40GB GPUs. To move toward that, change things in this order: More data. A 1 MB corpus caps everything else. A subword tokenizer. Byte pair encoding shortens sequences, so each token carries more meaning. A bigger model. Raise n_layer, n_head and n_embd, and lengthen block_size. A GPU. train.py picks CUDA or MPS automatically. I haven't run those paths. Don't do step 3 alone. Hoffmann et al. trained over 400 models and found that model size and training tokens should scale together: double the parameters, double the tokens. Their 70B model, Chinchilla, beat the 280B Gopher on a range of tasks using the same compute and 4× more data. The paper is Training Compute-Optimal Large Language Models. Tests python -m pytest -q......................... [100%] 25 passed in 12.23sThe suite covers the tokenizer round trip, batch shifting, the parameter-count formula, tied weights, the causal mask, agreement with PyTorch's attention, overfitting a single batch, reproducible sampling, greedy decoding, the learning-rate schedule, the weight-decay split, checkpoint loading, a two-step training run on a tiny corpus, the sampler's rejection of characters it has never seen, and the download script's fallback when Python can't verify certificates. I also broke the code on purpose. Deleting the masked_fill line from the attention layer makes exactly two tests fail: the leak test and the PyTorch-equivalence test. The other 23 still pass, which shows those two guard that line. Gotchas I hit while building it Split contiguously. Sample validation windows from the tail of the file, not from a shuffle of the whole text. Decay matrices only. Weight decay on biases and LayerNorm gains does nothing useful. The dim() &gt;= 2 split handles it. Crop the context. generate must slice to block_size, or long generations crash on the position table. The test asks for 40 new tokens with a block size of 16. Seed your sampler. A torch.Generator passed to generate makes samples reproducible, and one test checks exactly that. Don't hard-code a start token. My first version seeded sampling with a newline and crashed on any corpus that had none, which the tiny-corpus test now covers. Save the vocabulary. The checkpoint stores the character list next to the weights. Without it, ids can't be turned back into text. Download and run The project is tinygpt.zip. It contains every file listed above, the trained checkpoint in out/, and the training log. unzip tinygpt.zip cd tinygpt pip install -r requirements.txt python get_data.py python sample.py --prompt \"ROMEO:\" python -m pytestRetraining with python train.py writes to out/ and overwrites the shipped checkpoint. Copy out/ somewhere first if you want to keep it. FAQ Is this really an LLM? It's a language model built the same way as one. The \"large\" is about scale, and this is roughly 150 times smaller than the 124M GPT-2 model. Can I train it on my own text? Yes. Run python train.py --data yourfile.txt. The vocabulary comes from the file. With the default 128-character context the file needs at least 1,300 characters. Shorter files stop with a ValueError that says how many tokens are missing. Why does sampling fail on some prompts? The tokenizer only knows characters from the training text. sample.py stops and lists the ones it doesn't know: python sample.py --prompt \"ROMEO: ☃\"prompt has characters the model has never seen: ['☃']Calling encode yourself raises a KeyError instead: from data import CharTokenizer tok = CharTokenizer.from_text(\"hello\") try: tok.encode(\"hello!\") except KeyError as e: print(\"KeyError:\", e)KeyError: '!'Why does get_data.py fail with CERTIFICATE_VERIFY_FAILED? Python installed from python.org on macOS often ships without root certificates, so it can't verify any HTTPS server. get_data.py detects this and downloads with curl instead, which keeps verification on. To fix Python itself, run Install Certificates.command from the /Applications/Python 3.x/ folder. Don't switch verification off to make the error go away. Why does pip install torch say \"No matching distribution\"? Usually PyTorch doesn't publish a wheel for your Python version and Mac chip. I checked Python 3.14 with pip install --dry-run: Apple Silicon has wheels, and the Intel Mac tag I tried (macosx_10_15_x86_64) has none. Run uname -m to see your chip (arm64 means Apple Silicon). I tested everything on Python 3.12.3. What is the \"Failed to initialize NumPy\" warning? PyTorch prints it when NumPy isn't installed. Nothing in tinygpt uses NumPy, and the sampler, the training script and the test suite ran fine without it in my check. pip install numpy silences the message, and the install commands above include it. Does it run on a GPU? The code picks CUDA or MPS if PyTorch finds one. I only tested it on CPU. Why is the output nonsense? The model is small and the data is tiny. Better output comes from more data, a bigger model, and longer training, in that order. References Vaswani et al., Attention Is All You Need, 2017. Radford et al., Language Models are Unsupervised Multitask Learners (GPT-2), 2019. Source for pre-LN placement and residual-scaled initialization. Hoffmann et al., Training Compute-Optimal Large Language Models, 2022. Andrej Karpathy, nanoGPT and char-rnn. PyTorch documentation, scaled_dot_product_attention.","contentHash":"sha256:1a2f224176dc5a6db671479890b17e837afc056cce5da8e9fc2c4df51837d968","authorName":"Pradeep Kumar","authorUrl":"https://zyvop.com/author/pradeep","authorSameAs":[],"category":"Tutorial","tags":[],"audience":"Software engineers and developers building applications with Tutorial","tone":"Instructional, practical, code-first","readingTimeMinutes":24,"wordCount":5291,"faqs":null,"primaryTopic":"Tutorial","publishedAt":"2026-10-01T09:47:07.559Z","updatedAt":"2026-10-01T10:18:15.899Z","canonicalUrl":"https://zyvop.com/build-a-small-llm-from-scratch-a-tested-gpt-in-pytorch-62ccl"}