NLP / LLMPythonNumPyTransformers
Transformers from first principles
To optimize for language models you should understand how they read. I rebuilt the core pieces by hand.
10
runnable labs
3
stages completed
2
languages (EN / 中文)
Why I did it
GEO advice online is often folklore. I wanted a grounded mental model of what a language model actually does with a page of text: how it breaks text into tokens, represents meaning as vectors, and decides which parts of the context to attend to. That understanding shapes practical choices — clear entities, self-contained passages, consistent naming.
What’s in it
Stage 1 — NLP foundations
lab1_tokenizer.py— turning text into tokenslab2_onehot.py— one-hot vectors and why they fail to capture meaninglab3_embeddings.py— dense embeddings, similarity and t-SNE visualisation
Stage 2 — Attention with Q/K/V
Implements the core formula and checks every matrix dimension by hand:
Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V
- Basic attention, with weight heat-maps
- Scaled attention — why dividing by √d_k keeps softmax gradients healthy
- Multi-head attention, comparing 1, 2 and 4 heads
Stage 3 — The Transformer block
- Sinusoidal positional encoding
- Layer norm vs batch norm, and residual connections
- Feed-forward networks, ReLU vs GELU
- A complete assembled Transformer block, plus a BERT / GPT / T5 architecture comparison
Stage 4 (applications such as classification, NER and fine-tuning with Hugging Face) is planned.
What I took from it for GEO
- Attention rewards clear, local context. Passages that state who, what and where in the same paragraph are easier to retrieve and quote than facts scattered across a page.
- Tokens aren’t words. Inconsistent brand and product naming fragments the signal. Pick one canonical name.
- Structure is a cheap hint. Headings, lists and schema.org markup give models and crawlers unambiguous anchors.