GeoRanker
← All work
NLP / LLMPythonNumPyTransformers

Transformers from first principles

To optimize for language models you should understand how they read. I rebuilt the core pieces by hand.

10

runnable labs

3

stages completed

2

languages (EN / 中文)

Why I did it

GEO advice online is often folklore. I wanted a grounded mental model of what a language model actually does with a page of text: how it breaks text into tokens, represents meaning as vectors, and decides which parts of the context to attend to. That understanding shapes practical choices — clear entities, self-contained passages, consistent naming.

What’s in it

Stage 1 — NLP foundations

  • lab1_tokenizer.py — turning text into tokens
  • lab2_onehot.py — one-hot vectors and why they fail to capture meaning
  • lab3_embeddings.py — dense embeddings, similarity and t-SNE visualisation

Stage 2 — Attention with Q/K/V

Implements the core formula and checks every matrix dimension by hand:

Attention(Q, K, V) = softmax(Q·Kᵀ / √d_k) · V
  • Basic attention, with weight heat-maps
  • Scaled attention — why dividing by √d_k keeps softmax gradients healthy
  • Multi-head attention, comparing 1, 2 and 4 heads

Stage 3 — The Transformer block

  • Sinusoidal positional encoding
  • Layer norm vs batch norm, and residual connections
  • Feed-forward networks, ReLU vs GELU
  • A complete assembled Transformer block, plus a BERT / GPT / T5 architecture comparison

Stage 4 (applications such as classification, NER and fine-tuning with Hugging Face) is planned.

What I took from it for GEO

  • Attention rewards clear, local context. Passages that state who, what and where in the same paragraph are easier to retrieve and quote than facts scattered across a page.
  • Tokens aren’t words. Inconsistent brand and product naming fragments the signal. Pick one canonical name.
  • Structure is a cheap hint. Headings, lists and schema.org markup give models and crawlers unambiguous anchors.