I outlined a five‑stage pipeline that I use to take a randomly initialized language model all the way to a reasoning‑capable system that can be aligned with human preferences.
Stage 0 – Randomly initialized LLM – You ask it “What is an LLM?” and get gibberish like “try peter hand and hello 448Sn”.
Stage 1 – Pre‑training – This stage teaches the LLM the basics of language by training it on massive corpora to predict the next token.
Stage 2 – Instruction fine‑tuning – In PFT: The user chooses between 2 responses to produce human preference data.
Stage 3 – Preference fine‑tuning (PFT) – It teaches the LLM to align with humans even when there’s no "correct" answer.
Stage 4 – Reasoning fine‑tuning – This is called Reinforcement Learning with Verifiable Rewards. GRPO by DeepSeek is a popular technique.