CertificateAdvanced

Architecting Large Language Models from Scratch: From Perceptrons to Post-Training

Master the inner workings of LLMs by building a GPT-2-style transformer from scratch. Learn neural network fundamentals, GPU optimization, attention mechanisms, pre-training pipelines, instruction tuning, and policy optimization using PyTorch.

Justin Angel
Justin Angel@justinangelai
Published Jun 28, 2026
Updated Jun 29, 2026
Architecting Large Language Models from Scratch: From Perceptrons to Post-Training
14h 23m
2Enrolled0Bookmarked
/ 5.0

Course Overview

Demystify Large Language Models (LLMs) by building, training, and optimizing a GPT-2-style transformer from first principles. This course bridges the gap between high-level API usage and low-level implementation, taking you on a journey from a single perceptron to advanced reinforcement learning techniques.

What You'll Learn

  • Fundamentals of Neural Networks: Build perceptrons, implement activation functions (ReLU, GELU, ReLU²), and construct Multi-Layer Perceptrons (MLPs) using PyTorch.
  • Training Mechanics & Optimization: Understand cross-entropy loss, backpropagation, and state-of-the-art optimizers like AdamW.
  • Deep Network Stability: Prevent training collapse using He/Xavier initialization, additive residual streams, and pre-norm RMSNorm/LayerNorm.
  • The Transformer Architecture: Implement causal self-attention, Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and assemble a complete generative transformer.
  • Data & Pre-Training: Clean raw web crawls, implement Byte Pair Encoding (BPE), build high-throughput sharded dataloaders, and run pre-training.
  • Post-Training & Alignment: Fine-tune models using LoRA/QLoRA, and align them with human preferences using direct policy optimization (SimPO).

Prerequisites

  • Strong proficiency in Python programming.
  • Comfort with basic programming concepts (no prior advanced math or machine learning experience required, as fundamentals are built from the ground up).
  • Access to Google Colab or a local GPU environment for training exercises.

Tools & Technologies

  • Frameworks: PyTorch, Hugging Face Transformers, Tiktoken
  • Hardware Acceleration: CUDA, Triton, GPU-performant coding
  • Analysis & Visualization: Microsoft Excel / Google Sheets (for visual math demos), Bertviz, Neuronpedia

Course Content

8Chapters24Lessons8Quizzes
1

Introduction to LLMs, Text Generation, and Architecture Exploration

2 lessons · 1h 6mQuiz
1

Autoregressive Text Generation and Sampling Heuristics

50:41
2

Reverse-Engineering GPT-2 and Mapping the Architecture

15:32
2

Neural Network Foundations: Perceptrons, Activations, and GPU Optimization

4 lessons · 1h 43mQuiz
1

Perceptrons: The Mathematical Engine of Deep Learning

15:56
2

Activation Functions: Injecting Non-Linearity

25:54
3

GPU Coding: Optimizing PyTorch Execution with Fused Kernels

21:53
4

Multi-Layer Perceptrons (MLPs) and Feed-Forward Networks

40:14
3

Training Mechanics: MLPs, Loss Functions, Backpropagation, and Model Persistence

4 lessons · 2h 14mQuiz
1

Building Multi-Layer Perceptrons and Feed-Forward Networks

40:14
2

Loss Functions: From Residuals to Cross-Entropy

25:56
3

Backpropagation and Optimization Loops

56:57
4

Saving and Loading Models: Serialization Formats

11:39
4

Stabilizing Deep Networks: Initialization, Residuals, Normalization, and Regularization

5 lessons · 2h 40mQuiz
1

Random Initialization & Preventing Training Collapse

36:29
2

Residual Connections & Signal Preservation

27:32
3

Normalization: LayerNorm, RMSNorm, and Pre-Norm Architectures

35:52
4

Regularization: Weight Decay, Gradient Clipping, and Dropout

40:27
5

SoftMax: Mapping Logits to Probabilities

20:13
5

Representing Text: Tokenization, Embeddings, and Positional Encodings

2 lessons · 1h 19mQuiz
1

Subword Tokenization and Byte Pair Encoding (BPE)

30:46
2

Token Embeddings and Positional Encodings

48:41
6

The Core Transformer: Causal Attention, Multi-Head Architectures, and Full Model Assembly

2 lessons · 1h 52mQuiz
1

Implementing Causal Self-Attention and Multi-Head Attention

58:27
2

Assembling the Full GPT-2 Transformer Block

53:52
7

Pre-Training Pipelines and LLM Evaluation Strategies

2 lessons · 1h 47mQuiz
1

Building High-Quality Pre-Training Pipelines with FineWeb-Edu

72:49
2

Scientific Evaluation of LLMs: Benchmarks, Lambada, and Perplexity

34:55
8

Post-Training Adaptation: Instruction Tuning, Reinforcement Learning, and Scaling

3 lessons · 1h 38mQuiz
1

Instruction Tuning, Alpaca Formats, and LoRA Fine-Tuning

40:25
2

Alignment and Direct Preference Optimization with SimPO

42:11
3

Scaling Up: Parallelism, FlashAttention, and KV Caching

16:22
PriceFree
LanguageEnglish
XP8640
This course includes
About the Channel
Justin Angel
Justin Angel
@justinangelai1.8K
1 Course2 Learners

AI, ML, LLMs.

Share

Related Courses