5:23
Why LLMs Need Reinforcement Learning (The Alignment Problem)
5:34
Reward Models Explained: Turning Preference into a Number - RL for LLMs
4:41
RLHF & InstructGPT: The Recipe Behind ChatGPT RL for LLMs
4:31
Policy Gradients & REINFORCE, from Scratch Rl for LLMs
4:32
PPO Explained: The Clip and the KL Leash RL for LLMs
4:10
RLAIF & Constitutional AI: When AI Grades AI RL for LLMs
4:54
DPO: Alignment Without Reinforcement Learning
4:28
Rejection Sampling & Best-of-N: The Simplest RL. Rl for LLMs
4:25
Process Reward Models: Grading Every Step RL for LLMs
4:45
GRPO: Reasoning RL Without a Critic
o1 & Test-Time Compute: Teaching Models to Think RL for LLMs
4:52
DeepSeek-R1 & RLVR: Reasoning from Pure RL Rl for LLMs
4:50
Agentic RL: Rewarding Multi-Step Tool Use RL for LLMs
4:26
RLHF vs DPO vs GRPO: Which Method, and When RL for LLMs