RSS Amplifier

Video feed

Reinforcement Learning of Large Language Models

youtube.comSource feed ↗12 videos

Dormant Last read · last published · next check
Read 3 days ago and current, but nothing has been published for 13 months.

Written by

Latest videos

Saves to your Watch queue, to pick up on another day or another device.

[UCLA RL-LLM] Chapter 2.4: In-context learning and instruction fine-tuning

Play

[UCLA RL-LLM] Chapter 2.3: Transformers II (modern transformers updates and sampling methods)

Play

[UCLA RL-LLM] Chapter 2.2: Transformers I (BERT, GPT-1)

Play

[UCLA RL-LLM] Chapter 2.1: NLP foundations, language modeling, RNNs

Play

[UCLA RL-LLM] Chapter 3.2: Reinforcement learning with verifiable rewards (RLVR)

Play

[UCLA RL-LLM] Chapter 3.1: Reinforcement learning from human feedback (PPO, DPO)

Play

[UCLA RL-LLM] Chapter 1.5: AlphaGo, test-time compute, and expert iteration

Play

[UCLA RL-LLM] Chapter 1.4: Deep policy gradient methods (PPO, GRPO)

Play

[UCLA RL-LLM] Chapter 1.3: Deep policy gradient methods (A3C)

Play

[UCLA RL-LLM] Chapter 1.2: Deep policy evaluation

Play

[UCLA RL-LLM] Chapter 1.1: MDP foundations, imitation learning, and value iteration

Play

[UCLA RL-LLM] Chapter 0: Course outline and prologue

Play