Topic · policy gradient methods
policy gradient methods
The 32 most recent episodes and tracks on this topic.
Saves to your Watch queue, to pick up on another day or another device.
Pick anything below and it plays in the bar at the foot of the window — and keeps playing while you go on browsing the directory.
- [UCLA RL-LLM] Chapter 2.4: In-context learning and instruction fine-tuningReinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 2.3: Transformers II (modern transformers updates and sampling methods)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 2.2: Transformers I (BERT, GPT-1)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 2.1: NLP foundations, language modeling, RNNsReinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 3.2: Reinforcement learning with verifiable rewards (RLVR)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 3.1: Reinforcement learning from human feedback (PPO, DPO)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 1.5: AlphaGo, test-time compute, and expert iterationReinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 1.4: Deep policy gradient methods (PPO, GRPO)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 1.3: Deep policy gradient methods (A3C)Reinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 1.2: Deep policy evaluationReinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 1.1: MDP foundations, imitation learning, and value iterationReinforcement Learning of Large Language ModelsNotes
- [UCLA RL-LLM] Chapter 0: Course outline and prologueReinforcement Learning of Large Language ModelsNotes
- Outlook and Research Insights (Safe, Edge and Meta Reinforcement Learning - Lecture 14, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Further Contemporary RL Algorithms (TRPO, PPO - Lecture 13, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Deterministic Policy Gradient Methods (Lecture 12, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Stochastic Policy Gradient Methods (Lecture 11, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Value-Based Control with Function Approximation (Lecture 10, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- On-Policy Prediction with Function Approximation (Lecture 09, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Function Approximation with Supervised Learning (Lecture 08, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Planning and Learning with Tabular Methods (Lecture 07, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Multi-Step Bootstrapping (Lecture 06, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Temporal Difference Learning (Lecture 05, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
- Lecture 01 IntroductionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 02 Markov Decision ProcessesCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 03 Solving known MDPsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 04 Solving Known MDPsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 05 Monte Carlo MethodsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 06 Temporal Difference MethodCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 07 Neural Networks Architectures for RLCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 08 Function Approximation for PredictionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 09 Value FunctionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
- Lecture 10 Policy Gradient MethodsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
This playlist:.m3u.plsAll the feeds behind it
