RSS Amplifier

Topic · policy gradient methods

policy gradient methods

The 32 most recent episodes and tracks on this topic.

Saves to your Watch queue, to pick up on another day or another device.

Pick anything below and it plays in the bar at the foot of the window — and keeps playing while you go on browsing the directory.

  1. [UCLA RL-LLM] Chapter 2.4: In-context learning and instruction fine-tuningReinforcement Learning of Large Language ModelsNotes
  2. [UCLA RL-LLM] Chapter 2.3: Transformers II (modern transformers updates and sampling methods)Reinforcement Learning of Large Language ModelsNotes
  3. [UCLA RL-LLM] Chapter 2.2: Transformers I (BERT, GPT-1)Reinforcement Learning of Large Language ModelsNotes
  4. [UCLA RL-LLM] Chapter 2.1: NLP foundations, language modeling, RNNsReinforcement Learning of Large Language ModelsNotes
  5. [UCLA RL-LLM] Chapter 3.2: Reinforcement learning with verifiable rewards (RLVR)Reinforcement Learning of Large Language ModelsNotes
  6. [UCLA RL-LLM] Chapter 3.1: Reinforcement learning from human feedback (PPO, DPO)Reinforcement Learning of Large Language ModelsNotes
  7. [UCLA RL-LLM] Chapter 1.5: AlphaGo, test-time compute, and expert iterationReinforcement Learning of Large Language ModelsNotes
  8. [UCLA RL-LLM] Chapter 1.4: Deep policy gradient methods (PPO, GRPO)Reinforcement Learning of Large Language ModelsNotes
  9. [UCLA RL-LLM] Chapter 1.3: Deep policy gradient methods (A3C)Reinforcement Learning of Large Language ModelsNotes
  10. [UCLA RL-LLM] Chapter 1.2: Deep policy evaluationReinforcement Learning of Large Language ModelsNotes
  11. [UCLA RL-LLM] Chapter 1.1: MDP foundations, imitation learning, and value iterationReinforcement Learning of Large Language ModelsNotes
  12. [UCLA RL-LLM] Chapter 0: Course outline and prologueReinforcement Learning of Large Language ModelsNotes
  13. Outlook and Research Insights (Safe, Edge and Meta Reinforcement Learning - Lecture 14, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  14. Further Contemporary RL Algorithms (TRPO, PPO - Lecture 13, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  15. Deterministic Policy Gradient Methods (Lecture 12, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  16. Stochastic Policy Gradient Methods (Lecture 11, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  17. Value-Based Control with Function Approximation (Lecture 10, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  18. On-Policy Prediction with Function Approximation (Lecture 09, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  19. Function Approximation with Supervised Learning (Lecture 08, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  20. Planning and Learning with Tabular Methods (Lecture 07, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  21. Multi-Step Bootstrapping (Lecture 06, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  22. Temporal Difference Learning (Lecture 05, Summer 2023)Reinforcement Learning Course: Lectures (Summer 2023)Notes
  23. Lecture 01 IntroductionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  24. Lecture 02 Markov Decision ProcessesCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  25. Lecture 03 Solving known MDPsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  26. Lecture 04 Solving Known MDPsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  27. Lecture 05 Monte Carlo MethodsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  28. Lecture 06 Temporal Difference MethodCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  29. Lecture 07 Neural Networks Architectures for RLCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  30. Lecture 08 Function Approximation for PredictionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  31. Lecture 09 Value FunctionCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes
  32. Lecture 10 Policy Gradient MethodsCMU: 2018 Fall: 10-703 Deep Reinforcement Learning and ControlNotes