Mainak's PMRF Tutorials
Publishes 2 feeds
Jadavpur University, 2025: Introduction to Reinforcement Learning
10 posts · theirs
Jadavpur University: Foundation_Math_forML_Autumn23
10 posts · theirs
Lately
Session 21: Actor Critic based Policy Gradient, Safe RL, Planning, DYNA, Curriculum Learning
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 20: Deep Neural Networks, MLP, Backpropagation, Policy Gradient, REINFORCE
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 19: Asynchronous Q learning, Classification in ML, MLE, Logistic and Softmax Regression
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 18 Synchronous Q-learning, Model-free, based, tabular, with Linear Fn. Approx., Convergence
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 17: Off-Policy Evaluation of TD0 with linear function Approximation, Emphatic TD0
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 16 γ contraction, Banach's Fixed Point Theorem, How far is it far from the intended optimal
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 15 TD(0) convergence proof (contd), Point of Convergence of TD(0) (linear function approx.)
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 14: TD0 with linear function approximation, Glimpse at Stochastic Approximation Algorithm(1)
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 13: Function Approximation in RL, Policy Evaluation, SGD Monte Carlo, TD(0) Implementation
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 12: On Policy vs Off Policy Algorithms, Importance Sampling, Model-free Q learning, SARSA
Jadavpur University, 2025: Introduction to Reinforcement Learning ·
Session 10: Gradient descent, why it works, Linear and Logistic regression, ML estimate
Session 9: Introduction to convex functions, Jensen’s, Holder’s inequality, Minkowski, Lagrangian
Everything on this page was read from markup Mainak's PMRF Tutorials published — a rel="me" link, an h-card, or the feed’s own author element. Nothing was inferred from anywhere else. To correct or remove it, get in touch. Machine-readable: JSON
