# mdp (blogs) — RSS Amplifier

Recent posts from the 2 feeds in the RSS Amplifier directory that cover mdp.

Page: <https://rssamplifier.com/topics/mdp/blogs>  
Feed: <https://rssamplifier.com/topics/mdp/blogs.md>

---

## [ލޭ ހަދިޔާކުރުމުގައި ދީލަތިވުމަކީ ކޮންމެ މީހަކަށްވެސް އޮތް ފުރުސަތެއް: މެޑަމް ސާޖިދާ](https://tnn.mv/Ley-hadhiyaa-Kurumugai-Dheelathi-Vumakee-Konme-Meehakah-ves-Oiy-Furusatheh%3AMadam-Sajidha)

_2026-06-16 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Ley Hadhiyaa Kurumugai Dheelathi Vumakee Konme Meehakah Ves Oiy Furusatheh:Madam Sajidha

## [ޕޭޕަލްގެ ހިދުމަތާއިއެކު ބައިނަލްއަގުވާމީ މައިދާނުގައި ފައިތިލަ ސާބިތުކުރުމުގެ ފުރުސަތު ހުޅުވިއްޖެ: ރައީސް](https://tnn.mv/Paypal-ge-Hidhumathai-eku-bainal-Aguvaamee-maidhaanugai-Faithila-saabithu-kurumuge-Furusathu-Hulhuvejje%3ARaees)

_2026-06-15 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Paypal Ge Hidhumathai Eku Bainal Aguvaamee Maidhaanugai Faithila Saabithu Kurumuge Furusathu Hulhuvejje:Raees

## [ސަރުކާރުން މިވަނީ ރައްޔިތުންގެ ބޮލުގައި އެޅުނު ދަރަނި ކުޑަކޮށްދީފައި، ދެން ފެންނާނީ ތަރައްގީގެ ސްޕީޑް ބާރުވާތަން: ރައީސް](https://tnn.mv/Sarukaarun-Mivanee-Rayyithunge-Bolugai-Elhunu-Dharani-Kudakohdheefai%2Cdhen-fennaanee-Thaarhgeege-speed-Baaruvaathan%3ARaess)

_2026-06-15 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Sarukaarun Mivanee Rayyithunge Bolugai Elhunu Dharani Kudakohdheefai,dhen Fennaanee Thaarhgeege Speed Baaruvaathan:Raess

## [އެމްޑީޕީއަކީ &quot;ފަނާކުރަނިވި&quot; ފިކުރެއް،ޕީއެންސީގެ ފިކުރަކީ ގުޑުވާނުލެވޭނެ ވަރުގަދަ އަސާސެއް: ރައީސް މުއިއްޒު](https://tnn.mv/MDP-Akee-Fanaa-Kuruvanivi-Fikureh-%2CPNC-Ge-Fikurakee-Guduvaanuleveyne-Varugadha-Asaaseh%3ARaees-Muizzu)

_2026-06-15 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

MDP Akee Fanaa Kuruvanivi Fikureh ,PNC Ge Fikurakee Guduvaanuleveyne Varugadha Asaaseh:Raees Muizzu

## [Off-Policy Policy Evaluation](https://cruxponent.com/post/off_policy_eval/)

_2026-06-15 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

For its own sake or as part of a policy iteration scheme, evaluating policies is an important matter. This post is concerned with off-policy policy evaluation, or evaluating policies with data generated by others. The ambition is to provide a smooth progression towards the Retrace estimator and some of its extensions, with a focus on each operator&rsquo;s properties and stochastic approximation…

## [ހިދުމަތްތައް ޒަމާނީކޮށް ހަރުދަނާކުރުމުގެ އަމާޒުގައި ސިވިލް ސަރވިސް ކޮންފަރެންސް ފަށައިފި](https://tnn.mv/Hidhumayhthah-Zamaaneekoh-Harudhanaa-kurumuge-Amaazugai-Civil-Service-Conference-Fashaifi)

_2026-06-13 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Hidhumayhthah Zamaaneekoh Harudhanaa Kurumuge Amaazugai Civil Service Conference Fashaifi

## [ބޭއިންސާފުން ކަންކަން ކުރާނަމަ ސަރުކާރުގެ މައި އޮފީސްތައް ހިފާނަން : ނަޝީދު](https://tnn.mv/Bey-insaafun-kankan-Kuraanama-Sarukaaruge-Mai-office-thah-Hifaanan-%3A-Nasheedh)

_2026-06-13 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Bey Insaafun Kankan Kuraanama Sarukaaruge Mai Office Thah Hifaanan : Nasheedh

## [މި ދަނޑިވަޅުގައި ނަޝީދުގެ ތަޖުރިބާއަކީ ސަރުކާރު ޖަވާބުދާރީ ކުރުވަން މުހިންމު ހަތިޔާރަކަށް ވެގެންދާނެ: ސޯލިހު](https://tnn.mv/Mi-Dhandivalhugai-Nasheedh-Ge-Thajuribaa-akee-sarukaaru-Javaabudhaaree-Kuruvan-Muhinmu-Hathiyaarakah-Vegendhaane%3A-Solih)

_2026-06-13 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Mi Dhandivalhugai Nasheedh Ge Thajuribaa Akee Sarukaaru Javaabudhaaree Kuruvan Muhinmu Hathiyaarakah Vegendhaane: Solih

## [ރައީސް ނަޝީދު އުފެއްދެވި ދަ ޑިމޮކްރެޓްސް އުވާލައިފި](https://tnn.mv/Raees-Nasheed-ufehdhevi-The-Democrats-uvaalaifi)

_2026-06-03 · އާއިދާ އަބްދުލް ހަކީމް · ޓީއެންއެން &#45; : ސިޔާސީ_

Raees Nasheed Ufehdhevi The Democrats Uvaalaifi

## [އަލަށް އިންތިހާބުވި ކައުންސިލަރުން ހުވާކުރައްވައިފި](https://tnn.mv/Alah-Inthihaabuvi-counciler-In-Huvaakurahvaifi)

_2026-05-17 · azla74609@gmail.com · ޓީއެންއެން &#45; : ސިޔާސީ_

Alah Inthihaabuvi Counciler In Huvaakurahvaifi

## [CLI-first decentralized GPU compute (Sponsored)](https://crawlproof.com/a/PPtj1I0djGWS)

_2026-05-17 · **Sponsored**_

Pay workers or run your GPU for FFmpeg transcode and AI inference.

## [ހާއްސަ މަޖިލީހުގެ މެމްބަރުކަން ކުރެއްވި ރާމިޒް އަވަހާރަވުމާއި ގުޅިގެން ރައީސް ތައުޒިޔާ ވިދާޅުވެއްޖެ](https://tnn.mv/Khaassa-majilis-ge-member-kan-kurehvi-Ramiz-avahaara-vumaai-gulhigen-Raees-Thauziyaa-vidhaalhuvejje)

_2026-05-09 · އާއިދާ އަބްދުލް ހަކީމް · ޓީއެންއެން &#45; : ސިޔާސީ_

Khaassa Majilis Ge Member Kan Kurehvi Ramiz Avahaara Vumaai Gulhigen Raees Thauziyaa Vidhaalhuvejje

## [Post-Training is a Contextual Bandit](https://cruxponent.com/post/rl_4_llm/)

_2026-04-29 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

Coming from a control background, applying RL to text generation is hardly intuitive. Is there an actual, non-degenerate dynamical system at play here? What is the concrete novelty behind the shiny LLM post-training algorithms? This post is an attempt to answer those questions by providing a semiformal derivation of the celebrated GRPO algorithm through a contextual bandit lens. $\\quad$ The focus…

## [Control in (Generally) Regularised MDPs](https://cruxponent.com/post/regularised_mdp/)

_2025-12-05 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

This post explores the theory of regularised MDPs beyond entropic regularisation (which we covered in an older post). We will introduce convex regularisation of the classical Bellman operators and study the induced regularised policy iteration algorithms. On the way, we will tie some links with several popular algorithms. This post is mostly a good excuse to refresh some convex optimisation…

## [Information Theory Cheat-Sheet](https://cruxponent.com/post/inf_theory/)

_2025-10-13 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

Entropy, divergence, mutual information, etc. are central concepts in statistical machine learning. This post ties them together in a short collection of elementary information theoretic results. Below, we consider random variables $\\mathrm{X}, \\mathrm{Y}$ that take values in some discrete sets $\\mathcal{X}$ and $\\mathcal{Y}$. We denote, respectively, $p\_{\\tiny\\mathrm{X}}\\in\\Delta\_\\mathcal{X}$ and…

## [A PPO Saga](https://cruxponent.com/post/ppo_saga/)

_2025-08-12 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

For better or for worse, proximal policy optimisation (PPO) algorithms and its variants are dominating the RL landscape these days. This post aims at retracing their journey, from foundational concepts to LLM-savy innovations. We will start this saga on the theoretical trail, which we will progressively abandon to pay closer attention to algorithmic aspects. $\\quad$ We rely on standard notations…

## [Variational Inference in POMDPs](https://cruxponent.com/post/vi_and_pomdp/)

_2025-06-20 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

The goal of this post is to explore from first principles the learning of belief models in partially observable MDPs. We will start with a quick refresher on variational inference, and apply it to state estimation in POMDPs. Specifically, we will derive the update rule used to train Dreamer-like models. Variational Inference In this post, we are interested in latent variable models. We consider…

## [Successor States and Representations (2/3)](https://cruxponent.com/post/successor_2/)

_2025-05-03 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

In this second post of this series, we take a break from successor measures to focus on successor features . We will first review the use of a generalised policy improvement mechanism that can efficiently leverage the successor features of existing policies to enable zero-shot transfer to new tasks. We will then discuss the generalisation to universal successor features approximations, allowing…

## [Average Reward Control (1/2)](https://cruxponent.com/post/mdp_average/)

_2025-04-09 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

Thanks to its relative simplicity and conciseness, the discounted approach to control in MDPs has come to largely prevail in the RL theory and practice landscape. Departing from the myopic nature of discounted control, we study here the average-reward objective which focuses on long-term, steady state rewards. To start gently, we will limit ourselves to establishing Bellman equations for policy…

## [Successor States and Representations (1/3)](https://cruxponent.com/post/successor_1/)

_2025-02-24 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

The main promise of unsupervised RL is test-time adaptation to newly specified reward functions. This requires a systemic untangling of reward and dynamics in traditional RL tools. In this first post of a short series, we see how this can be done via the concept of successor states and successor representations. We will focus on policy evaluation, leaving control for a follow-up. One over-arching…

## [Stream Torrents and IPTV Instantly (Sponsored)](https://crawlproof.com/a/79MLei2BTklc)

_2025-02-23 · **Sponsored**_

Search, index, and play music, movies, books, and live TV in your browser.

## [Oldies but goodies: Optimal State Estimation](https://cruxponent.com/post/optimal_filter/)

_2025-01-16 · l.faury@hotmail.fr (Louis Faury) · CruxPonent_

This post is interested in state estimation in HMMs: filtering, prediction and smoothing. We will introduce state estimation as the solution of an optimisation problem, and prove the celebrated recursive updates for each inference use-case. A special attention will be given to HMM filters (and how they easily generalise to the celebrated Kalman filters). $\\quad$ The reader interested about…

