# td (blogs) — RSS Amplifier

Recent posts from the 5 feeds in the RSS Amplifier directory that cover td.

Page: <https://rssamplifier.com/topics/td/blogs>  
Feed: <https://rssamplifier.com/topics/td/blogs.md>

---

## [Opening Reception: Beauty in Tension](https://www.focusonvictoria.ca/events/event/3611-opening-reception-beauty-in-tension/)

_2026-08-20 · Ongoing & Upcoming Events_

Art Rental + Sales at the Art Gallery of Greater Victoria presents Beauty in Tension Show & Sale, featuring artworks by Andrea Soos, Alanna Sparanese, and Candis Zimmerman. Join us for the opening reception on Thursday, September 17 to meet the artists, and enjoy refreshments. Admission to the Massey Sales Gallery is always free! Beauty in Tension runs from September 17 to November 8, 2026.

## [The Avenue Gallery - Feature Artist - Shell Cooper](https://www.focusonvictoria.ca/events/event/3610-the-avenue-gallery-feature-artist-shell-cooper/)

_2026-08-20 · Ongoing & Upcoming Events_

The Avenue Gallery is thrilled to introduce Shell Cooper, a South African-born, British Columbia-based abstract painter whose work is shaped by the landscapes and memories of Canada, Costa Rica, and her South African origins. Working intuitively in acrylic, Shell creates large-scale layered paintings that draw on botanical forms and remembered landscapes while exploring the emotional terrain of…

## [Reinforcement Learning Part 14: Baselines, the Advantage Function, and Actor-Critic](https://shawnhymel.com/3705/reinforcement-learning-part-14-baselines-the-advantage-function-and-actor-critic/)

_2026-08-19 · ShawnHymel · Shawn Hymel_

In the previous post, we ended with a straightforward application of the policy gradient in the REINFORCE algorithm, which proves to be a useful stepping stone in our deep reinforcement learning (RL) journey. Here, we updated the parameters of the policy approximator (often a neural network) using this formula: Notice that we are using the \[…\]

## [Radical Regalia: Elizabeth Carefoot at Gage Gallery](https://www.focusonvictoria.ca/events/event/3609-radical-regalia-elizabeth-carefoot-at-gage-gallery/)

_2026-08-19 · Ongoing & Upcoming Events_

Visit REGALIA: an art show celebrating traditional, sacred and ceremonial dress and objects. See Elizabeth Carefoot’s work exploring regalia through stitch, fabric and paint, poking gentle fun of the things people wear to show off their status. Wearing regalia is a powerful, non-verbal communicator that shows respect for the occasion and a sense of control for the wearer. Donning regalia brings a…

## [Four Friends Having Fun](https://www.focusonvictoria.ca/events/event/3606-four-friends-having-fun/)

_2026-08-15 · Ongoing & Upcoming Events_

For the 11th year, this popular art event is back! Karen Wilk, Linda Butcher, Lois Kissinger, and Shirley Sarens are setting up their painting studios at the Art Sea Gallery and spending a week there together ... welcoming visitors, painting, sharing stories, laughing and showing off artwork. Enjoy a walk by the sea and stop in for some art, fun, and smiles. ArtSea Gallery, 9565 Fifth Street,…

## [The Avenue Gallery - Feature Artist - Carolyn Houg](https://www.focusonvictoria.ca/events/event/3605-the-avenue-gallery-feature-artist-carolyn-houg/)

_2026-08-14 · Ongoing & Upcoming Events_

Carolyn Houg's sculptures are whimsical and full of personality, often placing subjects in unexpected situations. Each piece invites us to pause and reflect on the shared connections between all living beings. Carolyn studied Fine Arts in Montreal and Edmonton, working in oils, acrylics, printmaking and concrete before falling in love with clay. Now based on Vancouver Island, she brings decades of…

## [Reinforcement Learning Part 13: Policy Gradient Causality Trick and REINFORCE](https://shawnhymel.com/3648/reinforcement-learning-part-13-policy-gradient-causality-trick-and-reinforce/)

_2026-08-11 · ShawnHymel · Shawn Hymel_

In the previous post, we showed how we can substitute our usual ε-greedy policy with a parameterized approximation (often a neural network), we then derived the policy gradient theorem required to optimize this approximator function, and demonstrated how it can be estimated using Monte Carlo sampling. At the very end, we pointed out that the \[…\]

## [Avery Raquel ~ 2026 Slow Motion Tour](https://www.focusonvictoria.ca/events/event/3604-avery-raquel-~-2026-slow-motion-tour/)

_2026-08-10 · Ongoing & Upcoming Events_

Ocean Entertainment Worldwide, in association with A.J.E. Global Productions, presents eight limited engagement concerts starring Canada's Avery Raquel ~ Live in Concert. Witness the evolution of one of Canada’s most formidable young musical forces. At just 24 years old, Billboard-charting recording artist and multi-instrumentalist Avery Raquel takes the stage for a mesmerizing live experience…

## [Mixture of Experts (MoE): How Transformers Scale Without Activating Everything.](https://chizkidd.github.io//2026/08/10/mixture-of-experts/)

_2026-08-10 · Chizoba Obasi blog_

Mixture of Experts (MoE) is one of the main techniques used to scale modern language models without making every token pay the full computational cost of the model. The basic idea is surprisingly simple: instead of sending every token through one enormous feed-forward network, we split it into many smaller expert networks and only activate a few experts for each token. These notes walk through the…

## [The Avenue Gallery - Feature Artist - Air & Earth Design by Heidi von der Gathen](https://www.focusonvictoria.ca/events/event/3603-the-avenue-gallery-feature-artist-air-earth-design-by-heidi-von-der-gathen/)

_2026-08-06 · Ongoing & Upcoming Events_

Heidi von der Gathen creates Air & Earth Design, a contemporary jewellery collection inspired by organic forms and natural materials. Her handcrafted pieces combine clean lines, unique finishes and sculptural elements to create jewellery with a refined sensibility. Timeless rather than trend driven, these are pieces designed to express individual style and be enjoyed for years to come. View Air &…

## [Ship software peacefully (Sponsored)](https://crawlproof.com/a/r1webHNnzqel)

_2026-08-06 · **Sponsored**_

Connect your repo and deploy instantly — Railway handles config, scaling, and monitoring.

## [How Attention Became Efficient & Scalable: KV Caching, MQA, GQA, MLA, and Sparse Attention.](https://chizkidd.github.io//2026/08/05/attention-efficient-scalable/)

_2026-08-05 · Chizoba Obasi blog_

Attention mechanisms have evolved considerably to make transformer inference faster and more memory-efficient. These notes trace that evolution: from vanilla self-attention , through KV caching , to memory-saving variants like MQA , GQA , and MLA , and finally to sparse attention methods like SWA and Deepseek Sparse Attention (DSA) . Table of Contents Self-Attention Masked (Causal) Self-Attention…

## [The Avenue Gallery &#x2013; Feature Artist &#x2013; Bruce Edmundson](https://www.focusonvictoria.ca/events/event/3602-the-avenue-gallery-%E2%80%93-feature-artist-%E2%80%93-bruce-edmundson/)

_2026-07-30 · Ongoing & Upcoming Events_

Bruce Edmundson began wood carving in the early 1990s while working in the forest industry doing timber development and silviculture work. Interest in high quality wood was a natural development of working in the BC forests. An especial interest was the locally exotic or unusual. The natural result is burls. These abnormal growths found on most trees demonstrate - to the extreme - the…

## [Love to Sing!!!!!! The Victoria Arion Male Choir Invitation to September 2026 Open Houses](https://www.focusonvictoria.ca/events/event/3601-love-to-sing-the-victoria-arion-male-choir-invitation-to-september-2026-open-houses/)

_2026-07-27 · Ongoing & Upcoming Events_

A community of male-voice singers who sing for health and joy: we are the Victoria Arion Male Choir. Our history goes back to 1893 when the Arion Club of Victoria gave its first public concert. Our mission is to maintain the great tradition of male-voice choral music, and to share that music through two or more public concerts each year. The choir also performs at senior residences, and the…

## [The Avenue Gallery &#x2013; Feature Artist &#x2013; Minori Takagi](https://www.focusonvictoria.ca/events/event/3600-the-avenue-gallery-%E2%80%93-feature-artist-%E2%80%93-minori-takagi/)

_2026-07-24 · Ongoing & Upcoming Events_

Born in Shizuoka, Japan, Minori Takagi brings traditional Japanese glasswork to life with a modern, West Coast twist. Since moving to Vancouver in 2006, she has developed a signature style of "wearable art," blending ancient lampworking techniques with contemporary design inspired by the city's natural beauty. Her award-winning jewellery, handcrafted from both soft and borosilicate glass, is known…

## [Reinforcement Learning Part 12: The Policy Gradient](https://shawnhymel.com/3632/reinforcement-learning-part-12-the-policy-gradient/)

_2026-07-24 · ShawnHymel · Shawn Hymel_

In the previous post, we introduced the breakthrough concept of combining deep learning and reinforcement learning (RL). Instead of recording estimated Q-values in a table, which is intractable for large or continuous state spaces, we approximated those Q-values using a neural network. This simple act spawned the current generation of deep RL, paving the way \[…\]

## [Reinforcement Learning Part 11: Deep Q-Networks (DQN)](https://shawnhymel.com/3588/reinforcement-learning-part-11-deep-q-networks-dqn/)

_2026-07-14 · ShawnHymel · Shawn Hymel_

Previously, we looked at how Q-learning used off-policy temporal difference (TD) updates to converge on an optimal policy. This reinforcement learning (RL) algorithm works surprisingly well, but it requires the environment to have relatively small, discrete state and action spaces. The Q-table can quickly grow to intractable sizes with environments that have a large number \[…\]

## [Reinforcement Learning Part 10: Q-Learning](https://shawnhymel.com/3580/reinforcement-learning-part-10-q-learning/)

_2026-07-08 · ShawnHymel · Shawn Hymel_

One of the biggest breakthroughs in reinforcement learning (RL) occurred in 1989 with Chris Watkins’s paper, Learning from Delayed Rewards. In it, he proposed Q-learning, which decouples the experience gathered from the policy update. In other words, the agent can collect experience using one policy (called the behavior policy) while updating a different policy (called \[…\]

## [Reinforcement Learning Part 9: TD(λ) and Eligibility Traces](https://shawnhymel.com/3513/reinforcement-learning-part-9-td%ce%bb-and-eligibility-traces/)

_2026-07-02 · ShawnHymel · Shawn Hymel_

In the previous post, we saw how temporal difference (TD) learning updated value predictions in the middle of an episode rather than waiting to the very end, like we do with Monte Carlo (MC) methods. If you recall, the TD(0) algorithm updates value estimates after every step using a single reward in order to bootstrap \[…\]

## [Reinforcement Learning Part 8: Temporal-Difference (TD) Learning](https://shawnhymel.com/3481/reinforcement-learning-part-8-temporal-difference-td-learning/)

_2026-06-23 · ShawnHymel · Shawn Hymel_

Temporal Difference (TD) learning is one of the foundational concepts in reinforcement learning (RL). It combines the notion of updating estimates before the final outcome is known, similar to how dynamic programming (DP) works, with the notion of learning directly from experience, like we saw with the Monte Carlo (MC) methods in part 7. MC \[…\]

## [Reinforcement Learning Part 4: Expected Return, Value Functions, and Bellman Equations](https://shawnhymel.com/3350/reinforcement-learning-part-4-expected-return-value-functions-and-bellman-equations/)

_2026-05-26 · ShawnHymel · Shawn Hymel_

In the previous post, we defined a policy, provided the foundational concept of a Markov Decision Process (MDP), and talked about trajectories. We’re going to combine these concepts with the idea of future discounted returns to create value functions. We now introduce two new concepts: state-value function (given by V(s)) that attempts to estimate the \[…\]

## [Open-source Screen Sharing with Control (Sponsored)](https://crawlproof.com/a/38fBPv6GaCde)

_2026-05-26 · **Sponsored**_

WebRTC screen sharing with simultaneous mouse & keyboard control; viewers join in browser.

## [Reinforcement Learning Part 3: Policies, Markov Decision Processes (MDPs), and Trajectories](https://shawnhymel.com/3328/reinforcement-learning-part-3-policies-markov-decision-processes-mdps-and-trajectories/)

_2026-05-19 · ShawnHymel · Shawn Hymel_

In the third part of this reinforcement learning (RL) series, we’re going to give a formal definition for a policy and then conceptualize how actions and states play out in a trajectory. While we discussed rewards and returns in the previous post, we’re going to see how Markov Decision Processes (MDPs) provide the underlying foundation \[…\]

## [Reinforcement Learning Part 2: Rewards, Returns, and the Discount Factor](https://shawnhymel.com/3322/reinforcement-learning-part-2-rewards-returns-and-the-discount-factor/)

_2026-05-12 · ShawnHymel · Shawn Hymel_

In this second post on reinforcement learning (RL), we build on the introduction from part 1 by revisiting the idea of a reward and building up to the idea of discounted returns. Recall that the goal of RL is to maximize the rewards earned by the agent over time. We’re going to discuss three main \[…\]

## [Policy Gradient Methods: REINFORCE, Actor-Critic, and the Policy Gradient Theorem (S&B Ch. 13)](https://chizkidd.github.io//2026/05/07/rl-sutton-barto-notes-ch013/)

_2026-05-07 · Chizoba Obasi blog_

Almost all the algorithms/methods covered so far have been action-value methods (except gradient-bandit algorithms, Section 2.8 ). Action-value methods learn the values of actions and then derive the policy thereafter to select actions based on their estimated action values. Here, we explicitly learn a parametrized policy that can select actions without consulting a value function. A value…

## [What is Reinforcement Learning?](https://shawnhymel.com/3316/what-is-reinforcement-learning/)

_2026-05-05 · ShawnHymel · Shawn Hymel_

Reinforcement learning (RL) is a field of study within machine learning (ML) concerned with developing intelligent agents that take actions in dynamic environments in order to maximize their rewards. RL has gained a lot of popularity in the past few years, most notably in robotics, where big-name companies are using it to create robust locomotion \[…\]

## [Transformer Architecture Explained: Self-Attention, Encoders, and Decoders](https://chizkidd.github.io//2026/04/17/transformers/)

_2026-04-17 · Chizoba Obasi blog_

Transformers are a sequence-to-sequence model : given an input sequence, produce an output sequence. Architecture: an Encoder processes the input; a Decoder generates the output autoregressively. \\\[\\text{(En) "I am sorry"} \\xrightarrow{\\text{Encoder}} \\xrightarrow{\\text{Decoder}} \\texttt{\<start\>}\\ \\text{Je suis désolé}\\ \\texttt{\<end\>}\\\] Autoregressive : the decoder generates one token at a time,…

## [SAM 2 Explained: Meta's Promptable Visual Segmentation Model](https://chizkidd.github.io//2026/04/17/sam-2/)

_2026-04-17 · Chizoba Obasi blog_

Meta’s unified model for promptable image and video segmentation. A foundation model for solving promptable visual segmentation in images & videos . Built a data engine to collect the largest video segmentation dataset to date. Model : Simple transformer architecture with streaming memory for real-time video processing. Trained on a wide range of tasks: video segmentation and image segmentation.…

## [Muon Optimizer Explained: Newton-Schulz Orthogonalization Beyond Adam](https://chizkidd.github.io//2026/04/04/muon-muonclip/)

_2026-04-04 · Chizoba Obasi blog_

Muon stands for M oment U m O rthogonalized by N ewton-Schulz and was invented by Keller Jordan . The key idea: Instead of applying Adam-style per-element adaptive updates to model parameters, Muon orthogonalizes the momentum matrix before using it as the update direction. Table of Contents Adam Optimizer Matrix Orthogonalization Newton-Schulz 5 Iteration Muon QK-Clip Multihead Latent Attention…

## [Inkcast: Turn Any EPUB or PDF into an Audiobook in Your Browser](https://chizkidd.github.io//2026/03/16/inkcast/)

_2026-03-16 · Chizoba Obasi blog_

Earlier this year, I decided to force myself to read more. Not a New Year’s resolution, because those never last. The reason is that growing up as a child and young teenager, reading often felt like punishment. My mum required my siblings and me to read a certain number of pages from a designated book every day throughout elementary school. Missing a day meant mandatory punishment. In boarding…

## [Eligibility Traces Explained: TD(λ), Sarsa(λ), and the λ-Return (S&B Ch. 12)](https://chizkidd.github.io//2026/03/13/rl-sutton-barto-notes-ch012/)

_2026-03-13 · Chizoba Obasi blog_

Eligibility traces are one of the basic mechanisms of RL that unify and generalize TD and Monte Carlo (MC) methods. TD methods augmented with eligibility traces produce a family of methods spanning a range from MC methods at one end ($\\lambda = 1$) to one-step TD (TD(0)) methods at the other end ($\\lambda = 0$). With eligibility traces, MC methods can be implemented online and on continuing…

## [The Deadly Triad in RL: Off-Policy Learning with Function Approximation (S&B Ch. 11)](https://chizkidd.github.io//2026/03/09/rl-sutton-barto-notes-ch011/)

_2026-03-09 · Chizoba Obasi blog_

Let’s discuss the extension of off-policy methods from the tabular case (Ch. 6 & 7) to function approximation. We’ll explore the convergence problems, the theory of linear function approximation, the notion of learnability, and stronger convergence off-policy algorithms. Off-policy learning with function approximation has 2 challenges: Finding the target of the update. The off-policy distribution…

## [Web Coding With AI, Live (Sponsored)](https://crawlproof.com/a/jdmRuWwojDND)

_2026-03-08 · **Sponsored**_

Collaborative screen sharing with simultaneous remote control, open source.

## [Semi-Gradient Sarsa and the Average Reward Setting in RL (S&B Ch. 10)](https://chizkidd.github.io//2026/03/09/rl-sutton-barto-notes-ch010/)

_2026-03-09 · Chizoba Obasi blog_

Let’s dive into the control problem now with parametric approximation of the action-value function $\\hat{q}(s, a, \\mathbf{w}) \\approx q\_{\*}(s, a)$, where $\\mathbf{w} \\in \\mathbb{R}^d$ is a finite-dimensional weight vector. We’ll focus on semi-gradient Sarsa , the natural extension of semi-gradient TD(0) to action values and to on-policy control. We’ll look at this extension in both the episodic…

## [A Reflection: The Will to Change by bell hooks](https://matthewscheffel.com/posts/the-will-to-change/)

_2026-02-26 · Matthew Scheffel_

I have heard of patriarchy since I was a child, but bell hooks&rsquo; book mapped that word onto my own experience. It has been an invisible force hollowing me out. Its agents have been family, friends, and media - all promoting its myths and enforcing its standards. It begins with shame and humour - shame for the uninitated, humour for the initiated. It hollows out boys into good little soliders…

## [Changelog](https://matthewscheffel.com/changelog/)

_2026-01-28 · Matthew Scheffel_

&#xA; &#xA; td;dr:&#xA; &#xA; &#xA; This page is auto-generated using my git commit log &#xA; &#xA; &#xA; &#xA; &#xA; 2022-Q4 &#xA; &#xA; Footnotes heading &#xA; Pizza dough - more footnotes &#xA; Pizza dough &#xA; Content edit, add Yew article &#xA; draft game idea text &#xA; Add privacy policy for NeuralNote &#xA; Colour tweak &#xA; Regenerate content following edits &#xA; Content edit &#xA;…

## [Square Root N Sampling](https://matthewscheffel.com/posts/square-root-n-sampling/)

_2025-04-12 · Matthew Scheffel_

&#xA; &#xA; td;dr:&#xA; &#xA; &#xA; The \\(\\sqrt{N}\\) sampling technique is invalid, absolute sample size is what matters&#xA; &#xA; &#xA; &#xA; My introduction to the technique &#xA; In the building automation world installation jobs require technicians to validate that installed devices are fit for purpose, i.e.: &ldquo;commissioning&rdquo;. It&rsquo;s a slow, and therefore expensive process. At…

## [building slow: great bones, terrible agony.](https://matthewscheffel.com/software-development/building-slow/)

_2025-01-27 · Matthew Scheffel_

I&rsquo;ve been working on the project for years now. It is the most ambitious thing I have ever made. I&rsquo;m proud of it, but I&rsquo;m far from done and it has been grueling. Each time I overcome a mountain of difficulty I am rewarded with the sight of the next mountain I must climb. It feels as though I am always on the cusp of &ldquo;rolling downhill&rdquo;, if only I complete just another…

