RSS Amplifier

Podcast

Shamsher's AI PM Brief

AI PM Briefs are inspired by my LinkedIn posts. Since it's hard to find posts from three months ago on LinkedIn, I decided to put them here on Substack. This will help me and others quickly learn what AI Product Managers need to know.

aipmbriefs.substack.comRSS feed ↗20 episodes

Written by

Latest episodes

Understanding Voice Agent Sessions: The Building Block of Conversational AI

Voice AI is becoming a core part of customer support, virtual assistants, and enterprise automation.

Why Every AI Agent Needs a Secure Sandbox

LLMs have made AI agents incredibly capable.

The Hardest Part of Product Management Is Not Building the Product

Many people think a Product Manager’s primary responsibility is building products.

How GPUs Talk to Each Other (And Why It Matters More Than You Think)

When people think about AI performance, they usually think about which GPU they are using.

How to Improve GPU Utilization?

GPU utilization is all about one thing:

From Engineer to AI Product Manager: The Art of Unlearning

One of my mentors gave me an important piece of advice when I was moving from an engineering role to product management:

Why Old LLM Systems Waste GPU Memory and How vLLM Fixes It

When people talk about running Large Language Models (LLMs), they usually focus on GPUs.

Why Is GPU Compute So Expensive Per Hour, Even for AI/ML Engineers?

If you’ve ever looked at GPU cloud pricing and thought

If AI Made Building Easy, Does Distribution Become the Hard Part?

I attended #ProductCon (Online) yesterday, and one slide from Elena Verna really stuck with me.

LLM Inference Explained: What Product Builders Need to Know

Over the last two years, I’ve spent a lot of time learning and experimenting with LLMs.

What is The Secret to Faster LLM Inference

Whenever I use an AI chat tool, I always wonder: how does it generate text so quickly?

Why 70% of GPU Spend Is Wasted, and How to Fix It

GPU Costs Are Skyrocketing, and Most of It Is Wasted

Why AI Agents Struggle Today, and How Memory Can Fix It

Without memory, your AI agent is just a stranger.

Token-Based Pricing in AI: What You’re Really Paying For

Agentic AI uses 10–100x more Tokens. Rough rule: 1,000 tokens ≈ 750 English words

Why LLM Inference Costs More Than Training

What Can You Do About It? Read if you want to save Inference Cost

Why Apple Could Dominate the On-Device AI

The on-device AI market (including Apple) is expected to hit $36.6 billion by 2030

How to Run vLLM on Apple M4 Mac Mini

vLLM is a powerful inference engine for LLMs. It’s designed to serve models faster and more efficiently.

How I Ran RamaLama on My Raspberry Pi

LLMs on CPU Made Easy, No GPU Needed with RamaLama.

Why is RamaLama the safest way to run a Local LLM?

No Dockerfile. No dependencies mess. Just run your models. Locally. Securely.

Why the Apple M4 Mac Mini Is a Perfect Machine for Running Local LLMs

If you're exploring running AI models locally, especially Large Language Models (LLMs) with a few billion parameters, the Apple M4 Mac mini might be one of the best desktop machines to get started with.