HAMi-core and KAI Scheduler: GPU Sharing Moves from Allocatable to Governable
GPU sharing is moving from soft allocation to governance—scheduling decisions and runtime isolation can finally be reconciled.
Recent content in Blog on Jimmy Song
GPU sharing is moving from soft allocation to governance—scheduling decisions and runtime isolation can finally be reconciled.
Code is cheap; consensus is the new scarce good.
HAMi moves from cluster to desktop with Olares.
From Anji bamboo weaving to industrialization
A GPU explainer for Kubernetes veterans new to AI. Maps token, model, training, inference, Transformer, Tensor Core, HBM, and KV cache to concepts you already know.
From GPU utilization to productive GPU-hours.
A practical AI Infra review of Agentic AI reliability, covering a five-dimension framework, fault tolerance, recovery, observability, and hybrid architecture design.
From GPU hardware, Kubernetes scheduling, inference engines to token cost — understanding the 8-layer observability architecture for modern AI infrastructure.
How I built a personal AI infrastructure using ChatGPT, OpenClaw, Obsidian, GitHub, Lark, GLM-5.1, and a Mac mini M4.
The Linux Foundation's Tokenomics Foundation signals a shift: tokens are becoming a core resource in the AI era, much like CPUs in the cloud era.
AI Native Landscape has moved to landscape.jimmysong.io with 600+ curated open-source projects, AI skill search support, and a call for community contributions.
Observations on the evolution of AI infrastructure control planes, focusing on HAMi v2.9, GPU scheduling, and Kubernetes resource models.
At KubeCon EU 2026, I witnessed Kubernetes' anxiety and transformation in the AI era. This article explores the challenges and future opportunities for Kubernetes in the age of AI.
KubeCon Europe 2026 Day One: How Kubernetes is adapting to the AI infrastructure wave and the evolution of the GPU resource layer.
A systematic upgrade to HAMi’s website and docs, improving community visibility, content structure, search, and usability.
On the eve of GTC 2026, rethinking whether AI is becoming the new infrastructure from NVIDIA's AI Five-Layer Cake, the rise of agent runtime, to AI-native infrastructure.
A CTO/VP view on open GPU scheduling: CDI, Kubernetes DRA, virtualization data planes, ecosystem governance, and lock-in risk.
A curated collection of AI learning resources we removed from the AI Resources list: awesome lists, courses, tutorials, and cookbooks. These educational materials deserve their own spotlight.
Before ChatGPT and TensorFlow, there was Hadoop, Kafka, and Kubernetes. This post honors the traditional open source infrastructure that became the foundation of today's AI revolution.
Observations from my first month at Dynamia: From cloud native to AI Native Infra, why this direction is worth investing in, and the key issues and opportunities in compute governance.
Exploring how Spec becomes the governable core asset in Agent-Driven Development (ADD) and the trend toward control-plane engineering systems.
Comparing Miaoyan, Zhipu, and Shandianshuo voice input methods for developers: speed, stability, command capabilities, and cost models.
How technical standards and data sovereignty shape AI open source paths and infrastructure competition in the global AI era.
Joining Dynamia as Open Source Ecosystem VP to drive AI-native infrastructure ecosystem development, transforming compute from hardware consumption to core asset.
A hands-on experience with Verdent's standalone Mac app, exploring how parallel AI agents, isolated workspaces, and task-oriented workflows change real-world development.
A look back at the major changes in 2025: shifting from Cloud Native to AI Native Infrastructure, AI tool ecosystem, and major website improvements.
Manus's acquisition by Meta sparked polarized opinions. This article explores the butterfly effect in AI applications and key lessons for entrepreneurs on growth strategies.
Beijing and Shanghai's open source plans reveal opportunities and challenges for China's AI infrastructure, balancing technology and governance.
In 2025, software engineering shifts from code-centric to runtime and cost governance. AI and Agents move complexity to runtime, compute, and budget layers, reshaping engineering value.
Explores why AI Agents need Kubernetes infrastructure and how Agent orchestration, MCP services, and AI gateways enable production-ready AI architectures.
Comprehensive introduction to the AI Open Source Landscape's positioning, interface, scoring model, and data mechanisms to help developers efficiently discover quality AI projects.
2026 AI's turning point: not models, but infrastructure, agentic runtimes, GPU efficiency, and new organizational forms.
From an engineering and organizer's perspective, real changes at COSCon'25: AI as the default backdrop, discussions returning to engineering issues, and Chinese open source entering a long-term phase.
An analysis of Block's Goose project, why it became one of the first Agentic AI Foundation (AAIF) projects, and what this means for Agentic Runtime and the evolution of AI-Native infrastructure.
How ARK uses cloud-native architecture and declarative runtime to drive engineering adoption of multi-agent systems and shape the Agentic Runtime ecosystem.
Lunary, an open-source project in the AI DevTool space, suddenly deleted its GitHub repo, exposing the instability of commercial open source projects.
An analysis of the background, strategic urgency, differences and division of labor between Agentic AI Foundation (AAIF) and CNCF/CNAI, and its significance for the AI Native era.
Bun's acquisition by Anthropic marks the first time a general-purpose language runtime is integrated into a large model engineering system, revealing a structural trend for AI-native runtimes.
Analyzing Ark from architecture, semantics, community activity, and engineering paradigms to reveal its impact on 2026 AI Infra trends and the ArkSphere community.
Analysis of McKinsey's Ark project: architecture, CRDs, control plane, design paradigms, production readiness, and implications for ArkSphere and AI infrastructure.
AI's real turning point is moving from using AI tools to building AI systems. Why the era of AI engineering hasn't begun, and the developer opportunity in the next three years.
A practical Antigravity setup guide for developers who want a VS Code-style AI IDE, including marketplace switch, AMP and CodeX installation, and workflow tuning.
An analysis of the Cloudflare global outage on November 18, 2025, exploring implicit assumptions, automated configuration pipelines, and systemic risks in modern infrastructure.
A decade of cloud native evolution, a look ahead to AI-Native Platform engineering, technical layers, and key changes. KubeCon NA 2025 signals a new era.
Based on months of deep usage, this article analyzes how NotebookLM helps me learn new technologies, read complex documents, generate teaching outlines, and shares future improvement expectations.
An analysis of Helm 4's core changes, including Server-Side Apply, WASM plugin system, kstatus status model, reproducible builds, and content hash caching, with a timeline review of Helm's history.
Kimi K2 Thinking's open source marks China's entry into thinking models. This article reviews its technical approach and compares it with Claude and Gemini.
Exploring how Gateway API Inference Extension brings model-aware inference traffic control through InferencePool, InferenceObjective, and metrics-driven routing.
A comparison of TRAE SOLO and VS Code (Copilot, Agent HQ) via the AI Engineering Entity framework, focusing on automation, collaboration, model transparency, and engineering roles.
Analysis of closed-source model acceleration and open-source ecosystem response, exploring core engineering contradictions and infrastructure evolution.