Platform engineering has matured from a buzzword into a recognised discipline over the last three years. Here is an honest assessment of where the practice actually is in 2026 — what's working, the gaps that practitioners are glossing over, and the three shifts that will define platform engineering over the next two years.
Agentic AI workloads have a threat surface that standard Kubernetes security models weren't designed for: wide egress requirements, dynamic tool invocation, and the potential for autonomous actions that cross trust boundaries. Here is the security architecture — network policies, RBAC, resource quotas, and audit logging — that actually contains them.
After several years of running a 70-person infrastructure engineering team across India, the UK, the US, and Singapore, here is the operating model that works: hiring principles, team topology, decision protocols, on-call design, and the two things that make or break distributed engineering teams at scale.
The IaC tool debate has moved on from religious arguments to practical questions: which tool for which problem, at which team maturity level, with which operational constraints? Here is the honest comparison — strengths, real failure modes, and when each tool is the wrong choice for enterprise platform teams.
The patch-to-exploit window is now hours. A 30-day patch cycle is no longer defensible. Here is the complete pipeline — CVE ingestion, Ansible Lightspeed playbook generation, staging validation, ring-based production deploy — that gets critical patches out in under 4 hours without breaking production.
The OTel setup in tutorials is not the setup that works at high transaction volumes in regulated environments. Here is the collector architecture, sampling strategy, cardinality budget, and compliance constraints that shape a production deployment for financial workloads.
AI agents are genuinely useful in incident response — for runbook execution, log triage, and communication drafting. They are not useful as autonomous decision-makers. Here is the practical setup that works: where to put the AI, where to keep the human, and the specific failure modes to design around.
The gap between how engineers see cloud costs and how finance sees them is the source of most FinOps failures. Here is the reporting structure, the metrics, and the framing that closes that gap — built from the experience of explaining infrastructure spend to finance teams who have zero tolerance for technical abstraction.
GPU infrastructure has a completely different operational profile from CPU-based compute. Before your first AI workload hits production, here is what every platform engineer needs to understand about GPU scheduling, memory architecture, multi-tenancy, and cost management.
70 engineers across four countries. The failure modes nobody warns you about and three practices that actually work at distance: writing culture as infrastructure, explicit decision protocols, and deliberate relationship investment.