0:00
-0:33
Over the past few weeks, I have been doing some soul-searching regarding telecom network automation. If you have read my recent LinkedIn posts, you know I am constantly looking for the line between actual operational reality and the endless stream of industry marketing.
The market is flooded with promises of “zero-touch operations” and fully autonomous networks. Yet, anyone who has worked with real-world telecom infrastructure knows that dropping an AI model on top of a fragmented legacy stack does not suddenly make the network autonomous.
Recently, my former Google colleague Gabriele Di Piazza, who leads Product Management, Alliances and Architectures at Blue Planet, shared some deep technical blueprints and operational data from their field deployments. Reviewing those materials gave me a chance to step back, compile my key learnings, and synthesize what actually works versus where the industry is fooling itself.
When you go into the rabbit hole of automating a multi-vendor, multi-domain network, it is clear to me that most operators are doing network automation wrong.
Sure, they are not failing because of a lack of effort, investment, or technical skill. They are struggling because they are automating in isolated silos. Essentially, they are applying next-generation intelligence to fragmented, legacy workflows.
If you are evaluating or building an automation roadmap today, here are the 10 basic non-negotiables that determine whether you build a genuinely autonomous network or just a faster set of problems.
If you want to move from isolated scripts to a functioning autonomous network, these are the ten structural requirements you cannot skip.
Automating isolated network domains, whether radio access, optical, or IP routing, creates intelligent silos. If every domain executes local optimizations without coordination, cross-domain policies can conflict.
TM Forum architectures emphasize layer decoupling, where resource domains execute local actions while a cross-domain intelligence layer coordinates overall operational outcomes. Without that coordination to reconcile business intent across access, transport, and core, domain-level automation simply accelerates policy clashes.
Data fragmentation is one of the biggest bottlenecks in telecom operations. You cannot deploy AI agents if your inventory is static, out of sync, or fragmented across legacy databases. AI needs a continuously reconciled view of what actually exists and its current state. Field deployments report up to a 95% reduction in capacity reporting time through automated, federated inventory.
Using model-driven discovery and standardized catalog and inventory interfaces such as TMF634, TMF638, and TMF639, operators can map physical, virtual, and cloud assets so AI agents act on live network state rather than outdated records.
Imperative scripting, with rigid playbooks defining step-by-step commands, becomes brittle as firmware, configurations, and topologies change. Autonomous operations require declarative, intent-based automation.
The operator defines the target outcome, while orchestration translates that intent into resource actions across access, core, and transport and continuously reconciles the network toward the desired state. Field deployments report reductions of up to 80% in service activation times.
Off-the-shelf Large Language Models parse text well, but they lack native understanding of network topology, service dependencies, and radio context. To give AI real operational reasoning, operators need an OSS knowledge graph that represents how network resources and services relate.
This graph maps both “North-South” service-to-resource relationships and “East-West” multi-vendor network adjacencies. Advanced implementations can use Graph Neural Networks to analyze IP topology, link utilization, and traffic-engineering constraints such as Segment Routing and IGP metrics. This topological context allows agents to perform more accurate root-cause analysis rather than guessing.
Adding a conversational AI widget to a legacy OSS dashboard is not agentic automation. Operationalizing AI agents across a complex, multi-vendor environment requires a dedicated three-layer architecture. At the foundation, agentic tooling exposes OSS knowledge, telemetry, and network capabilities through governed APIs and open interfaces such as the Model Context Protocol.
Above it, an agentic core manages the agent lifecycle, model access through secure LLM gateways, orchestration, policy, and Agent-to-Agent communication. Finally, agentic channels embed that intelligence into inventory, orchestration, assurance, and other operational workflows, making reasoning part of how the network is run rather than another isolated portal.
The skepticism surrounding AI-driven network control is justified by real-world outage data. Industry reports indicate that 80% of serious outages are preventable, stemming directly from process and configuration errors, while faulty software changes have been linked to hundreds of millions of lost user-hours.
Telcos cannot scale automation without Configuration and Change Management. This governance layer enforces real-time drift detection, validates live configurations against compliance baselines, and maintains a strict audit trail of every change, whether executed by a human engineer, an imperative script, or an AI agent.
In modern software development, code is tested before it hits production; high-impact network changes need the same guardrail. Before an automated system changes routing, capacity, or critical configuration, it should validate the action against a network digital twin.
Using route optimization and analysis tools, operators can run predictive “what-if” simulations. If an AI agent proposes changing routing costs or shifting traffic off an optical path, the twin first simulates the impact to verify that latency, packet loss, and enterprise SLAs remain within acceptable limits.
Traditional operations wait for an alarm before an engineer opens a ticket and investigates. In dynamic multi-domain networks, that is too slow. Assurance must evolve into a proactive, closed-loop engine that uses real-time telemetry to detect anomalies, predict degradation, and understand service impact before customers are affected.
When a risk is identified, assurance can trigger orchestration workflows to reroute traffic or adjust resources. In 5G slicing deployments, field data show a 75% reduction in root-cause identification time and an 85% reduction in mean time to resolve slice disruptions.
This is stronger because the number is now specific and defensible, rather than making 85% sound universal.
Viewing network automation merely as a back-office cost-cutting exercise misses its commercial value: service velocity and Network-as-a-Service monetization. Lumen now reports more than 3,000 NaaS customers, with new customer adoption growing 22% quarter-on-quarter as enterprises shift toward programmable connectivity for cloud and AI workloads.
Automated, intent-based orchestration enables this commercial model by compressing service delivery from days or months to hours. In one European Tier 1 5G slicing deployment, AI-driven order management reduced activation from months to hours while delivering 80% cost savings, transforming infrastructure into a programmable commercial asset.
“Level 4 Autonomous Networks” cannot be purchased off a vendor price sheet. Under TM Forum’s Autonomous Network Levels, Level 4 marks the shift from human-defined automation toward intent-driven, predictive, closed-loop decision-making. Operators are increasingly reaching this level in specific domains and operational scenarios rather than across the entire network at once.
Operators must progress systematically: starting with data unification, moving to domain automation and closed-loop assurance, and ultimately extending autonomous reasoning across the service lifecycle. Skipping the foundational data, orchestration, and governance layers simply automates the problems underneath.
Autonomous networking is not a fixed destination or a product you buy off a shelf. And adding more AI does not make a fragmented network more autonomous.
If your network strategy skips the hard work of unifying data, governing change, and moving from scripts to intent-based orchestration, you are not modernizing the network. You are digitizing your legacy debt. Put AI on top of a fragmented OSS architecture, and you simply automate chaos at a higher speed.
The operators that win the next decade will not be those that bought the most AI. They will be the ones who built the operational foundations to make AI useful: live inventory, cross-domain orchestration, governed change, digital twins, and closed-loop assurance working as one coherent system. If your automation roadmap starts with an LLM, you are probably starting in the wrong place.
Where are you seeing operators get stuck? I would be interested to hear which of these ten foundations is proving hardest to get right.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.