Graph-centric agentic intelligence

Graphs encode relationships. In a world where AI agents need to reason about complex systems, not just retrieve information, graph structure provides the scaffolding for causal inference, dependency tracking, and compositional reasoning. Unlike tabular or unstructured data, graphs preserve the topology of real-world systems: what connects to what and how influence flows. For networks, this is not a metaphor. Networks are natively graphs. Every device, link, open communication channel, and service dependency is a vertex or edge in a structure that can span millions of elements. The question facing operational teams today is not whether to represent networks as graphs but how to let AI agents reason over that structure autonomously, adaptively, and at speed. This post traces how graph intelligence evolved across scientific research on networks, from passive modeling to active agentic reasoning, and demonstrates agentic reasoning through cascaded graph analytics for root cause identification. For the solution architecture and a guide to common routines (a runbook), see our companion post, “Beyond correlation: Finding root causes using a network digital twin graph and agentic AI”. Graphs for networks The earliest use of graphs for networks was to represent topology — to model physical connectivity so that operators could compute paths and isolate failures in seconds rather than hours. Knowledge graphs and ontologies enhanced topological models by adding semantics, defining, for instance, what a "cell", an "SLA breach", or an "escalation procedure" is and how these concepts relate to each other across vendors and network generations. This enabled machine-readable operations and multivendor interoperability. Once the graph carried both structure and semantics, the next step was encoding causation. Alarm correlation graphs, which map relationships among network-generated alerts about phenomena like lost connections or network breaches, introduced hierarchical relationships between network events. Walking upstream through the graph condensed tens of thousands of raw alarms into a single causal chain within minutes, a significant advance over manual correlation but one that required predefined alarm hierarchies. Dependency graphs removed that requirement: autogenerated in real time from software-defined-networking (SDN) and network functions virtualization (NFV) controllers, they enabled Bayesian fault localization at 95% accuracy in under 30 seconds with no manually authored rules. Causal subgraphs answered not just where but why. Live alarms converted into directed acyclic graphs revealed the propagation paths of key performance indicators (KPIs), giving operators evidence rather than predictions. Graph neural networks (GNNs), which produce vector representations of graph nodes for use in downstream applications, learned what topology and causal structure alone could not show: hidden intercell dependencies, data patterns that would have emerged had failed nodes not stopped reporting, and spatiotemporal dynamics. These were sequential scientific advances, each addressing a limitation of its predecessor. Today they converge. A single network can be simultaneously represented as a topology graph, enriched with ontology semantics, annotated with temporal KPIs, and reasoned over by GNNs, all coordinated by an agentic layer that selects the right graph tool for each failure pattern. The graph went from a passive data model to an active agentic reasoning substrate. Root cause analysis Root cause analysis is the natural proving ground for graph intelligence. When a network fails, finding the root cause can take hours. For complex multilayer failures, remediation in traditional network operations centers (NOCs) averages four to five hours and can extend to days. The bottleneck is not engineering expertise; it is cognitive overload. Human operators cannot correlate hundreds of alarms, configuration files, and telemetry data across thousands of nodes faster than customer impact accumulates. Traditional approaches rely on temporal correlation: if alarm A precedes alarm B, the system infers that A caused B. This heuristic fails in complex topologies where failures propagate through multiple parallel paths, polling intervals (i.e., intervals between regular system state queries) make timing unclear, and the true root cause may generate no alarm at all. To address the problem of root cause analysis in complex networks, we designed and developed an approach combining cascaded graph algorithms with agentic-AI execution. The graph algorithms provide mathematical precision in the analysis of network topology; the agentic layer provides adaptive intelligence for applying those algorithms correctly across diverse failure scenarios. We demonstrated our approach with NTT DOCOMO at the Mobile World Conference (MWC) earlier this year, achieving root cause analysis in minutes on commercial networks. Three pillars underpin our design: graph modeling, graph-centric analytics, and agentic orchestration. 1. Graph modeling: The digital twin We represent the network as a continuously synchronized graph whose vertices are devices with attributes and whose edges represent connections. This “digital twin” of a physical or software-defined network ingests network dependencies, live alarms, and KPIs from multiple data sources across network segments and layers, transforming them into a topology-aware data structure that all analytics operate on. 2. Graph-centric analytics: The cascaded pipeline Against this graph, the system orchestrates a three-stage cascade of graph algorithms, each stage narrowing the search space for the next: Stage 1, decomposition Based on the number and strength of connections, we identify the most-connected parts of the topology. When failures sever links, the network may fragment into disconnected subgraphs, which allows us to immediately localize analysis to boundary nodes and prevent wasted computation on unaffected regions. This narrows the candidate set from thousands of nodes to hundreds. Stage 2, clustering Within the affected components, community detection algorithms (Louvain or label propagation) group nodes that frequently interact or share dependencies. This process reveals functional groupings and distinguishes failures affecting a single cluster from those propagating across cluster boundaries. The number of candidates narrows from hundreds to tens. Stage 3, centrality ranking Within the identified clusters, a suite of centrality algorithms ranks candidates according to the likelihood that they are the root cause. Standard centrality algorithms measure structural importance: PageRank determines which node is most central within the subgraph; degree centrality determines which node has the most connections; and closeness determines which node is nearest to everything. These answers are static; they don't change when a failure occurs. For root cause analysis, the question is fundamentally different: which node is most important relative to this particular failure? The alarming nodes define the reference frame. Every centrality measure must be recomputed, not against the full graph, but against the failure set. We apply this principle, conditioning on the alarm set, across three centrality algorithms in the suite. Personalized PageRank works by performing a random walk along network paths, evaluating the number and importance of each node’s connections. Our version of the algorithm seeds the random walk from nodes that are issuing alarms (“alarming nodes”). The ranking component of the algorithm concentrates on common ancestors of alarming nodes, tracing faults upward in hierarchical topologies. Our algorithm for computing degree centrality counts only edges to alarming nodes: a gateway with 50 total connections but zero to alarming nodes scores zero; a switch with five alarm connections scores five. The algorithm thus identifies the hubs of star topologies, or subgraphs with one central node. Alarm-relative closeness measures the average distance to alarming nodes only; the geometric center becomes irrelevant, and the node closest to the failure cluster ranks first. This metric surfaces peripheral failures in mesh topologies, or decentralized subgraphs where every node connects to every other node. Topology-aware selection determines which variant to emphasize for each incident. The agentic layer classifies each affected subgraph (hierarchical, star, or mesh) based on graph metrics including degree distribution, diameter, and hierarchical depth, then selects the algorithm combination accordingly. Real networks rarely conform to a single topology type: a subgraph may be hierarchical in one region and star-shaped in another, and the density of connections, the number of alarming nodes, and the presence of redundant paths all influence which combination produces the most accurate ranking. 3. Graph-based agentic orchestration AI agents select and compose graph algorithms adaptively based on failure characteristics. The first step is always to query the incident knowledge base: if the incoming alarm pattern matches a stored incident with high confidence, agents apply the prescribed remedy directly. When no match exists, agents invoke the full graph-driven cascade. Before invoking the cascade, the agent performs a complexity triage. This involves three properties: the number of affected nodes, resolved against the network digital twin; the topological spread of those nodes across connected components; and the semantic clarity of the fault signature — whether the alarm types, severities, and temporal ordering unambiguously implicate a single device or link. Once the cascade produces a ranked candidate list, agents build a failure subgraph by expanding from top candidates to dependency neighbors up to two or three hops away. For each candidate, the agent consults the alarm timeline, incident knowledge base, and runbooks. The resulting root cause determination carries a confidence score weighting temporal evidence, topology-pattern match, centrality scores, and correlated alarm count. It either opens a new trouble ticket or adds to an existing one, with NOC feedback for continuous learning.The pipeline also includes an AI assistant available on demand, allowing network engineers to query the digital twin interactively, explore alternative hypotheses, or request deeper analysis at any point during or after an automated investigation. Path forward There are many natural ways to extend this approach. From a graph intelligence perspective, graph neural networks and spatiotemporal deep learning can augment the deterministic cascade, fusing learned patterns with algorithmic precision. From the agentic-pipeline perspective, two research programs suggest themselves: Graduated autonomy: Autonomy should be earned rather than granted by default during dynamic composition. An agent can progress from advisory recommendations to supervised execution, bounded action within approved guardrails, and eventually end-to-end remediation for well-understood, reversible, low-blast-radius incidents. Promotion should depend on predefined evaluation criteria and shadow-mode results; performance drift should reduce the permitted autonomy. Authenticated agent identity, least-privilege access, machine-readable action policies, audit records, rollback, and observability keep that authority measurable and reversible. Self-learning agents: Agents learn from repeated interactions across sessions. When a resolution pattern proves repeatable, the agent identifies the underlying procedure and proposes it as a candidate skill. The user retains full governance: no skill becomes active until explicitly approved. Over time, the agent builds a library of validated, human-sanctioned skills drawn directly from operational experience. Acknowledgments: Dheeraj Oruganty, Pooja Chikkala, Yuki Miyazaki, Dai Kurosawa, Kazuma Iwamoto, Vijay Veggalam

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论