Cisco Data Fabric: Solving for Visibility, Economics, and AI Readiness

Platform Sancha Norris

Key takeaways

  1. Cisco Data Fabric powered by Splunk unifies access to distributed data, improving visibility across security, IT, and network environments without requiring data movement or ETL.
  2. The Machine Data Lake lowers data costs by retaining full-fidelity data in open formats, while higher-value data can be promoted to faster tiers when needed.
  3. Context, governance, and AI-ready data enable trusted agentic operations, with agents using organizational knowledge and controls to investigate, reason, and act.

The organizations with the most operational data often have the least ability to act on it. The data sits in separate systems and multiple tools with incompatible schemas and query languages, so reasoning across it is the bottleneck, not collecting it. Oxford Economics puts the cost of that fragmentation at roughly $900,000 per hour of downtime or about $600 billion a year.

econ-use.png

Three key problems define the crisis: visibility, economics, and AI readiness. More tooling will not solve them but a new architecture will. Cisco Data Fabric powered by the Splunk Platform is a redesigned modern architecture that addresses all three. It brings query semantics to data wherever it lives, retains everything in open formats at commodity cost, and adds the context that humans and agents need to reason.

Visibility With Reach and Discoverability

Organizations cannot correlate across security, IT, and network domains for two reasons. They cannot reach data trapped in incompatible systems, and analysts cannot act on data they do not know exists.

Federated Search solves reach. It brings Splunk query semantics to data wherever it lives across Amazon S3, Azure Blob, and Snowflake, with no movement and no ETL, queried through SPL2 in either SPL or SQL. The Machine Data Lake covers everything else, retaining all ingested data in open formats so it stays searchable regardless of age or tier. Data once discarded for cost is preserved and queryable on demand.

Discoverability is the other half. An analyst or an agent cannot search for what they do not know exists, so the Context Layer catalogs every source, maps how entities relate across domains in a knowledge graph, and indexes operational history for semantic search. Reach plus discoverability is what makes visibility real.

Economics That Aligns With Known Value

Data grows 30 to 50 percent a year, faster than budgets, which forces a binary choice: index everything at premium cost, or delete it and accept blind spots.

The Machine Data Lake removes the need for a binary choice. Raw data lands in open formats at commodity object storage cost as an immutable source of truth. Only data that proves valuable is promoted to faster, indexed tiers, and that decision waits until the value is known instead of being guessed at ingest. Organizations pay for performance, not for retention, and retaining a terabyte of raw data costs a fraction of indexing it. Because schema-on-read applies structure at query time, reprocessing history becomes a query rather than a re-ingestion project.

AI Readiness That Holds Up in Production

Agents reasoning over partial history produce partial conclusions. An organization that has not solved its underlying data architecture cannot use AI to solve it.

The same architecture that fixes visibility and economics produces AI-ready data. That data has to meet five conditions: complete, contextual, accessible, trustworthy, and high quality. The Machine Data Lake supplies completeness by keeping full-fidelity history rather than only what was affordable to index. The Context Layer supplies context, turning raw events into meaning an agent can reason over. Open formats and the Model Context Protocol keep the data accessible to any model or agent. AI-Powered Data Management keeps it clean as sources change, using automated field extraction and self-healing pipelines that repair schema drift before it breaks downstream consumers. Lineage back to the immutable raw tier makes every output traceable, which is what trust depends on.

econ-2.png

Building agents on this data foundation does not require a data-science team. Agent Launchpad lets security and IT teams create custom agents by stating the objective in natural language and invoking them from Splunk searches and alerts with no coding required. Under central role-based access control, you can build the agent harness with your own data, expertise, and processes embedded in agent skills and MCP tools.

The custom agents can run on hosted foundation models purpose-trained for specific tasks, the Foundation AI Security Model for security work and the Cisco Deep Time Series Model for forecasting, which are more accurate on operational telemetry than general models. For broader reasoning, the platform adds open-weight frontier models like GPT-OSS and integrates third-party frontier models from OpenAI, Anthropic, and Google. The result is faster agentic operations, built by the people who run the environment.

The Splunk MCP server allows users to build agents outside of Splunk for more complicated workflows and agentic ecosystems.

Here’s How It Works

A typical investigation runs as a loop. A model flags an anomaly in network traffic. The Cisco Deep Time Series Model tests whether it is a genuine deviation or normal seasonal variation. A foundation model weighs it against recent security events and known threat patterns. The agent queries the knowledge graph for every affected service, host, and user, then runs federated searches across Splunk indexes, the Machine Data Lake, and external stores to gather corroborating evidence. It closes by synthesizing a resolution or escalating to an analyst with the full picture. Each step is only as reliable as the data feeding the one before it, which is why the foundation matters more than any single model.

econ-3.png

Governance That Extends to the Agents

Autonomous agents raise the stakes. When agents act on data without a human in the loop, ungoverned access becomes a direct path to data exposure and unauditable decisions. Governance therefore extends to the agents themselves, not only the people. Access control operates at the field level, so reaching a dataset does not expose every field within it. Agents act inside the same boundaries as the people they work for, high-impact actions pass through approval gates, and every access, inference, and action is recorded.

Trust also depends on knowing where data came from. The Context Layer captures complete lineage for every record, where it originated, how it was transformed, and who has accessed it, so any conclusion an agent reaches can be traced to its source. Sensitive data is masked or tokenized at ingest before it enters the platform, and because federation queries data in place, regulated data can stay in its home region instead of being copied to a central store.

The Foundation Only You Can Own

The operational environment is shifting from human speed to machine speed. What separates organizations will not be how much data they produce. It will be how much of it they can reason over at the moment it matters, without pre-classifying, pre-moving, or pre-structuring it first. Cisco Data Fabric powered by the Splunk Platform turns the data estate from an operational liability into an AI-ready foundation that only the organization itself can own.

Since our announcement last year, we have steadily delivered all the major capabilities of the Cisco Data Fabric powered by the Splunk Platform. Read the announcement blog.

For more details, read the Cisco Data Fabric whitepaper.

Related Articles

From Registry With Love: Malware Registry Abuses
Security
13 Minute Read

From Registry With Love: Malware Registry Abuses

The Splunk Threat Research Team explores the common Windows Registry abuses leveraged by current and relevant malware families in the wild and how to detect them.
SNARE: The Hunters Guide to Documentation
Security
6 Minute Read

SNARE: The Hunters Guide to Documentation

Discover the SNARE framework for effective threat hunting documentation.
Building Large-Scale User Behavior Analytics: Data Validation and Model Monitoring
Security
6 Minute Read

Building Large-Scale User Behavior Analytics: Data Validation and Model Monitoring

Splunk's Cui Lin explores fundamental techniques to validate data volume and monitor models to understand the size of your own UBA clusters.