Making Machine Data Easier to Onboard, Prepare and Trust with AI-Powered Data Management
Platform Michelle Corpora , Yogesh SontakkeKey takeaways
- AI-Powered Data Management helps teams turn raw, changing machine data into trusted, usable signals faster.
- Guided Onboarding with Auto-Schematization (Auto-Schema) reduces onboarding guesswork by giving admins an opinionated path from intent to search-ready data.
- Self-Healing Pipelines help teams maintain data quality over time by detecting Common Information Model (CIM) compliance drift and surfacing AI-assisted remediation recommendations.
Every investigation, detection, dashboard, and AI-assisted workflow depends on one thing: data that teams can trust. But as environments grow more distributed, the data behind those experiences gets harder to manage. New applications, cloud services, security tools, infrastructure, and network devices constantly generate machine data, and each new source can introduce new formats, missing fields, inconsistent mappings, and pipeline changes that require expert attention.
Machine data is the digital exhaust generated by the systems that keep modern organizations running. It includes logs, metrics, traces, events, and other machine-generated signals from applications, infrastructure, cloud services, networks, and endpoints. This data is incredibly valuable, but only when teams can turn it into trusted, usable context.
That manual work matters. If teams do not onboard data correctly, they can miss important signals. If searches do not extract fields consistently, dashboards become harder to trust. If Common Information Model (CIM) compliance drifts over time, security content that depends on normalized data can produce incomplete results even while data continues to ingest.
This is why the new AI-Powered Data Management features are such an important part of Cisco Data Fabric powered by the Splunk Platform. Cisco Data Fabric powered by the Splunk Platform helps organizations bring together data, context, and action across security, observability, and networking. These AI-Powered Data Management features help teams get the right data to the right place for the right purpose, with less manual toil and more confidence in the quality of data.
Splunk recently announced the general availability of new Splunk Platform innovations supporting the Cisco Data Fabric powered by Splunk Platform. Here, we take a closer look at AI-Powered Data Management features and how it helps teams move from raw, changing machine data to trusted operational context.
Why Data Management Is an AI Readiness Problem
AI is only as useful as the data and context available to it. Raw telemetry alone is not enough. Teams and AI systems need data that is structured, governed, relevant, discoverable, and accurate enough to support decisions.
For many Splunk administrators and data teams, that readiness work still depends on specialized expertise and manual steps:
- Diagnosing pipeline issues when data changes
- Monitoring whether data remains compliant over time
- Creating or updating configuration artifacts
As one Splunk administrator at a renowned higher education institution put it:
That captures the real problem: the work is not just about ingesting more data. It is about making data usable and keeping it usable as environments change.
A More Guided Path from Raw Data to Search-Ready Signals
Guided Onboarding with Auto-schematization (Auto-schema) is now generally available as an AI-powered experience in Splunk Data Management. It helps administrators plan, parse, structure, and prepare data for the Splunk platform faster while keeping admins in control of the workflow.
The experience includes two primary workflows.
- Plan data onboarding reduces the overwhelm of choosing from many possible onboarding paths. The workflow gathers context through guided questions, then recommends a strategy, deployment considerations, and a focused task list so admins can move from "where do I start?" to a clear plan faster.
- Schematize custom data helps admins work from representative sample events toward search-ready structure. The workflow can identify patterns in custom data, recommend mappings, generate field extraction logic, and help produce outputs such as add-on packages or SPL2 templates based on the destination and scenario.
For customers that manage complex security data, the value is speed and confidence. A Splunk admin at Atea, a European IT infrastructure company, said,
AI does not replace the administrator. The workflow gives administrators a faster starting point, clearer recommendations, and reviewable outputs they can inspect and edit before using.
Then explore the Schematize custom data walkthrough to see how representative events can become recommended mappings and reviewable configuration outputs.
Streamlining Field Extraction
Manual field extraction is one of the most specialized and time-consuming parts of preparing custom data. Teams often need to inspect raw events, identify patterns, write regular expressions, test results, and repeat the process as data changes.
Automated Field Extraction currently remains in Controlled Availability and is designed to help reduce the regex burden by using AI to analyze data sources and suggest fields for extraction during ingestion and processing. The result is less manual setup, faster usable data, and cleaner signals for downstream security, observability, and AI workflows.
Keeping Data Useful After It Is Onboarded
Onboarding is only the beginning. Data sources evolve. Vendors change event formats. New fields appear. Existing fields disappear. Configuration files drift. A source can continue sending data while the normalized fields used by downstream security content become incomplete or incorrect.
Self-Healing Pipelines are now generally available with Splunk Ingest Monitoring 1.4.0. It helps teams detect CIM compliance issues, understand likely root causes, and generate proposed configuration fixes using AI.
The workflow design operates around four steps:
- Detect compliance degradation. The system monitors data models for missing fields and incorrect values. Admins can configure alert rules at the data model, dataset, and sourcetype levels so monitoring reflects the needs of their environment.
- Analyze the likely root cause with AI. When an alert triggers, the AI analysis service examines sample events and relevant configuration files to identify the likely cause of the issue.
- Review proposed changes. Recommendations are presented in a side-by-side diff view so admins can inspect proposed updates to props.conf, transforms.conf, or both. Admins can also preview how data would look with the proposed configurations applied before deploying anything.
- Deploy through an add-on overlay. After review, admins can download a ready-to-install add-on overlay package that layers corrected configuration on top of the existing add-on without modifying the original files.
The non-destructive model matters. It makes proposed fixes easier to review, manage, and reverse while preserving the integrity of the base add-on. It also keeps the administrator at the center of the workflow: AI proposes, the admin reviews, and the organization retains control over what they deploy.
Early customer feedback has reinforced the value of applying AI to pipeline monitoring and remediation. As one Splunk Solutions Architect at a global access solutions company shared, the "product concept is valuable for proactively detecting CIM compliance issues across data sources."
Turning Better Data Management into Better Outcomes
Consider a security team onboarding telemetry from a new identity system. The source includes useful authentication events, but the logs are custom, the field names do not map cleanly to existing content, and the team needs the data to support CIM-based detections in Splunk Enterprise Security.
With Guided Onboarding, the administrator can plan the onboarding path and identify the steps required to bring the source into the Splunk platform. With Auto-schema, they can use representative events to generate recommended mappings and reviewable configuration outputs. With Automated Field Extraction, teams can accelerate the field extraction work that often slows custom onboarding. Once the source is in production, Self-Healing Pipeline can monitor for CIM compliance drift and surface AI-assisted remediation recommendations if the data changes.
The result is a more resilient data lifecycle. Teams can onboard data faster, reduce manual setup and maintenance, and improve the quality of the data that powers investigations, detections, dashboards, and AI-assisted workflows.
One Data Foundation for Cisco Data Fabric
AI-Powered Data Management is one part of the broader Cisco Data Fabric journey.
Splunk Data Management helps control what data enters and moves through the fabric. Machine Data Lake provides a cost-effective environment for retaining full-fidelity machine data at scale. Catalog makes data easier to discover, understand, and govern. Federated Search helps users and AI agents query data where it lives without unnecessary movement or duplication.
Together, these capabilities help organizations move beyond data chaos. They support a more deliberate data strategy: onboard data faster, keep it accurate, retain it economically, find it when they need it, and activate it with the right context and controls.
For AI, that foundation is critical. Agentic operations require more than models and prompts. They require operational data that can be understood, trusted, governed, and connected to action. AI-Powered Data Management helps build that foundation by reducing manual friction at the point where raw data becomes usable intelligence.
Together, these capabilities help customers reduce manual toil, improve data quality, and build a stronger AI-ready operational data foundation for Cisco Data Fabric powered by the Splunk Platform.
Availability
Guided Onboarding with Auto-schema is generally available for eligible Splunk Cloud Platform customers in selected regions and requires Splunk Cloud Platform version 10.4 or higher.
Self-Healing Pipelines is generally available with Splunk Ingest Monitoring 1.4.0 and is automatically activated on eligible Splunk Cloud Platform deployments through a gradual rollout.
Automated Field Extraction remains in Controlled Availability.
Related Articles

Behind the Walls: Techniques and Tactics in Castle RAT Client Malware

High(er) Fidelity Software Supply Chain Attack Detection
