Feature Overview
AI-Powered Data Management is a suite of GenAI-powered capabilities within Splunk Data Management that helps administrators onboard, schematize, and operate data pipelines faster and with less manual effort. Using an agentic framework powered by LLMs (Large Language Models) and domain-specific tools, AI-Powered Data Management guides administrators through complex data onboarding workflows, automates schema generation and CIM mapping, and continuously monitors pipeline health — regardless of their depth of experience with Splunk configuration or data modeling.
The agentic framework refers to a set of specialized AI agents that can orchestrate domain-specific behaviour based on administrator actions and data context. An orchestration layer interprets intent, plans workflows, invokes the right tools, and synthesizes results.
Model Overview
AI-Powered Data Management is powered by a leading LLM provided by a trusted cloud vendors, enabling seamless integration and strong performance for structured reasoning, code generation, and conversational guidance. The model is complemented by deterministic algorithmic processing for tasks requiring precise, reproducible results — such as event clustering, configuration syntax validation, and pipeline execution testing. Together, the LLM and algorithmic components form a hybrid AI system where the LLM handles reasoning, generation, and natural language interaction, while deterministic methods handle validation and verification.
Model Evaluation and Performance
AI-Powered Data Management is evaluated through a multi-layered protocol. At the individual agent level, each capability is tested against representative datasets — for example, auto-schematization is evaluated against well-known data sources where correct CIM mappings, field extractions, and Technology Add-ons are known. At the end-to-end level, the full workflow is tested from sample upload through artifact generation, with results compared against expert-validated outputs. Through iterative refinements to prompt engineering, tool design, agent workflow, and validation loops, AI-Powered Data Management demonstrates its potential for driving value in real-world evaluation scenarios. In-product feedback loops are implemented to further improve quality over time.
Data Sources for Model Training
AI-Powered Data Management does not train or fine-tune foundation models on customer data. The LLMs used are accessed as managed API services and are not modified by Splunk. Domain-specific knowledge — such as CIM model definitions, Splunk configuration best practices, and validated architecture patterns — is provided to the models at inference time through system prompts, tool integrations, and retrieval mechanisms. Splunk uses information from customers' interaction with AI-Powered Data Management in accordance with Specific Terms for Splunk Offerings and Documentation.
Data Privacy and Security
Data that is directly relevant to the generation of a meaningful response — such as sample log data uploaded by the administrator, onboarding questionnaire answers, pipeline metadata, and tool outputs — is sent to the third-party LLM service, which generates responses. Sample data is limited in size and is explicitly provided by the administrator for the purpose of schematization; production customer data at rest in indexes, metric stores, or trace stores is not accessed or sent to the LLM. All workflow state, session data, and generated artifacts are strictly tenant-isolated with no cross-tenant data sharing. Data that is not necessary to generate a relevant response stays within the Splunk compliance boundary for the customer's underlying Splunk Cloud Platform offering. Data categories relevant to AI-Powered Data Management are described in Data sharing and use. Additionally, the Specific Terms for Splunk Offerings set forth Splunk's data practices relating to all data collected, generated, or used by AI-Powered Data Management, including our practices relating to your opt-out preferences.
Fairness
AI-Powered Data Management results — including generated schemas, CIM mappings, Technology Add-ons, SPL2 pipelines, and pipeline remediation recommendations — are unique to each customer's data and environment and should be reviewed for accuracy and appropriateness by a qualified administrator prior to deployment. Human-in-the-loop review is built into critical workflow stages to support this.