From Data at Rest to Data in Action: Machine Data Lake and Catalog
Platform Piyush Revuri , Teja GattupalliKey takeaways
- Retain More, Spend Less — Economically retain full-fidelity machine data in a fully managed Splunk experience.
- Discover and Activate Faster — Use Catalog to find, understand, and selectively activate data across Machine Data Lake, Splunk indexes, and supported federated datasets.
- Turn Data into Action — Promote only the relevant data for investigation or analysis while maintaining governance and controlling costs.
The machine data that looks routine today may be the missing piece in tomorrow’s breach investigation, service disruption, or compliance review. Yet keeping every byte in a high-performance tier is neither practical nor necessary. Machine Data Lake and Catalog change that tradeoff: retain more full-fidelity data economically, find what matters quickly, and activate the right data when the business needs it.
We recently announced the general availability of new Splunk Platform innovations supporting Cisco Data Fabric. Here, we take a closer look at two of those innovations—and how they help turn machine data into agentic action.
A More Flexible Approach to Machine Data
Machine data continues to grow across an increasingly distributed and hybrid data estate. Yet not all of it has the same immediate value or performance requirements.
Frequently searched, latency-sensitive data may belong in a Splunk index. Many organizations place rarely used telemetry in external, low-cost object storage or data lakes. While economical, those approaches can introduce separate routing, metadata, governance, and retrieval workflows. Managed by Splunk and open by design, Machine Data Lake brings economical, petabyte-scale retention into a fully managed experience while letting organizations choose how selected data is activated. But economical retention solves only part of the problem. Organizations also need a reliable, simple, and self-service way to find, understand, and activate what they have retained.
From Discovery to Activation with Catalog
Catalog makes data across Cisco Data Fabric easier to find, understand, and use through relevant context, creating a discovery experience with governed access across Splunk indexes, Machine Data Lake, and supported federated datasets. Users can begin with a keyword search across dataset names, owners, types, descriptions, and field names. For data in Machine Data Lake, users who need to investigate more deeply can narrow the results by source, sourcetype, host, and time range.
Before running a search or starting a promotion job, users can inspect dataset coverage, including event volume, time coverage, and field schema to understand what a dataset covers and how its data is structured. This helps them determine whether it can answer their question before consuming additional resources.
Once the right data is identified, users can promote a precise slice rather than activating an entire raw table. Depending on the work, that slice can be promoted to a Splunk index for interactive SPL investigation or an Analytics Table for large-scale analysis. A promotion can capture a fixed period, or keep selected data current through streaming promotion to a Splunk index
Turning Historical Telemetry into Evidence
Consider a security team investigating suspicious activity involving a privileged account. Current identity and endpoint events are searchable in Splunk, but they reveal only part of the story. The analyst also needs older network telemetry from the period when the account may first have been compromised.
That telemetry is retained in Machine Data Lake. Using Catalog, the analyst finds the relevant raw table, confirms its fields and time coverage, and narrows the data to the affected hosts, sources, sourcetypes, and time window.
Because the goal is an active investigation, the analyst promotes that focused slice to a Splunk index, where it becomes searchable with SPL and available to native workflows such as dashboards, saved searches, and alerts.
The analyst can now connect historical network activity with current identity and endpoint signals, reconstruct the compromise path, and carry the findings into containment and stronger detections. Historical telemetry retained because it might matter has become evidence that changes the response.
One Governed Journey from Data to Action
Together, Machine Data Lake and Catalog advance the data, context, and action journey behind Cisco Data Fabric powered by the Splunk Platform.
At the data layer, Machine Data Lake provides an economical environment for landing and retaining full-fidelity machine data. At the context layer, Catalog makes datasets easier to find, evaluate, and understand. Promotion connects that context to the action layer, where selected data becomes available for investigation and analysis.
Once promoted, users can search Splunk indexes with SPL and incorporate results into investigations, dashboards, saved searches, and alerts. They can also use Analytics Tables for large-scale analytical and machine-learning workloads.
This self-service journey operates within existing Splunk role-based access controls, so users see only the datasets and actions their roles permit. Analysts gain a faster path to relevant data, while administrators maintain governance without serving as the manual intermediary for every request. By promoting only the sources, hosts, and time ranges with a defined purpose, organizations can act on the data that matters while preserving the economic value of the lake.
Retain Broadly. Activate Deliberately.
Machine Data Lake helps organizations retain more full-fidelity data without assigning all of it to the same cost and performance tier. Catalog makes that data discoverable, understandable, and selectively actionable. Together, they offer a better balance of readiness, economics, and choice: preserve more full-fidelity data, discover its context when questions arise, and activate it in Splunk or open analytics workflows when it matters.
Availability: As of August 4, 2026, Machine Data Lake and Catalog are generally available for Splunk Cloud Platform 10.5.2605.5 in US-East, Frankfurt, Tokyo, and Sydney. Discovery and promotion of data in Machine Data Lake are available through Catalog in those regions.
Related Articles

Staff Picks for Splunk Security Reading February 2022

