Simplify OpenTelemetry Operations with the New Fleet Management UI

Observability Courtney Gannon

Key takeaways

  1. Monitor OpenTelemetry agents and collectors in one place with a new Fleet Management UI that shows health, versions, configurations, and status.

  2. View health, configuration, and diagnostic data together to identify telemetry collection issues and troubleshoot them more quickly.

  3. Centralized visibility and APIs help teams manage large OpenTelemetry deployments, reduce configuration issues, and operate with greater confidence.

Managing OpenTelemetry at scale starts with answering a few deceptively simple questions: Which agents and collectors are running? Are they healthy? Which versions and configurations are they using? And when something goes wrong, what is causing the problem?

In our introduction to OpenTelemetry Fleet Management, we shared how Splunk Observability Cloud uses the Open Agent Management Protocol (OpAMP) to establish a centralized control plane for observability agents and collectors. This foundation gives teams a consistent way to retrieve inventory, health, version and configuration information across distributed environments.

Now, we’re making that information easier to explore and act on with a new Fleet Management experience in the Splunk Observability Cloud user interface.

From Fleet Management APIs to an Integrated UI

The first release of OpenTelemetry Fleet Management gave customers APIs for retrieving detailed inventory and effective configuration information. These APIs enable teams to incorporate fleet health checks into scripts, CI/CD pipelines, incident-response workflows and other automation.

The latest update brings that operational visibility directly into Splunk Observability Cloud. Teams can now use a dedicated UI to inspect registered OpenTelemetry Collectors and supported instrumentation agents without first creating their own API integrations.

The experience is available from Data Management > Fleet Management, giving observability teams a central place to understand the software responsible for collecting telemetry across their environments—and to investigate problems when collection is disrupted.

See Your OpenTelemetry Fleet in One Place

Modern organizations can operate thousands of agents and collectors across Linux, Windows and Kubernetes environments. Without centralized inventory, it can be difficult to determine whether every component is present, healthy and configured as expected.

The Fleet Management UI helps teams answer those questions by surfacing important details about registered clients, including:

This consolidated view helps teams identify unhealthy clients, find outdated versions and investigate inconsistencies without connecting to individual hosts or requesting access to application source code.

Move From Detection to Diagnosis

Knowing that a Collector is unhealthy is useful. Understanding why it is unhealthy is what helps teams restore telemetry quickly.

Fleet Management correlates inventory, configuration and health information with internal Collector metrics to provide additional diagnostic context. Instead of treating fleet status and Collector performance as separate sources of information, teams can use the combined view to investigate how a Collector is behaving and narrow down the likely source of a problem.

For example, a team investigating missing or delayed telemetry might need to determine whether the issue is caused by:

By correlating the Collector’s identity, version and effective configuration with its internal operational metrics, Fleet Management gives teams a more direct path from symptom to probable cause.

This reduces the need to move between disconnected tools, manually match a metric to a particular deployment or inspect hosts one at a time. It also helps teams distinguish an isolated Collector problem from a configuration or capacity issue affecting a larger part of the fleet.

Troubleshoot in Context

Collector issues are rarely solved by looking at a single signal. A high-level health state can indicate that something is wrong, but teams often need several pieces of context to understand what changed and what to do next.

The Fleet Management experience brings that context together so teams can ask questions such as:

This contextual approach helps observability teams reduce investigation time and get closer to root cause before they connect to the underlying infrastructure.

It can also improve collaboration during an incident. Instead of handing another team a generic report that “telemetry is missing,” the observability team can provide more specific evidence about the affected component, configuration and stage of the collection pipeline.

Understand the Configuration That Is Actually Running

In a distributed environment, the intended configuration and the configuration running on an agent or Collector are not always the same.

A deployment might be updated only partially. An environment variable could override a default. A configuration file might differ between two otherwise similar hosts. When this happens, the resulting configuration drift can create gaps in telemetry or make troubleshooting unnecessarily difficult.

Fleet Management reports each client’s effective configuration—the settings the client is actually using. Bringing this information into the UI makes it easier to:

This gives central observability teams greater visibility while reducing their dependence on application teams for routine checks.

Support for Collectors and Instrumentation Agents

Fleet Management is designed to provide a unified inventory for the components that make up an OpenTelemetry deployment.

That includes the Splunk Distribution of the OpenTelemetry Collector, whether deployed in agent or gateway mode, as well as supported OpenTelemetry instrumentation agents. This broader inventory is important because reliable telemetry depends on the health and configuration of the entire collection path—not just one component.

Using OpAMP, registered clients report identity, capabilities, health, errors and effective configuration to the Fleet Management control plane. Telemetry continues to travel over the normal OpenTelemetry Protocol (OTLP) data path; management communication uses a separate OpAMP channel.

Separating these paths allows teams to monitor and manage their collection components without disrupting the flow of application and infrastructure telemetry.

A Better Foundation for Operating OpenTelemetry at Scale

The new UI is more than a presentation layer. It represents an important step toward making OpenTelemetry operations accessible to a broader group of users.

Developers and automation teams can continue using the Fleet Management APIs for programmatic workflows. At the same time, platform engineers and observability administrators can use the UI for interactive investigation, fleet reviews, troubleshooting and day-to-day operational checks.

Together, these interfaces help organizations:

This is especially valuable for centralized observability teams that are responsible for telemetry collection standards but do not directly control every application or service.

Getting Started

To appear in Fleet Management, supported agents and Collectors must be configured to connect to the Splunk Observability Cloud Fleet Management service through OpAMP. The OpenTelemetry Collector can also provide the management path between supported language agents and the service.

Once the clients are enrolled and reporting, go to Data Management > Fleet Management in Splunk Observability Cloud to explore your inventory and investigate fleet health.For current prerequisites and configuration instructions, see Manage OpenTelemetry agents and collectors.

What’s Next

Centralized visibility is a fundamental requirement for operating observability agents at enterprise scale—but visibility alone is not enough. Teams also need to understand why collection components are unhealthy and how those issues affect the telemetry pipeline.

By bringing fleet status, versions, effective configurations and internal Collector metrics into a correlated experience, Splunk Observability Cloud helps teams move from detecting a problem to diagnosing its likely cause faster.

With APIs for automation and a UI for interactive operations and troubleshooting, OpenTelemetry Fleet Management helps teams understand their telemetry fleet, resolve Collector issues more efficiently and operate OpenTelemetry with greater confidence. These releases lay a strong foundation towards support for OpAMP based remote configuration and lifecycle management operations in the future.

Related Articles

Key Findings From a Recent Study on Data Management in the Modern Security Operations Center
Security
4 Minute Read

Key Findings From a Recent Study on Data Management in the Modern Security Operations Center

Learn about cloud storage preferences, data cost challenges, and best practices for optimizing your SOC's security posture and cost efficiency.
Splunk Named a Leader in Gartner SIEM Magic Quadrant for the Fifth Straight Year
Security
2 Minute Read

Splunk Named a Leader in Gartner SIEM Magic Quadrant for the Fifth Straight Year

Gartner's 2017 Magic Quadrant for Security Information and Event Management names Splunk a leader for the fifth straight year
Using Splunk to Detect Abuse of AWS Permanent and Temporary Credentials
Security
7 Minute Read

Using Splunk to Detect Abuse of AWS Permanent and Temporary Credentials

In this blog, the Splunk threat research team shows how to detect suspicious activity and possible abuse of AWS Permanent and Temporary credentials.