Splunk Observability Cloud offers a better view
REPAY needed to expand visibility so they could see the entirety of their operational environment. Splunk Observability Cloud fit the bill.
By standardizing on OpenTelemetry (Otel), REPAY creates a single source of truth across the entire organization. Trace IDs can match the flow of requests across their services, so teams can see the entirety of their environment — from gateway requests to specific database queries that caused delays.
“We can see the full picture, which we didn’t have before,” said Wolfe. “Having that end-to-end visibility is a massive change for us.”
OpenTelemetry has also fundamentally changed how the team collaborates. When a trace identifies an anomaly, the team looks beyond their own components to inferred services and related dependencies, clearing a path for greater cross-collaboration.
“We can chase the red dots, dig into what’s working, and what’s not,” said Wolfe. “Somebody might see something that doesn’t look right, and they’ll share the trace. Because the traces include inferred services and related database requests, it’s not just one team’s problem — it’s a collective effort to understand what happened. It removes the ‘it’s not my service’ finger-pointing you might see elsewhere.”
Dashboards open windows to context and transparency
Splunk’s single pane of glass approach allows Wolfe’s team to monitor system health using simple, intuitive “stoplight” dashboards that focus their attention on high-level indicators like latency and error rates. Now the team can quickly determine if the system is healthy without getting bogged down in complex, noisy graphs.
“We’re tracking error rates that tell us if the system is performing as expected,” said Wolfe. “I think the ability to see an issue, respond to it, and take action, has greatly improved.”
Wolfe said that the simplified dashboards have also been a game-changer for batch processing — an even bigger engineering challenge than day-to-day payments — especially during peak payment periods like the 1st and 15th of the month, when payments are typically due and transaction volumes surge. At these times, clients often upload tens of millions of payment transactions that need to simultaneously be processed, settled, and reconciled, creating “bursty” traffic that suddenly strains the infrastructure.
“As volume goes up, transaction time ticks up because of system load,” Wolfe said. “We’ve got millions of requests hitting our API, so we have to look at traffic through a different lens. But then we can pick out signals that are specific to those peaks, for those specific types of days. That’s exciting.”
That’s where Splunk Observability Cloud comes in, helping the team identify where to adjust components, visualize the waterfall of requests, see the retry mechanisms in action, and ensure algorithms are working as designed. “It’s about ensuring that a massive batch load doesn’t impact the rest of our ecosystem,” he said, adding, “we optimize and scale things out and then that load comes back down.”
Using Splunk dashboards in Observability Cloud, the company identified and optimized SQL queries and endpoints to reduce transaction latency by 30%. “The goal is to continue to refine and improve transaction processing times and error rates. We’re looking for ways that we can alleviate friction and have a higher success rate in general,” Wolfe said.
With that in mind, REPAY is developing client-facing dashboards to give clients direct access to their performance metrics, creating a new level of transparency and trust. “We want to share observability metrics with our clients so that they can see how things are working independently of us,” Wolfe said. “When we show a client their own performance graphs in Splunk, they often tell us, ‘Oh I love Splunk!’ There is a certain level of trust in the industry for the tool — it gives them comfort.”
AI Assistant, an expert on call
If Splunk Observability Cloud opens the door to a new world for REPAY, Splunk AI Assistant in Observability Cloud is the expert docent guiding them through it.
“Knowing the intricacies of every single system is impossible,” Wolfe noted. The “Splunk AI Assistant allows us to ask questions about our metrics and data correlations in plain language. For example, I don’t have to activate a service in Trace Analyzer. I can just enter a question in human language. Now I get answers from large amounts of data with minimal work. That’s an accelerator.”
This efficiency helped them triage 50% faster. AI Assistant was particularly useful for avoiding false positive rabbit holes, Wolfe said. “Those discussions where our team works for weeks trying to chase a problem and understand the root cause — those days are long behind us.”
In addition to digging into root cause, AI Assistant can proactively determine if an error rate on its dashboard is worth investigating at all. The team recently put it to the test during a potential production incident that triggered an alarm about a spike in errors. In the past, the organization would have called “all hands on deck,” pulling engineers away from their work to manually investigate logs. Instead, the team relied on AI Assistant to query the error spike and quickly found that a non-production endpoint accidentally deployed to the production environment and was failing because it couldn’t reach the resources it needed.
“It was a simple mis-deployment,” said Wolfe. “AI Assistant told us exactly what was happening in seconds. We didn’t have to get on a call, we didn’t have to notify the entire company, and we were able to resolve the issue immediately.”
Looking ahead, the company will lean even harder into its AI strategy, actively exploring new capabilities, including AI Agent Monitoring and advanced predictive alerting, to further reduce noise and focus on service-breaking issues. Wolfe said the goal is not only to have AI flag issues, but to propose solutions and act on them.
“We are at an inflection point,” Wolfe said “We want to focus on performance and speed. We’re competing against companies ten times our size, and while they have different budgets, we have the agility to win.”
REPAY clients win with Splunk
By standardizing on Splunk Observability Cloud and integrating the Splunk AI Assistant, REPAY has fundamentally transformed its operational culture. But more than that, by being able to respond more quickly, troubleshoot more accurately, and better act on the data, Wolfe and his team are also transforming the client experience when it comes to critical payments.
With Splunk, they now have the confidence to know exactly what is happening in their environment — and they have the right tools to keep their clients moving forward.
“We’ve excelled because we can show the clients the problem, and the mitigation steps to resolve whatever it is,” said Wolfe. “As we continue to improve transactional efficiency and reduce error rates, the clients just end up winning. They get a better product from us. Without a tool like this, I think you’re just in the dark.”