Why the Smartest Enterprise AI Strategy Starts with Tokenomics and Not Giant Models

CIO Office Amin Karbasi VP and Chief AI Scientist Cisco
The most capable AI model isn’t necessarily the right model for every workload. As AI usage scales, leaders should match model capacity to business value, cost, performance and risk.

For much of the AI boom, progress was defined by pure scale, with more data, more compute, and increasingly capable models. Organizations built their AI strategies around access to that growing capability.

Today, that calculus is changing. Defaulting to the most advanced frontier model for every task can add unnecessary cost, latency, and resource demands. Organizations are no longer just managing bits and bytes. They are managing intelligence.

Tokenomics is why that shift matters. Routine work flows best to smaller models, while frontier models deliver real ROI when reserved for high-complexity tasks where their cost is justified. Every AI interaction consumes resources, and different levels of model capability carry vastly different price tags. In agentic systems, a single business task can trigger multiple model calls, tool interactions, and retries, compounding token usage before the work finishes.

The next phase of AI maturity belongs to leaders who treat intelligence as a strategic resource and allocate it with precision. For technical leaders, the challenge is no longer whether to access advanced intelligence, but how to allocate it across the enterprise to drive business value.

The hidden costs and limitations of frontier AI models

Selecting the right AI model is as much an economic decision as it is a technical one. Many organizations default to the most advanced models, assuming more reasoning equals better outcomes. However, that isn’t always the case. Routine tasks like structured data extraction — including parsing standard invoices or categorizing high-volume support tickets — don’t require the massive reasoning capabilities of a frontier model. Using them there is an inefficient use of resources. Frontier models offer broad capabilities and advanced reasoning across a wide range of tasks, excelling at complex problems, orchestration, and work that span multiple domains. But that breadth isn’t necessary for every enterprise workload.

One common misconception is that frontier models — the massive general-purpose models developed by a handful of labs — are the only path to innovation. While these models are undeniably powerful for generic complex reasoning, such as synthesizing high-level industry trends with global geopolitical data, they are often the wrong tool for the high-frequency workflows that power daily operations, like analyzing millions of network packets to flag security threats.

Renting versus owning your AI intelligence

When organizations rely exclusively on external APIs, they are essentially renting their intelligence.

They also create potentially dangerous vendor dependencies across access restrictions, price fluctuations, product roadmaps, and policy changes, alongside broader market shifts like global economic tides, technology trends and other elements completely out of their control.

For a CTO, renting intelligence is a major trade-off. While they might gain speed to market, they take on risks around uptime, regulatory compliance, and regional availability while staying tethered to a third party. They also risk exposing sensitive customer telemetry or losing the proprietary data that creates competitive advantage.

When organizations build a core product on intelligence they don’t own, they can unintentionally surrender rights or access to external providers. This opens the door for compliance violations and forces teams to add yet another layer of scrutiny over who truly controls their core workflows.

As AI consumption grows, the cost of intelligence will grow with it, often exponentially. CTOs who treat AI as a black box utility for every task will continue to struggle with rising costs and architectural fragility.

To manage intelligence, organizations will need to shift their approach from renting to owning. Renting intelligence works well for commodity tasks like basic coding assistance or generic content generation, as well as complex multi-domain orchestration. But building a durable competitive advantage requires sustainable AI ownership. That means taking strong, open-source models and training them on proprietary data, unique workflows, and domain-specific edge cases. For example, instead of relying on a general-purpose model that understands security in the abstract, an organization can fine-tune an open model on its own historical threat intelligence database and proprietary network logs. By shaping models around internal data procedures, organizations can create specialized systems that directly reflect your business, whether for vulnerability detection or network observability.

How agentic AI workflows explore token costs

Most organizations have a plethora of advanced intelligence capabilities at their disposal. The question is whether they should use those capabilities for every task. That’s where they can apply the Pareto Principle, or 80/20 rule.

Roughly 80% of tasks in a typical enterprise are routine, repetitive, or domain specific. Because they don’t require massive reasoning capabilities, they run faster and cheaper on compact, specialized models. The remaining 20% of tasks require complex, cross-domain reasoning, and genuinely warrant a frontier model to drive them.

However, economic inefficiency strikes when organizations force that 80% of routine work through expensive frontier models. In agentic systems, agents and sub-agents call each other repeatedly to complete a workflow, causing token consumption to multiply rapidly. For instance, a single customer support request might trigger a router agent to classify the issue, a search agent to query internal databases, a reasoning agent to analyze logs, and a final agent to synthesize the response. If each step in this chain relies on a high-end frontier model, you are paying for multiple rounds of expensive, redundant reasoning for a single routine task. For organizations paying premium per-token pricing for every step, operational costs quickly become unsustainable.

To control token count and manage budget, technical leaders are moving toward a tiered architecture:

  1. A foundational layer of highly optimized, smaller models (aka Tiny AI or SMLs) that handle the bulk of operational tasks.
  2. An orchestration layer featuring a smart routing system that directs complex, high-stakes queries to frontier models only when necessary.

Ultimately, the vast majority of routine business functions can be completed with smaller, more efficient models that consume fewer resources while delivering comparable outcomes.

Why domain-specific small language models outperform giants in production

The most effective AI model is the one that completely understands your specific domain.

That’s where a specialized AI leaves generic models behind, solving narrow problems like vulnerability detection far better than general-purpose options. These models are small enough to run on-premises, within a private data center, or even a laptop, yet powerful to perform almost any specific task.

When it comes to security and sovereignty, smaller, more focused models will have an upper hand. Specialized AI models are increasingly proving their value in SOC workflows where tasks are often narrow but mission-critical.

Consider what it takes for real-time log correlation. A frontier model might struggle to interpret the unique, proprietary naming conventions and traffic patterns specific to an organization’s network architecture. In contrast, a specialized model, fine-tuned on historical incident data, can identify a true breach signal amid millions from routine events — all while keeping sensitive telemetry entirely within a secure perimeter.

Ultimately, the decision to use a third-party API versus an in-house model is a strategic business calculation, not just a technical one. Organizations should weigh their specific risk tolerance against requirements around data sovereignty, privacy, and deployment. For teams in highly regulated sectors, exposing proprietary security logs or confidential customer data to an external vendor carries considerable risk. Deploying smaller, specialized models inside a private environment eliminates data transit risks and satisfies strict sovereignty requirements while maintaining the required high-performance reasoning.

Looking ahead, the economics of AI will increasingly reward organizations that match the right level of intelligence to the right task. Rather than maximizing model sophistication, the greatest return will likely come from accurately aligning model capability with business need.

A four-step framework for enterprise AI model selection

As tokenomics becomes a larger operational concern and priority, leaders need clear principles to balance cost, performance, and impact across their AI portfolios.

To start, CTOs and technology leaders will need rigorous evaluation to measure model performance against their own criteria, not just public benchmarks. Future assessments will also include understanding of domain specialization, which entails investment in the data and training pipelines so organizations can make models unique to their own business needs. And finally, it will require architectural discipline, balancing the use of frontier models with the deployment of specialized, internal models to optimize cost and performance.

Use these four questions to determine when to route work to smaller specialized models and when a frontier model is warranted:

  1. Is this a core differentiator? If the AI capability provides a unique competitive advantage (e.g. proprietary security threat detection), it should be “owned” through fine-tuning and internal data. If the capability is a commodity (e.g. basic summarization of generic coding assistance), it is a candidate for “renting” via a frontier API.
  2. What is the data sensitivity threshold? Does this workflow involve proprietary logs, confidential customer data, or regulated IP? If the answer is yes, the data sovereignty requirements dictate an in-house or provide deployment model, regardless of performance benchmarks.
  3. What is the scale and frequency of the task? Is this a “high frequency workflow” that runs millions of times a day? If so, the cost of a frontier model will likely be prohibitive. These tasks are best served by smaller, highly optimized models that offer predictable costs and lower latency.
  4. Does the task require world knowledge or domain expertise? Frontier models excel at broad, cross-domain reasoning, or world knowledge. However, if the task requires deep, narrow expertise — such as interpreting specific network protocols or internal business logic — a smaller, specialized model trained on your domestic data will almost always outperform a generalist.

Building enterprise AI you actually own

Looking ahead, intelligence is no longer synonymous with massive size. AI architecture now requires deliberate choices about model routing, deployment, and cost control, with frontier and specialized models each serving distinct roles.

Instead, the future of enterprise AI will evolve with more specialized models that are faster, cheaper, and more secure than their frontier counterparts. These models will allow organizations to deploy AI at scale, without soaring costs or loss of control associated with large-scale frontier models.

In the era of tokenomics, the biggest model is not always the smartest choice. The winning strategy is the one that aligns with your business needs, protects your data, and puts the control of your intelligence firmly back in your hands.

To learn more about tokenomics and how it can be applied to your organization, please subscribe to the Perspectives by Splunk monthly newsletter.

No results