Connect with us

Business

Why every company wants an AI model router right now 



No one likes a surprise sky-high bill. But that’s exactly what many companies have faced this year as they deploy popular AI coding agents like Claude Code and Codex for increasingly long-running, autonomous tasks. Unlike a chatbot conversation, these agents can work for hours, repeatedly calling frontier models and quietly racking up millions of tokens. A developer can go to lunch and return to find that his agent has spent thousands on the inference, or output of the model. The result? Sticker shock.

A new study illustrates the impact: 62% of organizations said an unexpected AI expense materially altered a business decision over the past year. Among them, 40% said the issue required board-level escalation, 33% implemented emergency spending freezes, and 25% delayed or canceled an AI initiative outright.

That’s why AI model routers, software that  allows organizations to choose the AI model for the right task at the right cost, have suddenly become one of the hottest areas in enterprise tech. Rather than sending every request to the most expensive frontier model, companies can either manually define or automatically choose the model that offers the best combination of cost, speed, and performance for each step of an agent’s work. Companies say this kind of intelligent model routing can reduce inference costs by double-digit percentages – in some cases up to 30%. 

A slew of startups have rushed into the model router space, though they’re taking different approaches. Companies like OpenRouter, which has reportedly been in acquisition talks with Stripe at a valuation of up to $10 billion, provide a marketplace and unified gateway to hundreds of AI models, while Not Diamond automatically routes requests to the model best suited for each task. Others, including LiteLLM, let enterprises build and manage their own routing infrastructure. Large vendors such as Salesforce and Databricks are also building routing capabilities into their AI platforms. Cursor, Ramp and Meta are reportedly working on their own model routers, and even video startup Runway has launched one. 

Token-hungry agents have driven demand

Token-hungry agents have driven demand

OpenRouter co-founder and chief operating officer Chris Clark told Fortune that the vast demand for output from tools like Anthropic’s Claude Code, which was released in mid-2025, has driven the demand. Through 2024 and 2025, the C-suite was “pounding the table” for companies to adopt AI, he explained, but it wasn’t until this year that it all started falling into place because harnesses and agentic flows began advancing capabilities beyond chat. “It’s an AI model using tools, calling out to different systems, taking action,” he said. But agents are also more token-hungry, which for companies paying per-token, became more and more expensive. 

“No one had any budgets in place,” Clark said. “It was sort of this maximalist attitude.” 

But for OpenRouter, it’s been good for business. “Our durable belief is that AI tokens “will be a massive line item for every business’ operating expense,” he said, which will lead nearly every company to seek relief. 

Not every task requires the smartest—or most expensive—AI model, Clark explained, so there are real benefits to being choosy. “We try to talk to people about the idea of intelligent saturation: When a new frontier model shows up and you switch that agent to the newest, smartest model, does the performance improve or not?” he explained. Many tasks are already ‘intelligence-saturated,’ meaning if agents use the latest and greatest, smartest model, the performance doesn’t improve because an older model can already tackle that task. 

Still, most companies today simply default to the most powerful model for everything, which becomes wasteful, said Tomás Hernando Kofman, CEO of Not Diamond, which works with large enterprise clients such as SAP. However, he added that choosing the wrong small model at the wrong time can also be expensive because smaller models can take longer to get the same amount of work done. That’s where an automated model routing system can help. 

Not Diamond likens the problem of automating this to robotics. Just as a robot must coordinate thousands of small actions to complete a task, an AI coding agent needs different models at different stages of a long-running workflow. Rather than optimizing each prompt in isolation, its routing system predicts the best model and reasoning level for the next step based on the complexity of the request, the conversation’s history, and other signals. 

More than cost control

But while enterprises are adopting AI model routers today to rein in soaring inference costs, Salesforce president and chief architect David Ward argues that’s only the first phase.

“Let’s solve the big bill,” he said. “That’s certainly one priority, but there are other priorities that need to come into place as well.”

Over time, Ward expects routing decisions to expand beyond model selection to include trust, compliance, governance, and measurable business outcomes. Rather than simply choosing the right model, routers could help determine what enterprise data an agent can access and whether a fine-tuned open-source model is sufficient for a particular task.

“It’s not just which model for which job; it’s which tool, which skill gives me the measured outcome I want,” he said. “The industry will be having those conversations.” 

Florian Douetteau, co-founder and CEO of enterprise AI platform Dataiku, said the recent access restrictions to Anthropic’s Fable model also prompted enterprise customers around the world to rethink their dependence on a single AI provider. Rather than tying critical business processes to one frontier model, companies increasingly want the flexibility to switch providers—or fall back on increasingly capable open-weight models—if access changes.

“That’s another reason to have a router type of mindset,” he said.

Douetteau recalled one Fortune 500 CIO describing the anxiety of discovering late on a Friday that a model underpinning a critical business process might not be available on Monday. 

“That’s the practical thing that triggered people to start looking at different ways to procure AI,” he said.  

AI model routing isn’t easy

Still, even as a rash of AI companies rush into the routing space to serve enterprise demands, it’s not easy, said OpenRouter’s Clark. 

“I know a lot of companies are trying this – we’ve seen some sort of launch from just about every company under the sun over the past few weeks,” he said, pointing out that OpenRouter launched in 2023. “But I think those companies and their customers will discover this is a harder problem to get right than they anticipate.” 

While it may appear to be little more than a unified API connecting different AI models, he said running a production-grade routing platform requires deep partnerships with model providers, constant monitoring of thousands of model endpoints, rapid response to outages and specification changes, and ongoing validation of factors ranging from data policies to GPU configurations. 

“It’s not a thin software wrapper,” he said. “It’s a deep partnership.”



Source link

Continue Reading

Copyright © Miami Select.