The AI ​​race is no longer about the biggest model



TL;DR

The assumption that the biggest AI model wins is breaking down, with companies now choosing models by task, cost, and control rather than by ranking. Driving this are model bills running into the millions per month, the rise of model routing, and agents specialized in specific tasks, which Gartner expects in 40% of enterprise applications by the end of 2026, up from less than 5%. If the capacity is to commodify, the margin passes to whoever makes the cheapest inference.

For years, the industry was based on one assumption: that the biggest model wins. That belief is now collapsing, CNBC reports.

Companies are choosing models by task, cost and control rather than by reference position. The border continues to matter, but it is no longer the only thing that is bought.

The reason is not romantic. On an enterprise scale, model invoices run into millions of dollars per month.

The rise of good enough

The operating principle is now the cheapest model that surpasses the quality bar. Buyers have discovered that most tasks do not need a border system.

Model routing has emerged to automate that judgment, sending each request to the model that best suits it. A summary paper and a multi-step reasoning paper no longer go to the same place.

Specialized, industry-specific models are filling the rest of the gap. Gartner expects 40% of enterprise applications to incorporate task-specific AI agents by the end of 2026, up from less than 5% a year earlier.

Why did the bills force this?

The economy stopped working. Prices per token have plummeted, yet Enterprise AI bills have tripled anywaybecause agent tools consume many more tokens per task.

Buyers took notice. Palo Alto Networks CEO Nikesh Arora has said Token prices must drop by up to 90%. for adoption at scale.

Some companies stopped waiting and started rationing. a wave of ‘Token Minimization’ Causes Companies to Limit Employee AI Spending total.

Where Value Moves Next

If capacity is commodified, the margin migrates to whoever manages it cheaper. Inference optimization has quietly become one of the most valuable layers of AI infrastructure.

Cheap, open models sharpen the point. Chinese models approach US border labs at a fraction of the price, limiting what anyone can charge for merely competent production.

This is uncomfortable for the scale thesis. Hundreds of billions in capital expenditures were justified on the premise that larger models would still be decidedly better, and buyers are now voting otherwise.

None of this means that the frontier models are finished. It means the industry is discovering that most work is boring, and boring work doesn’t need the most expensive tool in the shop.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *