Can Tech Companies Adopt More Cost-Effective AI Models?
The rise of AI has been based on a core belief: bigger models have more power, and the strongest models lead the way. However, the industry is about to uncover what happens when this assumption starts to unravel.
Rising costs have prompted users to explore smaller, more budget-friendly models. This cost-conscious approach to model selection is a relatively new development, and its impact on the industry is still unclear, yet the effects are expected to be significant.
A prediction made most clearly by Brian Armstrong, co-founder of Coinbase, indicates that a large portion of tasks will shift to these more economical models.
“[D]emand for intelligence is nearly endless, yet 80% of workloads will rely on models that cost 99% less in the next 12 to 18 months,” Armstrong shared on X. “The remaining 20% will still utilize cutting-edge models where peak performance is necessary.”
If Armstrong’s prediction holds true, it could signal a major transformation for the AI industry.
Traditionally, most AI firms have concentrated on quality, which has often meant leaning towards the most advanced models available. If tasks previously managed by high-end models can now be successfully executed by cheaper alternatives, it would mark a significant shift in AI’s economic landscape. Additionally, a considerable share of the resulting savings would impact leading labs, creating financial challenges for companies like OpenAI and Anthropic as they near their IPOs.
This could potentially usher in a revolutionary transformation within the industry, centered around a crucial question: Are companies ready to embrace smaller models?
Preliminary evaluations suggest that, if implemented properly, less expensive models can be deployed without compromising on quality. A recent assessment by the legal AI tool Harvey demonstrated a reduction in inference costs by a factor of three while maintaining quality. This test, conducted in partnership with the inference platform Fireworks AI, combined Claude Opus and Fireworks’ GLM 5.1, utilizing Opus for the most demanding tasks. The result was a marked decrease in server time and costs overall.
“Quality is paramount, and it will always be within the legal sector,” remarked Gabe Pereyra, co-founder of Harvey, during an interview with TechCrunch. “However, the definition of quality is shifting from merely using the most powerful model for every task to identifying the best model that provides the correct answer efficiently.”
This trend is often framed as a competition between major labs and Chinese or open-weight models, but that viewpoint overlooks the broader perspective. The real divide is not between proprietary and open models, but rather between larger and smaller models. Transitioning from GPT-5.5 to DeepSeek’s V4 Flash can save on costs, yet moving to GPT-5.4-mini may yield comparable results.
Currently, a price war is unfolding between in-house inference from large labs and independently serviced open-weight models. In terms of the broader small versus large debate, the specific type of small model may not be as crucial.
While it may seem obvious — one should avoid using excess computational power — it runs counter to the prevailing scaling-first mindset in the industry. Driven by historical lessons, labs have heavily invested in creating the most compute-intensive models possible, pushing the frontiers of AI capabilities. With costs largely covered by investors, clients had little reason to opt for anything other than the most advanced options available.
As token prices rise and subsidies fade, users are facing cost pressures for the first time. It remains uncertain whether these new financial limitations will lead enterprise users to smaller models. They might just as easily reduce costs by making fewer calls, cutting back on context usage, or discontinuing less viable deployments.
However, if it turns out that most deployments could operate just as effectively on smaller models, it could significantly dampen the soaring demand for inference and raise pressing questions about the financial rationale for training cutting-edge models.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.


