The Digital Data Design Institute at Harvard is now the Harvard Business School AI Institute.

Open Weights and the New Economics of Inference

For months this column has circled a question: as AI agents do more of your work, who controls the meter? In April, Uber disclosed that it had used up its entire 2026 AI budget in four months, after an agentic coding tool spread to roughly 5,000 engineers faster than finance had modeled, and it capped spending at $1,500 per engineer per tool while its COO openly questioned the return. [1] Weeks later, Tesla capped its own engineers at $200 per week on outside AI tools, after some had been spending thousands weekly. [2] Uber and Tesla are not outliers; they are early public examples of a problem now landing on CFO desks everywhere. This is why we’re looking at this issue this month. In May we called it “the end of the subsidy era”, and this month the bill has arrived.

Token Bills Have Exploded as Two Forces Collided

Agentic AI multiplied token consumption by orders of magnitude, since a single autonomous task can chain thousands of model calls, just as vendors shifted from flat subscriptions to usage-based pricing. With frontier models charging in the tens of dollars per million output tokens, costs that once looked like a rounding error now compound into six- and seven-figure invoices.

The Open-Weight Answer

Leaders are responding on two fronts. The first is orchestration: instead of routing every request to a top-tier model, a routing layer sizes up each task and sends routine work to smaller, cheaper models. New frameworks like RouteLLM prove that optimizing model selection per query can slash costs by up to 3.66x while preserving 95% of the response quality, [3] and 37% of enterprises now run five or more models. [4]

The second front is open-weight models, systems whose parameters you can download and run yourself. Here the price gap is startling. Alibaba’s Qwen-Turbo runs about five cents per million input tokens, while Zhipu’s GLM 5.2 vies with a leading frontier model on coding benchmarks at roughly one-fifth the cost, and is winning rapid adoption among developers. [5] Chinese open models have climbed to represent 45% of all traffic on Open Router, up from under 2% in 2024, [6] and Alibaba’s Qwen has now passed one billion cumulative downloads, overtaking Meta’s Llama as the world’s most-downloaded open model family. [7] They are powerful and inexpensive, but for regulated or sensitive data the prudent path is to run them on your own servers or a domestic cloud, so nothing leaves your control. Western options are maturing fast too: Google’s Gemma has passed 400 million downloads, [8] NVIDIA’s Nemotron has roughly doubled from more than 50 million downloads this spring [9] to 100 million by early July [10] and now anchors a new coalition that includes France’s Mistral, [11] and Microsoft’s Phi line delivers comparable quality on many tasks at a fraction of frontier cost. [12]

The Sovereignty Catch

There is a complication leaders cannot ignore: governments now treat the most capable models as strategic assets. In June, the U.S. Commerce Department restricted foreign-national access to one leading vendor’s most powerful models. [13] Weeks later, Beijing signaled it may curb overseas access to China’s top models, including open-weight releases. [14] The models in your stack, in other words, may be shaped as much by policy as by price. Sourcing AI is starting to resemble sourcing any critical input: subject to controls, and worth diversifying.

This Month’s Action Item for Leaders

As I’ve said before in these memos: route, pilot, and govern. To route, put an orchestration layer in front of your workflows so routine tasks go to cheap models. To pilot, run a controlled test of one or two open-weight models on a high-volume, low-sensitivity workload, hosted on infrastructure you control, and measure quality and cost against your incumbent. To begin to govern, write a simple model policy that classifies workloads by data sensitivity and specifies which models are approved for each.


One Other Thing

Watch model availability become a boardroom question. As Washington and Beijing both draw lines around their best systems, “Which models are we allowed to use?” shifts from an engineering detail to a supply-chain risk that belongs on the enterprise risk register, beside cloud concentration and single-vendor dependence.


Next Month at a Glance

The meter is running and next month we’ll ask what it buys. As boards press for proof of return, a real ROI reckoning is underway, and the picture is more balanced than the headlines suggest. Many organizations are capturing genuine, measurable value, especially where AI is pointed at bounded, high-volume problems such as customer service, fraud detection, and software development. Others are struggling, though often the value is real and only the measurement is missing. In September, we will discuss how to tell the difference in your own firm.


Meet the Author

Mike Grandinetti artistic headshot

Mike Grandinetti is an Executive Fellow at Harvard Business School. He’s a serial tech entrepreneur, board member, AI & innovation consultant, VC EIR, and award-winning professor in the practice. A former Silicon Valley engineer and McKinsey consultant, Mike has been C-Suite leader roles across 8 tech startups, resulting in 2 NASDAQ IPOs and 7 strategic exits.

He’s led senior executive workshops for Berkeley, Brown, Carnegie Mellon, Columbia, Cornell, Harvard P&ED, NYU & Oxford. He’s been a senior advisor and organizing team member for the MIT CIO Symposium for a decade.

https://www.linkedin.com/in/mikegrandinetti
www.mikegrandinetti.com


Sources

[1] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/

[2] https://electrek.co/2026/07/02/tesla-caps-employee-ai-spending-200-week/

[3] https://arxiv.org/abs/2406.18665

[4] https://www.digitalapplied.com/blog/llm-model-routing-2026-cost-quality-optimization-engineering-guide

[5] https://www.technology.org/2026/07/02/zhipus-glm-5-2-rivals-opus-4-8-on-coding-benchmarks-at-a-fifth-of-the-cost/

[6] https://www.digitalapplied.com/blog/openrouter-rankings-april-2026-top-ai-models-data

[7] https://www.opensourceforu.com/2026/07/alibabas-qwen-crosses-one-billion-downloads-eclipsing-metas-llama/

[8] https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/

[9] https://blogs.nvidia.com/blog/nemotron-3-nano-omni-multimodal-ai-agents/

[10] https://aidailybrief.beehiiv.com/p/ai-costs-are-surging-and-the-cheap-model-fix-might-not-last

[11] https://nvidianews.nvidia.com/news/nvidia-launches-nemotron-coalition-of-leading-global-ai-labs-to-advance-open-frontier-models and https://x.com/NVIDIAAI/status/2074252047151452238

[12] https://www.practicallogix.com/small-language-models-eat-the-edge-the-32x-cost-disruption-reshaping-enterprise-ai/

[13] https://www.csis.org/analysis/department-commerce-restricted-access-anthropics-latest-models-what-comes-next

[14] https://qz.com/beijing-china-ai-model-export-restrictions-070726