Open-weight AI models processed 56% of all tokens routed through Vercel's AI Gateway in August, marking the first time these models have accounted for a majority of monthly token volume, according to Vercel's September report published Thursday. The shift represents a dramatic acceleration from December 2025, when open-weight models handled just 7% of traffic, climbing to 13% by April before rising every subsequent month. Despite commanding more than half of usage, open-weight models captured only 14 cents of every dollar spent through the gateway, while Anthropic's proprietary Claude models took 64% of total spending.

The average price per token across Vercel's AI Gateway dropped 23.2% in August, the third straight monthly decline. Among teams processing more than 10 million tokens in both July and August, the median cost per token fell 7.6%. Anthropic has maintained its dominance on the spending side, never dropping below 61% of gateway expenditure in any month since December 2025, with its models occupying the top two positions by spend throughout that period. Within Anthropic's portfolio, Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the less expensive Opus 5 climbed to 22.5%. Google saw its overall share of token volume plummet from 30% to 5%, with the decline in Gemini 3 Flash alone accounting for 22 of those 25 percentage points, as more than three-quarters of the volume lost by that model moved to providers including OpenAI and Anthropic.

Vercel CEO Guillermo Rauch said the trend will likely continue, noting that "enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic." The report's authors observe that "lab loyalty doesn't follow brand, it follows model profile, and consistency wins," pointing to how 90% of teams using Fable reduced their usage, with more shifting those workloads to Opus 5 than to any other model. When Z.ai launched GLM-5.3-Flash, the new model was processing three times the daily volume of GLM-5.2 within five days.

The mechanics behind the spending gap are straightforward: open-weight models from developers like DeepSeek, Moonshot AI and Z.ai generally cost less to run than proprietary models from US frontier labs, meaning high token volume doesn't translate to equivalent revenue. Vercel's data suggests that when a new model preserves what users valued in its predecessor, the lab retains its customers, but when it doesn't, those customers fill the need through other providers. Opus 5 ultimately gained almost twice as much usage as Fable lost, which Vercel attributes to the newer model handling similar workloads at roughly half the price, allowing Anthropic to keep the dollars even as customers shifted toward a cheaper option within its own lineup. The AI Gateway sits between applications and underlying model providers, routing requests while tracking usage and costs across tens of trillions of tokens each month, giving Vercel visibility into which models its customers are actually running in production. Rauch declared August 22 a "record day for open weight share of tokens," when these models accounted for 62% of traffic, viewing the milestone as just an early indication of where usage is heading. The pattern suggests developers are willing to cross lab boundaries when a replacement fails to meet the same needs on capability and price, with loyalty tied to model performance rather than brand name. Companies betting on proprietary model ecosystems may find themselves caught between defending premium pricing and losing ground to cheaper alternatives that deliver comparable results. The shift also signals that enterprise AI budgets will increasingly flow through whoever can balance cost and capability, not necessarily whoever invented the technology first.