OpenAI wants to substantially reduce its API rates. That is what anonymous sources tell The Wall Street Journal. The move is said to be driven in part by anticipated price reductions from rival Anthropic, the maker of the Claude language model.
The API market for large language models has grown increasingly competitive in recent years. Developers and companies that integrate language models into their products pay per unit of text processed, expressed in tokens. Pricing is one of the primary competitive factors in that segment, alongside model performance and availability.
OpenAI has not yet published specific new rates. The reporting is based on sources who wish to remain anonymous; OpenAI has not provided an official comment to The Wall Street Journal.
Anthropic gains ground with enterprise customers
Anthropic has recently been gaining ground with business customers through Claude, partly due to strong performance on coding tasks and a relatively favourable price-to-quality ratio. This has made the company a serious challenger to OpenAI in the enterprise segment, where many customers evaluate multiple models simultaneously and switch providers on a regular basis.
According to The Wall Street Journal's sources, OpenAI expects Anthropic to lower its own rates further. By acting pre-emptively, OpenAI aims to prevent developers and businesses from switching. In practice, price differences between providers become apparent quickly: at high volumes, even small per-token rate differences add up fast on the monthly bill.
API costs feed directly into budgets
For software companies and developers that rely on language models on a structural basis, API rates are a direct line item. The more intensively an application uses a model, the greater the weight of the token price in total infrastructure costs. This makes price reductions relevant to a broad range of customers, from small startups to large enterprises with high usage volumes.
OpenAI has already cut its rates several times in recent years, partly as a result of more efficient model architectures and cheaper compute. At the same time, the company reported substantial losses in 2024, driven in part by high infrastructure costs. Further rate reductions increase pressure on the margin per processed token, unless larger usage volumes partially offset that effect.
Broader downward pressure across the market
OpenAI and Anthropic are not the only players reducing prices. Google has also made its Gemini models cheaper via the API over recent months, and open-source alternatives such as Meta's Llama models can additionally be self-hosted by companies, eliminating API costs altogether. That combination of competing providers and self-hosted options keeps pressure high across the entire market.
For the Dutch and broader European startup scene, this is a relevant development. Many AI companies here build on top of the APIs of major US providers. Lower rates reduce the barrier to building and scaling products, and lessen the reliance on subsidies or external funding to cover infrastructure costs. For investors and founders keeping a close eye on unit economics, every downward step in token pricing feeds directly into the business case.