Microsoft CEO Satya Nadella has publicly criticised AI labs such as OpenAI and Anthropic for using public data at scale to train their models while simultaneously prohibiting other parties from imitating those models through a technique known as distillation. Nadella calls this pattern a "reverse information paradox" and argues that companies using AI services are effectively paying twice: once with money and once with their own proprietary knowledge.
The remarks come from a CEO whose company works closely with OpenAI while also selling infrastructure that enables other businesses to set up their own AI learning processes. Nadella thus positions Microsoft as an advocate for greater autonomy for enterprise AI users, a stance that directly benefits his own product portfolio.
What the reverse information paradox entails
Nadella references the economist Kenneth Arrow, who described the information paradox as the problem that you can only assess the value of information after you already know it. Where Arrow was concerned with the difficulty of selling knowledge, Nadella identifies a mirror-image situation at AI labs. They accumulate knowledge from the general public and from paying customers, yet do not allow others to do the same with their models.
In practice, this involves two distinct behaviours. First, labs such as OpenAI and Anthropic train their foundation models on large quantities of publicly available text and data, invoking fair use. Second, those models learn through interactions with enterprise customers, meaning those customers inadvertently contribute to the improvement of commercial models. At the same time, the terms of service of these labs prohibit third parties from using their models as training data for their own models, a technique known as distillation.
Distillation allows parties to achieve strong performance with relatively limited resources by training a smaller model on the output of a larger one. The core of Nadella's objection is that the major labs ban this practice while having themselves used similar methods to develop their own models.
Broader criticism from the business world
Nadella is not alone in his objections. Alex Karp, CEO of Palantir, stated that enterprise customers are "furious" about token-based pricing models and the risks of sharing intellectual property with AI providers. David Sacks, former White House AI adviser and tech investor, accused Anthropic of an "observe-copy-expand" pattern, in which the company benefits from others' work while accusing competitors of the same behaviour.
That accusation directed at competitors also comes from within Anthropic itself. CEO Dario Amodei previously complained about Chinese model makers allegedly stealing his company's work. Anthropic accused Alibaba of a large-scale distillation attack on its models. Elon Musk previously criticised Anthropic's data collection practices. Media companies including the BBC have also sued AI labs for using their material without permission.
Google DeepMind finds itself in a position comparable to that of OpenAI and Anthropic: that lab too builds on the work of third parties while protecting its own models.
Microsoft's own interest in the debate
Nadella's case for greater autonomy for enterprise AI users cannot be separated from Microsoft's commercial position. Microsoft sells infrastructure that allows companies to manage their own AI learning processes independently of the major model labs. If businesses become convinced that they need their own learning infrastructure to maintain control over their data and knowledge, they are more likely to procure it from Microsoft.
At the same time, Microsoft is one of the largest investors in and partners of OpenAI, one of the labs Nadella is now criticising. That dual position makes his remarks notable: he is calling out a practice of a partner while himself benefiting commercially if companies decide to become less dependent on those very same labs.
What this means for the European and Dutch AI scene
The debate over distillation and data ownership is also relevant for European companies and policymakers. European businesses working with American AI services face the same questions Nadella raises: to what extent are they contributing to the improvement of models over which they have no control, and what are the risks to their proprietary knowledge?
For Dutch and European founders and investors, the discussion underscores the importance of clear contractual arrangements regarding data use with AI providers. Policymakers in Brussels and The Hague are already examining how training data and copyright relate to one another, including through the AI Act and ongoing court cases. Nadella's remarks add an extra dimension to that discussion: they show that even within the American tech sector, the boundaries of fair use and reciprocal data use remain contested.