Microsoft executive called OpenAI's data collection the largest labor theft in human history
Newly unsealed court documents reveal that Microsoft and OpenAI executives discussed AI training on news content far more bluntly behind closed doors than in public. The New York Times reported on the filings, which came out of the lawsuit the paper itself filed against the companies.
Microsoft's director of Applied Science, Dr. Brent Hecht, compared OpenAI's work to the "largest theft of labor in human history," warning it could trigger a "doom loop." OpenAI executive Nick Turley described the situation as an "existential threat to publishers."
The documents surfaced as part of a lawsuit The New York Times filed against OpenAI and Microsoft back in 2023. Only select snippets of the internal communications became public, stripped of surrounding context.
Beyond the quotes, the filings also lay out the technical side of how the data got collected. According to the report, OpenAI and its partners bypassed website paywalls, built training datasets from millions of documents, and stripped copyright notices from that material.
Several similar lawsuits have already ended in favor of AI companies, though judges have pointed out that those rulings stem from the lack of settled law around generative models.
The New York Times case is seen as central to determining whether AI training falls under "fair use," the legal doctrine that allows copyrighted material to be used without permission in cases like parody or journalism.
Microsoft, for its part, distanced itself from its own employees' comments:
Microsoft's position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers' journalism.
Other exchanges are less dramatic but just as telling. When an employee told OpenAI president Greg Brockman about a new paywall workaround, he reportedly replied "ah nice."
An internal Microsoft document from 2023 stated that "millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions."
Executives also weighed in on the risks to the news industry itself. Hecht described large models as "a product that destroys its supply chain," while an OpenAI engineer wrote to colleagues back in 2023 that "no matter how prominently we show the links, users won't click."
Turley echoed that sentiment, calling AI products "largely substitutive" to journalism.
- Gamer set GPT-6 Astra to play through Fallout 3 – the AI reached the ending in 59 hours and started paranoidally checking saves
- AI GPT-6 Astra mined a diamond in Minecraft overnight by directly controlling a computer
- GPT-6 Astra was trained on one hundred thousand Nvidia video cards – Jensen Huang is ready to allocate 400 thousand video cards for the next model