Microsoft Called AI Scraping ‘the Largest Theft of Labour in Human History.’

Microsoft Called AI Scraping the Largest Theft of Labour in Human History

The New York Times has been in court with OpenAI and Microsoft since December 2023 over the use of its journalism to train generative models. The brief the publishers filed on Thursday, 17 September 2026, quotes Microsoft and OpenAI staff directly on how the training sets were assembled, what the products do to news traffic, and what follows for the writers of the work those models learned from.

What Microsoft’s own papers said

Brent Hecht is Microsoft’s director of applied science. In a January 2023 internal memo, as quoted in the publishers’ brief, he wrote that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions.” He added that the practice could amount to “the largest theft of labour in human history.” The memo itself has not been released in full.

A year later, in a January 2024 presentation also quoted in the filing, Hecht described what Microsoft’s own traffic data was showing. Copilot’s “answer engine,” compared with ordinary Bing search, cut click-through rates to The New York Times domain by as much as 93 percent. He called the pattern a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

The same document, as cited by the publishers, put it this way: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

A separate Microsoft paper, again as quoted, warned of a “real risk” that generative systems could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

Microsoft’s reply, given to Reuters, is that Hecht’s remarks “reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.”

What Nadella said under oath

Satya Nadella, Microsoft’s chief executive, sat for a deposition earlier this year. According to the brief, he said “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training.” He also said that if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

The filing also quotes him on what a chatbot does to a publisher’s page. Talking to one, he agreed, “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

Microsoft told Reuters that Nadella’s testimony does not undercut its defence in the case.

What OpenAI’s own people wrote

Nick Turley, who runs ChatGPT, wrote internally that publishers face an “existential threat” from products that are already “largely substitutive” and “will get more and more substitutive as they get better.”

Greg Brockman, OpenAI’s president, described the models as “excellent at news,” “particularly good at predicting text of news articles,” and “very good at any news task.”

The brief also records a short exchange. OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall.” Brockman replied: “ah nice.”

How large the copies are, on the publishers’ figures

The filing puts numbers on the datasets for the first time in public. OpenAI’s mid-training sets, it says, contain more than 91,692 copies of works published by The New York Times, the Daily News and the Center for Investigative Reporting. A dataset drawn from Common Crawl is said to include more than two million documents from nytimes.com alone.

The companies, the publishers allege, also moved data between themselves. “OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products.” Microsoft, the same passage says, sent training data the other way through projects named Taxi and Mango. Project Mango, on the publishers’ count, holds copies of at least 160,903 unique works from the news plaintiffs.

The brief further alleges that copyright notices were stripped from training text because researchers “wouldn’t want model outputting” those notices to users. Steven Lieberman, counsel for the New York Daily News, said in a statement: “The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.”

Those are the publishers’ figures. The underlying datasets have not been released with the brief.

Where the case stands

The suit is three years old. It sits in the Southern District of New York before Judge Sidney Stein. The publishers’ live claims include direct copyright infringement, vicarious infringement and the removal of copyright-management information. Summary judgment papers went in at the start of September. A decision on those motions has not been issued.

On 1 September the Trump administration filed a brief in the same case in support of the position that training on copyrighted work can be fair use.

Fair use is the American rule that sometimes allows copyrighted work to be used without a licence. One test is whether the new use takes the market for the old one. The publishers say these quotes go to that test. OpenAI and Microsoft say the training is transformative. Judge Stein has not ruled.

Germany has already tried the training question on different facts. In July the Munich Regional Court held that Suno infringed GEMA repertoire by training without a licence, including for training that took place in the United States. The same court found against OpenAI in 2025 over memorised song lyrics. Both judgments are on appeal.

Author: Andy Samu

See Also:

What Does NVIDIA’s Hugging Face Bid Mean for Europe’s Open-Source Bet? – MRKT3.0

Did OpenAI Build a Hacker, Then Lose It for a Week? – MRKT3.0

Share this article

Latest news

Subscribe to our newsletter

More News