MikhbarMIKHBAR
Artificial Intelligence

Microsoft filings call AI scraping ‘largest theft’

Newly unsealed court filings in The New York Times’ copyright lawsuit reveal internal Microsoft and OpenAI warnings about the effects of scraping news content for AI training. The documents also describe alleged efforts to bypass paywalls and remove copyright notices from datasets.

Microsoft filings call AI scraping ‘largest theft’

Internal warnings exposed in copyright case

New information unsealed in The New York Times’ copyright lawsuit against OpenAI and Microsoft shows that senior employees at both companies privately described the risks of using news content to train and power artificial intelligence products. The filings were submitted by news plaintiffs seeking summary judgment, and much of the newly disclosed material comes from The Times’ own brief rather than the underlying exhibits, which remain sealed.

In a January 2023 internal memo, Microsoft Director of Applied Science Brent Hecht reportedly called the scraping of news content “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” Another document allegedly said that widespread scraping made “a complete mockery of the idea of ‘fair use,’” a legal doctrine that can permit some unlicensed uses of copyrighted works.

Executives warned of a threat to publishers

The filings also describe warnings that generative AI products could damage the economic foundations of the news industry. An internal Microsoft presentation written by Hecht in January 2024 characterized declining traffic from AI products as a “doom loop” that could hurt both publishers and the performance of AI models by weakening the web’s supply of content.

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers,” the document said, according to the filing. It added that this was the situation Microsoft had created for its large-language-model business and its “content supply chain.” A separate Microsoft document identified a “real risk” that generative AI could significantly disrupt the employment of the people who produced the data used to train foundation models.

At OpenAI, Nick Turley, the company’s head of ChatGPT, allegedly wrote that publishers faced an “existential threat” from products such as the chatbot. The filing says Turley described those products as “largely substitutive” and predicted that they would become more substitutive as they improved. OpenAI President Greg Brockman was also quoted as describing the models as “excellent at news.”

Traffic data supports substitution claims

The Times’ filing points to Microsoft data showing that Copilot’s answer engine reduced click-through rates to The New York Times’ domain by as much as 93% compared with traditional Bing search. Ars Technica reported that Microsoft recorded declines of 83% to 93% for some news plaintiffs and 51% to 94% for others.

Microsoft CEO Satya Nadella also testified in a deposition earlier this year that paywalled material should be licensed by anyone using it for grounding or training. He said that if he had known OpenAI had scraped and trained on paywalled information, he would have invoked Microsoft’s right to require OpenAI to retrain its models.

Nadella further acknowledged that chatbots can substitute for news websites by providing information directly on an AI platform rather than sending users to the underlying source. OpenAI employees reportedly expressed similar concerns. One software engineer wrote that users would not click links if the chatbot already supplied the information, while Turley allegedly said there was “no good reason to click” in that situation.

Filings detail alleged scraping practices

The unsealed material describes how the companies allegedly obtained and shared news content. According to The Times’ filing, OpenAI delivered its entire GPT-3 training dataset to Microsoft, which used it to evaluate how OpenAI’s models could be integrated into commercial products. Microsoft also provided training data to OpenAI through initiatives called Project Taxi and Project Mango.

The companies allegedly assembled Project Mango data into a training dataset containing copies of at least 160,903 unique works from news publishers. OpenAI’s mid-training datasets alone reportedly contained more than 91,692 copies of works published by The New York Times, Daily News and the Center for Investigative Reporting. A dataset derived from Common Crawl allegedly included more than 2 million documents from the Times’ website.

The filings further allege that OpenAI employees developed a way to circumvent the Times’ paywall without detection. When researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman allegedly replied, “ah nice.” The documents also describe efforts to remove copyright notices from training data because researchers did not want models to output those notices to users.

Companies dispute plaintiffs’ interpretation

The disclosures are significant because they appear to conflict with parts of OpenAI and Microsoft’s fair-use defenses. The plaintiffs argue that internal statements and traffic data show AI products can substitute for original reporting rather than transform it, while verbatim or near-verbatim outputs may further reduce demand for news sites.

Microsoft defended its products as transformative fair use and said Nadella’s testimony reflected broad observations about changing patterns in how people find and consume information, not conclusions about the copyright issues before the court. A Microsoft spokesperson also said Hecht’s documents represented one employee’s individual perspective, were not legal analysis and did not express the company’s views.

OpenAI and Microsoft did not respond to TechCrunch’s requests for comment, while OpenAI did not immediately respond to Ars Technica. Steven Lieberman, counsel for the New York Daily News and other plaintiffs, said the newly revealed evidence showed that both companies knew their conduct was wrong. The court has not resolved the broader question of whether using copyrighted material to train AI models is lawful.

Sources

  • TechCrunchMicrosoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
  • Ars TechnicaMicrosoft exec called AI scraping the “largest theft of labor in human history”