Newly unsealed court documents have revealed concerns inside Microsoft and OpenAI about the impact of artificial intelligence on journalists, publishers and other workers whose content is used to train AI models.
The documents were released as part of the ongoing copyright lawsuit brought by The New York Times against Microsoft and OpenAI. The case focuses on whether the companies unlawfully used copyrighted articles and other material to train their AI systems.
One of the strongest statements came from Brent Hecht, a Microsoft director of applied science. In an internal document, Hecht described the large-scale use of online content for AI training as the “largest theft of labor in human history.” In another document, he called it an “astonishing theft of unprecedented proportions.”
Microsoft has said that Hecht’s comments represented his personal views and were not the company’s legal position.
Internal Concerns Over AI Training
AI companies train large language models on huge amounts of data collected from the internet. This data can include news articles, books, websites and other forms of written work.
Publishers have argued that their content was used without permission to build AI systems that can then provide information directly to users. They say this could reduce the number of people visiting the original websites and hurt the revenue that supports journalism.
The newly unsealed documents provide more details about concerns raised inside Microsoft and OpenAI over this issue. Microsoft documents reportedly warned that generative AI could “significantly disrupt the employment” of people who created the material used to train AI models.
The documents also describe what Microsoft called a possible “doom loop.” The concern was that AI systems could reduce traffic and revenue for news publishers, leading to less investment in journalism. That could eventually mean less high-quality material is available for future AI models to learn from.
OpenAI Executive Called AI Products “Substitutive”
The court documents also include comments from OpenAI executives about the effect of AI on news organizations.
Nick Turley, who leads the ChatGPT team, reportedly wrote that OpenAI’s products were “largely substitutive” for publishers. He also said that this effect would increase as AI systems became more capable.
Microsoft CEO Satya Nadella made a similar point during a deposition. According to the court filing, Nadella acknowledged that people could receive information directly through chatbots instead of visiting the original publisher’s website.
OpenAI President Greg Brockman had also written internally that AI models were particularly good at handling news-related tasks. In one 2020 message, he noted that the model could accurately predict text from a New York Times article.
These statements are now being used by the publishers to support their argument that Microsoft and OpenAI understood that AI products could compete directly with news organizations.
Microsoft Data Shows Drop In Publisher Traffic
The court filing also points to Microsoft’s own data on how AI-generated answers affected traffic to news websites.
According to the filing, Microsoft’s data showed an 83% to 93% decline in click-through rates for content from The New York Times and Daily News in some comparisons. Data cited for Ziff Davis websites showed declines ranging from 51% to 94%.
Publishers argue that these figures demonstrate the financial risk of AI systems answering users’ questions without requiring them to visit the original source.
Microsoft’s internal documents reportedly described this as a problem for both publishers and AI companies because the technology could weaken the businesses producing the content used to train AI models.
Questions Over Paywalled Content
The unsealed material also raises questions about how some copyrighted content was obtained.
The documents describe allegations that OpenAI employees looked for ways to access content behind The New York Times’ paywall. According to the filing, OpenAI researcher Nick Ryder told Brockman about a way to get around the Times’ paywall, and Brockman responded positively.
The filing also describes projects involving Microsoft and OpenAI that exchanged or collected large amounts of web content for use in AI-related work.
One Microsoft project, known as Project Taxi, provided OpenAI with a large collection of webpages gathered for Bing. Another effort, called Project Mango, involved Microsoft developing and operating a crawler to collect webpage content for OpenAI.
These claims are part of the publishers’ arguments in the lawsuit and have not themselves established that the companies violated copyright law.
Microsoft Defends Its Position
Microsoft has pushed back against the interpretation of the documents. A company spokesperson said Hecht’s statements reflected one employee’s individual perspective and were not a legal analysis or Microsoft’s official position.
Microsoft has also maintained that its use of copyrighted material for AI falls under fair use, a legal principle that allows certain uses of copyrighted works without permission.
The company has argued that AI training is different from simply republishing copyrighted material and that its products do not unlawfully reproduce publishers’ work. OpenAI has similarly defended the use of training data under fair-use principles.
Topics
- Microsoft



