Sun, 20 Sept 2026
In the News

Microsoft executive brands AI data harvesting as unprecedented labor theft

UnbarNewsUpdated 17 Sept 2026· 2 min read

Unsealed court documents reveal Microsoft officials called OpenAI's content scraping a historic theft of labor, raising fresh legal and ethical concerns.

Microsoft executive brands AI data harvesting as unprecedented labor theft

New court filings that were recently unsealed indicate a senior Microsoft leader described the practice of training large language models on copyrighted material as "the largest theft of labor in human history." The internal memorandum, cited by TechCrunch, specifically targets OpenAI’s approach to building its datasets.

According to the filings, Microsoft executives warned that both Microsoft and OpenAI had systematically scraped pay‑walled articles from The Times and other publishers, converting the text into massive training corpora. The documents say the companies were aware that such extraction could "gut" the business models of news outlets that rely on subscription revenue.

The memo also noted that the partnership between Microsoft and OpenAI, which includes a multibillion‑dollar investment and exclusive cloud hosting, could be jeopardized if the alleged data‑theft claims lead to litigation. Internal emails reportedly urged caution, suggesting that the fallout could damage Microsoft’s reputation and expose it to liability under emerging copyright frameworks.

Background: The controversy sits amid a wave of lawsuits filed by news organizations, authors, and artists accusing AI developers of violating copyright by using their works without permission. In the United States, the legal question hinges on whether the mass ingestion of publicly available text constitutes “fair use” or an infringement of the creators’ exclusive rights. Recent cases, such as the Authors Guild v. OpenAI and News Corp v. Google, have underscored the tension between rapid AI development and the protection of intellectual property. Courts are still grappling with how existing copyright law applies to machine‑learning training data, and regulators in the EU and U.S. are considering new rules that could require explicit licensing.

If the allegations in the unredacted filings hold up, Microsoft could face pressure to renegotiate its data‑sharing arrangements with OpenAI or to implement stricter safeguards for copyrighted content. Industry observers note that the outcome may set a precedent for how tech giants source data for generative AI, potentially reshaping the balance between innovation and the rights of content creators.

The next steps involve a series of hearings where both parties will likely present technical evidence about how the data was collected and processed. Until a court renders a decision, the dispute highlights the broader challenge of aligning AI advancement with legal and ethical standards that protect the labor of journalists and other creators.

This report is based on original reporting by TechCrunch. Read the original source →

#Microsoft#OpenAI#AI ethics#copyright#legal