Microsoft executives criticize AI scraping as historic labor theft in filings
Microsoft executives labeled AI scraping as "the largest theft of labor in human history," revealing internal concerns about using paywalled content for AI training. This legal battle with the New Yoโฆ
Microsoft executive called AI scraping โthe largest theft of labor in human history,โ as newly unredacted court filings show the tech giantโs internal warnings about its partnership with OpenAI and the use of paywalled NewโฏYorkโฏTimes content. The documents, released in a lawsuit filed by the Times, reveal that Microsoftโs senior staff had been aware of the legal and ethical risks of using copyrighted text to train large language models.
The issue matters because AI models such as GPTโ4 rely on vast amounts of text from the web, including articles behind paywalls. The Times and other publishers have sued Microsoft and OpenAI for copyright infringement, arguing that the companies scraped millions of words without permission. Microsoft is a major investor and cloud partner of OpenAI, and it supplies the computing power that powers the model. The filings suggest that the partnership was built on the assumption that data scraping would be tolerated, but the companyโs own lawyers flagged the practice as โtheftโ and warned that publishers could be โgutโedโ if they did not grant licenses.
In the unredacted documents, a memo from a senior Microsoft executive describes the โlargest theft of laborโ and outlines a plan to negotiate with publishers. The memo cites the Timesโ 2021 lawsuit, which sought damages of $2โฏbillion, and notes that Microsoft had already begun to explore licensing agreements. The filings also show that Microsoftโs legal team had drafted a letter to the Times offering a settlement, but the Times rejected it. The documents reveal that Microsoft was preparing to defend the use of scraped content in court, arguing that it was part of a broader โopenโsourceโ approach to AI development.
The next steps in the dispute will likely involve court hearings that could force the tech industry to rethink how it sources training data. If the lawsuit succeeds, Microsoft and OpenAI could face significant financial penalties and be required to pay royalties to publishers. Regulators in the United States and Europe may also scrutinise AI training practices, potentially leading to new dataโprotection rules. For the Times, a victory could set a precedent that protects publishersโ rights and limits the ability of AI firms to use copyrighted text without compensation. The outcome will shape how the AI industry balances innovation with respect for intellectualโproperty rights.
Read Full Story at TechCrunch โ

