Friday, 18 September 2026NewsWorldBusinessTech
Latest

Microsoft Exec Called AI Web Scraping Largest Labor Theft in History

A senior Microsoft executive privately warned that web scraping to train artificial intelligence models amounts to the largest theft of labor in human history, according to court documents unsealed Thursday, September 17, 2026, in a Manhattan federal copyright lawsuit brought by news publishers.

The revelation emerges from an ongoing landmark copyright infringement battle in Manhattan Federal Court, where major news organizations—including The New York Times, the New York Daily News, and the Orange County Register—are locked in legal combat with tech giants Microsoft and OpenAI. The newly unredacted filings expose sharp internal friction within Microsoft over the data collection practices fueling the generative AI boom, directly challenging the industry’s reliance on the fair use doctrine.

Internal Warnings and the Doom Loop at Microsoft

Brent Hecht, Microsoft’s Director of Applied Science, wrote an internal memo in January 2023 predicting that the industry-wide practice of scraping the open web would provoke intense public backlash. According to the court records, Hecht warned that millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions, noting further that it could constitute the largest theft of labor in human history.

Hecht’s internal paper trail extended into a follow-up presentation in January 2024. In that deck, he cautioned that Microsoft’s own Copilot answer engine was cannibalizing the referral traffic it was supposed to build upon. Internal Microsoft data showed click-through rates to New York Times pages plummeting by as much as 93% compared with traditional Bing search results. Hecht labeled this dynamic a doom loop, explaining that it is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’

Disputing Fair Use and Documenting Scale

Tech companies have long defended large-scale web scraping as a transformative use permitted under copyright law. However, legal experts note that courts weigh both the transformative nature of a product and its economic impact on the market for original works, as well as the good faith of the copying parties. The newly unsealed documents directly target both pillars of that defense.

Steven Lieberman, an attorney leading the Daily News’ defense, argued that the disclosures prove the defendants understood the unfairness of their behavior while simultaneously shielding the documents behind confidentiality requests. The filings also reveal the sheer volume of material ingested: OpenAI’s mid-training datasets contain more than 91,692 separate copies of articles from The Times, the New York Daily News, and the Center for Investigative Reporting, while a separate Common Crawl-derived dataset held over 2 million documents pulled directly from nytimes.com.

Bypassing Paywalls and Horse-Trading Content

Beyond standard training sets, the unsealed material details technical methods used to acquire restricted material. In one chat exchange cited in the filings, OpenAI researcher Nick Ryder explained a method for bypassing the Times’ paywall to CEO Greg Brockman. Brockman responded with ah nice, a casual two-word reply that plaintiffs’ lawyers intend to leverage to challenge claims of good-faith operation.

Microsoft Exec Called AI Scraping the Largest Theft of Labor in History
Photo: startupfortune.com

Furthermore, the motion alleges that the tech companies copied billions of web pages not just for training, but for horse trading-like deals, selling the harvested content to each other rather than licensing it from news organizations. Specific figures unsealed in the filings point to more than 3.9 million copies taken from The Times and 7.3 million copies taken from the Orange County Register, the New York Daily News, and sister publications including the Chicago Tribune, Orlando Sentinel, Denver Post, and Mercury News.

Corporate Response and What Comes Next

Microsoft has pushed back against the characterization of Hecht’s statements. A company spokeswoman stated that his remarks reflected one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views, maintaining that Copilot is not a substitute for professional journalism and that its transformative uses align with copyright law.

Microsoft Executive Exposed As Internal Filings Reveal AI Theft Scandal

The newspapers’ pending motion seeks summary judgment on acquisition-, training-, and distribution-based infringement. Plaintiffs are also pursuing sanctions against OpenAI for allegedly destroying evidence and concealing its capability to locate stolen news articles within its framework. The consolidated litigation—which also includes claims from The Authors Guild and best-selling writers—proceeds toward further discovery, expert testimony, and trial proceedings in Manhattan Federal Court.

Top 5 AI News: Microsoft exec called AI scraping 'the largest theft of labo 🔥