DigestAI news desk

Cut through the AI noise.

Policy & Regulationupdated 5 min read

Microsoft scientist calls AI training the largest theft of labor

Dr. Brent Hecht, Microsoft’s Director of Applied Science, described AI model training as “the largest theft of labor in human history” in a filing for the consolidated OpenAI copyright lawsuit before Judge Sidney H. Stein. Hecht’s comments are presented as his individual view, not a legal analysis, and Microsoft later said the remarks reflect only one employee’s perspective.

2 sources

Key points

  • Brent Hecht called AI training “the largest theft of labor” in the OpenAI copyright case filing.
  • Click‑through rates fell 87‑93% (Times), 83‑91% (Daily News) and 51‑94% (Ziff Davis) when users saw Copilot vs Bing.
  • OpenAI’s training data include over 91,692 copies of plaintiffs’ works, with 6,552‑66,780 copies per source and 160,903 in the Mango set.

The plaintiffs – including The New York Times, the Daily News Group, the Center for Investigative Reporting and Ziff Davis – cite internal Microsoft data showing click‑through rates dropping dramatically when users see Copilot results versus traditional Bing: 87%‑93% for the Times sites, 83%‑91% for Daily News, and 51%‑94% for Ziff Davis. The brief also lists the volume of copyrighted material in OpenAI’s training sets, noting more than 91,692 copies of plaintiffs’ works overall, with WebText2 containing at least 6,552 Times articles, 18,609 Daily News pieces and 66,780 Ziff Davis items, plus a Mango dataset of at least 160,903 unique plaintiff works and a New York Times Annotated Corpus of over 1.8 million articles.

The United States filed a statement of interest arguing that training on copyrighted text is fair use, while OpenAI calls the lawsuit an “undeserved payday.” The case could determine whether large‑scale data scraping for foundation models constitutes copyright infringement, shaping future AI development and publisher rights.

The story so far

4 episodes →
  1. Microsoft scientist calls AI training the largest theft of laborthis story
Full story fromthenextweb.com · by Ana Maria Constantin · via Search: MicrosoftOpen source ↗

A Microsoft scientist called AI training the largest theft of labor

thenextweb.com · 21 September 2026

A 92-page brief unsealed on Thursday opens by quoting the same Microsoft employee twice. Dr Brent Hecht, the company’s Director of Applied Science, wrote that people would come to see large models hoovering up their work as “an astonishing theft of unprecedented proportions”. In a second document he called it perhaps the “largest theft of labor in human history”.

Hecht is a research scientist. He is not a board member and he is not the chief executive, which matters as the phrase travels.

The News Plaintiffs’ combined summary judgment brief sits in the consolidated OpenAI copyright litigation before Judge Sidney H. Stein. Redactions covered most of its damaging passages until Thursday.

Jason Kint of the trade body Digital Content Next spotted the unredacted version. George Hammond and Stephen Morris reported it that day for the Financial Times, Scott Nover and Gerrit De Vynck for the Washington Post, Jason Koebler for 404 Media, and Karen Weise and Mike Isaac for The New York Times, which is also a plaintiff.

The quotations below come from the filing itself.

The plaintiffs are not only The New York Times. They include the Daily News group and the Center for Investigative Reporting. They also include Ziff Davis, whose titles run to CNET, ZDNET, PCMag and Mashable. The people suing are, in other words, this publication’s direct peers.

The doom loop, in Microsoft’s own numbers

The brief quotes a Microsoft document describing what the company had built. Its AI content strategy, the document says, “has started a ‘doom loop’” that will damage both its own models and the entire web. The document calls the position unusual. An end product now threatens the economic foundations of its own suppliers.

The figures behind it are Microsoft’s own. The brief cites company data on click-through rates falling 87% to 93% for the Times’s websites, 83% to 91% for the Daily News group’s, and 51% to 94% for Ziff Davis’s. The comparison is Copilot against traditional Bing search.

Hecht wrote that document soon after this lawsuit was filed. The systems work, he explained, by replicating the patterns in their training content. There is no way of passing economic value back down the supply chain. That, he wrote, “necessarily threatens the economic stability of those who create the content”. A ruling for his own employer, he added, would arguably “make a complete mockery of the idea of ‘fair use’”.

Other Microsoft material in the brief goes further. One document warns of a “real risk” that the technology could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained”. Another puts it in eight words: “LLMs are a product that destroys its supply chain.”

‘Users won’t click’

The sharpest line in the filing belongs to an OpenAI software engineer the brief does not name. “No matter how prominently we show the links, users won’t click.”

Nick Turley, who runs ChatGPT, wrote that publishers face an “existential threat”. The products, he said, are “largely substitutive, period” and will get more so as they improve. Jack Clark, OpenAI’s policy director, put it another way. The company’s work would increasingly mean building systems that substitute for the labour of the people who define a society’s culture.

Satya Nadella agreed under oath that chatbots had substituted for going to the underlying source. Internal OpenAI documents call ChatGPT “the modern newsstand”.

What they knew they were training on

OpenAI’s VP of Research gave the brief a sentence its lawyers will enjoy. “We train our networks to memorize the training data. That’s their objective.” By November 2019 the company was worrying internally about regenerating copyrighted works.

The brief puts numbers to the copying. OpenAI’s mid-training datasets hold more than 91,692 copies of the plaintiffs’ works. WebText2 held at least 6,552 from the Times, 18,609 from the Daily News group and 66,780 from Ziff Davis.

OpenAI separately obtained the New York Times Annotated Corpus, over 1.8 million articles from 1987 to 2007, under a licence limited to non-commercial research.

Around 2017, Greg Brockman wrote that he was “deeply motivated by the gazillions” he hoped to make from commercialising OpenAI’s technology.

Then there is the exchange the coverage keeps returning to. Nick Ryder told Brockman about “a hack to get around nytimes paywall”. Brockman replied: “ah nice.” OpenAI’s own store lists Custom GPTs named “Bypass Paywall” and “Remove Paywall”. Nadella testified that “anything that is paywalled should be licensed”.

Horse trading

The brief gives a whole section the title “Horse Trading”. Microsoft scraped for Bing, the filing says, then passed that content to OpenAI. OpenAI handed Microsoft the entire GPT-3 training dataset. Two joint initiatives, Taxi and Mango, moved data back, and the Mango dataset alone held at least 160,903 unique plaintiff works.

What the government told the same court

Sixteen days before the unsealing, the United States filed its own statement of interest in the same litigation. It argues that training on copyrighted text is fair use. TNW covered it at the time, and the 20-page filing reads differently now.

Its core argument is that training and output are separate uses. Copying a work to train a model reveals nothing to the public, so it cannot substitute for the original. Market harm counts only where the output is substantially similar to the source. General competition does not qualify.

It also frames licensing as an antitrust problem. Only the biggest technology companies could afford the fees, it argues, which would become “large subsidies for old mainstream media companies”.

What the brief does not prove

This is the plaintiffs’ document. Lawyers building a case chose every quotation in it, and the exhibits underneath are still sealed.

Microsoft has answered the central quotation. Hecht’s words “reflect one employee’s individual perspective, are not a legal analysis”, the company told the Financial Times. It said Nadella’s testimony spoke to broad principles rather than reaching a conclusion on the copyright questions. OpenAI did not respond to the FT. It has previously called the case an attempt at “an undeserved payday at the expense of progress that benefits everyone”.

The defendants have not conceded the legal point either. The brief quotes Microsoft’s own economist expert, Dr Tucker, drawing the distinction their lawyers will rely on. Search engines return a ranked list of links. Grounded models synthesise information from retrieved sources. Whether that synthesis counts as substitution in the sense copyright recognises is the question in front of the judge.

Courts have so far leaned towards the AI companies. Anthropic settled its book piracy case for $1.5bn rather than test it, and publisher suits against Google over Gemini are still running.

Europe is asking the opposite question

Brussels spent this month asking publishers whether Google’s AI opt-out works, on the assumption that a publisher should be able to refuse. Washington has told a court that letting them refuse would be an antitrust harm.

The doom loop has a payroll attached on both sides of the Atlantic. Reach, publisher of the Daily Mirror, cut 220 editorial jobs this month and pointed at AI summaries eating its traffic.

Judge Stein now has both documents. One says the training was transformative and the harm is not the kind copyright recognises. The other is the defendants, in their own words, calling it theft. The summary judgment ruling settles which reading the law accepts.

This text was published by thenextweb.com and written by Ana Maria Constantin. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

2sources
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Policy & Regulation

All →

Related stories