DigestAI news desk

Cut through the AI noise.

Policy & Regulation2 min read

OpenAI and Microsoft allegedly knew AI training used copyrighted books

Newly unsealed court filings in the Authors Guild lawsuit reveal internal comments suggesting OpenAI and Microsoft were aware that their models were trained on copyrighted material. An August 2019 OpenAI note bluntly states, “We trained GPT-3 on pirated stuff! No sharing that!” and a May 2020 testimony by OpenAI policy director Jack Clark warned that “the better we do on GPT‑X, the more worried…

1 source

Key points

  • OpenAI’s August 2019 note admitted training GPT‑3 on pirated material.
  • Jack Clark warned in May 2020 that AI could replace genre‑fiction authors on Amazon.
  • Microsoft knew of OpenAI’s LibGen usage as early as April 2019 and planned its removal in 2022.

The documents also show that hiring of Tarun Gogineni in 2022 was meant to improve model writing quality with the expressed goal of “creating a machine that would supplant human authors.” Gogineni described authors’ complaints about stolen datasets as “acceptable economic disruption” and predicted “the death of the reader.” Microsoft’s awareness of OpenAI’s use of LibGen dates back to April 2019, and a 2022 memo from VP of Research Bob McGraw discussed removing LibGen files. The filings come as the case moves toward summary judgment.

Full story from publishersweekly.com · via Search: MicrosoftOpen source ↗

Unsealed Files Show Open AI/Microsoft Knew Copying Was Illegal and Could Hurt Authors

publishersweekly.com · 20 September 2026

Loading the full article…

This text was published by publishersweekly.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
OpenAIMicrosoftAuthors GuildGPT-3GPT-XJack ClarkTarun GogineniBob McGraw

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Policy & Regulation

All →

Related stories