{"version":1,"type":"story","url":"https://digestai.news/story/openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books","json":"https://digestai.news/story/openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books.json","markdown":"https://digestai.news/story/openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books.md","slug":"openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books","headline":"OpenAI and Microsoft allegedly knew AI training used copyrighted books","summary":"Newly unsealed court filings in the Authors Guild lawsuit reveal internal comments suggesting OpenAI and Microsoft were aware that their models were trained on copyrighted material. An August 2019 OpenAI note bluntly states, “We trained GPT-3 on pirated stuff! No sharing that!” and a May 2020 testimony by OpenAI policy director Jack Clark warned that “the better we do on GPT‑X, the more worried genre fiction authors will become about us substituting for them on Amazon.”\n\nThe documents also show that hiring of Tarun Gogineni in 2022 was meant to improve model writing quality with the expressed goal of “creating a machine that would supplant human authors.” Gogineni described authors’ complaints about stolen datasets as “acceptable economic disruption” and predicted “the death of the reader.” Microsoft’s awareness of OpenAI’s use of LibGen dates back to April 2019, and a 2022 memo from VP of Research Bob McGraw discussed removing LibGen files. The filings come as the case moves toward summary judgment.","keyPoints":["OpenAI’s August 2019 note admitted training GPT‑3 on pirated material.","Jack Clark warned in May 2020 that AI could replace genre‑fiction authors on Amazon.","Microsoft knew of OpenAI’s LibGen usage as early as April 2019 and planned its removal in 2022."],"whyItMatters":"The allegations could reshape legal expectations for AI training data, impacting publishing, media, and future AI liability across the industry.","category":{"slug":"policy","name":"Policy & Regulation","url":"https://digestai.news/category/policy"},"entities":{"companies":["OpenAI","Microsoft","Authors Guild"],"models":["GPT-3","GPT-X"],"people":["Jack Clark","Tarun Gogineni","Bob McGraw"]},"firstPublishedAt":"2026-09-20T17:00:00Z","updatedAt":"2026-09-20T17:00:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"publishersweekly.com","title":"Unsealed Files Show Open AI/Microsoft Knew Copying Was Illegal and Could Hurt Authors","url":"https://publishersweekly.com/pw/by-topic/digital/copyright/article/101300-unsealed-files-show-open-ai-microsoft-knew-copying-was-illegal-and-could-hurt-authors.html","publishedAt":"2026-09-20T17:00:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"OpenAI and Microsoft allegedly knew AI training used copyrighted books\", 20 September 2026, https://digestai.news/story/openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books","publisher":"Digest AI","title":"OpenAI and Microsoft allegedly knew AI training used copyrighted books","datePublished":"2026-09-20T17:00:00Z","url":"https://digestai.news/story/openai-and-microsoft-allegedly-knew-ai-training-used-copyrighted-books"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}