An active legal and regulatory clash over whether large AI models may lawfully ingest and reproduce copyrighted news and journalism has framed recent debates between publishers and tech firms throughout 2023–2026.
The structural rules at issue derive from the Copyright Act of 1976 (17 U.S.C.), the fair-use doctrine codified in 17 U.S.C. §107, and the Digital Millennium Copyright Act of 1998, plus pivotal court precedents such as Campbell v. Acuff‑Rose Music, Inc. (U.S. Supreme Court, March 8, 1994) and the Authors Guild v.
Unredacted excerpts filed in The New York Times’ copyright suit against OpenAI and Microsoft show executives at the companies privately describing large-scale AI scraping of news content in stark terms.
The newly revealed material includes a Microsoft executive calling the scraping effort “the largest theft of labor in human history,” and it quotes OpenAI leaders warning that their models posed an “existential threat” to publishers and journalists, details TechCrunch extracted from the filings (per techcrunch).
The filings describe specific practices plaintiffs say undercut news revenue: systematic paywall bypassing, mass scraping of articles and removal of copyright notices from scraped content (per techcrunch).
The Times frames those facts as direct evidence that OpenAI and Microsoft relied on copyrighted journalism to train models and to displace publishers’ value — a legal strategy meant to weaken OpenAI’s fair-use defense (per techcrunch).
OpenAI and Microsoft have defended their work in other public statements, but the unredacted filings shift the factual record toward internal acknowledgments reported by The New York Times and summarized by TechCrunch (per techcrunch).
Why now: the suit has moved into phases where previously redacted internal communications are being disclosed to the court, and those communications now provide contemporaneous characterizations of how company leaders viewed massive news-data ingestion (per techcrunch).
The filings do not quantify how many articles were scraped or give detailed timelines of specific scraping operations in the TechCrunch excerpt, so the court will likely test whether these internal descriptions meet the legal thresholds for copying and market harm asserted by The New York Times (per techcrunch).