Economic Analysis and Competition Policy Research

Home   •   About   •   Analytics   •   Videos

Unless Courts Get the Copyright Analogy Right, AI Is Positioned to Destroy Newspapers

Newspapers already survived one extraction cycle by dominant tech platforms. Can they afford another?

Local news in America has been dying for fifteen years, and national outlets are only barely doing better. Facebook and Google spent the 2010s and early 2020s absorbing the digital advertising revenue that used to fund local newsrooms. Ad dollars followed eyeballs because clickbait captures more attention than deeply researched long form writing, to platforms, and those platforms captured the value of news content—headlines, snippets, links driving engagement—without paying newsrooms anything close to what that content was worth to their own engagement metrics.

The result is well documented: newsroom employment cut by roughly half since 2008, more than two thousand local newspapers closed, and a genuine crisis of “news deserts“—counties with no local paper at all, tied by researchers to lower voter turnout, weaker government accountability, and even higher municipal borrowing costs. Meta’s own decisions to deprioritize news content in the Facebook feed, after years of encouraging publishers to build their distribution strategy around the platform, accelerated the collapse for outlets that had nowhere else to turn.

Generative AI is positioned to do the same thing to what’s left of the news business—only faster, and with less friction. The ad-platform extraction at least sent some traffic back to publishers: a reader clicking a Facebook link still landed on the newspaper’s site, saw its ads, sometimes converted to a subscription. A chatbot that synthesizes the news directly in the answer box closes that loop entirely. As explained on this site by Dilan Alma, there’s no click-through when the AI just tells you what happened. Search referral traffic to publishers has already begun measurably declining as AI-generated summaries answer queries directly in search results, and the mechanism is structurally identical to the one at issue in the copyright suits now working through the courts: the value of the reporting gets extracted into a product that competes with, rather than routes traffic to, the outlet that produced it.

A well-worn defense

AI defendants argue that requiring them to compensate newspapers and other creators for training content—and ruling against fair use—would put the United States behind China in AI development. Whatever one makes of that argument, licensing may add friction to AI development, and some of the damages sought by plaintiffs may prove excessive. But China does not have the independent local-news ecosystem that the United States does. If preserving American technological leadership requires accelerating the collapse of American journalism, we should at least acknowledge the tradeoff we are making. The United States could win the AI race and still lose something essential: an independent information system capable of holding American institutions accountable. The argument China will win if we make AI companies pay for inputs assumes that AI capability is the only strategic asset that matters. It isn’t.

This is why the copyright doctrine question in The New York Times Co. v. Microsoft Corp. and OpenAI, Inc. isn’t academic. Every time a court hands down a ruling in one of the AI training-data cases, someone reaches for Google LLC v. Oracle America, Inc. (2021) as the governing analogy. The instinct is understandable—it’s the most recent Supreme Court fair use decision involving a trillion-dollar tech platform, it involves code and interoperability, and Google won. For AI defendants, that combination is irresistible as a rhetorical shield. But the analogy doesn’t hold here, and if courts import Oracle’s outcome without its reasoning, the fair-use factor that’s actually built to ask “does this harm the market for the original” gets steamrolled before it can do its job—in a market that was already cut in half by the last extraction cycle. Local newsrooms, which operate on thinner margins and less brand loyalty than the Times, are the least equipped to survive a second wave of value extraction dressed up as fair use.

Following the precedent

Oracle sued Google over Android’s use of 37 Java API packages—specifically, the declaring code (the method headers and organizational structure programmers use to call pre-written functions) rather than the implementing code (the actual instructions that do the work). Google had written its own implementing code from scratch; what it reused was the API’s naming and organizational structure, because that structure was what millions of Java programmers already knew. The Supreme Court, in an opinion by Justice Breyer, assumed without deciding that the API declaring code was copyrightable, and held that Google’s use was nonetheless fair use as a matter of law.

Three things about the Court’s reasoning matter for the comparison to come:

First, the material copied was a system of interoperability, not expressive content. The Court repeatedly stressed that declaring code is “different” from other computer code because it’s bound up with uncopyrightable ideas and functional necessity—it’s the “method[] of operating” a system, closer to the buttons on a QWERTY keyboard than to a novel. Justice Breyer’s opinion draws directly on the merger doctrine and the idea/expression dichotomy: when there’s essentially one way to name a function that does a particular job in a way third-party programmers can use, the “expression” contained in that name is thin to the point of vanishing.

Second, the purpose of the copying was interoperability and reimplementation, not substitution. Google used the declaring code so that the millions of programmers who already knew Java could write for Android without relearning a new system. The Court treated this as “a new platform” that expanded—rather than substituted for—the market Oracle originally served with Java SE, which was desktop and enterprise computing, not mobile.

Third, and most important for what follows: the Court’s market-harm analysis found that Android did not act as a market substitute for the Java platform Oracle actually sold, and that Oracle in fact benefited from Java’s increased ubiquity even as it lost the Android licensing deal it wanted. The fourth fair use factor—the effect on the market for the copyrighted work—cut for Google because Sun/Oracle’s own licensing business model contemplated exactly this kind of platform-to-platform reuse, and the record did not show Android cannibalizing Java SE’s actual market.

Oracle v. Google, in short, is a case about (1) functional, non-expressive interoperability code, (2) reused for a new purpose, (3) in a market the copyright owner wasn’t actually competing in with the copied product. Not one of those conditions is satisfied in the news-copyright AI cases.

An entirely different animal

The New York Times v. Microsoft/OpenAI turns on a different kind of copying, for a different purpose, aimed at a market in which the plaintiff obviously does compete.

The material is thickly expressive, not functional. The Times’s complaint centers on millions of its articles—investigative reporting, essays, criticism, feature writing—ingested into training corpora and, the complaint alleges, capable of being reproduced by ChatGPT in outputs that are sometimes verbatim or near-verbatim. This is nothing like an API’s declaring code. There is no merger-doctrine argument that says there’s only one way to write a Times investigative feature; journalism is exactly the kind of expressive work around which copyright’s core protections were built. The functional/expressive line that did nearly all the work in Breyer’s opinion runs the opposite direction here.

The purpose is training a substitute product, not achieving interoperability with an existing standard. Google reused Java’s API structure so third parties could keep using skills they already had, on a new platform that did not replace the old one. OpenAI’s use of Times content, by contrast, trains a model that competes directly with the Times for the attention of readers seeking news and analysis—a chatbot that can summarize the day’s events, answer questions about them, and (per the complaint) sometimes reproduce the reporting itself, without the reader ever clicking through to the Times’s website. That is substitutional in exactly the way Android was found not to be. The Times doesn’t need OpenAI to reimplement its API to reach a new audience; OpenAI needs the Times’s reporting to make its product useful, and having gotten the benefit, routes attention away from the source rather than toward it.

The market-harm analysis also runs the other way. Oracle’s licensing business model contemplated derivative platform use; Sun had a history of encouraging just this kind of reuse before its litigation strategy changed. The Times’s business model is the diametric opposite: subscriptions and advertising, both of which depend on readers visiting the Times’s own product rather than getting the substance of its reporting mediated through a chatbot. If a generative AI tool can answer “What did the Times report about Trump’s press conference” accurately enough that the reader never needs to see the original, the third factor—market substitution—points toward infringement in a way it never could for Android and Java SE, because there was no evidence Android displaced Java SE licensing revenue, and there is very direct concern that AI summarization displaces subscription and pageview revenue for news publishers.

The scale and mechanism of copying also differ. Google reused a defined set of 37 API packages, disclosed and specific. Training-data cases involve ingestion of enormous, largely undisclosed corpora, often including outputs that can reproduce close paraphrases or verbatim excerpts of the underlying works—the “memorization” problem that has become a live discovery fight in nearly every one of these suits, including Bartz v. Anthropic and Thomson Reuters v. Ross Intelligence. That’s a different fact pattern than “we rewrote the implementation and kept your naming conventions.”

Put simply: Oracle v. Google is a case where the Court found thin/functional expression, transformative new-platform purpose, and no market substitution. The Times case, and most of the newsroom-plaintiff AI suits, present the opposite of all three. Citing Oracle as controlling authority in the news-copyright context is not applying precedent—it is borrowing the outcome of a case whose reasoning cuts against the party invoking it.

Shishene Jing is a former antitrust attorney with the Federal Trade Commission Bureau of Competition Technology Enforcement Unit, where she worked on an investigation into the artificial intelligence industry. She clerked for Judge Jed S. Rakoff on the Southern District of New York and Chief Judge Robert A. Katzmann on the Second Circuit.

Share this article:
Share this article:
Facebook
Twitter
LinkedIn

Subscribe now to get email updates about The Sling

Related Articles

© 2026 The Sling