OpenAI and Microsoft asked a federal judge on September 4 to end, without a trial, the copyright case The New York Times filed against them in December 2023 -- and the number at the center of their argument is small enough to state in one sentence. Sampling 20 million ChatGPT conversation logs turned over in discovery, OpenAI's own expert found 24 instances of verbatim reproduction of a plaintiff's article, the longest running 29 and 43 words -- a reproduction rate of 0.00012%. The publishers on the other side of the same filing are working from a very different number: they say 10.8 million of their works were copied, not at the moment a chatbot answers, but earlier, when the model was built.
The motions land in In re: OpenAI, Inc., Copyright Infringement Litigation (MDL No. 3143), the multi-district litigation docket in the Southern District of New York that consolidates the Times' suit with more than a dozen others -- Daily News, Ziff Davis, the Center for Investigative Reporting, The Intercept -- all before U.S. District Judge Sidney H. Stein. It's the same docket the Justice Department filed a statement of interest in two days earlier, arguing training itself is transformative fair use; the Friday deadline that piece flagged was this filing.
How the case reached a summary-judgment fight
- Dec 27, 2023 — The Times sues OpenAI and Microsoft in the Southern District of New York.
- Mar 26, 2025 — Judge Stein denies most motions to dismiss, letting the core copying claims proceed.
- 2025 — The Daily News, Ziff Davis, CIR, and Intercept suits are folded into the same docket as MDL No. 3143.
- Sep 2, 2026 — The DOJ files a statement of interest backing OpenAI's fair-use defense.
- Sep 4, 2026 — OpenAI, Microsoft, and the publisher plaintiffs each file for summary judgment.
What OpenAI's own filing says it found
24 (verbatim passages OpenAI's own expert found across 20 million sampled ChatGPT conversations) is the entire evidentiary base for OpenAI's argument that the product itself doesn't infringe, whatever happened during training. Its motion argues that number is too small to support a claim that ChatGPT's outputs are the problem, whatever a court eventually decides about training itself. The company also leans on timing: it says its ChatGPT-User crawler agent was disclosed in March 2023, more than a year before most plaintiff publishers moved to block it via robots.txt in April 2024, and argues the crawling before that point was impliedly licensed. On the separate DMCA claim, OpenAI's filing says 92.7% of the outputs publishers flagged as unauthorized reproductions have an obvious, traceable source rather than evidence of copyright-management-information stripping.
What the publishers say a court should still weigh
The plaintiffs' motion starts from a different place entirely: not what a chatbot's answer looks like today, but what went into building it. Their filing asserts 10.8 million works were copied across four stages they describe separately -- acquisition, training, "grounding" (retrieval at answer time), and output -- of which OpenAI's verbatim-count defense above addresses only the last. (The Times' own count inside that combined figure is more specific: 6,030,928 individual articles, including a New York Times Annotated Corpus of 1.8 million pieces the paper says OpenAI used without a license. The remaining plaintiffs' works make up the rest of the 10.8 million.) The publishers also point to conduct the reproduction-rate framing doesn't capture: custom GPTs built on the models -- named in the filing as News Summarizer, Ace, and NYTimesGPT, plus a separate paywall-bypass tool -- and a crawl-to-referral ratio they calculate at 1,500 to 1 by June 2025, meaning OpenAI's crawlers hit their sites roughly 1,500 times for every reader ChatGPT sent back.
What each headline number in this case actually covers
- 10.8M · all five plaintiff groups
- Works the publisher coalition says were copied
Includes: Every article the five plaintiff groups assert was used at any of the four alleged infringement stages
Excludes: Any court finding -- this is the plaintiffs' own claimed count, not an adjudicated one - 6,030,928 · the Times alone
- The Times' own asserted article count inside that total
Includes: Times-only articles, including its 1.8M-piece Annotated Corpus
Excludes: Daily News, Ziff Davis, CIR, and Intercept works - 24 · OpenAI's discovery sample
- Verbatim passages OpenAI's expert found
Includes: Output-stage reproduction only, in a 20-million-conversation sample from discovery
Excludes: Any copying alleged at the acquisition or training stage - 0.00012% · computed
- Reproduction rate implied by 24 in 20 million
Includes: The instances-per-sampled-conversation rate, expressed as a percentage
Excludes: Any weighting for how often a plaintiff's specific work appeared in the sample
OpenAI's brief adds a market-harm argument the copyright statute treats as its own fair-use factor: the Times' digital advertising revenue rose 20.7% to $114 million in the second quarter of 2026, and subscriptions passed 12 million on the way to a stated target of 15 million by 2027 -- numbers OpenAI cites as evidence the paper isn't losing readers or revenue to ChatGPT, whatever the training-stage copying amounted to. The Times' suit was never built on a claim that it's losing money today, though; it argues AI-generated answers substitute for the visit itself, a harm that would show up as suppressed growth rather than an outright decline -- a claim a revenue number rising is unable to disprove or confirm from the outside.
What a ruling either way actually changes
- A ruling that training itself is fair use would end their strongest claim without a jury ever seeing the evidence compiled over two years of discovery.
- Summary judgment on the core question would resolve the single largest legal exposure either company carries from its training data -- in this case, and as precedent for others.
- A docket-wide fair-use finding here would settle the core legal question before any other case gets this far, shifting their leverage toward a licensing deal instead of a lawsuit.
- His March 2025 order already found the Times' core copying claims strong enough to survive a motion to dismiss -- summary judgment is a different, more fact-dependent test than that was.
Twenty-four sentences and ten million works are both real figures in the same case file -- they just aren't measuring the same thing.
Both sides asked for oral argument; as of this filing, Judge Stein hasn't scheduled one. The two outcomes on record elsewhere in AI copyright litigation sit at opposite ends of what could happen next: in a separate authors' case, Judge Vince Chhabria granted Meta summary judgment on training-stage fair use in June 2025, largely because the plaintiffs there couldn't show market dilution -- the exact question OpenAI's revenue numbers above are aimed at. Anthropic chose not to test that question in court at all; it settled a related books-piracy claim for $1.5 billion, a deal a federal judge approved on July 20, 2026. A finding that ends the case, or a number large enough to make every other AI company's general counsel read this docket closely -- that's the range Stein is now working inside.
- OpenAI and Microsoft moved for summary judgment September 4 in the consolidated NYT-led copyright docket.
- OpenAI's expert found 24 verbatim passages sampling 20 million ChatGPT logs -- a 0.00012% rate.
- Publishers count 10.8 million works they say were copied across acquisition, training, and output stages.
- The DOJ's September 2 fair-use brief already sits in the same case record.
- Caveat: a low output-reproduction rate doesn't resolve whether training itself infringed -- that's still Judge Stein's call.