An unredacted version of The New York Times' own summary-judgment filing, unsealed Sept. 17 in the Southern District of New York against OpenAI and Microsoft, quotes a January 2024 internal memo from Brent Hecht, Microsoft's director of applied science, describing the use of publishers' content to train AI models as "an astonishing theft of unprecedented proportions" and, in the same document, "the largest theft of labor in human history." The case is part of a consolidated docket, No. 1:25-md-03143, that folds the Times' own suit (No. 1:23-cv-11195) together with more than a dozen others against OpenAI and Microsoft, all before Judge Sidney H. Stein -- confirmed directly against OpenAI's own amended answer in the underlying docket. The quote is the Times' own evidence, submitted in support of its motion asking Judge Stein to rule on liability before trial -- it is an allegation the plaintiff is making using the defendant's own words, not yet a finding by the court.
Microsoft's response, given to reporters after the filing became public, is that Hecht's memo "reflected one employee's perspective, not company policy." That's the frame worth holding onto through everything below: every internal quote in this filing is somebody's candid assessment at a specific moment, not a company's official position, and the case will ultimately turn on conduct and contracts, not on how bluntly any one employee once described what the conduct amounted to.
The Times isn't relying on Hecht's memo alone. The unsealed filing also quotes Nick Turley, OpenAI's head of ChatGPT, telling colleagues that products built this way "are largely substitutive, period," and that publishers face an "existential threat" -- and a since-departed OpenAI policy director, Jack Clark, warning internally that the company's systems would "increasingly lead to us creating systems that substitute for the labor" of the very creators whose work trained them. On the numbers side, the Times cites internal estimates putting more than 91,692 copies of its own, the Daily News' and the Center for Investigative Reporting's articles inside OpenAI's mid-training datasets, a separate Common Crawl-derived set with more than 2 million documents from nytimes.com alone, and a third dataset -- internally called Project Mango -- holding at least 160,903 unique works. The filing also cites Microsoft's own internal research finding that Times click-through rates fell by as much as 93% when readers got answers from AI-powered search instead of clicking into the article.
The Times' unsealed numbers, and what each one claims to measure
- 91,692 · copies
- Times/Daily News/CIR articles in OpenAI's mid-training data
Includes: The Times' own count, drawn from internal OpenAI records produced in discovery
Excludes: Independent verification -- no third party has audited the dataset itself - 2M+ · documents
- nytimes.com pages in a Common Crawl-derived training set
Includes: Pages scraped by the third-party Common Crawl project, then incorporated into an OpenAI training corpus, per the filing
Excludes: Whether every page was paywalled or full-text at time of scraping - 160,903 · unique works
- "Project Mango" dataset, per the filing
Includes: The Times' characterization of an internal OpenAI dataset name and count surfaced in discovery
Excludes: OpenAI's own description of what Project Mango was built for, which is not quoted in current reporting - 93% · CTR drop
- Times click-through rate from AI-powered search, per Microsoft's own internal research
Includes: Microsoft's own study, cited in the filing, of referral traffic when users get answers via AI search instead of clicking through
Excludes: Whether this figure covers all AI search products or Microsoft's own Copilot/Bing specifically -- current reporting doesn't specify
There's a real tension inside Microsoft's own cited internal material that the filing itself surfaces, whether or not Microsoft intended it to. In the same period Hecht was calling the practice theft internally, CEO Satya Nadella said in a 2026 deposition that "anything that is paywalled should be licensed by anyone who wants to use it" -- a position that, taken at face value, argues against Microsoft's own products having trained on paywalled Times content without a license in the first place. (A company's deposed CEO stating a licensing principle and a company's own director privately calling the practice theft aren't necessarily contradictory -- both can be true if the licensing Nadella describes simply didn't happen for the specific content the Times is suing over. But it does mean Microsoft's public defense and its own executives' recorded statements aren't yet telling a fully consistent story.)
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers.” — internal Microsoft document, January 2024, as quoted in the Times' unsealed filing
This unsealing lands two weeks after a Statement of Interest filed by the Justice Department in the same consolidated docket, which told Judge Stein that training large language models on copyrighted text is legally transformative fair use and warned that ruling otherwise would cede AI dominance to "foreign adversaries." The two filings aren't answering the same question -- the DOJ's brief argues the legal fair-use standard should favor OpenAI regardless of what any employee said internally, while the Times' newly unsealed material goes to a different, earlier question: whether OpenAI and Microsoft's own people understood, at the time, that what they were doing displaced the outlets whose work they were using. A court can find training legally transformative and still weigh a defendant's own contemporaneous statements about intent and harm when deciding related claims, like willfulness -- the two documents pull in different directions without technically contradicting each other.
In re: OpenAI, Inc. Copyright Infringement Litigation (MDL 3143)
- 2025 — Twelve-plus suits against OpenAI and Microsoft, including the Times', consolidated before Judge Stein.
- Feb 6, 2026 — Stein sets aside a magistrate's order compelling disclosure of OpenAI's privileged attorney communications.
- Sep 2, 2026 — DOJ files a Statement of Interest backing OpenAI's fair-use position.
- Sep 4, 2026 — Both sides file cross-motions for summary judgment.
- Sep 17, 2026 — The Times' unredacted filing is unsealed, revealing the internal quotes and dataset figures above.
- Pending — Judge Stein's ruling on summary judgment.
The gap between filing and a first substantive ruling on liability isn't unusual for federal litigation of this scope and complexity -- but it does mean the internal statements unsealed this month have been sitting in discovery, known to both sides, well before a wider public saw them this week.
- Now have OpenAI and Microsoft's own internal language on the public record ahead of a summary-judgment ruling, regardless of how a court ultimately weighs it.
- Face internal statements that undercut a 'we didn't realize the harm' framing, even though neither company has been found liable for anything yet.
- A ruling either way sets precedent for every other publisher-versus-AI-lab suit now sitting in the same or comparable dockets.
A grant of summary judgment wouldn't need a trial to settle liability -- it would mean Judge Stein finds no genuine factual dispute large enough to require one, and rules on the legal question of infringement directly from the paper record already in front of him. That's a higher bar for either side than simply looking sympathetic: OpenAI's own amended answer to the Times' complaint, filed in the same docket in December, denies the core infringement allegations outright, which is exactly the kind of live factual disagreement that ordinarily sends a case to trial rather than resolving it on the papers. Whether the volume and specificity of the unsealed internal statements is enough to close that gap, on copying alone, is the question Stein's ruling will actually have to answer.
None of the unsealed material decides the case. Summary judgment turns on whether the underlying conduct -- copying, and the specific uses made of the copies -- amounts to infringement as a matter of law, not on which side's internal emails read worse in a headline. But a Microsoft director's own words are now sitting in the same docket as the DOJ's fair-use argument, and Judge Stein will have to weigh both when he rules on cross-motions neither side has withdrawn.
- An unsealed Sept. 17 filing quotes a Microsoft director calling AI training 'an astonishing theft.'
- The Times says OpenAI's training data held over 91,000 copies of its and partner outlets' work.
- Microsoft says the quote reflects one employee's opinion, not the company's official position.
- The filing backs the Times' own motion for summary judgment, filed in the same docket Sept. 4.
- Caveat: Judge Stein has not ruled -- these are the plaintiff's characterizations, not a court finding.