P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Unredacted OpenAI filings: an internal email is evidence of what one person wrote, not a ruling

On 17 September 2026 the news publishers suing OpenAI and Microsoft filed a 92 page brief with its redactions lifted, quoting staff at both firms calling AI training theft. It is counsel's argument citing a sealed record, not a finding by any court.

By Parminder Kumar Sharma · · 22 min read

Editorial illustration for the briefing: Unredacted OpenAI filings: an internal email is evidence of what one person wrote, not a ruling

Six words, 995 days, and no ruling

On 17 September 2026 the news publishers suing OpenAI and Microsoft filed a public version of their combined summary judgment brief with the redactions lifted wherever nobody had asked for sealing. The first sentence quotes a Microsoft executive on what large models are doing: an "astonishing theft of unprecedented proportions". That is six words. The New York Times Company filed its complaint on 27 December 2023, which makes 17 September 2026 the 995th day of the case, and on day 995 no court had decided anything about fair use in it.

The brief is 92 pages, filed as Dkt 1977-1 in In re OpenAI, Inc., Copyright Infringement Litigation, 25-md-3143 (SHS) (OTW), in the Southern District of New York. Its text cites the News Plaintiffs' Rule 56.1 statement of undisputed facts at least 291 times, by paragraph number. Almost every phrase that made the headlines is one of those citations: SF1437 for the astonishing theft line, SF1652 for the "largest theft of labor in human history", SF1672 for the "doom loop". The documents behind those numbers, the emails and internal papers themselves, were filed as exhibits to declarations that the public docket marks SEALED.

None of that makes the quotations false. It changes what they are. A damaging phrase in an internal email is admissible evidence of what one employee wrote on one day. It is not a finding of fact, it is not a concession by the company, and it is certainly not a ruling on the four fair use factors. Two things are worth separating before going further, because the coverage runs them together: a copyright claim about copying, and an economics argument about referral traffic. They appear in the same brief and they are not the same argument.

What was actually filed, and by whom

The multidistrict litigation was opened on 11 April 2025 before Judge Sidney H. Stein, with Magistrate Judge Ona T. Wang handling pretrial matters. The news side of the MDL consolidates five actions: The New York Times Company, the Daily News Plaintiffs (eight regional titles including the Chicago Tribune and the Orlando Sentinel), Ziff Davis, the Center for Investigative Reporting, and The Intercept. The news plaintiffs moved for summary judgment as Dkt 1728. Both sides filed public redacted versions of their briefs on 17 September 2026: the plaintiffs at Dkt 1977-1, Microsoft at Dkt 1975, the OpenAI defendants at Dkt 1980, with OpenAI's revised Rule 56.1 statement at Dkt 1979.

This matters for reading the coverage. The 17 September batch is not a disclosure event forced on the companies by a journalist or a leak. It is the routine unsealing step in a cross motion timetable, and the same step produced the defendants' briefs, which contain the numbers that cut the other way. Any account that quotes only one of the four documents filed that day is quoting one side of a summary judgment fight.

Read from the docket and the filings themselves, In re OpenAI, 25-md-3143 (SDNY), 19 September 2026

Fixed on the public recordNot established by these filings
Named employees of Microsoft and OpenAI wrote the quoted words, as characterised by plaintiffs' counselThat any company adopted those words as its position, or that they are accurate
Cross motions for summary judgment are fully briefed before Judge SteinAny judicial finding on fair use, infringement or damages in this case
The plaintiffs assert over 10.8 million works; The Times alone asserts 6,030,928 articlesThat all of them, or most of them, were copied. Defendants say the plaintiffs' own experts put the trained-on figure at 6.3 million
The United States filed a Statement of Interest on 1 September 2026 supporting fair use for trainingThat the court will follow it. A section 517 statement binds nobody

Who said what, and in what capacity

Four people carry most of the weight in the coverage. Their roles are given in the brief itself, and the kind of statement each made is different in every case. A line from a strategy paper, an answer given under oath in a deposition, a message in an internal chat and a sentence written by counsel are four different evidential animals.

Quotes as they appear in the News Plaintiffs' brief, Dkt 1977-1, with the paragraph of the sealed Rule 56.1 statement each is cited to

Words quotedAttributed by the brief toWhat kind of statement
"astonishing theft of unprecedented proportions" (SF1437)Dr Brent Hecht, Microsoft Director of Applied ScienceA prediction in an internal document about what others will soon think, not a conclusion of his own. See below
"largest theft of labor in human history" (SF1652)The same Microsoft executiveQuoted in the introduction with no surrounding sentence given in the public version
"are largely substitutive, period" (SF1473)Nick Turley, OpenAI's Head of ChatGPTAn internal written statement about OpenAI's own products
conversing with chatbots "has substituted" for going to the source (SF1432)Satya Nadella, Microsoft chief executiveDeposition testimony under oath, as summarised by opposing counsel

The same care is owed to the rest. Greg Brockman, OpenAI's co-founder and president, is quoted saying he was "deeply motivated by the gazillions" (SF630) and, of a colleague's message about a paywall workaround, "ah nice" (SF791). Those are chat messages. Satya Nadella is quoted agreeing under oath that "anything that is paywalled should be licensed by anyone who wants to use it" for grounding or training (SF794), and saying that had he known OpenAI had trained on paywalled material he would have required OpenAI to retrain its models (SF796). That is testimony, and it is testimony that cuts against his own co-defendant.

A four step chain showing how a quoted phrase reaches a reader. Step one, the underlying internal documents, filed as exhibits to sealed declarations, are not public. Step two, the News Plaintiffs' Rule 56.1 statement supplies the numbered paragraphs the brief cites at least 291 times. Step three, the public redacted brief at Dkt 1977-1 is written by plaintiffs' counsel. Step four, news coverage quotes it. A closing panel records that no court has ruled on fair use in this case.
Drawn from the docket and filings in In re OpenAI, 25-md-3143 (SDNY)

Five places the copying is said to happen

The brief does not ask the court to find that "AI training" infringes. It asks for judgment at five separate points, and the distinction matters enormously for anyone trying to work out their own exposure. The five are acquisition (obtaining the articles, in some cases from behind paywalls or from third party scrapers), training (pre-training, mid-training and post-training copies), grounding (Microsoft fetching article text from the Bing index at query time to feed the model), output (what the user is shown), and what the plaintiffs call "horse trading", the two defendants supplying each other with copies for money or other consideration.

Some of these are far stronger than others. The grounding and output claims run against Microsoft only. The training claim is the one most people mean when they argue about fair use, and it is the one where the defendants have the most law on their side. Mixing them produces the usual confusion: a ruling that grounding is not fair use would say very little about training, and a ruling that training is fair use would say nothing about a product that reproduces an article to a reader.

A left to right pipeline of five boxes labelled Acquisition, Training, Grounding, Output and Horse trading, the five points at which the news plaintiffs allege copying occurred. A band below, headed reach of a typical supplier copyright indemnity, marks acquisition, training and grounding as not covered, output as covered subject to conditions, and horse trading as not a customer matter. A footnote records that Anthropic's terms reach training data claims.
Stages from the News Plaintiffs' brief, Dkt 1977-1; indemnity reach from the suppliers' published terms

A few of the plaintiffs' concrete figures are worth holding on to, because they are the sort of thing a court can measure. Mid-training datasets are said to contain 91,692 copies of the plaintiffs' works (SF1188-89, 1202, 1205, 1221). At least two million Copilot conversations are said to have grounded on plaintiff website content, including 34,970 unique URLs for The Times and 35,202 for the Daily News group (SF1234, 1237, 1252). On the DMCA claim, the brief tabulates copies from which at least one piece of copyright management information was removed: 3,462,149 for Ziff Davis, 3,002,796 for The Times, 2,638,497 for the Daily News group.

The numbers the defendants put in the same record

If the plaintiffs' brief is read alone it looks overwhelming. Read alongside Dkt 1975 and Dkt 1980, filed the same day, it looks like a contested record with genuine disputes of fact on both sides, which is exactly what summary judgment briefing is supposed to look like. The defendants' central move is to shift from what people wrote to what the products measurably do.

From the OpenAI brief (Dkt 1980) and the Microsoft brief (Dkt 1975), 17 September 2026

What the defendants measuredThe figure they report
Verbatim regurgitation of asserted works in a sample of ChatGPT conversation logs24 instances in 20 million logs, a rate of 0.00012 per cent, per OpenAI's expert
Longest verbatim regurgitations identified29 words and 43 words
Copilot conversations with a 16 word match between grounding content and output59,545 out of 8.2 million logs, which Microsoft's expert puts at 0.73 per cent
Copilot users researching current events at all1.3 per cent of end users, per Microsoft's survey expert
Asserted works with evidence of having been used in training6.3 million of over 10.8 million asserted, on the plaintiffs' own experts' figures

OpenAI also points at the plaintiffs' commercial performance rather than their traffic: The Times's revenues have increased since ChatGPT launched, it passed 12 million subscribers in 2025 against a goal of 15 million by the end of 2027, and its economic expert Dr Avi Goldfarb found no impact on visits to plaintiff websites from individual users adopting ChatGPT. Microsoft's argument is narrower and harder to wave away: since September 2023 any publisher has been able to exclude its pages from Copilot grounding with a NOARCHIVE meta tag, or limit the amount of text used with NOCACHE, and Microsoft says a decision not to use them amounts to an implied licence.

The doom loop is an economics argument, not a copyright claim

The second headline from these filings is a phrase from a Microsoft document: "Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'" (SF1672). It is a striking piece of internal strategy writing. It is not a cause of action.

Where it does legal work is inside fair use factor four, the effect of the use on the market for the copyrighted work, and inside the related question of what would happen if the conduct became widespread. That is why the brief pairs it with click-through data. The relevant numbers, from Microsoft's own records, compare the click-through rate of Copilot's answer engine with traditional Bing search.

Click-through rate reduction between Copilot and traditional Bing search, from Microsoft data as cited at SF1536 in Dkt 1977-1

Publisher groupReported reduction in click-through rate
The New York Times websites87 per cent to 93 per cent
Daily News Plaintiffs websites83 per cent to 91 per cent
Ziff Davis websites51 per cent to 94 per cent

Note what that table is and is not. It is a comparison between two Microsoft products, so it controls for a great deal. It is not a measurement of harm to the publishers, because it says nothing about how many people used each product. The brief is candid that the plaintiffs do not think they need to prove harm empirically, arguing that "common sense suffices". That is a legal position, not an evidential one, and it is the position a defendant will attack hardest.

The most interesting part of the traffic thread is that the defendants' own experts concede the direction of travel, while attributing it elsewhere. OpenAI's economic expert Dr Goldfarb is quoted as opining that the decline in referral traffic to The Times has been driven by a combination of factors including Google AI Overviews, and that Google's introduction of AI Overviews "may have depressed search referrals by 20 to 60 percent" for the Daily News group (SF1784-86). OpenAI's media expert Dr Sinnreich is quoted citing a 2026 Reuters Institute analysis showing Google Discover referrals falling from over 5 billion a month to fewer than 4 billion, and Search referrals from well over 3 billion to slightly more than 2 billion (SF1787).

Where the fair use question actually stands

There is no appellate ruling in the United States on whether training a generative model on copyrighted text is fair use. The two district court decisions everybody cites point in opposite directions and both are cited in these filings. In Bartz v. Anthropic, 787 F. Supp. 3d 1007 (N.D. Cal. 2025), Judge Alsup held that training on lawfully acquired books was fair use but that building a library from pirated copies was not; the resulting class settlement of 1.5 billion US dollars received final approval in July 2026. In Kadrey v. Meta, 788 F. Supp. 3d 1026 (N.D. Cal. 2025), Judge Chhabria found for Meta on the record before him while writing that plaintiffs "will often win" when works are used for substitutive AI products, and that news publisher cases present even stronger arguments against fair use. The news plaintiffs quote that second passage in their introduction. OpenAI quotes the Bartz language about transformativeness in its own.

The first federal appellate answer is pending. The Third Circuit heard argument in Thomson Reuters v. Ross Intelligence on 11 June 2026, reviewing a decision that training a legal research tool on Westlaw headnotes was not fair use. Whatever it holds will be the first binding appellate treatment of transformativeness and market harm in an AI training case, and it will land on Judge Stein's desk while these motions are under consideration.

What a ruling either way would settle, and what it would leave open

If the court holds training is fair useIf the court holds it is not
Settles the acquisition and training claims in this MDL, subject to appealExposes statutory damages across millions of asserted works and forces licensing for US training
Leaves the grounding and output claims against Microsoft untouchedLeaves open whether grounding and output are separately excused
Says nothing about copies obtained in breach of a licence, as in the Annotated Corpus allegationSays nothing about models already trained outside the United States
Does not change UK, EU or any other national lawDoes not change UK, EU or any other national law

None of this is the law in the United Kingdom

The United Kingdom has no fair use doctrine. It has fair dealing, which applies only to a closed list of purposes, and a set of specific exceptions. The exception that people reach for when discussing AI training is section 29A of the Copyright, Designs and Patents Act 1988, inserted on 1 June 2014. Its terms are narrow and worth reading rather than paraphrasing: a copy may be made by a person who has lawful access, "in order that a person who has lawful access to the work may carry out a computational analysis of anything recorded in the work for the sole purpose of research for a non-commercial purpose". Transferring that copy to another person, or using it for any other purpose, infringes. A contract term purporting to restrict a copy allowed by section 29A is unenforceable, but nothing in the section makes commercial training lawful.

The government confirmed the position on 18 March 2026, when the Department for Science, Innovation and Technology and the Department for Culture, Media and Sport published the Report on Copyright and Artificial Intelligence required by section 136 of the Data (Use and Access) Act 2025. Having consulted on four options, the government recorded that its originally preferred proposal, a broad data mining exception with an opt out, "is no longer the government's preferred way forward", and that "we will not introduce reforms to copyright law until we are confident that they will meet our objectives". The report lists the relevant existing exceptions as section 28A (temporary copies) and section 29A (non-commercial research data mining), and notes that both "contain a number of conditions and will only be applicable when those conditions are met".

The one English judgment on the point did not decide it. In Getty Images (US) Inc v Stability AI Ltd [2025] EWHC 2863 (Ch), handed down on 4 November 2025, Getty abandoned its training and development claim because, as Mrs Justice Joanna Smith recorded, "there is no evidence that the training and development of Stable Diffusion took place in the United Kingdom". The secondary infringement claim was dismissed: an "article" may be intangible, but a model "which does not store or reproduce any Copyright Works (and has never done so) is not an 'infringing copy'". Getty succeeded only on limited historic trade mark points about generated watermarks.

UK position as at 19 September 2026, from CDPA 1988 section 29A, the DSIT report of 18 March 2026 and Getty v Stability AI

What is settled in UK lawWhat is not settled
There is no fair use defence. Fair dealing applies only to listed purposesWhether any existing exception covers commercial model training at scale
Section 29A covers non-commercial research only, and requires lawful accessWhat "lawful access" means where terms of service or paywalls are involved
No broad text and data mining exception will be introduced for nowWhat will replace it. The government said it will gather evidence and consider alternatives
A model that stores no copies is not an infringing article for secondary infringementWhether training performed in the UK would infringe. No English court has decided it

What a UK buyer should read in the supplier's terms

For a UK business deploying these models, the litigation risk is not the one in the headlines. You are very unlikely to be sued for OpenAI's training decisions. You are exposed on a narrower and more realistic front: that something your deployment publishes, or that you distribute to customers, reproduces protected expression. Supplier indemnities are built around exactly that distinction, and the marketing names obscure it. "Copyright Shield" and "Customer Copyright Commitment" both mean, in the terms themselves, an output indemnity.

Published supplier terms read on 19 September 2026. Check the version that applies to your contract before relying on this

SupplierWhat the indemnity coversMain conditions and gaps
Microsoft (Product Terms, Customer Copyright Commitment)Third party IP claims "based on Customer's use or distribution of Output Content of a Covered Product"Five conditions, including not disabling content filters or metaprompt restrictions, not using output you should know is infringing, having rights in the input, and implementing all required mitigations. Trade mark claims excluded
OpenAI (Service Terms sections 1 and 3(b))Claims that the customer's use or distribution of Output infringes third party IP, for API and Enterprise or Business customersSix exclusions, including knowing or constructive knowledge of infringement, disabling citation or filtering features, modifying or combining output, and trade mark claims. Beta services excluded entirely
Anthropic (Commercial Terms section K)Claims that paid use of the Services, "which includes data Anthropic has used to train a model that is part of the Services", or outputs, violate third party IPSix exclusions broadly similar to the others. Notably the only one of the three whose wording reaches training data claims

Two details are worth checking rather than assuming. First, the caps. Commentary still repeats that these indemnities are capped at twelve months of fees. In the OpenAI Services Agreement effective 1 January 2026, clause 14.2 excludes "a party's indemnification obligations" from the twelve month cap, and clause 13.1 states that the service specific terms indemnity "is not subject to any liability cap". Anthropic's clause L.3.b likewise disapplies its caps to the indemnity. Second, control. All three agreements give the supplier control of the defence and choice of counsel, which means a claim against your company is defended on the supplier's strategy, not yours.

What to do, in the order worth doing it

Take this with you

For a UK organisation deploying commercial LLMs

  • Write down which of your AI uses produce text that leaves the organisation: published copy, customer-facing summaries, research notes sold on. That list, not the model choice, is your copyright exposure.
  • Pull the actual indemnity clause from your contract, not the marketing page. Note whether it covers output only, whether it is capped, and who controls the defence.
  • Check the five conditions. Confirm nobody has disabled content filters, system prompt restrictions or citation features in any production path, and record that you checked.
  • Map any use of retrieval augmented generation over third party sources. Grounding copies sit outside every output indemnity in the table above.
  • If you fine-tune or train on material you did not create, get advice on section 29A rather than assuming a research exception applies. It covers non-commercial purposes only.
  • If you are a publisher, set NOARCHIVE or NOCACHE where you want exclusion, and keep dated evidence that you set them. Microsoft is running an implied licence argument against publishers who did not.
  • Add a contract review trigger tied to the Third Circuit decision in Thomson Reuters v. Ross Intelligence and to any ruling by Judge Stein on these motions.
  • Stop treating US fair use commentary as guidance on UK exposure. The UK has no fair use doctrine and the government has said it will not legislate for one yet.

The question that exposes the gap

The strongest material in the plaintiffs' brief is not the word theft. It is the ordinary commercial detail: an OpenAI engineer writing that "no matter how prominently we show the links, users won't click" (SF1475), Nick Turley saying that once the chatbot answers there is "no good reason to click" through to the source (SF1525), a Copilot home page telling users that "instead of clicking through links, we can talk about whatever you're curious about" (SF1524). Those are product decisions, made deliberately, recorded in writing, and none of them require a judge to characterise anyone as a thief.

Which leaves the question worth putting to your own supplier, and to your own board. If the answer engine is designed so that users do not click, and the supplier's indemnity covers only what the model outputs, who is carrying the risk that the thing you deployed reproduces someone else's expression on a page with your name at the top? Nobody in these 92 pages is offering to carry it for you.

Key facts

Sources

  1. PrimaryNews Plaintiffs' Combined Summary Judgment Brief, In re OpenAI, Inc., Copyright Infringement Litigation, 25-md-3143 (SHS)(OTW), Dkt 1977-1, filed 17 September 2026. Read in full; source of every plaintiff-side quote and figure.US District Court, Southern District of New York (via CourtListener RECAP)accessed 2026-09-19
  2. PrimaryLetter from Davida Brook to Judge Stein, Dkt 1977, 17 September 2026, explaining that redactions were removed under paragraph 9.c of the Stipulated Omnibus Sealing Order (Dkt 1685).US District Court, Southern District of New York (via CourtListener RECAP)accessed 2026-09-19
  3. PrimaryOpenAI defendants' memorandum in support of summary judgment, Dkt 1980, 17 September 2026. Source of the regurgitation rates, asserted work counts and traffic evidence.US District Court, Southern District of New York (via CourtListener RECAP)accessed 2026-09-19
  4. PrimaryMicrosoft's memorandum in support of summary judgment in the news plaintiffs' consolidated cases, Dkt 1975, 17 September 2026. Source of the 8.2 million chat log analysis and the meta tag opt out argument.US District Court, Southern District of New York (via CourtListener RECAP)accessed 2026-09-19
  5. PrimaryStatement of Interest of the United States, Dkt 1682, filed 1 September 2026, arguing that training large language models on written works is fair use.US Department of Justice via CourtListener RECAPaccessed 2026-09-19
  6. PrimaryDocket for In re OpenAI, Inc., Copyright Infringement Litigation, 1:25-md-03143 (SDNY). Used for filing dates, entry numbers, sealed declaration listings and the assigned judges.CourtListener (Free Law Project)accessed 2026-09-19
  7. PrimaryCopyright, Designs and Patents Act 1988, section 29A, copies for text and data analysis for non-commercial research. Full current text checked.legislation.gov.ukaccessed 2026-09-19
  8. PrimaryReport on Copyright and Artificial Intelligence, published 18 March 2026 under section 136 of the Data (Use and Access) Act 2025. Source of the decision not to proceed with a broad exception.Department for Science, Innovation and Technology and Department for Culture, Media and Sportaccessed 2026-09-19
  9. PrimaryGetty Images (US) Inc and others v Stability AI Ltd [2025] EWHC 2863 (Ch), judgment of Mrs Justice Joanna Smith, 4 November 2025. Read for the abandoned training claim and the secondary infringement conclusion.Courts and Tribunals Judiciary of England and Walesaccessed 2026-09-19
  10. PrimaryMicrosoft Product Terms, universal licence terms for online services, including the Customer Copyright Commitment and its five conditions.Microsoftaccessed 2026-09-19
  11. PrimaryOpenAI Service Terms, sections 1 and 3(b), the API and Enterprise output indemnity and its six exclusions.OpenAIaccessed 2026-09-19
  12. PrimaryOpenAI Services Agreement, effective 1 January 2026, sections 13 and 14, indemnification and the limitation of liability carve out.OpenAIaccessed 2026-09-19
  13. PrimaryAnthropic Commercial Terms of Service, sections K and L.3, whose indemnity expressly extends to data used to train the model.Anthropicaccessed 2026-09-19
  14. Reported byCoverage of the same filings, used only as a pointer to the docket entry number.The Decoderaccessed 2026-09-19
  15. Reported byCoverage of the same filings. The page could not be retrieved during research and nothing in this briefing rests on it.The Vergeaccessed 2026-09-19

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.