P.K. SHARMA

Cyber security intelligence, AI governance, practitioner analysis

Astra for Law passed 54% of a US legal research test. The remaining 46% is the governance problem

OpenAI has combined GPT-6 Astra with a daily legal index spanning more than 230 million URLs. Its own benchmark shows a substantial gain, while also showing why every authority and proposition still needs professional review.

By Parminder Kumar Sharma · · 5 min read

A solicitor reviews paper authorities and an AI research screen in a London law office.

A legal configuration, not a new source of law

OpenAI describes Astra for Law as GPT-6 Astra combined with instructions for legal analysis and writing, a legal search index, governance settings and connections to specialist systems. The distinction matters. The model still generates an answer; the index supplies material that can ground it in authorities. Neither changes which court, legislature or regulator creates the law.

The search index covers US case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs, with sources added daily. OpenAI says its work with Free Law Project brings a case-law collection covering more than 99.9% of published US precedential case law into the experience. That is broad published coverage. It does not mean every filing, unpublished decision, local source, docket development or paid commentary is present. OpenAI explicitly positions the index as complementary to licensed products such as Thomson Reuters.

Initial availability is also narrow. Selected law firms receive Trusted Access in ChatGPT and Codex; the API is described as coming soon. This is a US legal-research product at launch, not a substitute for UK, EU or multi-jurisdiction research.

The 54% result is useful because OpenAI states the denominator

OpenAI tested the complete Astra for Law configuration on 200 US legal-research questions from a private validation set of Vals AI's Legal Research Bench. At the highest reasoning setting, it passed the overall correctness check on 54.0% of questions. GPT-6 Astra using web search alone passed 38.7%. OpenAI calls that a 40% relative improvement.

On case-law questions, the configured product found 24% more reference cases and retrieved up to 54% more relevant passages from correct opinions on an audited set. These measures support a specific conclusion: a curated legal retrieval layer improves research over general web search. They do not establish 54% accuracy for every matter, jurisdiction or legal task. The questions are US-focused, the validation set is private, and the published page does not provide a confusion matrix for missed authorities, wrong propositions or citation-treatment errors.

What the published evidence supports

Published measureResultWhat it means
Overall correctness54.0% on 200 questionsThe complete configuration passed just over half of the private US research set
Web-search comparison38.7%The legal index and configuration materially improved the same model
Reference cases24% moreIt found more benchmark reference authorities on case-law questions
Relevant passagesUp to 54% moreIt retrieved more target passages on the audited subset
Unanswered by the releaseNo public error distributionFirms still need matter-specific acceptance tests

The control boundary moves into the matter workspace

The product becomes more useful, and riskier, when it connects to a firm's own material. OpenAI announced 26 partner-built plugins. Its examples include saving a negotiation brief to an iManage matter file, bringing prior deals into a comparison through DeepJudge and connecting matter context from HighQ. Custom systems described by firms can analyse agreements, trace diligence findings to source documents and carry changes across an IPO filing.

At that point, the central questions are permissions and provenance. Can the assistant retrieve material from a matter that the user cannot otherwise open? Does a draft saved to the document-management system show which model, sources and instructions produced it? Can an ethical wall prevent one team's precedent from entering another team's context? Can a reviewer reconstruct the answer after the index and model have changed?

OpenAI says eligible firms can receive Zero Data Retention on the API and that ChatGPT Enterprise use is excluded from human review by default. It is working with Latham & Watkins on information permissions, ethical walls, client instructions and firm oversight. Those are relevant platform properties, but each firm still owns its access model, retention decisions, client commitments and professional duties.

Take this with you

Controls before a client matter goes live

  • Restrict the pilot to named practice groups, matter types and approved jurisdictions
  • Test retrieval against a lawyer-built set containing binding, adverse and superseded authorities
  • Require source links and passage review before any proposition enters client work
  • Apply document permissions and ethical walls before connecting matter repositories
  • Record model, index date, prompt, retrieved authorities, output and reviewer decision
  • Define what may be written back to the matter file and who approves it
  • Create a rapid correction route for wrong authority, stale law and confidentiality incidents

A small example shows where judgment remains

Suppose a lawyer asks whether a termination clause is enforceable after a recent appellate decision. A useful system can identify the governing jurisdiction, retrieve the decision, point to the relevant paragraphs, find contrary authority and compare the client's clause with the facts. The lawyer must still decide whether the decision is binding, whether it has been stayed or distinguished, whether the factual analogy is sound and whether the answer fits the client's risk appetite.

The same split appears in transactional work. Legora reported that an Astra-powered agent checked 41 documents in minutes, found four planted errors and completed roughly 50 more checks than the previous model. The useful output was a granular record for a professional to review. It was not an autonomous sign-off on the accounts.

Key facts

Sources

  1. PrimaryIntroducing Astra for LawOpenAIaccessed 2026-09-20
  2. PrimaryAstra for Law help articleOpenAIaccessed 2026-09-20
  3. PrimaryLegora reviews financial statements with GPT-6 AstraOpenAIaccessed 2026-09-20

Share this briefing

Know someone who owns this problem? Send it to them.

Related briefings

The briefing, in your inbox

Practitioner analysis of cyber and AI security news. No vendor noise.

One email per briefing. Unsubscribe any time.