Claude searched 1.9 billion protein clusters. The lab found where the claim must stop
Anthropic researchers report that 949 Claude sessions surfaced a previously undescribed family of array-associated reverse transcriptases in bacteriophages. The computational discovery is substantial, but the system's biological function and any biotechnology use remain unknown.
By Parminder Kumar Sharma · · 7 min read

The discovery is a family of biological systems, not a finished gene-editing tool
On 23 September 2026, Anthropic introduced a life-sciences research group and released a 40-page preprint describing array-associated reverse transcriptases, or ARTs. The team says Claude agents found the pattern while surveying reverse transcriptase loci across 1.9 billion protein clusters. Human scientists then extended the computational analysis and examined biological data from a bacteriophage infection experiment.
The finding is easy to overstate because ART arrays resemble CRISPR arrays. They are not reported as CRISPR-Cas systems. The paper found no nearby cas genes, and the researchers do not yet know ART's primary function. Anthropic says the architecture is interesting because other repeat-associated systems can be programmable, but the present evidence does not show that ART cuts, copies or edits a chosen target.
That boundary matters. The paper supports the discovery of a previously undescribed reverse-transcriptase family with a repeat array and a dedicated partner gene. It does not establish a treatment, a gene-editing platform or even the biological purpose of the system.
The campaign combined a fixed research brief with agents that could open follow-up work
The researchers did not ask one chatbot to inspect a database and return an answer. They built an agentic harness around Claude Code. A research brief became a sequence of analysis stages. Worker agents proposed and executed plans, supervisors reviewed the plans and results, and curators maintained a shared record. Supervisors could create additional tasks when an observation justified a new line of work.
Using Claude Mythos 5, the campaign assembled sequence-profile models, recovered about 198,290 reverse-transcriptase clusters, classified them into nine classes and sampled approximately 10,983 loci. It scored 3,564 recurring neighbouring protein families as possible partners. Sixteen families passed the agents' initial criteria, while a follow-up task promoted a seventeenth.
The full run comprised 119 tasks and 949 agent sessions. The preprint reports 77 agent-hours, 215.6 million tokens and 21.5 hours of wall-clock time without human intervention during the run. Those are measures of orchestration scale, not proof that every generated hypothesis was useful.
The funnel reported in the preprint
| Stage | Reported result | What it means |
|---|---|---|
| Search space | 1.94 billion protein clusters | The database scale the agents queried |
| Initial recovery | 198,290 RT clusters | Reverse-transcriptase candidates after filtering |
| Neighbour census | 3,564 partner families across 10,983 loci | Possible recurrent genomic relationships |
| Deep investigation | 16 promoted families plus one follow-up | Candidates selected for dedicated agent work |
| Campaign output | 19 reports | Sixteen partner-family reports and three new RT-lineage reports |
| Execution | 949 sessions, 21.5 wall-clock hours | Parallel agent work coordinated by the harness |
A rejected partner hypothesis created the path to the useful anomaly
ART was not the result the original partner-gene census was designed to find. Agents first selected the candidate because a reverse transcriptase appeared beside a phage RNA-polymerase gene. A worker later rejected that relationship as a meaningful dedicated partnership. Instead of ending the branch, it opened follow-up work on the reverse transcriptase itself.
The next step compared non-coding DNA upstream of related reverse transcriptases. One group had a median upstream region of only 27 base pairs; close relatives had a median of 940. A worker loaded one of those longer raw sequences into its context and noticed a repeated motif. Its own repeat-counting script found a locus with 14 copies of a 16-nucleotide repeat, separated by unique spacers around 100 to 200 nucleotides long.
The agent then tried to disprove its novelty by comparing the arrangement with retrons, CRISPR arrays and other reverse-transcriptase systems. The paper's transcript is important evidence: it shows the observation emerging from direct inspection of sequence data rather than appearing only in a polished retrospective narrative. It also shows uncertainty and self-correction, rather than a straight line from prompt to discovery.
The expanded analysis found 95 ART clusters and 28 detectable repeat arrays
After the autonomous campaign, the researchers used interactive Claude Science sessions and conventional analyses to define the family. They report 95 distinct ART reverse-transcriptase clusters at 90% identity in cultured jumbo phages and predicted viral contigs. Twenty-eight had a detectable repeat array upstream.
The systems share three features: the reverse transcriptase, a dedicated gene immediately downstream and an upstream non-coding repeat array. The arrays span approximately 0.3 to 4.1 kilobases and contain three to 21 copies of a repeat. The repeats are 15 to 49 nucleotides long, while their intervening spacers are 120 to 220 nucleotides and differ from one another.
That structure is distinct from a typical CRISPR array, where near-identical repeats alternate with much shorter spacers and nearby cas genes provide the machinery. The ART reverse transcriptases also have an unusually long amino-terminal region. The catalytic YxDD motif was retained in all 93 family members whose sequence covered that region, supporting their classification as active-looking reverse transcriptases rather than arbitrary sequence neighbours.
Why the authors call the arrays CRISPR-like without calling them CRISPR
| Feature | ART reported in the preprint | Typical CRISPR comparison |
|---|---|---|
| Repeated units | 15 to 49 nucleotides | Near-identical repeat units |
| Spacers | 120 to 220 nucleotides | About 30 nucleotides in the paper's comparison |
| Associated enzyme | Reverse transcriptase with long N-terminus | Cas machinery |
| Nearby cas genes | None found | Expected for CRISPR-Cas systems |
| Known function | Unknown | Adaptive defence and programmable targeting are established for characterised systems |
Laboratory and infection data confirmed expression, not the system's ultimate purpose
The strongest evidence after sequence discovery concerns expression. The team reanalysed published RNA sequencing from a Staphylococcus phage SA1 infection. ART-array RNA was abundant at early, middle and late infection stages. The signal appeared as discrete units rather than one undivided transcript, consistent with the array producing a repertoire of short RNAs.
Human researchers also performed the laboratory work. Anthropic states that its laboratory operates at BSL-1 and BSL-2 and does not handle pathogens capable of infecting humans. The company says Claude helped interpret results, but people conducted the experiments.
These results make ART more than a visual coincidence in a genome browser. The locus is present across a related family, retains catalytic features and produces RNA during infection. They still do not tell us what the RNA guides, what substrate the enzyme acts on, how the partner protein participates or whether the system can be engineered. Those questions require biochemical and structural experiments that the authors say are continuing.
Take this with you
Evidence that would move ART towards a biotechnology claim
- Identify the natural molecular substrate and product of the reverse transcriptase
- Show how the repeat-derived RNAs interact with the enzyme and partner protein
- Demonstrate sequence-specific activity rather than expression alone
- Reconstitute the system with defined components and appropriate controls
- Establish whether targeting can be deliberately reprogrammed
- Measure specificity, efficiency and unintended activity before discussing applications
The paper exposes enough of the agent process to ask reproducibility questions
The publication is unusually useful because it includes the research brief, task counts, selected transcripts and methods rather than presenting the model as a mysterious author. That makes it possible to distinguish the autonomous campaign from the later interactive analysis and human experimentation.
Several limitations remain. The authors work at Anthropic and used an Anthropic model and harness. The preprint had not completed peer review when released. A second team has not yet reported an independent rerun of the campaign, and the model's internal signals are not a substitute for biological validation. The funnel also shows how much proposed work failed: of 17 candidate partner families, only three were retained as previously unreported associations; fourteen were rejected or set aside.
That rejection rate is not a defect by itself. Discovery programmes expect many candidates to fail. It does mean that the useful unit is the complete system of search, review, experiment and correction. Counting generated hypotheses without measuring their survival through that system would reward volume rather than discovery.
Key facts
Sources
- PrimaryClaude discovers a novel enzyme system with CRISPR-like repeatsAnthropicaccessed 2026-09-27
- PrimaryAutonomous AI agents discover reverse transcriptases with tandem repeat arraysYoon et al.accessed 2026-09-27


