Brand Logo

Claude's Autonomous CRISPR-Like Enzyme Discovery, Explained

Aerin Kim

Written by

Aerin Kim

Anthropic's Claude autonomously found a novel CRISPR-like enzyme system in phage DNA across 949 agent sessions. Here is how the search, the agent, and the wet lab actually worked.

If you build or manage autonomous AI agents for anything longer than a single chat session, you already know the hard part is never getting a model to answer one question well. The hard part is keeping an agent coherent, honest, and actually useful across hundreds of sessions and millions of tokens, with no human watching every step. On September 23, 2026, Anthropic published a result that is one of the most concrete public case studies of that exact problem, in one of the least likely domains for it: molecular biology.

Claude, running unsupervised across a custom multi-session harness, searched a database of roughly 1.9 billion protein clusters, flagged a reverse transcriptase enzyme sitting next to a long, evenly spaced array of DNA repeats, and correctly recognized that the layout looked like nothing in the published literature [1]. Anthropic's scientists confirmed the pattern was real, ran it through their wet lab, and are calling the new system array-associated reverse transcriptases, or ART. Feng Zhang, one of the people most responsible for turning CRISPR into a real technology, reviewed the pre-print and called the finding "genuinely intriguing" [1].

This post is not a press release recap. It is a working explanation, for people who build agents, of what actually happened inside that 21.5-hour, 949-session, 215.6-million-token run [3], why a search space of 1.9 billion protein clusters is genuinely hard to comb by hand, why a CRISPR-trained human eye is structurally biased to miss exactly this kind of pattern, and what the discovery run's own architecture, checkpointing, hypothesis logging, human-gated escalation, teaches anyone building long-horizon research or verification agents outside biology entirely. Every code block below is a real, runnable illustration of the mechanics involved, not decoration, and every claim about the discovery itself is cited back to Anthropic's own post, its technical report, and independent science-desk coverage.

claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-hero-wet-lab-bench.png

Step 1: What Anthropic Actually Announced

In spring 2026, Anthropic formed a new life sciences research group and built a real molecular biology laboratory in the Bay Area, staffed by scientists whose prior careers focused on discovering and characterizing unusual proteins through computational genome mining [1]. The group's stated goal was narrow and testable: find out whether a general-purpose AI model, given only a high-level research prompt, can independently drive the kind of biological discovery that historically required a specialist noticing something odd in a mountain of DNA sequence data, the same way restriction enzymes, Taq polymerase, and CRISPR itself were all first noticed [1].

The team's first public result, described in the September 23 post "Claude discovers a novel enzyme system with CRISPR-like repeats," is exactly that kind of finding. Anthropic gave Claude one instruction: search a massive protein sequence database for interesting, previously uncharacterized reverse transcriptases (RTs), enzymes that copy RNA back into DNA. Beyond that initial prompt, the company's own account is direct about the division of labor: "Our involvement was limited to the initial prompt and the lab work, while Claude agents combed through the database, investigated the distinct RT families, and used their own judgment to identify interesting candidates" [1].

Independent outlets that reviewed the same announcement converge on the same core numbers. The Next Web reports the campaign ran for 21.5 hours without human input, across 949 agent sessions, consuming 215.6 million tokens, after being pointed at a database of roughly 1.9 billion protein clusters [3]. Interesting Engineering's coverage lands on the same order of magnitude, describing "950 Claude agents" searching DNA for "21 hours" [4]. Gizmodo and Quartz both corroborate the ART name, the three-part structure, and the fact that Anthropic still does not know what the system actually does biologically [6] [5].

That last point matters and is worth stating plainly before anything else: this is not a claim that Claude invented a new gene-editing tool. ART's function is unknown. What was found is a structural pattern, a reverse transcriptase, an adjacent gene of unknown function, and a repeat array, that has previously only ever been seen together in a small handful of programmable DNA-editing systems, CRISPR chief among them [1]. The scientific claim is narrow and testable. The engineering claim, that a fully autonomous multi-session agent swarm can independently surface a pattern like this out of billions of sequences with only a one-line research brief, is the part that is actually verifiable today and is the real subject of this post.

Step 2: Why 1.9 Billion Protein Clusters Is a Genuinely Hard Search Space

To understand why this result is interesting from an engineering standpoint, you have to understand what "1.9 billion protein clusters" actually means as a search problem, and why it is not something you point a single model context window at and read through.

Public protein sequence databases like UniProtKB have grown past hundreds of millions of raw entries, and even after deduplication, the reference cluster databases built on top of them, most notably UniRef, still contain hundreds of millions to low billions of non-redundant cluster representatives depending on the identity threshold used [16] [17]. A database on the order of 1.9 billion protein clusters is not a curated shortlist. It is the raw output of clustering essentially everything that has ever been sequenced from bacteria, archaea, viruses, and their mobile elements, at a resolution loose enough to still separate genuinely distinct protein families from one another.

No text search or keyword filter helps here, because the overwhelming majority of these clusters have no functional annotation at all. You cannot search for "reverse transcriptases with adjacent repeat arrays" the way you would search a well-labeled dataset, because almost nothing in this space is labeled. The only tractable approach is the one genome miners have used for two decades: cluster sequences by similarity, pull out family representatives, and inspect the ones that look structurally unusual, an approach known as genome mining, which historically produced most of the reverse transcriptase families discovered in bacteria over the past several years [1].

The tooling that makes billion-scale clustering possible at all

The reason a search like this is even computationally feasible is a specific line of bioinformatics tooling built explicitly for this scale. MMseqs2, developed by Martin Steinegger and Johannes Söding, filters sequence comparisons through three stages of increasing sensitivity and decreasing speed, a fast k-mer prefilter that eliminates 99.9% of irrelevant sequence pairs, an ungapped alignment pass that eliminates another 99% of what remains, and only then a full Smith-Waterman alignment on the roughly one in one hundred thousand pairs that survive both filters [18]. That funnel is what turns an all-against-all comparison of billions of sequences, which would be computationally impossible at brute force, into something that finishes in a practical amount of wall-clock time.

Here is what a real clustering pass over a protein database at this kind of scale actually looks like as a command-line pipeline, using MMseqs2's own easy-cluster workflow plus a Foldseek structural pass for candidates that clear the sequence-level filter:

bash
/code # Illustrative large-scale protein clustering pipeline, structurally representative # of how a genome-mining search over a database on the order of ~1.9 billion # protein clusters is actually scripted with real, widely used bioinformatics # tools (MMseqs2 for sequence clustering, Foldseek for structural follow-up). # This is a teaching example of the pipeline shape, not a reproduction of # Anthropic's exact internal harness or database. set -euo pipefail # 1. Build an MMseqs2 database from a raw FASTA of candidate protein sequences # (in a real run, this would be a reverse-transcriptase-focused subset # pulled from a much larger reference cluster database). mmseqs createdb candidate_proteins.fasta candidate_db # 2. Cluster at a moderate sequence identity threshold. easy-cluster runs the # full pipeline: k-mer prefilter -> ungapped alignment -> Smith-Waterman # on survivors, which is what makes clustering billions of sequences # computationally tractable in the first place. mmseqs easy-cluster candidate_db clustered_out tmp/ \ --min-seq-id 0.3 \ --cov-mode 1 \ --cluster-mode 2 \ --threads 32 # 3. Pull out one representative sequence per cluster. This is the # "one entry per protein family" reduction that turns a raw # multi-billion-sequence collection into a tractable number of # distinct families to actually inspect. mmseqs createsubdb clustered_out_cluster.tsv candidate_db cluster_reps # 4. For clusters that pass an initial anomaly filter (e.g. reverse # transcriptase domain present, but no matching characterized system), # run a structural search with Foldseek against a reference structure # database to check for remote structural relationships that pure # sequence identity would miss. foldseek easy-search cluster_reps.fasta \ reference_structures_db \ structural_hits.tsv \ tmp_foldseek/ \ --format-output "query,target,evalue,prob,alntmscore" # 5. Candidates with no strong structural or sequence match to any known, # annotated system are exactly the shortlist worth a closer look, either # by an automated anomaly scorer or, as in the ART campaign, an agent # reading the raw sequence context around each one. awk -F'\t' '$4 < 0.5' structural_hits.tsv > low_confidence_matches.tsv echo "Candidates with no confident structural match: $(wc -l < low_confidence_matches.tsv)"

Two things about that pipeline matter for understanding what Claude's agents were actually doing. First, sequence-level clustering alone gets you groups of similar proteins, not an explanation of what they do or whether they are interesting. Second, structural search tools like Foldseek, which encodes 3D structure as a sequence over a discrete structural alphabet so it can be searched with the same speed tricks as ordinary sequence search, let you find remote relationships that pure sequence identity would miss entirely, since protein structure is conserved far longer through evolution than raw sequence is [19]. Neither tool tells you a cluster is worth a human's time. That triage step, deciding which of thousands of structurally coherent clusters is actually novel and biologically interesting, is exactly the judgment call Anthropic's own writeup says Claude agents were making on their own [1].

claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-search-space-mining-cross-section.png

The funnel, in real numbers

The scale of the actual filtering that happened during the 21.5-hour run is worth laying out explicitly, because it is the clearest evidence that this was not a lucky single hit, it was a systematic funnel that happened to surface one genuinely novel result at the bottom:

StageCount / DurationWhat Happens
Starting search space~1.9 billion protein clustersFull non-redundant reference cluster database, almost entirely unannotated
Reverse transcriptases recovered~200,000Sequences matching RT domain signatures, gathered by autonomous agents
Candidate partner families scored3,564Distinct RT-adjacent gene families with no known characterized system
Human-readable reports filed19Final shortlist that cleared the campaign's own evidence and novelty bar
Novel system confirmed1 (ART)Verified in Anthropic's wet lab as a genuinely new RT + repeat array system
Total campaign time21.5 hoursFully autonomous, across 949 agent sessions and 215.6 million tokens

Going from roughly 200,000 recovered reverse transcriptase sequences down to 19 human-reviewable reports is a reduction of more than four orders of magnitude, achieved without a human touching a single intermediate result. For an individual expert scientist doing the equivalent survey by hand, Anthropic's own team estimates this kind of analysis normally takes weeks to months of dedicated work [1]. Compressing that into under a day is the headline efficiency claim. But the number that should actually catch an agent engineer's attention is 19, not 1.9 billion: the entire campaign produced a shortlist small enough that a team of maybe three or four human scientists could plausibly review every single report by lunch. That ratio, billions of raw candidates funneled down to a double-digit human review queue, is the actual design target for any agent system meant to operate at this kind of scale, in biology or anywhere else.

Step 3: The Biology a CRISPR-Trained Eye Is Built to Miss

To understand why this specific pattern, a reverse transcriptase sitting next to an evenly spaced repeat array, is both meaningful and easy for a human expert to walk past, you need a short tour of what reverse transcriptases in bacteria and phages actually do, because RT-associated systems are a much bigger and stranger family than most people outside microbiology realize.

CRISPR itself was first noticed not as a gene-editing tool but as an unexplained repeat sequence. Francisco Mojica spent years in the 1990s cataloguing short, regularly spaced repeats in archaeal and bacterial genomes before anyone understood their function, and it took until 2002 for Mojica and Ruud Jansen, working independently, to settle on the acronym CRISPR and to formally name the associated Cas genes sitting beside the repeats [9]. The full story of how that repeat pattern eventually became a gene-editing platform took another decade of independent groups piecing together the mechanism [10]. The point is not nostalgia, it is that the single most consequential discovery in modern molecular biology started exactly the way ART did: as an unexplained, evenly spaced repeat sitting next to a gene, noticed by someone doing systematic sequence comparison rather than looking for anything in particular.

Reverse transcriptases are not a CRISPR sideshow, they are their own sprawling family

Reverse transcriptases sit at the center of several distinct, independently evolved bacterial and phage systems, and this is exactly where a CRISPR-trained intuition becomes a liability rather than an asset.

Diversity-generating retroelements (DGRs) use an RT to convert a template repeat region into hypervariable target genes, generating up to roughly 10^13 protein variants from a single locus, a mechanism bacteriophages use to keep evolving new receptor-binding proteins fast enough to chase a host's changing surface receptors [11]. A 2024 survey found roughly 31,000 DGRs spread across more than 1,500 bacterial and archaeal genera [11].

Retrons pair an RT with a non-coding RNA to produce a branched RNA-DNA hybrid molecule (msDNA), and were shown only in 2020 to function as an anti-phage abortive-infection defense system, sacrificing the infected cell to stop a phage from spreading further through a bacterial population [[12]](https://www.cell.com/cell/fulltext/S0092-8674(20)31306-4). Retrons had been known since 1984 and sat unexplained for roughly 36 years before their defensive function was demonstrated.

Defense-associated reverse transcriptases (DRTs, sometimes called UG/Abi systems) were catalogued in a 2022 systematic computational survey that used exactly the genome-mining approach described above, sequence clustering across bacterial genomes, to identify 42 highly diverse groups of RT-linked defense systems, most paired with an adjacent gene of unknown function at the time of publication [13]. If that structure sounds familiar, an RT paired with a mystery partner gene, it should: it is the closest documented precedent to what ART turned out to be, and it is a useful reminder that this general shape of discovery, RT plus unexplained neighbor, has a real track record.

Group II introns, the most abundant class of mobile retroelements in bacteria, pair a self-splicing catalytic RNA with an RT that also carries maturase and DNA endonuclease activity, letting the whole element "retrohome" into new genomic locations via an RNA intermediate [14]. They are also believed to be evolutionary ancestors of the introns found throughout eukaryotic genomes, including our own.

Jumbo phages, the specific host context ART was found in, are bacteriophages with genomes larger than 200 kilobase pairs, historically rare (only 93 were known as of 2016) but increasingly common in recent metagenomic surveys, and notorious for carrying genuinely novel biological machinery not found in smaller, more conventional phages, including proteinaceous nucleus-like compartments that physically protect their DNA from host defenses [15]. Anthropic's own post confirms the RT underlying ART was first identified in a jumbo phage in prior studies, but that no one had previously noticed its adjacent repeat array or accessory gene [1].

Lay all five of these side by side and the actual scale of the search space comes into focus. A human expert's intuition is trained by years of reading the specific literature they specialize in. A CRISPR specialist's trained eye is tuned to recognize CRISPR-Cas repeat spacing and PAM-adjacent motifs. A DGR specialist's eye is tuned to variable-repeat asymmetry. A retron specialist is tuned to msDNA branch structures. None of those trained intuitions transfer cleanly to a jumbo phage RT sitting next to an unfamiliar repeat pattern that does not quite match any of the five known families, because specialization, by definition, means becoming extremely good at recognizing what you already expect and comparatively worse at flagging genuine anomalies outside your trained distribution.

claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-art-system-cutaway.png

A repeat-array detector, in real code

The actual pattern-recognition step, spotting a tandem repeat array by scanning raw DNA sequence, is not exotic machine learning. It is a well-understood computational biology technique: sliding-window autocorrelation to find periodic repeats, then measuring the spacing consistency between repeat units. Here is a self-contained, runnable version of the core logic an agent (or a script) would use to flag a candidate array in a raw sequence, structurally close to what a genome-mining pipeline would run automatically across every RT-adjacent region pulled out of the funnel described in Step 2:

python
/code # Self-contained tandem repeat array detector. # Real, runnable logic (pure Python, no external dependencies) for the core # pattern-recognition step described in Step 3: scanning a raw DNA sequence # for an evenly spaced tandem repeat array, the same class of pattern that # both CRISPR arrays and the ART system share. from dataclasses import dataclass from typing import List, Optional @dataclass class RepeatArrayCandidate: unit_length: int spacing: int repeat_count: int start: int consistency_score: float # 1.0 = perfectly evenly spaced def find_tandem_repeats(sequence: str, min_unit_len: int = 20, max_unit_len: int = 40, min_repeats: int = 3) -> List[RepeatArrayCandidate]: """Scan `sequence` for tandem repeat arrays using sliding-window autocorrelation: for each candidate repeat unit length, check how many consecutive windows of that length match closely, and how consistent the spacing between matches is. """ sequence = sequence.upper() candidates: List[RepeatArrayCandidate] = [] for unit_len in range(min_unit_len, max_unit_len + 1): i = 0 while i < len(sequence) - unit_len * min_repeats: unit = sequence[i:i + unit_len] positions = [i] j = i + unit_len while j <= len(sequence) - unit_len: window = sequence[j:j + unit_len] if _hamming_similarity(unit, window) >= 0.85: positions.append(j) j += unit_len else: break if len(positions) >= min_repeats: spacings = [positions[k + 1] - positions[k] for k in range(len(positions) - 1)] consistency = 1.0 - (_stdev(spacings) / unit_len) candidates.append(RepeatArrayCandidate( unit_length=unit_len, spacing=unit_len, repeat_count=len(positions), start=positions[0], consistency_score=max(0.0, consistency), )) i = positions[-1] + unit_len else: i += 1 # Keep only the highest-confidence, non-overlapping candidates. candidates.sort(key=lambda c: (c.repeat_count, c.consistency_score), reverse=True) return candidates[:5] def _hamming_similarity(a: str, b: str) -> float: if len(a) != len(b): return 0.0 matches = sum(1 for x, y in zip(a, b) if x == y) return matches / len(a) def _stdev(values: List[int]) -> float: if len(values) < 2: return 0.0 mean = sum(values) / len(values) variance = sum((v - mean) ** 2 for v in values) / len(values) return variance ** 0.5 def flag_for_review(candidate: Optional[RepeatArrayCandidate]) -> bool: """Decide whether a detected repeat array is worth escalating. Mirrors the kind of evidence threshold a triage step would apply before a candidate reaches a human-readable report. """ if candidate is None: return False return (candidate.repeat_count >= 4 and candidate.consistency_score >= 0.8) if __name__ == "__main__": # A synthetic example: a repeat unit of length 24, repeated 5 times # with near-perfect spacing, flanking a stretch of non-repetitive # sequence on either side. unit = "ACGTTGCATCGGATCCAGTTACGA" flank_left = "TTGACCGATCGGTATTCCGATTAGGCTA" flank_right = "GGCTTACCGATTGCAATCGGATTAGC" seq = flank_left + (unit * 5) + flank_right hits = find_tandem_repeats(seq) for hit in hits: print(hit, "-> escalate:", flag_for_review(hit))

Running that kind of scan is trivial computationally, a few milliseconds per candidate sequence. The genuinely hard part was never detecting the repeat once you are looking at the right few hundred base pairs. The hard part was being the one process, out of a search space of 1.9 billion clusters, that decided this particular obscure RT-adjacent region was worth pulling up and looking at in the first place.

Step 4: Why a Systematic Agent Notices What a Trained Eye Skips

This is the part of the story that separates "a computer ran a known algorithm faster" from something genuinely interesting about how the agent behaved. Anthropic's own account describes the moment of discovery almost like a lab notebook entry. While combing through the raw DNA sequence adjacent to an unusual RT family, one agent wrote: "[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!" [1]. It then did what a careful human scientist would do next: it counted the repeats, measured their spacing, compared the layout against known RT systems, searched the literature for any prior report of the same pattern, and only after that verification chain filed a written report for human review [1].

That sequence, notice, count, compare, search literature, write up evidence, is the actual mechanism worth studying, not a metaphor for it. It maps directly onto something every agent builder eventually has to design explicitly: a self-verification loop that runs before a low-confidence finding is escalated to a human, rather than an agent that reports every anomaly it stumbles across and burns its reviewers' trust with false positives.

A genuinely useful, slightly humbling benchmark result

Buried in the coverage of this release is a detail that is arguably more important for agent design than the discovery itself. The Next Web's reporting on the technical report states that when Claude's most capable models were shown DNA sequences directly and asked to identify the repeat array, they succeeded in at least 90% of attempts. But when those same models were given files and tools, and had to locate and process the sequence data themselves rather than having it handed to them, accuracy dropped to just 32% [3].

That is a 58-percentage-point gap between "recognize the pattern when it's in front of you" and "successfully retrieve and process the right data to get the pattern in front of you in the first place," using the exact same underlying models. It is a direct, quantified illustration of a lesson agent builders learn the hard way: raw model capability and tool-use reliability are not the same axis, and a swarm's overall success rate is gated far more tightly by its tool-use and data-handling reliability than by its underlying reasoning quality. This is also precisely why running 949 independent sessions rather than one long session mattered. A single session with a 32% success rate at correctly retrieving and framing the right slice of data would need to get lucky. Nearly a thousand independent attempts, each with its own tool-use trajectory, dramatically raises the odds that at least one of them threads the retrieval step cleanly and gets a genuinely novel candidate in front of the model's much stronger 90%-plus direct pattern-recognition ability.

claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-accuracy-benchmark-gauges.png

Here is a simplified, runnable sketch of the kind of harness you would actually build to measure that gap yourself, comparing a direct-input condition against a tool-mediated retrieval condition on the same underlying detection task:

python
/code # ILLUSTRATIVE benchmark harness sketching how you would measure the gap # between "recognize a pattern when handed the right data directly" versus # "correctly retrieve and frame that data yourself via tools", the same # structural gap reporting described a 90%+ vs 32% split on. This is a # teaching skeleton, not a reproduction of Anthropic's actual eval code. import random from dataclasses import dataclass, field from typing import Callable, List @dataclass class EvalResult: condition: str correct: int = 0 total: int = 0 @property def accuracy(self) -> float: return self.correct / self.total if self.total else 0.0 def run_direct_input_condition(model_call: Callable[[str], bool], test_sequences: List[str]) -> EvalResult: """Condition A: the raw sequence containing (or not containing) a repeat array is placed directly in the prompt. Tests pattern recognition only. """ result = EvalResult(condition="direct_input") for seq, has_repeat in test_sequences: result.total += 1 predicted = model_call(seq) if predicted == has_repeat: result.correct += 1 return result def run_tool_mediated_condition(model_call_with_tools: Callable[[str], bool], test_file_paths: List[str]) -> EvalResult: """Condition B: the model is given a file path or database handle and must locate, fetch, and correctly slice the relevant sequence itself before it can even attempt the same recognition task. """ result = EvalResult(condition="tool_mediated_retrieval") for path, has_repeat in test_file_paths: result.total += 1 predicted = model_call_with_tools(path) if predicted == has_repeat: result.correct += 1 return result def report(results: List[EvalResult]) -> None: for r in results: print(f"{r.condition}: {r.correct}/{r.total} " f"({r.accuracy * 100:.1f}% accuracy)") if len(results) == 2: gap = abs(results[0].accuracy - results[1].accuracy) * 100 print(f"Capability vs. tool-use gap: {gap:.1f} percentage points") if __name__ == "__main__": # Stand-in model callables. In a real eval these would call an actual # model API; here they simulate the reported accuracy gap for illustration. def fake_direct_model(_seq: str) -> bool: return random.random() < 0.90 def fake_tool_model(_path: str) -> bool: return random.random() < 0.32 direct_tests = [(f"seq_{i}", True) for i in range(200)] tool_tests = [(f"file_{i}.fasta", True) for i in range(200)] results = [ run_direct_input_condition(fake_direct_model, direct_tests), run_tool_mediated_condition(fake_tool_model, tool_tests), ] report(results)

The practical implication for anyone building a research or verification agent outside biology: do not assume that improving your model's raw reasoning score will fix a low end-to-end success rate if your actual bottleneck is tool reliability, data retrieval correctness, or context assembly. Measure the two separately, the way this benchmark implicitly did, before deciding where to spend engineering effort.

Step 5: Orchestrating 949 Agent Sessions Over 21.5 Hours

Anthropic's own description of its tooling is specific: the work ran inside Claude Science and Claude Code, "the same tools available to any scientist," plus, in their words, "sometimes with a harness of our own that coordinates many Claude sessions running in parallel" [1]. That custom harness, not any single model call, is the actual engineering artifact behind this result, and it is the same category of system Anthropic has documented publicly in a different context: their multi-agent research system, where an orchestrator agent decomposes a task and spawns subagents with their own context windows that report condensed findings back up [23]. Anthropic's internal evaluation of that architecture found a multi-agent setup with a Claude Opus lead and Claude Sonnet subagents beat a single-agent Opus baseline by 90.2% on an internal research benchmark [23], which is the same structural bet the ART campaign made at a much larger scale: many parallel, bounded sessions instead of one enormous one.

Three specific engineering problems have to be solved to make 949 sessions cohere into one 21.5-hour campaign instead of 949 unrelated conversations, and Anthropic's own engineering writing on context engineering names the same three techniques independently of this specific result: compaction, structured note-taking, and sub-agent decomposition with condensed handoff [24].

Context handoff between sessions. No single context window holds 215.6 million tokens of history, so each new session cannot simply "continue" the previous one's raw transcript. Instead, a session has to end by writing a compact, structured summary, what was tried, what was ruled out, what remains open, to persistent storage outside the model's context, and the next session has to load that summary rather than the full history [24].

Hypothesis logging as a first-class object. Anthropic's team notes that because Claude produces hypotheses "prolifically," the hypotheses themselves became an object of study: with hundreds to thousands of candidate reports from a single campaign, the team started analyzing what distinguished proposals worth testing from ones correctly set aside, and fed that back into the instructions given to future runs [1]. That is a durable, queryable hypothesis store, not an ephemeral chat log, and it is what let 3,564 candidate families get narrowed to 19 reports without any single session needing to hold all 3,564 in memory at once.

A stopping and escalation condition. A campaign that runs until a human manually kills it is not an autonomous system, it is an unsupervised one, and those are different things with very different risk profiles. Something in the harness has to decide, per candidate and in aggregate, when a finding has cleared enough independent verification steps (repeat count, spacing consistency, literature novelty check) to warrant flagging for human review, and separately, when the overall search has run long enough or hit diminishing enough returns to end the session budget.

Here is a simplified but structurally real skeleton of what that kind of checkpointed, multi-session orchestration loop looks like in Python, covering session handoff, a persistent hypothesis log, and an explicit escalation condition:

python
/code # Checkpointed multi-session agentic loop skeleton. # A structurally real illustration of how a ~949-session, 21.5-hour # autonomous research campaign is orchestrated: bounded sessions, persistent # context handoff between them, a durable hypothesis log, and an explicit # stopping / escalation condition. This is a teaching skeleton meant to be # adapted to a real agent SDK and a real task queue, not a drop-in production # system. import json import time import uuid from dataclasses import dataclass, field, asdict from pathlib import Path from typing import List, Optional STATE_DIR = Path("./campaign_state") STATE_DIR.mkdir(exist_ok=True) HYPOTHESIS_LOG = STATE_DIR / "hypothesis_log.jsonl" CHECKPOINT_FILE = STATE_DIR / "latest_checkpoint.json" MAX_SESSIONS = 1000 MAX_CAMPAIGN_HOURS = 24.0 ESCALATION_CONFIDENCE_THRESHOLD = 0.8 DIMINISHING_RETURNS_WINDOW = 50 # sessions with no new high-confidence hit @dataclass class SessionSummary: session_id: str candidates_examined: int high_confidence_hits: int compact_notes: str # what to carry forward, NOT the raw transcript timestamp: float = field(default_factory=time.time) @dataclass class CampaignCheckpoint: sessions_completed: int started_at: float total_candidates_examined: int total_high_confidence_hits: int sessions_since_last_hit: int carry_forward_notes: str def load_checkpoint() -> CampaignCheckpoint: if CHECKPOINT_FILE.exists(): return CampaignCheckpoint(**json.loads(CHECKPOINT_FILE.read_text())) return CampaignCheckpoint( sessions_completed=0, started_at=time.time(), total_candidates_examined=0, total_high_confidence_hits=0, sessions_since_last_hit=0, carry_forward_notes="No prior sessions. Begin broad survey of " "reverse transcriptase clusters lacking a " "characterized function.", ) def save_checkpoint(cp: CampaignCheckpoint) -> None: CHECKPOINT_FILE.write_text(json.dumps(asdict(cp), indent=2)) def log_hypothesis(session_id: str, candidate_id: str, confidence: float, evidence: str) -> None: entry = { "session_id": session_id, "candidate_id": candidate_id, "confidence": confidence, "evidence": evidence, "logged_at": time.time(), } with HYPOTHESIS_LOG.open("a") as f: f.write(json.dumps(entry) + "\n") def run_one_session(checkpoint: CampaignCheckpoint) -> SessionSummary: """Stand-in for a single bounded agent session. A real implementation calls an agent SDK with `checkpoint.carry_forward_notes` as the seed context, lets it explore a bounded slice of the search space, and returns a compact summary rather than its full raw transcript. """ session_id = str(uuid.uuid4())[:8] # ... real agent call happens here, scoped to a bounded slice of the # search space and seeded with `checkpoint.carry_forward_notes` ... candidates_examined = 40 # placeholder for real search output high_confidence_hits = 0 # most sessions find nothing notable summary = SessionSummary( session_id=session_id, candidates_examined=candidates_examined, high_confidence_hits=high_confidence_hits, compact_notes=( "Surveyed 40 candidate RT-adjacent regions. No repeat arrays " "above the consistency threshold. Ruled out 3 families as " "matching known DGR architecture." ), ) return summary def should_stop(cp: CampaignCheckpoint) -> Optional[str]: elapsed_hours = (time.time() - cp.started_at) / 3600 if elapsed_hours >= MAX_CAMPAIGN_HOURS: return f"Time budget exhausted ({elapsed_hours:.1f}h)" if cp.sessions_completed >= MAX_SESSIONS: return f"Session budget exhausted ({cp.sessions_completed})" if cp.sessions_since_last_hit >= DIMINISHING_RETURNS_WINDOW: return (f"Diminishing returns: {cp.sessions_since_last_hit} " f"sessions with no new high-confidence hit") return None def run_campaign() -> None: cp = load_checkpoint() while True: stop_reason = should_stop(cp) if stop_reason: print(f"Campaign ending: {stop_reason}") break summary = run_one_session(cp) cp.sessions_completed += 1 cp.total_candidates_examined += summary.candidates_examined cp.total_high_confidence_hits += summary.high_confidence_hits cp.carry_forward_notes = summary.compact_notes # compaction step if summary.high_confidence_hits > 0: cp.sessions_since_last_hit = 0 log_hypothesis( session_id=summary.session_id, candidate_id=f"candidate-{cp.sessions_completed}", confidence=ESCALATION_CONFIDENCE_THRESHOLD, evidence=summary.compact_notes, ) else: cp.sessions_since_last_hit += 1 save_checkpoint(cp) # crash-safe: a killed process resumes cleanly if cp.sessions_completed % 50 == 0: print(f"[{cp.sessions_completed} sessions] " f"{cp.total_candidates_examined} candidates examined, " f"{cp.total_high_confidence_hits} flagged for review") if __name__ == "__main__": run_campaign()
claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-agent-expedition-basecamps.png

Reading that skeleton, notice what it deliberately does not do: it does not let any single session accumulate unbounded state, it does not let a candidate get escalated to a human on the strength of one session's opinion alone, and it never lets the loop run forever on an open-ended "keep researching" instruction without a checkable stopping condition. Those three constraints are what separate a system that can run for 21.5 hours and 949 sessions without babysitting from one that quietly drifts off task, hallucinates a growing pile of false positives, or simply never terminates.

Step 6: The Human-in-the-Loop Wet Lab Gate

The most consequential design decision in the entire ART campaign is also the least flashy: Claude never touches a pipette. Every physical laboratory step, expressing the candidate protein in a standard laboratory strain, characterizing it biochemically and structurally, confirming the repeat array is actually transcribed into RNA, is performed exclusively by human scientists [1]. Claude's role stops at searching genomic datasets, generating hypotheses, analyzing candidates computationally, and helping interpret the resulting lab data once it exists [1] [6].

The lab itself is described as ordinary by molecular biology standards, Bay Area-based, operating only at the lower biosafety risk levels (BSL-1 and BSL-2), and explicitly not handling pathogens capable of infecting humans [1]. Anthropic's post even notes they experimented with using AI to accelerate physical lab work through an internal initiative called the Model Hardware Standard, and concluded that approach was "less conducive to the sort of ad hoc workflows" involved in this kind of exploratory molecular biology [1]. In other words, the boundary between computational agent and physical human is not an arbitrary safety theater line, it is a deliberate engineering judgment about where an AI agent's current reliability profile actually matches the task, and where it does not yet.

What a candidate has to clear before it reaches a human

For a 19-report shortlist to be trustworthy enough for scientists to spend real lab time on, each report needs a consistent, structured shape, not a free-form paragraph a reviewer has to parse from scratch every time. Here is a realistic schema for what that kind of human-review candidate report looks like, matching the structure implied by Anthropic's own description (proposed function, supporting evidence, comparison to known systems, confidence level):

json
/code { "$schema": "https://json-schema.org/draft/2020-12/schema", "title": "CandidateSystemReport", "description": "Structured shape for a single human-readable candidate report, the kind of artifact an autonomous genome-mining campaign hands to a human scientist for review before any lab work is scheduled.", "type": "object", "required": [ "candidate_id", "proposed_function", "supporting_evidence", "comparison_to_known_systems", "confidence", "session_id", "flagged_at" ], "properties": { "candidate_id": { "type": "string" }, "proposed_function": { "type": "string", "description": "One or two sentence hypothesis, e.g. 'Reverse transcriptase with an adjacent evenly spaced repeat array, structurally reminiscent of a CRISPR array.'" }, "supporting_evidence": { "type": "array", "items": { "type": "string" }, "minItems": 1, "description": "Each item is a discrete, checkable observation: repeat count, spacing consistency, absence of matching literature, etc." }, "comparison_to_known_systems": { "type": "array", "items": { "type": "object", "required": ["system_name", "similarity", "key_difference"], "properties": { "system_name": { "type": "string" }, "similarity": { "type": "string" }, "key_difference": { "type": "string" } } } }, "confidence": { "type": "number", "minimum": 0, "maximum": 1 }, "session_id": { "type": "string" }, "flagged_at": { "type": "string", "format": "date-time" }, "requires_wet_lab_verification": { "type": "boolean", "const": true } } }

And here is the triage logic, in practice a small script, that would sit between the agent swarm's raw output and the human review queue, enforcing that nothing reaches a scientist's desk without meeting a minimum evidence bar, while still logging every rejected candidate for later audit rather than silently discarding it:

python
/code # Triage script: sits between the raw agent-generated candidate reports and # the human lab-review queue. Nothing reaches a scientist's desk without # clearing a minimum evidence bar, and every rejected candidate is still # logged for later audit rather than silently discarded. import json from pathlib import Path from datetime import datetime, timezone MIN_CONFIDENCE = 0.75 MIN_EVIDENCE_ITEMS = 2 REVIEW_QUEUE = Path("./human_review_queue.jsonl") REJECTED_LOG = Path("./rejected_candidates.jsonl") def load_reports(path: Path): with path.open() as f: for line in f: if line.strip(): yield json.loads(line) def passes_triage(report: dict) -> bool: if report.get("confidence", 0) < MIN_CONFIDENCE: return False if len(report.get("supporting_evidence", [])) < MIN_EVIDENCE_ITEMS: return False if not report.get("comparison_to_known_systems"): # A candidate with no comparison against known systems has not # actually demonstrated novelty yet, just an unexplained pattern. return False return True def triage(raw_reports_path: Path) -> None: promoted, rejected = 0, 0 with REVIEW_QUEUE.open("a") as queue_out, \ REJECTED_LOG.open("a") as rejected_out: for report in load_reports(raw_reports_path): record = dict(report) record["triaged_at"] = datetime.now(timezone.utc).isoformat() if passes_triage(report): queue_out.write(json.dumps(record) + "\n") promoted += 1 else: record["rejection_reason"] = ( "confidence_below_threshold" if report.get("confidence", 0) < MIN_CONFIDENCE else "insufficient_evidence_or_comparison" ) rejected_out.write(json.dumps(record) + "\n") rejected += 1 print(f"Promoted to human review: {promoted}") print(f"Rejected (logged for audit): {rejected}") if __name__ == "__main__": triage(Path("./raw_candidate_reports.jsonl"))
claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-detection-pipeline-blueprint.png

This computational-to-physical handoff is also where the campaign's efficiency claim has to be read carefully. Compressing weeks of expert survey work into 21.5 hours of unsupervised search is a real result [1], but it only compresses the survey and triage stage. The actual biological verification, cloning the candidate, expressing the protein, running the biochemistry that confirmed the ART array is transcribed into a set of distinct short RNAs [1], still runs on ordinary wet-lab timescales, days to weeks, and Anthropic is explicit that understanding ART's actual function is still ongoing work [1]. The agent compressed the needle-in-a-haystack search. It did not, and could not yet, compress the biology.

To make that computational-to-physical pipeline concrete for anyone who wants to generate visual explainer material from it, here is a video generation prompt covering the full arc, from raw sequence data through to a scientist confirming a result at the bench:

A 12-second cinematic sequence in a clean 3D-rendered technical explainer style, matching the pastel-blueprint diagram aesthetic of the post's other visuals rather than live-action photography. Shot 1 (0-4s): a slow overhead dolly across a glowing abstract grid of billions of tiny colored points representing a protein cluster database, with soft ambient blue backlighting, as a single point near the center pulses and turns bright amber, a thin label fades in reading 'Candidate Flagged'. Camera holds a steady, smooth glide with everything in soft focus except the pulsing point, which stays crisp. Shot 2 (4-8s): a match-cut zoom into that amber point, which resolves into a clean vector-style cutaway of a DNA strand with an evenly spaced repeat array, thin white leader lines drawing in short labels 'Reverse Transcriptase', 'Partner Gene', and 'Repeat Array' one at a time, rendered with flat color fills and soft ambient light, no harsh shadows. Shot 3 (8-12s): a smooth dissolve to a real molecular biology lab bench, shallow depth of field, warm overhead lab lighting, a human hand in a nitrile glove carefully placing a labeled sample vial into a rack, with a soft focus pull from the vial to a small monitor in the background showing a simplified data readout, ending on a calm static hold. Consistent soft, ambient lighting throughout with no harsh flicker, no camera shake, no distorted or extra fingers on the hand, no garbled or illegible on-screen text beyond the three short labels named above, no watermark or logo.

Step 7: What Feng Zhang and Independent Reviewers Actually Said

Treat expert reaction to this announcement carefully, because the two most substantive outside reactions found in circulation say meaningfully different things, and reporting only one of them would misrepresent how the scientific community actually received this.

Feng Zhang, a professor at MIT and the Broad Institute and one of the people most responsible for developing CRISPR-Cas9 into a practical genome-editing tool [8], reviewed Anthropic's pre-print directly. His full quoted statement is worth reading in full rather than trimmed to a soundbite: "This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research" [1]. That is a real, specific, on-record endorsement of the finding's biological interest from one of the few people on Earth qualified to judge whether a CRISPR-adjacent repeat pattern is worth investigating, and it corroborates Anthropic's own framing without simply repeating it.

A more skeptical, and equally important, reaction came from Lucas Harrington, who trained under Nobel laureate Jennifer Doudna and co-founded Mammoth Biosciences, a CRISPR-focused biotech company, and who works directly in genome mining for new CRISPR-adjacent systems. According to reporting on his response, Harrington's core objection is not that the finding is wrong, but that the method behind it is not new: "the method has been around for decades," and "similar systems have been known since 2008" [7]. His sharper point is about what actually constitutes a discovery in this field: identifying a sequence pattern is meaningfully different from understanding what a system does, and he argues the harder, unfinished part, "figuring out what a system actually does," is exactly the part Anthropic has not yet demonstrated [7]. He further suggests that framing an early, functionally unresolved finding as a major discovery "isn't helpful" to the field [7].

Both reactions can be true at once, and holding them together is actually the more accurate read than picking one. Harrington is correct that genome mining as a method is not new: the Mestre et al. 2022 survey of defense-associated reverse transcriptases used essentially the same clustering-and-triage approach to identify 42 diverse RT-linked defense system groups, entirely through human-driven computational analysis [13]. Zhang is also correct that the specific pattern found, an RT with this exact repeat-array architecture, appears genuinely novel and worth investigating on its own scientific merits, independent of who or what found it. The part of this story that is actually new is not the discovery method. It is that the method was executed autonomously, end to end, from database to candidate shortlist, by an AI agent swarm given nothing but a one-line research brief, in under a day, at a scale and speed no individual human team plausibly matches by hand. That is the claim worth evaluating, and it is a narrower, more defensible claim than "Claude invented a new biological system," which is not what Anthropic actually said.

Step 8: How This Compares to Prior AI-for-Science Systems

ART is not the first time an AI system has meaningfully accelerated a stage of scientific discovery, and putting it next to its closest predecessors makes clear exactly what is genuinely new here versus what is an extension of an established trend.

DeepMind's AlphaFold changed structural biology by predicting protein structures at a scale no experimental method could match: the AlphaFold Protein Structure Database now covers structure predictions for more than 214 million protein sequences, up from an initial 2021 release covering roughly 350,000 [20], and DeepMind reports the database has been used by more than 3 million researchers across over 190 countries [21]. AlphaFold's contribution was prediction at unprecedented scale, turning an experimentally expensive measurement (structure) into a fast computational one. It did not generate novel hypotheses about what an unstudied gene does; it answered a well-defined question (what shape is this sequence) far faster and more broadly than crystallography or cryo-EM ever could.

Google's "AI co-scientist," built on Gemini, tackles a different stage of the pipeline: hypothesis generation itself. In a 2025 paper, the system used a multi-agent "generate, debate, and evolve" loop to propose novel research hypotheses across biomedical goals, and its proposals were independently validated in real wet-lab trials, including predicting drug repurposing candidates for acute myeloid leukemia and identifying epigenetic targets for liver fibrosis [22]. Like the ART campaign, it keeps a human expert in the loop for both direction-setting and experimental confirmation, and it explicitly frames itself as augmenting rather than replacing a scientist's judgment [22].

SystemPipeline StageWhat It Actually DoesHuman Role
AlphaFold (DeepMind)Structure predictionPredicts 3D protein structure from sequence at massive scale; 214M+ structures in its public databaseScientists query and interpret predictions; no autonomous hypothesis generation
AI co-scientist (Google)Hypothesis generationMulti-agent 'generate, debate, evolve' loop proposes novel research hypotheses from a stated research goalScientist sets the research goal and runs wet-lab validation of proposed hypotheses
Claude / ART campaign (Anthropic)Search, triage, and hypothesis generationAutonomously searches an unlabeled sequence database, decides what counts as anomalous, and files its own candidate reports from a one-line briefScientist gives the initial prompt and performs 100% of physical lab verification

Set against both of these, what the ART campaign actually contributes is not a new capability category, it is a demonstration that the search-and-triage stage of discovery, the stage that has historically consumed the most unglamorous human labor, systematically combing an enormous, mostly unlabeled sequence space for the rare structurally anomalous entry, can now run with meaningfully less continuous human steering than either AlphaFold's structure prediction or the AI co-scientist's hypothesis generation required. AlphaFold answers a question you already knew to ask. AI co-scientist proposes a hypothesis from a research goal you already articulated. The ART campaign's agents decided, on their own, which of billions of otherwise unremarkable database entries was worth a second look at all, from a starting instruction that specified almost nothing about what to look for beyond "interesting reverse transcriptases."

Step 9: Engineering Lessons for Long-Horizon Autonomous Research Agents

Strip away the biology entirely and the ART campaign is a case study anyone building autonomous agents for research, security analysis, code auditing, or any other open-ended search-and-verify task should study directly, because the constraints that made it work generalize far outside molecular biology.

Treat the search space as a funnel, and instrument every stage of it. The 1.9 billion to 19 reduction described in Step 2 only works because each stage, sequence clustering, structural similarity, anomaly scoring, candidate report generation, human triage, has its own explicit success criteria and its own measurable pass-through rate. An agent system with no visibility into where candidates are getting filtered, or why, cannot be tuned; a system where every stage logs its accept and reject counts can be debugged and improved stage by stage.

Separate raw capability from tool-use reliability, and measure them separately. The 90% direct-recognition versus 32% tool-mediated-retrieval gap from Step 4 is the single most transferable lesson here [3]. If your agent's end-to-end success rate is low, do not assume the fix is a better or bigger model. Isolate whether the failure is in reasoning over data it already has, or in correctly retrieving and framing the right data in the first place, before spending budget on the wrong half of the problem.

Run many bounded sessions instead of one unbounded one. Anthropic's own multi-agent research system work found a 90.2% improvement over a single-agent baseline on internal research evaluations [23], and the ART campaign's 949-session structure is the same bet taken further: many independent, checkpointed attempts each with their own tool-use trajectory outperform one session trying to hold everything in a single, ever-growing context window, a conclusion Anthropic's own context engineering guidance reaches independently, through compaction, structured note-taking, and sub-agent decomposition with condensed handoff [24].

Make hypotheses a durable, queryable object, not a transient thought. The fact that Anthropic's team started studying which of the campaign's own hypothesis reports were worth testing, and fed that pattern back into future instructions [1], is a meta-lesson: a well-designed long-horizon agent system produces an audit trail that its own operators can mine for patterns, improving the next run without retraining anything.

Put the physical or irreversible action behind a human gate, always. Every part of this campaign that could be run computationally, database search, clustering, scoring, report writing, ran autonomously. Every part that involved an irreversible physical action, expressing a protein, running a biochemical assay, stayed with a human scientist [1]. That is not a limitation of the system, it is the design choice that makes the autonomous part trustworthy enough to run unsupervised for 21.5 hours in the first place. The same principle applies directly to agents that can execute financial transactions, modify production infrastructure, or send external communications: keep the exploratory, reversible search space fully autonomous, and keep a human explicitly in the loop for anything that cannot be undone.

claude-crispr-like-enzyme-discovery-life-sciences-agent-2026-checkpoint-relay-diagram.png

Here is a compact, standalone version of the stopping-condition logic referenced in Step 5, isolated because it is the single piece of this architecture most teams skip when they first build a long-running agent loop, and the piece whose absence causes the most expensive failures, runaway token spend, silently looping agents, and false-positive floods that erode a human reviewer's trust in the system:

python
/code # Standalone stopping / escalation condition logic, isolated from the full # orchestrator in Step 5 because it is the piece most long-running agent # systems skip when first built, and the piece whose absence causes the # most expensive failures: runaway token spend, silently looping agents, # and false-positive floods that erode a human reviewer's trust. from dataclasses import dataclass from typing import Optional @dataclass class CampaignState: elapsed_hours: float sessions_completed: int token_budget_used: int token_budget_total: int sessions_since_last_high_confidence_hit: int false_positive_rate_last_50: float # fraction of escalations later rejected # Tune these per domain. The values below are illustrative, not the # specific thresholds used in any real published campaign. LIMITS = { "max_hours": 24.0, "max_sessions": 1000, "max_token_fraction": 0.95, "diminishing_returns_sessions": 50, "max_false_positive_rate": 0.6, } def evaluate_stopping_condition(state: CampaignState) -> Optional[str]: """Returns a human-readable stop reason, or None if the campaign should continue running. """ token_fraction = state.token_budget_used / state.token_budget_total if state.elapsed_hours >= LIMITS["max_hours"]: return f"Time budget exhausted: {state.elapsed_hours:.1f}h elapsed" if state.sessions_completed >= LIMITS["max_sessions"]: return f"Session budget exhausted: {state.sessions_completed} sessions" if token_fraction >= LIMITS["max_token_fraction"]: return f"Token budget nearly exhausted: {token_fraction * 100:.1f}% used" if (state.sessions_since_last_high_confidence_hit >= LIMITS["diminishing_returns_sessions"]): return (f"Diminishing returns: " f"{state.sessions_since_last_high_confidence_hit} sessions " f"since the last high-confidence hit") if state.false_positive_rate_last_50 > LIMITS["max_false_positive_rate"]: # This is the reviewer-trust guardrail: if more than 60% of recent # escalations are being rejected by humans, the escalation # threshold itself needs tightening before the campaign continues, # not more raw search volume. return (f"Escalation quality degraded: " f"{state.false_positive_rate_last_50 * 100:.0f}% of recent " f"escalations rejected on human review") return None if __name__ == "__main__": state = CampaignState( elapsed_hours=21.5, sessions_completed=949, token_budget_used=215_600_000, token_budget_total=250_000_000, sessions_since_last_high_confidence_hit=12, false_positive_rate_last_50=0.3, ) reason = evaluate_stopping_condition(state) print("Stop campaign:", reason or "No, continue running")

Common Mistakes When Reading or Replicating This Kind of Result

Treating "novel pattern found" as equivalent to "new technology invented." ART's function is still unknown, and Anthropic says so directly [1]. The programmable, CRISPR-like potential is a hypothesis worth testing, not a demonstrated capability. Reporting or repeating this as "Claude invented a new gene-editing tool" is a real distortion of what was actually announced, and several outlets covering this story have been careful to draw exactly this distinction [6].

Ignoring the benchmark gap between direct pattern recognition and tool-mediated retrieval. Teams that read this story and conclude "the model is smart enough for autonomous discovery" while skipping the 90% versus 32% detail [3] will build agent systems that look capable in a demo, where data is handed to the model cleanly, and fail in production, where the model has to find and correctly frame the data itself.

Assuming more agent sessions alone fixes a broken harness. 949 sessions did not succeed because 949 is a magic number. It succeeded because each session had a bounded scope, a structured handoff format, and an explicit escalation condition. Scaling session count without also building the checkpointing and hypothesis-logging infrastructure described in Step 5 just produces 949 sessions of the same failure mode, faster and more expensively.

Skipping the human gate on irreversible actions to "save time." The single design choice most responsible for this campaign being trustworthy at all was that nothing physical happened without a human scientist personally performing it [1]. Teams building autonomous agents for domains with real-world consequences, financial, medical, infrastructure, that skip an equivalent human gate to move faster are trading away the exact property that made this specific result credible.

Conflating genome mining's maturity as a method with the campaign's actual novelty. Lucas Harrington is right that clustering-and-triage genome mining is a decades-old technique with real precedent, including the 2022 DRT survey [13] [7]. Missing that distinction, and either overclaiming a brand-new scientific method or dismissively concluding "this is nothing new" because the method is old, both miss the actual engineering claim, which is about autonomy and scale of execution, not method novelty.

Production Best Practices for Long-Horizon Discovery and Verification Agents

Anyone building a system in this shape, autonomous search across a large, mostly unlabeled space, with human-gated escalation for high-stakes findings, should treat the following as close to a checklist, drawn directly from how this specific campaign was structured.

Instrument every stage of your funnel with explicit pass and reject counts, not just a final output count, so you can tell whether a low yield is coming from an overly strict early filter or a genuinely sparse search space. Separate your evaluation of raw model reasoning from your evaluation of tool-use and retrieval reliability, and run them as distinct benchmarks, the same way the 90% versus 32% split was measured for direct sequence input versus tool-mediated retrieval [3]. Design your context handoff format before you scale session count, since a poorly structured summary handed from session N to session N+1 compounds errors across hundreds of sessions rather than resolving them.

Log every hypothesis or candidate your system generates, including the ones it rejects, in a durable, queryable store, because the rejected set is often more informative about your system's calibration than the accepted set, exactly the pattern Anthropic's team describes studying [1]. Build an explicit, checkable stopping condition into every autonomous loop, whether that is a token budget, a diminishing-returns threshold on new findings, or a maximum session count, rather than relying on a human to notice the system should stop. And draw a hard, non-negotiable line between what your agent can do autonomously and what requires a human to physically or explicitly execute, the same line Anthropic drew between computational search and wet-lab work, and keep that line even when it slows the system down, because it is precisely what makes the fast part trustworthy.

For teams building their own automated content or research pipelines rather than biology pipelines specifically, the underlying orchestration pattern, bounded sessions, structured handoff, human review gates before anything irreversible publishes or executes, is the same one worth studying regardless of domain. Miraflow AI's own content generation workflow, taking a topic through script, visuals, video, and thumbnail in one automated pipeline at miraflow.ai, runs on a version of the same idea at a much smaller scale: automate the exploratory, reversible stages end to end, and keep a clear point where a human reviews the output before anything goes live.

Frequently Asked Questions

Did Claude actually discover a new CRISPR system? Not exactly. Claude found a novel enzyme system, called ART, with a structural layout (a reverse transcriptase, a partner gene, and an evenly spaced repeat array) that resembles CRISPR's architecture, but ART's actual biological function has not yet been determined [1]. It is a candidate for a new programmable DNA system, not a confirmed one.

How long did the autonomous search take, and did any human intervene during it? The campaign ran for 21.5 hours across 949 agent sessions and roughly 215.6 million tokens, with Anthropic's own account stating their involvement was limited to the initial research prompt and the subsequent lab work, not the search itself [1] [3].

What database did Claude search? A protein sequence cluster database of roughly 1.9 billion clusters, drawn from the kind of large-scale, non-redundant reference collections (comparable in spirit to UniRef) used across modern computational biology [3] [16].

Is this the same as Anthropic's Claude Code or Claude Research features? It uses the same underlying tools, Anthropic's post specifically names Claude Science and Claude Code plus a custom multi-session harness [1], but the orchestration layer coordinating hundreds of parallel sessions over 21.5 hours is purpose-built for this kind of long-horizon research campaign, closer in spirit to the architecture described in Anthropic's multi-agent research system writeup [23] than to a standard chat session.

Why does the 90% versus 32% accuracy gap matter so much? Because it isolates two different failure modes that are easy to conflate. A model can be extremely good at recognizing a pattern once the right data is in front of it (90%+) while being much less reliable at retrieving and framing that data correctly on its own using tools and files (32%) [3]. Any team building an autonomous agent should benchmark both halves separately, because they require different fixes.

Did any outside scientists dispute the finding? Yes, and the dispute is about framing, not the underlying pattern. Feng Zhang called the finding "genuinely intriguing" and worth further investigation [1]. Lucas Harrington, a genome-mining specialist and Mammoth Biosciences co-founder, argued the discovery method itself is not new and that presenting an early, functionally unresolved finding as a major discovery is premature [7]. Neither disputes that the repeat-array pattern is real.

Could this approach be applied outside biology? The underlying architecture, bounded parallel sessions, structured context handoff, a durable hypothesis log, and a human gate before any irreversible action, is domain-agnostic. It maps directly onto security research (searching large codebases or infrastructure for anomalies), literature review at scale, or any other search-and-verify task where the space is too large for exhaustive human review but the cost of a false positive reaching a human reviewer needs to stay low.

Conclusion

The most useful way to read Anthropic's ART announcement is not as a story about AI discovering biology, it is a story about what a well-instrumented, checkpointed, human-gated autonomous agent system can accomplish over a long horizon when the search space is genuinely too large for exhaustive human review. A 1.9-billion-cluster database is not something a person combs through by hand, and the specific pattern that got flagged, a reverse transcriptase beside an evenly spaced repeat array in a jumbo phage, sits exactly in the blind spot a specialist's trained intuition is structurally likely to miss, because specialization means becoming excellent at recognizing what you already expect [1] [11] [[12]](https://www.cell.com/cell/fulltext/S0092-8674(20)31306-4).

What actually made the 21.5-hour, 949-session run work was not raw model intelligence alone, it was the architecture around it: a funnel with instrumented stages, checkpointed session handoff, a durable hypothesis log, an explicit stopping and escalation condition, and, most importantly, a hard, unmoved line between what the agent could do computationally and what stayed exclusively in human hands [1]. Feng Zhang's endorsement of the finding's scientific interest and Lucas Harrington's insistence that method maturity and functional understanding both still matter are not actually in conflict, they are two halves of the same accurate picture: something real and worth investigating was found, using a method that is not new, executed with a degree of autonomy that is. Whether you build agents for biology, security, or content, that last part, the specific engineering discipline that made a nearly thousand-session autonomous run trustworthy enough to publish, is the part of this story worth studying closely.

References

  1. Claude discovers a novel enzyme system with CRISPR-like repeats — Anthropic
  2. Anthropic technical report / pre-print PDF: array-associated reverse transcriptases
  3. Anthropic says Claude found a new enzyme system with CRISPR-like repeats — The Next Web
  4. Claude scans 200,000 enzymes to uncover new CRISPR-like system hidden in phages — Interesting Engineering
  5. Anthropic's Claude AI discovers CRISPR-like enzyme system — Quartz
  6. Claude Found a Mysterious CRISPR-Like System, but Anthropic Can't Say What It's Capable of — Gizmodo
  7. Anthropic says Claude discovered a new enzyme system, but CRISPR researchers call it routine genome mining — The Decoder
  8. Feng Zhang — Wikipedia
  9. Mojica, F.J.M. (2016). The discovery of CRISPR in archaea and bacteria. The FEBS Journal
  10. Lander, E.S. (2016). The Heroes of CRISPR. Cell
  11. Diversity-generating Retroelements in Phage and Bacterial Genomes — Microbiology Spectrum
  12. [Millman, A. et al. (2020). Bacterial retrons function in anti-phage defense. Cell](https://www.cell.com/cell/fulltext/S0092-8674(20)31306-4)
  13. Mestre, M.R. et al. (2022). UG/Abi: a highly diverse family of prokaryotic reverse transcriptases associated with defense functions. Nucleic Acids Research
  14. Mobile Bacterial Group II Introns at the Crux of Eukaryotic Evolution — Microbiology Spectrum
  15. The biology of jumbo phages — Nature Communications
  16. Suzek, B.E. et al. (2007). UniRef: comprehensive and non-redundant UniProt reference clusters. Bioinformatics
  17. Suzek, B.E. et al. (2015). UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches. Bioinformatics
  18. Steinegger, M. & Söding, J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology
  19. van Kempen, M. et al. (2024). Fast and accurate protein structure search with Foldseek. Nature Biotechnology
  20. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences — PMC
  21. AlphaFold reveals the structure of the protein universe — Google DeepMind
  22. Gottweis, J. et al. (2025). Towards an AI co-scientist. arXiv
  23. How we built our multi-agent research system — Anthropic
  24. Effective context engineering for AI agents — Anthropic
  25. Effective harnesses for long-running agents — Anthropic