RSS Amplifier

Blog

LatchBio

Intelligence Infrastructure for Biology

blog.latch.bioSource feed ↗20 posts

Live Last read · last published · next check

Written by

Latest posts

Grok 4.6 is a Frontier Biology Model

Following yesterday's Grok 4.6 release, we ran it against our short-horizon biology tasks on benchmarks.bio. Across 1716 trajectories we generated, we find that Grok 4.6 sits roughly at Opus 5/GPT-Sol-5.6 level of intelligence while being far cheaper.

Open Source Classifiers Do Not Stop AI Bioweapon Generation

A few days ago Mistral released Shieldstral, its new safety classifier, without disclosing its performance on major biosecurity related tasks, such as viral engineering.

How Good is Opus 5 at Biology?

An in depth analysis across our benchmark suite

Surfacing Benchmark-Maxxing in Kimi-K3

Today we released results for Kimi-K3, an open-source LLM boasting GPT-5.6/Mythos-level coding-benchmark scores, across our short-horizon therapeutics and -omics benchmarks on benchmarks.bio.

VariantBench: An Agentic Benchmark for Genetic Variant Discovery and Interpretation

A verifiable benchmark for variant discovery, statistical genetics, and personal genomics

Why An Open Source Harness Outperforms Claude Code On Frontier Biology Tasks

In current generation models, behavioral priors introduce by harnesses such as Claude Code, and Codex cause substantial performance swings on our benchmarks.

Benchmarking AI Agents on Pathogen Genomic Surveillance

A verifiable benchmark for practical decisions about taxonomy, variants, AMR, source tracking, anomaly detection, and engineered sequences in pathogen genomic surveillance workflows.

Benchmarking Refusals in Agentic Biology

A paired benchmark for capability and caution in agentic biosecurity risk assessment

Latch MCP: Agent-Native Data Infrastructure

Access the Latch platform using Claude Code, Codex, or any agent

LatchBio acquires TwentyTwo to form Latch Biosecurity

Welcoming Harmon Bhasin, Evan Seeyave, and John Wang to the team.

Benchmarking AI Agents on Long-Horizon Single-Cell Biology

A verifiable benchmark for recovering scientific conclusions from raw single-cell data

Benchmarking AI Agents on Small-Molecule Preclinical Pharmacology

A verifiable benchmark for practical decisions about potency, mechanism, exposure, safety, and efficacy

EpiBench: AI agents still struggle with epigenomics analysis

Benchmarking frontier models on practical CUT&Tag/CUT&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows

Human Verification of SpatialBench

Two rounds of independent expert attempts define a verified subset of 115 spatial biology tasks and expose ambiguity in benchmark specification and grading.

Verifiable Benchmarking of Long-Horizon Spatial Biology

Evaluating whether AI agents can recover complex scientific conclusions from raw spatial biology data

scBench Updates: Opus 4.7, GPT 5.5, Gemini 3.1

Benchmarking frontier models on messy, real-world single cell data analysis

Agentic biology is shaped like software

Biology will not jump straight to autonomous AI scientists. Like software, it will first accelerate where work is executable, feedback-rich, and economically bottlenecked: data analysis.

New Frontier Models Are Faster, Not More Reliable, at Spatial Biology

Overall accuracy for GPT-5.5 and Opus 4.7 remains flat on SpatialBench. Scientist-reviewed trajectories reveal persistent gaps in assay-aware biological judgment.

scBench: Can AI Agents Analyze Real-World Single-Cell Data?

394 verifiable problems across six sequencing platforms: 10X Chromium, BD Rhapsody, CSGenetics, Illumina, MissionBio, ParseBio

SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

first contribution towards a broader family of benchmarks spanning major data types // 146 verifiable problems spanning 5 spatial platforms and 7 task categories