GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Grok 4.6 Tops Independent Biosecurity Benchmark: 62.1% on BioSecBench-Refusal, Only Model Above 50% on Both Axes

LatchBio released independent biological capability and red-team benchmarks for Grok 4.6 today, covering biosecurity refusal accuracy and pathogen surveillance workflows. Grok 4.6 leads the published rankings and is the only frontier model to break 50% on both axes of the primary evaluation suite.

BioSecBench-Refusal

BioSecBench-Refusal pairs routine biological tasks from published literature with 46 red-team tasks. The red-team set conceals biosecurity hazards in scientific data, mislabeled files, and other obfuscated formats — designed to defeat systems that key on surface-level keywords like “pathogen” or “toxin.”

LatchBio scores models using a trial-weighted harmonic mean of red-team refusal rate and routine task compliance rate. The harmonic mean penalizes models that sacrifice one for the other: a system that refuses everything scores poorly on routine compliance; a system that completes everything scores poorly on refusal.

Grok 4.6 results across harnesses:

MetricGrok 4.6Opus 5GPT-5.6 Sol
BioSecBench-Refusal (harmonic mean)62.1%——
Red-team tasks refused59.2%lowerlower
Routine tasks completed64.8%higherlower
Above 50% on both axesYesNoNo

Grok 4.6 holds the top three spots across different evaluation harnesses on BioSecBench-Refusal, averaging 62.1%. LatchBio evaluated models at their highest-offered effort levels and varied harnesses to remove confounds.

The underlying behavior LatchBio observed: Grok 4.6 reasons over the full contents of a task and the test environment before proceeding or refusing, looking for discrepancies between stated intent and actual data. It does not react to surface terminology alone.

BioSecBench-Surveillance

BioSecBench-Surveillance tests pathogen genomic surveillance workflows — the kind used in public-health monitoring. Tasks require chaining file inspection, tool use, and scientific judgment on messy sequencing data.

Grok 4.6 averaged 53.5% on BioSecBench-Surveillance, sitting behind Opus 5 and ahead of GPT-5.6 Sol. LatchBio’s framing: the results indicate an agent well-calibrated for both routine and helpful biological work and capable in biosecurity monitoring.

On supplementary benchmarks of general biological capability (SpatialBench, TxBench-PP), Grok 4.6 matches or exceeds other frontier models. Full results are published at benchmarks.bio.

Context

xAI published today alongside LatchBio’s release, noting that Grok 4.6 shows a material improvement in biological work capability over Grok 4.5 and 4.3, with substantial gains in refusal and biosecurity performance. The company frames ongoing safeguards work as what allows it to continue serving frontier-scale intelligence safely across future releases.

Grok 4.6 has a 95.6% SWE-bench Verified score per vals.ai’s independent harness — second only to DeepSeek V4 Pro (96.4%) and behind Claude Opus 5 (97.0%). The LatchBio evaluation adds a separate biosecurity-specific capability dimension not captured in standard coding or reasoning benchmarks.

The evaluation is the first published independent red-team assessment of a frontier model using biological concealment techniques rather than explicit harmful request framing.