GPT-56T 861 —
MUSE-SPK 837 —
GLM-52 792 -4.8%
GLM-5 779 -9%
GPT-56SC 774 -4.9%
CL-OP55X 767 -5.1%
GROK-46H 767 -5.1%
DSK-V4PH 767 -5.1%
QWEN-38X 751 -8.9%
KIMI-K3X 745 -0.1%
GPT-6A 742 -9.5%
GEM-38FH 712 +0.1%
CL-FAB5H 685 -5.9%
CL-OP5H 660 -6.1%
CL-OP5X 658 -5.9%
GEM-37FH 657 -5.9%
CL-OP46H 644 -6%
CL-OP55H 641 -6%
CL-OP47H 633 -6.4%
GPT-56S 592 -0.7%
GEM-36FH 583 —
CL-OP48H 580 —
GEM-35FH 567 —
CL-OP47 556 -0.4%
INKL 531 —
GPT-55H 525 —
GEM-31P 500 -0.4%
GEM-3P 496 —
CL-OP46 492 -0.2%
CL-OP48 485 -0.2%
GPT-55 408 —
GPT-56T 861 —
MUSE-SPK 837 —
GLM-52 792 -4.8%
GLM-5 779 -9%
GPT-56SC 774 -4.9%
CL-OP55X 767 -5.1%
GROK-46H 767 -5.1%
DSK-V4PH 767 -5.1%
QWEN-38X 751 -8.9%
KIMI-K3X 745 -0.1%
GPT-6A 742 -9.5%
GEM-38FH 712 +0.1%
CL-FAB5H 685 -5.9%
CL-OP5H 660 -6.1%
CL-OP5X 658 -5.9%
GEM-37FH 657 -5.9%
CL-OP46H 644 -6%
CL-OP55H 641 -6%
CL-OP47H 633 -6.4%
GPT-56S 592 -0.7%
GEM-36FH 583 —
CL-OP48H 580 —
GEM-35FH 567 —
CL-OP47 556 -0.4%
INKL 531 —
GPT-55H 525 —
GEM-31P 500 -0.4%
GEM-3P 496 —
CL-OP46 492 -0.2%
CL-OP48 485 -0.2%
GPT-55 408 —
← Back to feed

Harvard Physicist Found the Right Problems for Claude: 15 New Feynman Integrals, 36 Papers in 90 Days

The problem with using AI for science, Harvard particle physicist Matthew Schwartz argues, is impedance mismatch. LLMs can do a great deal, but asking them to work like a human collaborator wastes what they are actually good at. His fix: find “Claude-shaped problems.”

In an October 1 guest post on the Anthropic Science Blog, Schwartz described what he found when he stopped fighting the mismatch and built a toolkit around it.

BootLoops

Schwartz built BootLoops while working as a visiting researcher at Anthropic, using Claude Fable 5. The toolkit is an open-source harness for exact calculations in quantitative science — it is not an Anthropic product and is maintained by Schwartz on GitHub.

The origin was scattering amplitudes: the theoretical bridges between particle collision data and the underlying physics. At their core are Feynman integrals — multidimensional integrals whose hardest instances take a research group years. Schwartz identified the semi-numerical S-matrix bootstrap as especially well-suited to AI assistance. The method works by imposing physical constraints until only one answer remains, then using high-precision numerical checks (sometimes to 1,000 digits) to pin down remaining coefficients. It draws on mathematics, physics, and computer science that no single person has fully mastered, and every result is independently checkable by running two scripts.

That last property matters: Claude cannot reliably judge whether a scientific result is interesting, but it can execute long, technically correct computation chains, and anyone can verify the output without domain expertise.

What Happened

Schwartz started by asking Fable 5 to port existing scattering-amplitude software — spread across Wolfram Language, C++, Python, and Julia, some with no public code at all — into a unified framework. Then he asked it to find unsolved integrals it could compute.

The model initially confined itself to the simplest function class: logarithms. Schwartz pushed it to the next level, elliptic functions, which are substantially harder. Only a handful of elliptic Feynman integrals had ever been computed before, and none by the bootstrap method alone. The difficulty, Schwartz writes, was not that the methods wouldn’t work — it was that no one had tried, because the expertise required was spread across too many people for any one team to assemble.

Claude assembled it. As the toolkit grew, integrals fell one after another.

Final count in physics: 30 integrals BootLooped end to end — 15 reproductions of known results using a new method, and 15 that had never before been computed.

Cross-Disciplinary Reach

A structural property of mathematics is that the same equations appear across fields. BootLoops found this automatically. Bayesian evidence integrals in population genetics share form with Feynman integrals from physics. Finite-field methods for Feynman integral reduction apply to evolutionary biology. Claude made those connections; Schwartz then recruited domain experts to test whether the results were scientifically relevant, not merely technically correct.

In ecology, the toolkit built a demographic model of forest inventory species from US plot data, classifying each species on longevity, growth, and recruitment. In population genetics, it solved a 30-year-old integral expression for how natural selection shapes rare mutations, applied to the gnomAD catalog of human genetic variation.

Running the whole program meant managing: 36 manuscripts in 18 fields with 19 coauthors over three months, drawn from roughly 400 candidate problems. Schwartz ran multiple Claude Code sessions on Google Cloud virtual machines, one per project plus a master session to coordinate them, allocate compute, and validate results. Subagents ran computations in the background with intermediate results stored in markdown files.

Where Claude Failed

Schwartz is explicit about the failure modes. Claude could declare a task finished before resolving its central problem. It misjudged how long work would take. It would sometimes pursue long calculations when building a new tool would have been more effective.

More fundamentally: it could not evaluate whether a result was scientifically interesting. Domain experts — biologists, ecologists — were often impressed by the technical correctness but uncompelled by the science. The physicist still had to supply the intuition about which problems were worth solving. That division of labor is the central point.

The Model Limitations

BootLoops ran on Fable 5, not a specialized scientific model. Schwartz notes Fable 5’s classifiers interrupted individual sessions — the multi-agent architecture was partly a workaround, so that classifier blocks shut down one subagent rather than the whole session.

The harness is model-agnostic by design and will work with other frontier models. As the underlying models improve, Schwartz expects the harness to improve independently, through human interaction rather than lab-scale training runs.

BootLoops is available at bootloops.ai and on GitHub.