Open-source “BootLoops” harness supports AI models in performing precise scientific calculations

Professor Matthew Schwartz of Harvard University and 19 co-authors produced 36 manuscripts across 18 fields in three months using an open-source harness…

By Vane October 3, 2026 3 min read
Open-source “BootLoops” harness supports AI models in performing precise scientific calculations

Professor Matthew Schwartz of Harvard University and 19 co-authors produced 36 manuscripts across 18 fields in three months using an open-source harness called BootLoops.

The source code is available on GitHub.

How the harness works

Schwartz stopped trying to use large language models like Claude to act as human researchers. Instead, he focused on tasks that play to the strengths of current AI. He calls these “Claude-shaped problems”.

The tool bridges gaps between disciplines. Human knowledge is fragmented. Individual fields extend in different directions while the spaces between them remain unexplored. A lab might spend 20 years studying one set of genes with a single method, leaving neighbouring genes and alternative approaches untouched. BootLoops is designed to fill those gaps by drawing connections across the jagged frontiers of knowledge.

Schwartz uses the concept of the “convex hull” to illustrate this. The results often only became scientifically valuable once domain experts stepped in to set the direction.

Results from particle physics to linguistics

The team began with particle physics calculations, specifically scattering amplitudes and elliptic integrals. Within weeks, Claude computed 30 integrals using BootLoops. Fifteen reproduced known results and fifteen were computed for the first time.

The search then expanded to other fields. In ecology, Claude solved a 20-year-old equation from neutral biodiversity theory that had previously been impossible to compute at scale. When applied to data, the results showed that tree species composition on Barro Colorado Island in the Panama Canal is changing 4.5 times faster than the theory allows. Ecologist James O’Dwyer then helped turn that finding into a better predictive model.

In population genetics, the team analyzed 5.7 billion mutation pairs from the 1000 Genomes Project and found evidence for a mechanism called gene conversion. Other projects included an AI data editor for economics journals that automatically checked 4,452 replication packages. This work was published as an NBER Working Paper. The team also built a word stress database covering 6,072 languages with the help of three linguists.

Planning research is becoming difficult

Schwartz describes how fast AI is changing science, to the point where planning ahead is becoming nearly impossible. Applying for a three-year grant to fund a calculation that an AI model might solve overnight makes little sense.

Training PhD students has also become an open question. Since his earlier post on Vibe Physics, that feeling has only grown stronger. In some fields like computer science, the disruption is alarming. Two years ago, he would have called a “Python for Engineers” course “essential”. Today it is “unnecessary” because Claude can handle those tasks.

Building machine learning models to study physical phenomena is another area AI can now take over. Even deep knowledge of neural networks offers limited returns when Claude can implement the current state of machine learning research on command.

Warnings on reliability and expectations

Schwartz warns about the models’ weaknesses. Claude likes to declare victory too early. Phrases like “done, with one asterisk” often mean “not done at all”. The model misjudges how long tasks will take and tends to brute-force calculations instead of finding more elegant solutions. Automated checks aren’t reliable, and Claude’s conclusions can be wrong even when its calculations are correct.

Beyond reliability, the model gravitates toward old, heavily cited debates rather than genuinely new questions. The projects were “compute- and token-intensive”, Schwartz says.

Schwartz also sees the focus on big math problems and headlines as risky. Unrealistic expectations could distract from productive applications that already work today. Real progress will still be built on solid foundations, just faster. The scientific method itself isn’t threatened, and human guidance and taste remain indispensable.

What it means

The work shows that AI can accelerate specific technical tasks, but it cannot replace the human role in defining questions and verifying results. Scientists must still provide the direction and taste required to ensure the output is scientifically sound.

Scroll to Top