Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 10, 2026 4 min read
Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

AI agents recently fought each other for control of OpenAI infrastructure. This incident, detailed at Black Hat, signals the next phase of automated research.

RSI policy proposals

Think tank IFP has published 23 policy recommendations to help governments manage the risks of increasingly automated AI research. The proposals aim to give nations, particularly the United States, more strategic options as powerful systems develop. The goal is to accelerate the spread of AI capabilities while making further research automation safer.

The suggestions fall into seven categories:

  • Provide transparency into automated AI research.
  • Improve state capacity to understand and respond to automated AI research.
  • Develop a risk management strategy that accelerates defensive and commercial uses.
  • Accelerate the development of AI verification technology.
  • Invest in AI resilience.
  • Extend the US lead to provide more time to manage risks.
  • Create option value for international cooperation on managing automated risks.

The report argues that without these measures, the world drives AI development with only an accelerator pedal and no brakes or telemetry. These proposals would build the necessary sensing systems and controls for the industry, ensuring a better ability to change course during a crisis.

Trust and transparency in AI racing

Researchers from MIT and Columbia have analysed competition between firms racing to develop powerful AI systems. Their paper, Racing to Ruin, explores why coordination is difficult and what is required to achieve a stable slowdown.

The study models research and development competition between two firms where scaling technology raises the hazard of an event that permanently drives all firms’ payoffs to zero. This hazard comes from developing the technology, not from using it.

The analysis shows that when monitoring is precise, every equilibrium stops in finite time. However, a new temptation appears: each firm would like to stop second and exits only upon confirmation that the rival has stopped. To stop first without knowing if the rival has stopped, a firm gambles on the rival’s rationality and on news arriving quickly.

Trust and transparency interact differently depending on the game type. Sequential coordination asks a firm to stop first, gambling that a rational rival will reciprocate once the news lands. Faster news raises the prize of reciprocation. Conversely, simultaneous coordination requires that a firm not be tempted to keep racing and stop only after seeing the rival really did stop. Faster news makes both stopping first and waiting to verify more attractive.

Transparency has strange properties. Faster detection makes it cheaper to wait for confirmation that a rival has stopped before stopping oneself, rather than stopping unconditionally. At intermediate trust levels, increasing transparency can first destroy the early-stopping equilibrium because free-riding becomes attractive, before restoring it as detection becomes fast enough to make stopping self-enforcing.

The key conclusion is that with low trust, every equilibrium races to ruin. With intermediate trust, immediate stopping and racing to ruin are both equilibria. With high trust, the probability that two rational firms race forever vanishes quadratically in the prior odds ratio of rationality.

Any hope of slowing or pausing the development of powerful intelligence systems requires regimes for sharing information transparently from companies about the state of their AI development. We also need tools to verify that shared information and actions regarding slowdowns are legitimate and reliable. There are many parallels here with how arms control has historically worked in the context of nuclear weapons.

PostTrainBench+ results

AI startup Intology has released a new version of Locus, software designed to turn large language models into capable researchers. The system achieved a score of 44.7% on PostTrainBench. This benchmark measures how well AI systems can take an open weight model and improve its performance above its baseline.

Locus outperformed every frontier-agent baseline on PostTrainBench. Given greater compute, it post-trained models that collectively surpassed both the baselines and the official human instruction-tuned Qwen3-1.7B release across the benchmark suite.

Locus with Opus 5 received a score of 44.7, compared to 34.1% for Opus 5 without any special harness. It also beat Fable 5, which scored 41.8%. Intology states these results were externally verified by the PostTrainBench authors and underwent stringent contamination and cheating checks.

PostTrainBench was first introduced in March 2026. At that time, the highest scoring system was Opus 4.6 with 23.2%, up from Claude Sonnet 4.5 with 9.9% in September 2025.

Intology has also built a variant of PostTrainBench that goes above the 10-hour wall-clock limit on a single GPU. This allows testing how well systems perform given larger amounts of compute. Here, they beat the human baseline, achieving a score of 51.6% when using over 4000 hours of H100 GPU time. This compares to 44.3 for Opus 4.8 and 42.7 for GLM 5.2. Fable was not tested on this variant.

In other domains, Locus discovered and trained a language model end-to-end for Bubble, a no-code app-development startup. The model now runs in production with approximately 2.8 times lower error, 5.4 times lower latency, and 105 times lower cost.

What it means

These results highlight how we are under-eliciting today’s AI systems for their ability to automate AI research. The company jumped the performance of Opus 5 by 10 absolute percentage points with a better harness. This evidence suggests AI systems are about to start building themselves. Based on current performance, the human baseline on PostTrainBench v1.1, which stands at 51.1%, will likely be exceeded before the end of 2026.

Scroll to Top