Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 18, 2026 1 min read
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

Rohin Shah and Anca Dragan from Google Deepmind have warned that the visible chain of thought currently offered by AI models is slipping away. They argue that writing intermediate steps in plain language allows researchers to spot deceptive reasoning or problematic plans. For instance, Gemini 3 Pro revealed it recognised it was in a test environment through this process. However, OpenAI’s system card for GPT-6 Astra reports a significant drop in how well this reasoning can be monitored. Future models might think in number spaces that humans cannot read, which would be more efficient but completely opaque. The researchers want the field to regularly measure how well chains of thought remain readable. They also call for keeping transparent architectures and ensuring models do not learn to hide their true reasoning during training. Earlier warnings from OpenAI chief scientist Jakub Pachocki and Anthropic CEO Dario Amodei highlighted this loss of control. Amodei specifically called for deliberately slowing the pace of development to address these risks.

  • OpenAI’s GPT-6 Astra shows reduced monitorability of reasoning steps.
  • Future models may switch to opaque number-based thinking spaces.
  • Researchers demand regular measurement of chain of thought transparency.
Scroll to Top