In this article
Nathan Lambert and Tom Zick have launched Trillium Labs, a nonprofit dedicated to conducting high-risk AI research in public view rather than behind closed doors.
The founders argue that keeping advanced models locked inside private labs prevents the wider scientific community from scrutinising them effectively. Their approach involves publishing full details of experiments so peers can study and replicate the work.
The problem with secrecy
Lambert states that frontier AI companies currently operate in a way that hinders collective effort. He believes transparency regarding model construction and tuning is essential for reducing danger.
“Over the past few millennia, humanity has had the scientific method in our toolbox as a way to mitigate harms and build better futures,” Lambert tells WIRED. “The current closed trajectory of frontier AI development is taking us a step backwards.”
Most powerful models from firms like OpenAI and Anthropic remain accessible only via applications or application programming interfaces. This restriction often hides how the systems are built and how they behave.
Some organisations, particularly in China, offer models that users can download and run on local hardware. Xiaomi recently released live details of a major training run for one of its systems. Meanwhile, researchers at Stanford are pretraining the AI model Marin openly.
The industry is currently divided over which strategy offers the best safety. Powerful frontier models can now automate the discovery of software vulnerabilities and probe systems for weaknesses. Recent high-profile hacking campaigns have increased scrutiny on these capabilities.
Supporters of limited access argue that power must stay with a trusted few. Lambert and Zick contend that a shared understanding of risks benefits everyone.
What they have done before
Lambert previously worked at Ai2, a research lab known for publishing data and training methods alongside its models. He formerly worked at Hugging Face, runs a technical blog, and founded Truly Open Models, an initiative encouraging US companies to release more open systems.
Zick worked at Harvard University and helped Charles Schwab create policies around responsible AI.
The two met over Zoom during the COVID-19 pandemic while attending UC Berkeley as graduate students. They founded the new nonprofit after observing a disconnect between industry research and academic work. Professors and students often cannot replicate big company projects due to a lack of resources.
What the lab will study
Trillium Labs will initially focus on post-training, which involves fine-tuning large models after they have been built. Another key area is recursive self-improvement, a process where AI contributes to its own research.
The prospect of ongoing progress continuing indefinitely, potentially leading to a loss of human control, has alarmed many researchers. The issue gained mainstream attention earlier this month when an Anthropic researcher left the company and warned that recursive self-improvement could pose an existential threat to humankind.
The nonprofit will also examine reinforcement learning, a method that rewards a model for good results and punishes it for bad outcomes. This approach has made agents far more capable but also more inclined to do unexpected things. The team will study how this method shapes the character and behaviour of AI models, noting problems when a model becomes overly sycophantic.
“To understand something like how reinforcement learning scales in post-training, you need significant compute and a lot of careful experimentation,” Zick says. She adds that publishing details of how reinforcement training runs work could yield surprising insights as outside researchers scrutinise the work.
Funding and goals
The lab has raised an undisclosed sum from Schmidt Sciences, Halcyon Futures, and others. The founders state they aim to raise between $40 million and $100 million in total. They plan to spend $30 million on training over the next 18 months.
“I’m a massive fan of much more transparency than we currently have in R&D,” Tim Fist, director of emerging technology policy at the Institute for Progress, tells WIRED.
Lambert and Zick hope Trillium Labs will add nuance to the wider discussion on how best to build AI.
“We’re in an era of AI discourse dominated by a few world views,” Lambert says. “We believe that the scientific method and careful measurement of recent events is the best way to understand new behaviors of AI models.”




