2026 in LLMs (so far)

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase. We do…

By Vane September 28, 2026 3 min read
2026 in LLMs (so far)

Simon Willison closed his keynote at the WeAreDevelopers World Congress North America in San Jose with a chronological review of large language model developments so far in 2026. The full video is available on YouTube, and his annotated slides follow below.

The November turning point

Willison treats November 2025 as the true start of the year. That month saw the release of Claude Opus 4.5 and GPT-5.1. These were incremental updates to existing models, but they crossed a threshold where coding agents moved from unreliable to usable.

Claude Code had existed since February 2025, while Codex was younger. The new models paired with their agent harnesses improved from making frequent errors to being reliable enough for daily work.

A pelican on a bicycle

The speaker has used a specific prompt to test models for years: “Generate an SVG of a pelican riding a bicycle”. He describes it as the world’s stupidest benchmark, yet it remains a challenge because drawing pelicans is hard, drawing bicycles is hard, and pelicans cannot ride bicycles.

In November, Claude could not draw a bicycle properly. GPT-5.1 produced a slightly better bicycle frame and a better beak, but the results were still poor.

Warelay and the new year

Also in November, the first commit appeared in an obscure GitHub repository called Warelay. The commit added an MIT license file.

The December holidays followed, giving individual developers time to experiment with the new coding agent combinations. Many realised how much more these tools could do compared to before. By January, excitement grew as people began putting the technology into action.

A change of ambition

For every previous year, Willison’s New Year’s resolution was to take on fewer projects and focus on the most important tasks in existing ones. He has held this stance as long as he can remember.

This year he decided to go the other way. Since the old approach never worked, he chose to be more ambitious. He planned to take on as many new projects as he liked.

Willison noted he has a lot of plates spinning at the time of writing. He stated that the only way to find the limits of this technology is to keep pushing them until they break.

Predictions and reality

Willison shared predictions for the next year, three years, and six years on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal.

His LLM predictions were unambitious in hindsight. He said it would become undeniable that LLMs write good code. He considers that point reached now.

He predicted we would finally solve sandboxing. Around 40 of the 277 sessions at the conference touched on sandboxing or agent security, showing significant effort in that area.

He predicted a “Challenger disaster” for coding agent security. There has been a lot of noise around agent security this year, but the specific disaster he foresaw, where coding agents are hijacked and cause real-world economic damage, has not occurred.

He also made a joke prediction that the Pope would weigh in on the economic impact of LLMs.

Kākāpō parrots

Willison predicted that New Zealand’s Kākāpō parrots would have an outstanding breeding season. These are flightless nocturnal parrots that live in New Zealand. They look dumpy but are beautiful. There were only 236 of these parrots in the world at the start of the year.

Kākāpō only breed when Rimu trees have a big fruiting season. That event had not happened in four years, but the Rimu fruit looked excellent this year.

Deep Blue

On the podcast, Willison and Adam Leventhal coined the term “Deep Blue”. It describes the feeling of AI-induced ennui where software engineers become listless because the AI can do anything.

This theme ran throughout the year and was discussed by several speakers at the conference. Willison noted he has never had a year where everything changed so quickly and dramatically.

He stated a lot of his work this year involves coming to terms with those changes and what they mean for his profession.

AI mania

In January, Willison suffered from what he calls AI mania. This is not the same thing as the AI mania described in other contexts.

Scroll to Top