These startups are chasing the next big thing in LLMs

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 10, 2026 8 min read
These startups are chasing the next big thing in LLMs

MIT Technology Review’s What’s Next series examines industries, trends, and technologies to offer a first look at the future. Read the rest of the series here.

AI researchers at Google published a paper titled “Attention Is All You Need” in the summer of 2017. It introduced a new neural network type called a transformer. The design proved effective at processing long sequences of data, particularly text.

Nine years later, transformers power every major large language model available. “The entire AI industry is built on transformers,” says Justin Dangel, cofounder and CEO of the AI startup Subquadratic. “They are one of the most important innovations in the history of computer science, and they’ve changed the world.”

Transformers are starting to show their age. Recent advances in large language models, such as reasoning models and their ability to handle large amounts of input simultaneously, are not neat extensions of that core technology. They are workarounds that patch over fundamental flaws.

Scientists and engineers are now asking what comes next. Large language models are not going anywhere, but the way they get built is up for grabs. MIT Technology Review dubbed this future generation of models LLMs+ in this year’s list of the 10 things that matter in AI.

A wave of startups hopes to push the boundaries of this boomtown technology. Some will no doubt fail—but they have everything to play for and far less to lose than the companies at the front of the pack today.

Strength in numbers

The key strength of transformers lies in a mechanism called dense attention, which encodes the meaning of a block of text in a series of numbers. The process involves comparing every word, or part of a word known as a token, with every other word via a form of multiplication.

Dense attention can capture the meaning of text with remarkable accuracy. But as the length of that text grows, the number of computations needed to process it adds up fast. A document 10,000 words long might require a transformer to perform 50 million multiplications. That is the main reason large language models suck up so much power.

The costs are huge. OpenAI is set to spend $50 billion on computing this year, according to the company’s president, Greg Brockman. And the International Energy Agency predicts that the total amount of electricity consumed by data centers will double by 2030.

Transformers struggle with what many of the latest models are designed to do. Because they process text word by word, they are not great at keeping track of a lot of information at once. Their context window cannot get too large. If large language models are to carry out harder tasks, they will need to take in larger amounts of data: a whole library of documents, an entire code base, or in the case of agents, output from other large language models.

Reasoning models work by writing notes to themselves in a kind of scratch pad known as a chain of thought, then reading them back. This adds to the amount of data the system must manage.

As large language models get bigger and better, transformers have become a bottleneck. The technology’s key strength is now a limitation.

Here are four new ideas for how to solve the transformer problem—innovations that could change large language models for good, making them faster, far more efficient, and perhaps even smarter.

01: Rethinking attention

An obvious way to make large language models faster and cheaper is to tackle the problem head on and change the way attention works. Swapping out dense attention for a mechanism called sparse attention, which runs calculations on only some pairings of words in a block of text instead of all of them, can radically reduce the amount of computation large language models need to do.

Researchers have come up with plenty of sparse attention mechanisms over the years. The problem is that none of them were as good as dense attention at capturing meaning.

That might have changed. Subquadratic, a startup based in Miami, claims it has invented the first sparse attention mechanism that rivals top mainstream large language models on a handful of tasks, including search and coding. It is a huge claim, and some people in the industry remain skeptical.

Subquadratic says its model, SubQ, works by figuring out on the fly, for each piece of text it is given, which words matter and which do not. The company also claims that thousands have signed up to its waitlist and plans to make the model widely available soon.

Meanwhile, Manifest AI, a startup based in San Francisco, is coming at the problem from a different angle. Instead of changing how attention works, it is replacing it with something else.

It has developed a mechanism it calls power retention, which stores only the most relevant information for a given task and ensures that the amount of data a large language model has to keep track of does not blow up.

Attention mechanisms force large language models to keep track of everything in their context window. A sparse attention model, such as SubQ, throws out a lot of the individual words, but it still retains a rough picture of everything it has seen. In contrast, power retention works by providing the model with a rolling summary of its context window. As new information is added, less relevant information is dropped.

The basic principle of retention has been around for a decade. Manifest AI claims it has updated those techniques to build models that can stand up to transformer-based large language models for the first time.

The company says it is possible to adapt a transformer model into a power retention model with minimal retraining. To demonstrate this, it has turned an existing open-source coding large language model called StarCoder into a version that uses power retention, called PowerCoder. It has also released a model called Brumby, which it claims rivals some versions of Alibaba’s popular open-source model Qwen.

Manifest AI wants its power retention tech to become the go-to solution when large language models need to carry out tasks that involve processing huge amounts of data. There are many useful applications, Manifest AI’s cofounder and CTO, Carles Gelada, claimed in a video announcing his company’s technology last year—from analyzing videos that are hours long to building agents that can stay on task for weeks at a time.

02: Making models smaller and more flexible

Liquid AI, an MIT spinout based in Cambridge, Massachusetts, has not changed or ditched transformers fully but pairs them with its own tech, liquid neural networks, to build what cofounder and CEO Ramin Hasani calls LFMs, liquid foundation models.

Liquid AI’s models are far smaller and use less energy than most large language models. The firm builds models for car makers, including Mercedes, which run on the small chips inside vehicles. Its latest models can run on a Raspberry Pi, a low-powered hobbyist computer that costs $50.

Its models are available for free to any organization with an annual revenue less than $10 million. And they have proved popular: The company has racked up almost 34 million downloads, says Hasani.

Liquid neural networks were inspired by worm brains. They are an extension of another type of neural network that predates transformers, called convolutional networks. The key innovation is a mechanism that lets a model adapt its behavior to new information, so it can learn as it goes. That is not possible with transformers: Once a model is trained, its behavior is fixed.

Liquid AI’s first models were pretty basic but could fly drones or drive vehicles. With LFMs, the company is trying to scale up its technology to compete with mainstream large language models. Its new models match the performance of rivals four times bigger, including versions of Alibaba’s Qwen and Google’s open-source LLM Gemma.

A typical large language model is built from a stack of transformers wired together. Liquid AI’s recent LFMs are hybrid models made up of 20% transformers and 80% liquid neural networks.

That ratio was hit upon by another AI system that Liquid AI has built, which it uses to help design all its models. “It’s the core technology of our company right now,” says Hasani. This designer AI sifts through many different combinations of neural networks—liquid, convolutional, and more, as well as transformers—and comes up with designs that bolt different ones together to hit a sweet spot of performance and efficiency.

Hasani thinks transformers were just the beginning: “Your brain is an AGI system, you know, and it operates with 20 watts of power. How is it possible? We can get a lot more innovative.”

03: Generating text all at once

Almost all large language models produce their output one word at a time. It makes sense, because that is how people speak and write. But for computers, it is very inefficient.

It is faster and cheaper for large language models to generate text all at once—spitting out whole sentences or paragraphs in one shot. That is the approach taken by Inception, a startup based in Palo Alto, California, which is building large language models using a technique called diffusion.

Diffusion is better known as the technology that drives most image and video generation models. Diffusion models are trained to take a random grid of pixels, like the static on an old TV set, and turn it into an image. They do this by working on all the pixels at the same time, figuring out which need changing to make the static look more like a high-definition photo.

It turns out this process works on text too. Inception has trained its large language models to take a random string of words and turn it into sentences that make sense. Diffusion large language models still use transformers to encode meaning, but by producing whole blocks of text at once, they make transformers do more for less. “You’re still using a big transformer model, but you can predict many tokens at the same time,” says Inception’s cofounder and CEO, Stefano Ermon. “That’s why these models are so much faster and cost-efficient compared to what most other people are building today.”

The challenge was to take a technology designed for image generation and apply it to text. With images, if you need to change a blue pixel to a red one you can step through intermediate colors, says Ermon. That does not work with text.

What it means

Developers face a choice between spending billions on compute or adopting new architectures that reduce that need. Models like PowerCoder allow teams to process long documents without paying for massive context windows, while Liquid AI’s approach lets smaller teams run capable models on cheap hardware. The shift moves the industry from scaling up parameters to scaling down efficiency.

Scroll to Top