AMD acquires Taalas, a startup that bakes AI models directly into silicon

AMD has acquired Canadian startup Taalas to integrate its model-specific inference technology into its own accelerator roadmap. The Toronto-based firm, founded in…

By Vane August 7, 2026 1 min read
AMD acquires Taalas, a startup that bakes AI models directly into silicon

AMD has acquired Canadian startup Taalas to integrate its model-specific inference technology into its own accelerator roadmap. The Toronto-based firm, founded in 2023, embeds trained parameters directly onto silicon to achieve high speeds, though this design locks each chip to a single model architecture. A recent demonstration showed the hardware processing over 16,000 tokens per second per user with Llama 3.1-8B, significantly outperforming generic competitors. Google is reportedly developing similar dedicated silicon for its Gemini models, indicating a shift toward specialised hardware for large language models.

The acquisition allows AMD to offer this approach alongside its existing Instinct GPUs as a system-level solution. While standard regulatory approvals are pending, the deal expands AMD’s portfolio beyond flexible accelerators to include fixed-function chips designed for specific workloads. This move addresses the growing demand for low-latency inference in data centres where throughput matters more than model versatility.

* Taalas was in stealth mode until February 2024
* The demo chip ran Llama 3.1-8B at 16,000 tokens per second per user
* Google is reportedly working on comparable hardware for Gemini

Scroll to Top