Kog aims to extract significantly more processing power from standard datacentre GPUs rather than relying on custom silicon.
In this article
The French startup recently made headlines by demonstrating that extremely fast single-request decoding is possible on hardware enterprises already possess. The demo relied on AMD MI300X and NVIDIA H200 chips.
While some observers noted the technology did not yet extend to laptop GPUs, the potential for reducing inference costs and speed remains significant. CEO Gaël Delalleau told TechCrunch the company had generated 200 tangible business leads following the announcement.
Current applications
Early feedback suggests software engineering is the primary use case. Developers using tools like Claude Code often face waits of several hours for results. Anthropic acknowledges this speed is financially valuable, charging a premium multiple for its Fast Mode.
Kog targets professionals relying on AI for daily workflows who find these delays unacceptable. The startup also works with design partners generating games and apps via prompt. For these clients, faster output from the Kog Inference Engine (KIE) translates directly to increased revenue.
Market maturity
The company recognises the market is not yet fully developed. Prospective customers currently lack the resources to fine-tune small models. Consequently, Kog has shifted focus to accelerating the development of larger models to meet observed demand.
This creates a substantial challenge for the promise of 30x faster LLM inference. The initial demo achieved 3,000 tokens per second per request. This result was driven by the Laneformer 2B, an open-sourced model with approximately 2 billion parameters.
Delalleau believes the same approach functions effectively with larger LLMs, contradicting sceptics who view standard GPUs as unsuited for decoding. He argues newer hardware offers increased memory bandwidth that requires unlocking through better software.
Competitive landscape
Kog is not the only entity exploring software optimisation for standard hardware. ZML, another French firm, released hardware-agnostic software that bypasses Nvidia’s CUDA to enable fast inference across competing chips. Delalleau distinguishes Kog by comparing it to Stanford University’s Hazy Research, citing a deeper focus on GPU acceleration.
Founder background
Delalleau is not a researcher by trade. His first company, Stribe, launched in 2009 and shares no technical connection with Kog. However, his former co-founder Kamel Zeroual, now a venture capitalist, co-led the seed round through Varsity VC.
The startup’s methodology stems from Delalleau’s unique history. He studied solid-state physics at France’s École Polytechnique before working in offensive cybersecurity. He describes this background as fostering a mindset of understanding the fundamental laws governing GPUs to maximise their utility.
His experience as a four-time finalist in DEF CON’s CTF tournament reinforced this approach. He learned to reverse-engineer systems down to assembly language and binary code to achieve goals not originally intended by the design.
Operational limits
The downside of this method is that it is highly manual and time-consuming. For every new GPU architecture, the team dedicates several weeks or months to deep-dive engineering research. With a staff of 11 people, this restricts the number of chips Kog can support in the immediate future.
Long-term plans involve feeding this methodology into agent-based pipelines to broaden support for various chips and models. This strategy may offer sovereignty advantages as Europe builds its own capabilities, supported by Scaleway and French government programmes including Bpifrance and French Tech 2030.
Next steps
Kog must now prove the approach works on large language models to secure further funding. Delalleau expects to implement the first major model at 10x speed by September. Achieving this milestone will allow the company to demonstrate customer traction and pursue Series A financing.



