Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI

In this articleHugging Face releases 207 WebGPU kernels for local AIWhy focus on kernels?A kernel repository, not just a shaderLoading a kernel…

By Vane September 1, 2026 6 min read
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI


Hugging Face releases 207 WebGPU kernels for local AI

The team has launched

@huggingface/kernels

, a library for loading and running optimised WebGPU kernels directly from the Hugging Face Hub. The initial release contains 207 kernels available at huggingface.co/webgpu-kernels.

Each kernel is published as a complete, versioned package. Its interface, shader templates, correctness tests, benchmarks, and usage instructions are all hosted together on the Hub.

Alongside the kernels, the team is launching Fleet, an in-browser benchmarking suite. It runs tests on users’ hardware and scores the results. Beyond individual reports, Fleet allows the community to contribute performance data from devices that cannot be tested in a conventional lab. With permission, every run adds private evidence to help identify failures, improve kernel variants, and guide optimisation decisions across real-world hardware.

Why focus on kernels?

A model running in the browser breaks down into a sequence of GPU operations: matrix multiplications, normalizations, convolutions, attention primitives, quantization, and data-layout transformations. WebGPU provides a portable API for these operations, while WGSL serves as the common language for the shaders that execute them.

Portability does not guarantee performance. Two shaders can implement the same operation and produce identical output while behaving differently across different accelerators. Workgroup sizes, memory access patterns, vectorization, data types, and fusion strategies all affect speed. The best choice changes with the input shape, device, browser, and available WebGPU features.

Kernels form the foundational layer for fast browser inference. Higher-level runtimes are only as efficient as the operations they dispatch. By making these operations individually discoverable, testable, benchmarkable, and versioned, the team can improve the foundation independently while keeping a stable contract for the layers above it.

A kernel repository, not just a shader

Each kernel has its own repository and a kernel card. The card documents the operation’s semantics, inputs, outputs, attributes, supported data types, source files, and a ready-to-run

@huggingface/kernels

example.

For example,

ai.onnx.Add

implements elementwise addition with multidirectional broadcasting. It is one of the simplest operations in a neural network, used everywhere from residual connections to adding a bias. Its card documents the two inputs, the broadcasted output shape, supported data types, and the variants available for different shapes and devices.

Behind the card, the repository contains the artifacts needed to understand and evaluate the implementation:

  • manifest.json

    is the source of truth for the operation contract. It defines inputs, outputs, attributes, type constraints, and shape derivation rules.

  • metadata.json

    records the kernel identifier, digests, and provenance.

  • test.json

    contains correctness cases so an implementation can be checked against expected behavior.

  • bench.json

    contains benchmark and tuning cases that represent the workloads used to evaluate the kernel.

  • *.wgsl.jinja

    files contain the parameterized WGSL implementations used to produce shaders for a particular request and device.

This structure turns a shader into a reusable software artifact. The interface is inspectable without reading WGSL, correctness and performance cases travel with the implementation, and published versions can be loaded explicitly rather than depending on an unversioned file URL. The kernels also serve as reference implementations for developers building custom WebGPU kernels or integrating these operations into their own runtimes.

Loading a kernel from the Hub

Install the package from npm:

npm install @huggingface/kernels@preview

Running these kernels requires a browser with WebGPU support. WebGPU availability depends on the browser, operating system, GPU, and driver. You can check for it in JavaScript with

"gpu" in navigator

@huggingface/kernels

provides the bridge between a kernel repository and your application. Call

getKernel

with a Hub repository ID and a contract version, then invoke the returned function with typed input data and tensor shapes. Here is a small bias-add example:

import { getKernel } from "@huggingface/kernels";

const add = await getKernel("webgpu-kernels/ai.onnx.Add", { version: 1 });

const { c } = await add({
  a: {
    data: new Float32Array([1, 2, 3, 4, 5, 6]),
    shape: [2, 3],
  },
  b: {
    data: new Float32Array([10, 20, 30]),
    shape: [3],
  },
});

The second input is broadcast across the first dimension, producing an output with shape

[2, 3]

. The loader derives that output shape and logical data type from the manifest contract and the inputs, then allocates

c

automatically.

Addition on six floats is deliberately the smallest possible demo. At this size, the GPU round trip costs far more than the math. The point is the call pattern: it stays exactly the same for the heavyweight operations where optimised kernels actually pay off, such as matrix multiplication (

ai.onnx.MatMul

). Only the repository ID and the inputs change.

Even this elementary operation illustrates why kernels need variants. Equal-shape addition can use a direct vectorized path, while broadcasted inputs need different indexing logic. The published Add kernel includes variants for equal shapes, vectorized broadcasting, scalar processing, and general broadcasting. The runtime can select an implementation that fits the current call and device without changing the application-facing API.

The

version: 1

option selects version 1 of the published kernel contract. It is separate from an ONNX opset, an operator’s

since_version

, or a model revision. Keeping those concepts separate lets applications depend on a stable JavaScript-facing contract while kernel implementations evolve behind it.

How fast are the kernels?

To measure the difference optimised kernels make, the team put their collection head-to-head with ORT WebGPU on an Apple M4 GPU, using ONNX Runtime Web

1.30.0-dev.20260826-b1f76d586a

. They started with 1,756 test cases across all 207 operations and kept the 809 cases where both sides produced matching outputs and reliable timings.

Across those comparisons, the team’s kernels were 2.57x faster by geometric mean and 1.90x faster at the median, with 629 wins, 176 losses, and 4 ties. Here is a closer look at four familiar operations:

OperationCompared casesOur WebGPU KernelORT WebGPUSpeedup
Add50.064 ms0.227 ms3.52x
MatMul290.115 ms0.131 ms1.14x
Softmax120.114 ms0.240 ms2.11x
LayerNormalization60.061 ms0.135 ms2.22x

Some individual wins were much bigger. A particularly difficult bilinear Einsum case (

i,ij,j

with size 4096) ran in 0.136 ms with the team’s kernel versus 1,396 ms with ORT WebGPU: more than 10,000x faster. A row-wise CumSum over

[256, 4096]

was 301x faster, at 0.016 ms versus 4.784 ms. These are unusual cases rather than the speedups you should expect everywhere, but they show how much a specialised kernel can help when a general implementation hits a slow path.

The team timed the work done on the GPU itself, leaving out setup such as loading kernels, creating sessions, uploading inputs, compiling shaders, and reading outputs back. Very short workloads are naturally harder to measure, and small cases can benefit from the GPU cache, so these numbers are best read as a useful comparison rather than a promise for every application.

They are also results for individual operations, not complete models. Exact performance will change across GPUs and browsers, which is why Fleet is so important for building a broader picture.

The team is also working with the ONNX Runtime team to upstream these improvements so they can benefit the broader ONNX Runtime Web ecosystem.

From one device to a fleet

WebGPU performance varies across GPUs, browsers, and drivers, so results from one machine only tell part of the story. Fleet lets anyone run correctness and performance checks in the browser and see how the kernels behave on their hardware.

With consent, each run privately contributes evidence that helps the team spot device-specific failures, compare variants, and improve selection rules. The goal is simple: use broad, real-world coverage to make the kernels faster and more reliable for everyone.

Building a shared foundation for WebAI

The initial 207 kernels are a starting point, not the end state. Publishing kernels independently on the Hub gives the team a common place to inspect contracts, compare implementations, reproduce correctness checks, and improve performance without embedding every shader directly into every runtime.

The collection is also part of the Hub’s broader kernel ecosystem: on the Kernels page, the WebGPU kernels sit alongside kernels for CUDA, ROCm, Metal, and other platforms, and can be filtered, sorted, and explored like any other artifact on the Hub.

The pieces reinforce one another:

  • Kernel repositories define transparent, versioned operation contracts.
  • @huggingface/kernels

    makes those operations straightforward to load and run from JavaScript.

  • Fleet crowdsources real-world evidence across a much broader range of devices than a conventional benchmark lab can cover.
  • Every contributed run can reveal failures, guide tuning, improve variant selection, and help validate future kernel versions.
Scroll to Top