Open Source Hugging Face Blog

Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI

WebGPUkernelsHugging Facebrowser inference

Hugging Face's WebAI team is working to make browser inference as fast and user-friendly as possible, a multi-layer effort spanning model representations, runtime execution plans, and individual GPU operations. Today they released the first layer: @huggingface/kernels, a minimal library that downloads, prepares, and runs WebGPU kernels directly from the Hugging Face Hub, along with an initial collection of 207 kernels published in the webgpu-kernels organization. The collection covers operations used across a wide variety of machine learning architectures and workloads. Each kernel is a complete, versioned package including its interface, WGSL shader templates, correctness tests, benchmark cases, and usage instructions, all Apache-2.0 licensed.

They also launched Fleet, an in-browser GPU benchmarking and testing suite that runs and scores kernels on the user's hardware. With the user's consent, every run adds private evidence to the project, helping the team find failures such as incorrect results or pathologically slow cases, improve kernel variants, and make better optimization decisions across real-world GPUs.

The post explains why kernels are foundational: a model running in the browser becomes a sequence of GPU operations like matrix multiplications, normalizations, convolutions, attention primitives, quantization, and data-layout transformations. WebGPU offers portability across browsers, but portability doesn't guarantee performance—two shaders for the same operation can behave very differently across accelerators. Workgroup sizes, memory access patterns, vectorization, data types, and fusion strategies all affect performance, and the best choice can change with input shape, device, browser, and available WebGPU features. Higher-level runtimes can only be as efficient as this kernel layer beneath them.

Read original →

← Back to home