Open Source Hugging Face Blog

tokenizers v1: encode, decode and scaling, measured

tokenizersHugging Faceperformance benchmarksopen source

Tokenizers have not historically been the bottleneck in ML workflows: tokenization is compute-light compared with the heavy modeling in the rest of the pipeline. But that balance is shifting as models get faster and workloads scale. Training on massive datasets, serving many concurrent requests, or repeatedly processing long inputs can put enough pressure on the tokenizer that it starves the model of data. Hugging Face therefore focused heavily on performance for the upcoming v1 of tokenizers, aiming to keep tokenization light and scalable so GPUs never sit idle waiting for CPU tokenization to finish.

The article says v1 is faster than v0.23, often by tens of times, and credits the broader ecosystem. Tokenization is an active area of open-source work, with libraries such as gigatoken, tiktoken, kitoken, tokie, fastokens, wordchipper, ai-tokenizer, and many others each pushing what a fast tokenizer can be. Several ideas in the refactor reached the team because another project showed they were worth trying. Before this refactor, tokenizers was nowhere near its potential performance, so contributing may not have seemed worthwhile; with the refactor, Hugging Face hopes to make clear that tokenizers is a library worth contributing to. The post also thanks IBM, NVIDIA, and the ExecuTorch team for contributing patches and helping test across a wide range of hardware to broaden platform support.

For results, Hugging Face showcases the tokenizers v1 release candidate against other widely used alternatives. The benchmarks cover single-threaded performance, multi-threaded performance, scaling across threads, per-model comparison, per-language comparison, latency, decoding throughput, memory heap, and crate size. They are run from the tokbench repository, which includes a command to rerun the benchmarks on your own hardware.

On what v1 is: it will produce the same token IDs as v0.23. The goal was to preserve the output, API, vocabulary, and merge ranks while improving everything that can be improved, including breadth. The library remains general across tokenizer families rather than specializing on BPE, so v1 loads everything v0.23 loaded.

A tokenizer converts text into the list of integers a model reads, and tokenizers runs that conversion in four stages. Normalization applies operations such as lowercasing or Unicode normalization to the raw text. Pre-tokenization splits the text into smaller pieces called pre-tokens. The model turns each pre-token into tokens and maps them to IDs in its vocabulary. Post-processing adds any special tokens the model expects. The model stage is where most of the work described in the article happens. The excerpt's final line begins to state that eight of the ten model families measured in the article use byte pai (truncated).

Read original →

← Back to home