Baseten on Hugging Face Inference Providers ๐ฅ
Hugging Face announced that Baseten has joined its Inference Providers ecosystem, expanding serverless inference options on the Hub's model pages. Baseten is an AI infrastructure platform offering serverless AI, training, and a catalog of frontier models, with support for model types from LLMs to text-to-speech. Initially, the integration covers conversational and text-generation tasks, providing access to open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2, with more tasks planned.
In the Hugging Face UI, users can set custom API keys for providers or leave requests routed through HF (with charges applied to their HF account). Providers can be ordered by preference, and model pages display compatible providers sorted accordingly. The integration is available in Python via huggingface_hub >= 1.26.1 and in JavaScript via @huggingface/inference. Example code shows using DeepSeek V4 Flash through Baseten with a Hugging Face token, where the request is routed automatically to Baseten.
Baseten-hosted models also plug directly into agent harnesses such as Pi, OpenCode, Hermes Agents, and OpenClaw without extra glue code. This streamlines the developer experience by letting users choose Baseten as a provider across the Hub, SDKs, and agent frameworks.