Model Releases Hugging Face Blog

NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

NVIDIAtabular datafoundation modelin-context learning

Tabular data is the backbone of enterprise machine learning—customer records, transactions, sensor logs, claims, and orders—and predicting churn, default, demand, or price from such tables is among the most common industrial ML tasks. For two decades these problems have been handled well by gradient-boosted trees, but the lifecycle barely changes: every new question requires collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that learns each task from scratch. Large language models demonstrated an alternative through in-context learning, where a pretrained model solves a new task from a few prompt examples without updating weights; the same idea applies to tables, since a model pretrained on millions of tables can read a labeled table as context and predict labels for new rows directly.

Today NVIDIA is releasing Kumo Tabular (GitHub, Hugging Face), an open foundation model for tabular classification and regression and part of the NVIDIA Kumo Structured model collection. Given a table with labeled rows and the rows to predict, it returns class probabilities or numeric predictions in a single forward pass. It requires no training, no tuning, and no feature engineering. It was pretrained only on artificial data, comes in three sizes ranging from 28M to 215M parameters, runs through NVIDIA's open-source library, and is released under the OpenMDW-1.1 license for commercial use. It ranks first on four benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench. Model code is at https://github.com/NVIDIA/structured-data-models and weights at https://huggingface.co/nvidia/Kumo-Tabular.

Kumo Tabular is a Transformer built around the structure of a table, utilizing column, row, and in-context attention as introduced in TabICL and TabPFN. To predict a label, it must do three things: understand what each value means within its column, understand how the columns of a row interact, and relate the context rows with existing labels to the query rows with unknown labels.

The model achieves this through cell and row embeddings. In cell embedding, a group of cells becomes a token; numerical and categorical values pass through Fourier features—sines and cosines of learned frequencies—with separate weights for each type, while missing values need no imputation and are treated specially. Every token in the context also receives a label embedding. In row embedding, each row is turned into an embedding by alternating two kinds of attention multiple times, starting with column attention.

Read original →

← Back to home