In this article
NVIDIA Kumo Tabular predicts labels without training
NVIDIA has released an open foundation model for tabular data that predicts classification or regression outcomes in a single forward pass. The system requires no training, no tuning, and no feature engineering. It operates on labeled rows to predict new ones, running on an open-source library under the OpenMDW-1.1 license. The model is available in three sizes, ranging from 28 million to 215 million parameters. It currently ranks first on four benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench.
Model Code: https://github.com/NVIDIA/structured-data-models
Model Weights: https://huggingface.co/nvidia/Kumo-Tabular
Why tabular data needs a new approach
Enterprise machine learning relies heavily on tabular data. Customer records, transactions, sensor logs, claims, and orders all live in tables. Predicting churn, default, demand, or price from these tables remains one of the most common tasks in industry. For two decades, this work has used gradient-boosted trees. The process has barely changed. Every new question requires collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that learns each task from scratch.
Large Language Models showed a different way of working with new tasks. Given a few examples in the prompt, a pretrained model solves the task without updating a single weight. This is in-context learning, and it applies to tables just as well as to text. A model pretrained on millions of tables can read a labeled table as its context and predict the labels of new rows directly.
NVIDIA Kumo Tabular is an open foundation model for tabular classification and regression. It takes a table with labeled rows and the rows you want predictions for, returning class probabilities or numeric predictions in a single forward pass.
How the model processes tables
Kumo Tabular is a Transformer built around the structure of a table. It uses column, row, and in-context attention. To predict a label it performs three actions: (1) understand what each value means within its column, (2) understand how the columns of a row interact, and (3) relate the context rows with existing labels to the query rows with unknown labels. The system handles this through specific embedding and attention mechanisms.
Cell Embedding converts a group of cells into a token. Numerical and categorical values pass through Fourier features, sines and cosines of learned frequencies, with separate weights for each type. Missing values require no imputation and are treated specially. Every token in the context receives a label embedding.
Row Embedding turns each row into an embedding by alternating two kinds of attention multiple times. Column attention looks down a single column and learns what a value means in the distribution of its column. For example, it determines whether a 42 is typical or extreme via induced self-attention. Its cost grows linearly with the number of rows. Row attention looks across the tokens of a single row and learns how features interact, using rotary positions to distinguish columns. Four learnable [CLS] tokens join each row and act as the final readout. After this row compression, the cost of the final stage no longer depends on the number of columns.
In-context Learning uses a final Transformer operating on the row embeddings. Context rows attend to each other, while query rows attend to context rows only. Each prediction depends only on the context and on the row itself, not on which other rows are scored alongside it. Because the context never looks at the queries, its keys and values are computed once and can be reused for follow-up predictions. Query rows use Test-GQA, which shrinks the cache that every prediction reads. A head turns each query row into class probabilities for classification and 999 quantiles for regression, from which a point prediction and an uncertainty estimate follow.
Length-aware Attention Temperature adjusts softmax attention as the number of keys grows. Attention that is sharp over a few hundred rows can dissolve over tens of thousands, which is exactly the situation when a table at inference is much larger than a typical training table. Kumo Tabular scales every query by a temperature that grows with the logarithm of the number of keys. The coefficient is learned separately for each attention head. The result is attention that stays sharp as tables grow longer or wider.
Training on synthetic data
Kumo Tabular is pretrained entirely on artificial tables. Each training table is sampled from a Structural Causal Model in six steps. The process first draws a configuration for the whole table, from its size and task to its mechanisms and missingness. A random causal graph then links hidden variables, evaluated from root to leaf via randomly drawn functions at every node. Some nodes become numerical or categorical columns, one becomes the target, and the rest stay hidden. Post-processing correlates groups of columns, clips outliers, and injects missing values. A quick tree-ensemble check discards any table without a learnable signal. Because the generator is a procedural sampler rather than a trained model, it produces an endless supply of tables, each with a new graph and new mechanisms.
Real-world tables are messy, so the generator includes imperfections. Values go missing in several patterns, some features are coarsened so that duplicate rows may disagree on their label, some categorical columns carry many levels, and regression targets can be heavy-tailed. A model that has seen millions of such tables learns to handle these imperfections without any cleanup.
On every artificial table, the model sees most of the rows with their labels as context and learns to predict the labels of the remaining rows. It uses a cross-entropy loss for classification and a quantile loss for regression. Classification and regression are trained as separate models. Training runs in three stages, similar to TabICLv2. The first and longest stage uses tables of 1,024 rows and up to 100 columns to teach the model what tables look like. The second stage varies the context from 400 to 10,240 rows, and the third extends it to 60,000 rows. In total, Kumo Tabular-Small/Medium/Large saw about 35/71/137 million artificial tables.
The training recipe and artificial data generators will be released soon.
Performance results
All three Kumo Tabular sizes ran with default settings against the full TabArena leaderboard. The leaderboard spans tuned gradient-boosted trees, AutoGluon, and the latest tabular foundation models. Kumo Tabular ranks first overall with an ELO of 1950 while running 17 faster than LimiX-2 under a uniform single RTX 6000 Pro evaluation setup. Across all three model sizes, the system establishes a new state-of-the-art on the accuracy-efficiency Pareto front.
Evaluation on BeyondArena reached an ELO of 1418 with an Improvability score of 7.78%, placing first on the leaderboard. On TALENT, it achieves the top overall ranking across classification accuracy, classification log-loss, and regression RMSE, with average ranks of 6.67, 3.98, and 4.22. On ScoringBench, a benchmark for predictive distributions, Kumo Tabular-Large and Medium rank first and second on average rank.
Limitations
Kumo Tabular works on numerical and categorical columns only. Text, images, or timestamps can be turned into features via built-in pre-processing recipes. A single forward pass covers up to 10 classes, which the library extends to any number of classes with error-correcting output codes. Accuracy may degrade on tables far beyond the training ranges or when the query rows come from a different distribution than the context rows. As with any predictive model, validate accuracy and calibration on your own held-out data before deployment.
Getting started
Kumo Tabular runs via NVIDIA’s newly released GPU-native library for structured-data-models. The library downloads the weights from the Hub on first use and provides the preprocessing, ensembling, and many-class handling used in evaluations. The following code takes a pandas.DataFrame to a prediction:
import sdm # structured-data-models
# Tensorize tabular data:
table = sdm.TableTensor.from_pandas(pd.load_csv(...), device="cuda")
na_mask = table["target"].isnan()
model = sdm.models.KumoTabular(device="cuda")
pred = model(
# In-context examples (features/targets):
x_context=table[~na_mask].drop_columns("target"),
y_context=table[~na_mask, "target"],
# Prediction examples (features):
x_query=table[na_mask].drop_column("target"),
)
Licensing
Kumo Tabular is released under the OpenMDW License Agreement, version 1.1. NVIDIA believes Trustworthy AI is a shared responsibility. When downloaded or used in accordance with the terms of service, developers should work with their supporting model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Report model quality, risk, security vulnerabilities, or NVIDIA AI concerns here.
- Model Code: https://github.com/NVIDIA/structured-data-models
- Model Weights: https://huggingface.co/nvidia/Kumo-Tabular
Acknowledgements
We thank David Holzmüller for contributing significant ideas and ablations to Kumo Tabular. We thank Vignesh Kothapalli for his help on Kumo Tabular during his internship.




