NVIDIA Kumo Tabular 是 NVIDIA Kumo 结构化模型系列的一部分,是一个面向表格数据的开放基础模型,现已在 Hugging Face 上提供。给定一个带有标注行的表格,它能够在单次前向传播中预测新行的标签,无需训练、无需微调、也无需特征工程,即可同时处理分类和回归任务。该模型仅在人工数据上进行预训练,提供三种尺寸(28M 至 215M 参数),通过我们的开源库运行,并在 OpenMDW-1.1 许可下发布以用于商业用途。它在 TabArena、BeyondArena、TALENT 和 ScoringBench 四个基准测试中均排名第一。
转向表格基础模型
表格数据是企业机器学习的骨干。客户记录、交易、传感器日志、索赔和订单都存储在表格中,从中预测流失率、违约、需求或价格属于工业界最常见的机器学习任务之一。在过去二十年中,这项工作一直由梯度提升树完成,且效果良好。然而,围绕这些模型的开发生命周期几乎没有变化。每一个新问题都需要收集标签、进行特征工程、搜索超参数、验证模型,并部署一个对表格本身一无所知、从头学习每个任务的模型。
大型语言模型展示了处理新任务的不同方式。在提示中提供少量示例后,预训练模型无需更新任何权重即可解决该任务。这就是上下文学习(in-context learning),它同样适用于表格:在数百万张表格上预训练的模型可以将带标注的表格作为其上下文,直接预测新行的标签。
今天,我们发布 NVIDIA Kumo Tabular(GitHub、HuggingFace),这是一个用于表格分类和回归的开放基础模型。给定一个带有标注行以及你需要预测的行组成的表格,Kumo Tabular 会在单次前向传播中返回类别概率或数值预测。
Kumo Tabular 的工作原理
Kumo Tabular 是一个围绕表格结构构建的 Transformer,利用了 TabICL 和 TabPFN 中引入的列注意力、行注意力和上下文注意力。为了预测标签,它必须完成三件事:(1) 理解每个值在其列中的含义;(2) 理解一行中各列之间的交互关系;(3) 将带有现有标签的上下文行与带有未知标签的查询行关联起来。Kumo Tabular 通过以下方式实现这一点:
单元格嵌入(Cell Embedding):一组单元格成为一个 token。数值型和类别型值通过傅里叶特征、学习频率的正弦和余弦函数,每种类型拥有独立的权重。缺失值无需插补,而是被特殊处理。最后,上下文中的每个 token 都会接收一个标签嵌入。
行嵌入(Row Embedding):随后,我们通过交替使用两种类型的注意力多次将每一行转换为嵌入。列注意力沿单列向下查看,学习一个值在其列分布中的含义,例如,通过诱导自注意力判断 42 是典型值还是极端值。因此,其成本随行数线性增长。行注意力跨越单行的 token 查看,学习特征如何交互,并使用旋转位置信息区分各列。四个可学习的 [CLS] token 加入每一行,作为该行的最终读取输出。经过这一行压缩步骤后,最终阶段的成本不再依赖于列数。
上下文学习:最终的 Transformer 作用于行嵌入。上下文行彼此进行注意力交互,而查询行仅对上下文行进行注意力交互。因此,每个预测仅依赖于上下文和该行本身,而不依赖于与其一起评分的其他行。由于上下文从不查看查询,其键(keys)和值(values)只需计算一次,即可用于后续的预测。查询行利用 Test-GQA 技术,缩小了每次预测所读取的缓存。每个注意力头将每个查询行转换为分类的类别概率以及回归的 999 个分位数,从而得出点预测和不确定性估计。
长度感知的注意力温度:随着键数量的增加,Softmax 注意力的分布会变得更加分散。在几百行上表现尖锐的注意力,在数万行时可能会变得弥散,而这正是推理时的表格远大于典型训练表格时所面临的情况。Kumo Tabular 因此根据键数量的对数来缩放每个查询的温度系数,且该系数针对每个注意力头单独学习。其结果是,随着表格变长或变宽,注意力依然保持尖锐。
Kumo Tabular 是如何构建的
Kumo Tabular 完全在人工表格上进行预训练。每个训练表格通过以下六个步骤从结构因果模型(SCM)中采样生成:
我们首先为整个表格绘制配置,涵盖其大小、任务、机制和缺失情况。随后,一个随机的因果图连接隐藏变量,这些变量从根节点到叶节点通过在每个节点随机抽取的函数(例如线性映射、小型神经网络、树或高斯过程)进行评估。某些节点成为数值型或类别型列,其中一个成为目标变量,其余则保持隐藏状态,类似于真实数据背后未被测量的原因。后处理步骤对列组进行相关性调整,裁剪异常值,并注入缺失值,同时通过快速的树集成检查丢弃任何缺乏可学习信号的表格。由于生成器是一个过程采样器而非训练过的模型,它源源不断地产生具有新图和全新机制的表格。
现实世界的表格杂乱无章,因此我们在生成器中融入了更多的不完美特性。数值以多种模式缺失,部分特征被粗化,导致重复行可能在标签上存在分歧;某些类别列包含大量层级,回归目标可能呈现重尾分布。见过数百万此类表格的模型能够处理这些不完美之处,而无需任何数据清洗。
在每一个人工表格上,模型将大部分带有标签的行作为上下文,并学习预测剩余行的标签,其中分类任务使用交叉熵损失,回归任务使用分位数损失。分类和回归作为独立的模型进行训练。与 TabICLv2 类似,训练过程分为三个阶段。第一个也是最长阶段使用包含 1,024 行和多达 100 列的表格,教导模型理解表格的结构。第二阶段将上下文行数从 400 变化至 10,240 行,第三阶段将其扩展至 60,000 行,同时仍保持最多 100 列。总体而言,Kumo Tabular-Small/Medium/Large 分别经历了约 3500 万/7100 万/1.37 亿个人工表格的训练。
我们的训练配方和人工数据生成器即将发布。
性能表现
我们在默认设置下,将三种尺寸的 Kumo Tabular 与完整的 TabArena 排行榜进行了对比,该排行榜涵盖了调优后的梯度提升树、AutoGluon 以及最新的表格基础模型。Kumo Tabular 以 1950 的 ELO 评分位居整体榜首,在统一的单张 RTX 6000 Pro 评估设置下,其运行速度比 LimiX-2 快 17 倍。在所有三种模型尺寸中,Kumo Tabular 在准确率-效率帕累托前沿上确立了新的最先进水平:
我们还在 BeyondArena、TALENT 和 ScoringBench 上对 Kumo Tabular 进行了评估。在 BeyondArena 上,Kumo Tabular 达到了 1418 的 ELO 分数,可改进性得分为 7.78%,在排行榜上位居第一。在 TALENT 上,它在分类准确率、分类对数损失和回归 RMSE 方面均取得了整体排名第一的成绩,平均排名分别为 6.67、3.98 和 4.22。在预测分布基准测试 ScoringBench 上,Kumo Tabular-Large 和 Kumo Tabular-Medium 的平均排名分别位列第一和第二。
局限性
Kumo Tabular 仅适用于数值型和分类型列,而文本、图像或时间戳可以通过内置的预处理配方转换为特征。单次前向传播最多可覆盖 10 个类别,该库通过纠错输出码(error-correcting output codes)将其扩展至任意数量的类别。如果表格数据远远超出训练范围,或者查询行与上下文行来自不同的分布,准确率可能会下降。因此,与任何预测模型一样,在部署之前应使用自己保留的数据验证准确性和校准情况。
演示
Kumo Tabular 通过 NVIDIA 新发布的面向结构化数据的 GPU 原生库运行。该库在首次使用时从 Hub 下载权重,并提供我们在评估中使用的预处理、集成和多数类处理功能。以下代码即可实现从 pandas.DataFrame 到预测的转换:
import sdm
table = sdm.TableTensor.from_pandas(pd.load_csv(...), device="cuda")
na_mask = table["target"].isnan()
model = sdm.models.KumoTabular(device="cuda")
pred = model(
x_context=table[~na_mask].drop_columns("target"),
y_context=table[~na_mask, "target"],
x_query=table[na_mask].drop_column("target"),
)
使用 Kumo Tabular 开始构建
Kumo Tabular 根据 OpenMDW 许可协议 1.1 版本发布。NVIDIA 认为可信 AI 是共同责任,我们已制定政策和实践以支持广泛 AI 应用的开发。在下载或按照服务条款使用时,开发者应与他们的模型支持团队合作,确保该模型满足相关行业和用例的要求,并解决不可预见的产品滥用问题。请在此处报告模型质量、风险、安全漏洞或 NVIDIA AI 相关问题。
致谢
我们感谢 David Holzmüller 为 Kumo Tabular 贡献的重要思想和消融实验。我们感谢 Vignesh Kothapalli 在实习期间对 Kumo Tabular 提供的帮助。
NVIDIA Kumo Tabular, part of the NVIDIA Kumo Structured model collection, is an open foundation model for tabular data now available on Hugging Face . Given a table of labeled rows, it predicts the labels of new rows in a single forward pass, with no training, no tuning, and no feature engineering, for both classification and regression. It was pretrained only on artificial data, comes in three sizes (28M to 215M parameters), runs through our open-source library , and is released under the OpenMDW-1.1 license for commercial use. It ranks first on the four benchmarks TabArena , BeyondArena , TALENT and ScoringBench .
The Shift to Tabular Foundation Models
Tabular data is the backbone of enterprise machine learning. Customer records, transactions, sensor logs, claims, and orders all live in tables, and predicting churn, default, demand, or price from them is among the most common machine learning tasks in industry. For two decades, this work has been done with gradient-boosted trees, and it has worked well. But the lifecycle around those models has barely changed. Every new question means collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that knows nothing about tables in general and learns each task from scratch.
Large Language Models showed a different way of working with new tasks. Given a few examples in the prompt, a pretrained model solves the task without updating a single weight. This is in-context learning , and it applies to tables just as well as to text: a model pretrained on millions of tables can read a labeled table as its context and predict the labels of new rows directly.
Today, we are releasing NVIDIA Kumo Tabular ( GitHub , HuggingFace ), an open foundation model for tabular classification and regression. Given a table with labeled rows and the rows you want predictions for, Kumo Tabular returns class probabilities or numeric predictions in a single forward pass.
How Kumo Tabular Works
Kumo Tabular is a Transformer built around the structure of a table, utilizing column, row and in-context attention as introduced in TabICL and TabPFN . To predict a label it has to do three things: (1) understand what each value means within its column, (2) understand how the columns of a row interact, and (3) relate the context rows with existing labels to the query rows with unknown labels. Kumo Tabular achieves this as follows:
Cell Embedding: A group of cells becomes a token. Numerical and categorical values pass through Fourier features, sines and cosines of learned frequencies, with separate weights for each type. Missing values need no imputation and are treated specially. Finally, every token in the context receives a label embedding.
Row Embedding: We then turn each row into an embedding by alternating two kinds of attention multiple times. Column attention looks down a single column and learns what a value means in the distribution of its column, e.g. , whether a 42 is typical or extreme, via induced self-attention. Its cost therefore grows linearly with the number of rows. Row attention looks across the tokens of a single row and learns how features interact, with rotary positions to tell columns apart. Four learnable [CLS] tokens join each row and act as the final readout of a row. After this row compression, the cost of the final stage no longer depends on the number of columns.
In-context Learning: A final Transformer operates on the row embeddings. Context rows attend to each other, while query rows attend to context rows only. Each prediction therefore depends only on the context and on the row itself, not on which other rows are scored alongside it. Because the context never looks at the queries, its keys and values are computed once and can be reused for follow-up predictions. Query rows utilize Test-GQA, which shrinks the cache that every prediction reads. A head turns each query row into class probabilities for classification and 999 quantiles for regression, from which a point prediction and an uncertainty estimate follow.
Length-aware Attention Temperature: Softmax attention spreads out as the number of keys grows. Attention that is sharp over a few hundred rows can dissolve over tens of thousands, which is exactly the situation when a table at inference is much larger than a typical training table. Kumo Tabular therefore scales every query by a temperature that grows with the logarithm of the number of keys, with a coefficient learned separately for each attention head. The result is attention that stays sharp as tables grow longer or wider.
How Kumo Tabular was Built
Kumo Tabular is pretrained entirely on artificial tables. Each training table is sampled from a Structural Causal Model (SCM) in the six steps shown below:
We first draw a configuration for the whole table, from its size and task to its mechanisms and missingness. A random causal graph then links hidden variables, evaluated from root to leaf via randomly drawn functions at every node ( e.g. , linear maps, small neural networks, trees or Gaussian processes). Some nodes become numerical or categorical columns, one becomes the target, and the rest stay hidden, like the unmeasured causes behind real data. Post-processing correlates groups of columns, clips outliers, and injects missing values, and a quick tree-ensemble check discards any table without a learnable signal. Because the generator is a procedural sampler rather than a trained model, it produces an endless supply of tables, each with a new graph and new mechanisms.
Real-world tables are messy, so we built more of their imperfections into the generator. Values go missing in several patterns, some features are coarsened so that duplicate rows may disagree on their label, some categorical columns carry many levels, and regression targets can be heavy-tailed. A model that has seen millions of such tables learns to handle these imperfections without any cleanup.
On every artificial table, the model sees most of the rows with their labels as context and learns to predict the labels of the remaining rows, with a cross-entropy loss for classification and a quantile loss for regression. Classification and regression are trained as separate models. Similarly to TabICLv2, training runs in three stages. The first and longest stage uses tables of 1,024 rows and up to 100 columns and teaches the model what tables look like. The second stage varies the context from 400 to 10,240 rows, and the third extends it to 60,000 rows, still with up to 100 columns. In total, Kumo Tabular-Small/Medium/Large saw about 35/71/137 million artificial tables.
Our training recipe and artificial data generators will be released soon.
Performance
We ran all three Kumo Tabular sizes with default settings against the full TabArena leaderboard, spanning tuned gradient-boosted trees, AutoGluon, and the latest tabular foundation models. Kumo Tabular ranks first overall with an ELO of 1950 while running 17 faster than LimiX-2 under a uniform single RTX 6000 Pro evaluation setup. Across all three three model sizes, Kumo Tabular establishes a new state-of-the-art on the accuracy-efficiency Pareto front:
We also evaluated Kumo Tabular on BeyondArena , TALENT and ScoringBench . On BeyondArena, Kumo Tabular reaches an ELO of 1418 with an Improvability score of 7.78%, placing first on the leaderboard. On TALENT, it achieves the top overall ranking across classification accuracy, classification log-loss, and regression RMSE, with average ranks of 6.67, 3.98, and 4.22. On ScoringBench, a benchmark for predictive distributions, Kumo Tabular-Large and Medium rank first and second on average rank.
Limitations
Kumo Tabular works on numerical and categorical columns only, while text, images, or timestamps can be turned into features via built-in pre-processing recipes. A single forward pass covers up to 10 classes, which the library extends to any number of classes with error-correcting output codes. Accuracy may degrade on tables far beyond the training ranges or when the query rows come from a different distribution than the context rows, so, as with any predictive model, validate accuracy and calibration on your own held-out data before deployment.
Demo
Kumo Tabular runs via NVIDIA's newly released GPU-native library for structured-data-models . The library downloads the weights from the Hub on first use and provides the preprocessing, ensembling, and many-class handling used in our evaluations. The code below is all it takes to go from a pandas.DataFrame to a prediction:
import sdm
table = sdm.TableTensor.from_pandas(pd.load_csv(...), device= "cuda" )
na_mask = table[ "target" ].isnan()
model = sdm.models.KumoTabular(device= "cuda" )
pred = model(
x_context=table[~na_mask].drop_columns( "target" ),
y_context=table[~na_mask, "target" ],
x_query=table[na_mask].drop_column( "target" ),
)
Start Building with Kumo Tabular
Kumo Tabular is released under the OpenMDW License Agreement, version 1.1 . NVIDIA believes Trustworthy AI is a shared responsibility, and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their supporting model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report model quality, risk, security vulnerabilities, or NVIDIA AI concerns here .
Acknowledgements
We thank David Holzmüller for contributing significant ideas and ablations to Kumo Tabular. We thank Vignesh Kothapalli for his help on Kumo Tabular during his internship.
首次收录 · 2026-09-30 · 10.62 分