TabKDE: Simple and Scalable Tabular Data Generation with Kernel Density Estimates. Alishahi, M., Zheng, Y., Wang, J., Yeh, C. M., & Phillips, J. M. May, 2026. arXiv:2605.17642 [cs.LG]
TabKDE: Simple and Scalable Tabular Data Generation with Kernel Density Estimates [link]Paper  doi  abstract   bibtex   
Tabular data generation considers a large table with multiple columns – each column comprised of numerical, categorical, or sometimes ordinal values. The goal is to produce new rows for the table that replicate the distribution of rows from the original data – without just copying those initial rows. The last 4 years have seen enormous progress on this problem, mostly using computational expensive methods that employ one-hot encoding, VAEs, and diffusion.
@misc{alishahi_tabkde_2026,
	title = {{TabKDE}: {Simple} and {Scalable} {Tabular} {Data} {Generation} with {Kernel} {Density} {Estimates}},
	shorttitle = {{TabKDE}},
	url = {http://arxiv.org/abs/2605.17642},
	doi = {10.48550/arXiv.2605.17642},
	abstract = {Tabular data generation considers a large table with multiple columns – each column comprised of numerical, categorical, or sometimes ordinal values. The goal is to produce new rows for the table that replicate the distribution of rows from the original data – without just copying those initial rows. The last 4 years have seen enormous progress on this problem, mostly using computational expensive methods that employ one-hot encoding, VAEs, and diffusion.},
	language = {en},
	urldate = {2026-06-15},
	publisher = {arXiv},
	author = {Alishahi, Meysam and Zheng, Yan and Wang, Junpeng and Yeh, Chin-Chia Michael and Phillips, Jeff M.},
	month = may,
	year = {2026},
	note = {arXiv:2605.17642 [cs.LG]},
	keywords = {Computer Science - Machine Learning, WG: Observable},
}

Downloads: 0