OASIS: Observation-Aware Simulation-Based Inference via Distributional Matching. Farahi, A., Zhou, C., & Vashistha, R. 2026. Version Number: 1
OASIS: Observation-Aware Simulation-Based Inference via Distributional Matching [link]Paper  doi  abstract   bibtex   
We introduce OASIS, a simulation-based inference framework for scientific settings where observations are distorted by measurement error, selection effects, and other survey-specific transformations. In many real applications, simulators generate latent, noiseless quantities, while the data are observed only after passing through a complex observational pipeline. Standard simulation-based inference methods often ignore this distinction, comparing observations to idealized simulator outputs or relying on low-dimensional summaries that can miss important structure. OASIS addresses this mismatch by explicitly embedding the observation model into the simulator and performing inference directly at the level of observed-data distributions. The method constructs a pseudo-posterior by reweighting prior samples according to a maximum mean discrepancy (MMD) loss between the empirical distributions of the observed data and forward-simulated observations, thereby avoiding both handcrafted summaries and learned neural surrogates. We provide theoretical guarantees for Monte Carlo consistency, convergence of the empirical pseudo-posterior to its population counterpart, and posterior concentration on the MMD-identified parameter set, with consistency for the true parameter under correct specification and identifiability. In controlled errors-in-variables regression experiments, OASIS delivers robust parameter recovery and well-calibrated uncertainty under heterogeneous and non-Gaussian measurement noise. We then demonstrate the method on a realistic cosmological application involving galaxy cluster observations across multiple wavelengths, in which latent physical properties are linked to observables through nonlinear scaling relations, heteroscedastic errors, selection functions, and incomplete coverage.
@misc{farahi_oasis_2026,
	title = {{OASIS}: {Observation}-{Aware} {Simulation}-{Based} {Inference} via {Distributional} {Matching}},
	copyright = {Creative Commons Attribution 4.0 International},
	shorttitle = {{OASIS}},
	url = {https://arxiv.org/abs/2606.22572},
	doi = {10.48550/ARXIV.2606.22572},
	abstract = {We introduce OASIS, a simulation-based inference framework for scientific settings where observations are distorted by measurement error, selection effects, and other survey-specific transformations. In many real applications, simulators generate latent, noiseless quantities, while the data are observed only after passing through a complex observational pipeline. Standard simulation-based inference methods often ignore this distinction, comparing observations to idealized simulator outputs or relying on low-dimensional summaries that can miss important structure. OASIS addresses this mismatch by explicitly embedding the observation model into the simulator and performing inference directly at the level of observed-data distributions. The method constructs a pseudo-posterior by reweighting prior samples according to a maximum mean discrepancy (MMD) loss between the empirical distributions of the observed data and forward-simulated observations, thereby avoiding both handcrafted summaries and learned neural surrogates. We provide theoretical guarantees for Monte Carlo consistency, convergence of the empirical pseudo-posterior to its population counterpart, and posterior concentration on the MMD-identified parameter set, with consistency for the true parameter under correct specification and identifiability. In controlled errors-in-variables regression experiments, OASIS delivers robust parameter recovery and well-calibrated uncertainty under heterogeneous and non-Gaussian measurement noise. We then demonstrate the method on a realistic cosmological application involving galaxy cluster observations across multiple wavelengths, in which latent physical properties are linked to observables through nonlinear scaling relations, heteroscedastic errors, selection functions, and incomplete coverage.},
	language = {en},
	urldate = {2026-07-06},
	publisher = {arXiv},
	author = {Farahi, Arya and Zhou, Conghao and Vashistha, Ritwik},
	year = {2026},
	note = {Version Number: 1},
	keywords = {Computation (stat.CO), Data Analysis, Statistics and Probability (physics.data-an), FOS: Computer and information sciences, FOS: Physical sciences, Instrumentation and Methods for Astrophysics (astro-ph.IM), Methodology (stat.ME), WG: Explainable},
}

Downloads: 0