19–23 Oct 2026
Lanthieri Mansion, Vipava
Europe/Ljubljana timezone

Trustworthy species distribution modelling: reliability audits for inference from imperfect observations

20 Oct 2026, 11:45
15m
Lanthieri Mansion, Vipava

Lanthieri Mansion, Vipava

Glavni trg 8, Vipava, SI 5271, Slovenia

Speaker

Kristian Miok (FRI UL)

Description

Machine-learned models of species distributions are built, more often than not, on observations the modeller did not collect. Aggregated archives pool records from many contributors, instruments and eras, and with them a wide range of positional and observational quality. Our work asks one question across this pipeline: how far can the resulting predictions be trusted, and where exactly do they fail? It treats it as a problem of reliability propagation: uncertainty and contamination in the observations travel through the fitted model into the ecological conclusions drawn from it.
Occurrence records are routinely filtered by positional quality before modelling, a step almost universally treated as neutral hygiene. It is not. Filtering is neutral only if quality is unrelated to the predictors; when it is not, discarding low-quality records removes them non-randomly across predictor space, the distribution the model learns from shifts, and cleaning becomes a source of bias. We formalise this as a covariate-shift problem and develop a diagnostic that detects, without ground truth, whether and along which axes filtering displaces the data. Applied to a curated global freshwater-crayfish database (115,191 records) and to open-aggregator Odonata records, filtering shifts the observed distribution significantly and coherently; two defensible corrections move the model in opposite directions and bracket a truth that is not identifiable from the data; and stricter filtering induces more bias, not less.
The consequence does not stay small. Mean per-cell changes in predicted suitability of 0.03–0.05 become changes of 7–52% in predicted range area, the quantity conservation decisions actually use. Error is amplified, not absorbed, on the way downstream.
I will place this within a wider reliability layer under construction: auditing the observations that enter, calibrating the predictions that leave, where miscalibration proves spatially structured, concentrating in particular parts of the river network, and propagating the residual uncertainty into derived quantities, together with an emerging R implementation designed to sit alongside existing modelling stacks rather than replace them.
The lesson generalises beyond ecology: wherever a data-quality flag correlates with the predictors, cleaning is a modelling decision with consequences, not a hygiene step taken before the modelling begins.

Author

Co-authors

Presentation materials

There are no materials yet.