arXiv Science⌕ Search

arXiv · 2609.38586

Fraud Detection via Bayesian Positive-Unlabeled Learning with Gaussian Processes and Multilayer Networks

Abstract

Tax fraud remains a central challenge for public revenue authorities worldwide, imposing fiscal losses estimated to reach up to 1 trillion euros annually in the EU alone. Fraudulent firms intentionally manipulate reported figures to conceal their activity. Business networks can provide complementary information and reveal fraud that covariates alone may miss. We propose a Bayesian positive-unlabeled (PU) classification framework that combines firm-level covariates with multilayer network information. As a policy relevant case study, we apply the framework to firms in a regulated segment of the Greek energy market, a sector exposed to excise tax evasion. Confirmed fraud labels exist only for audited cases, while the remaining observations are unlabeled rather than verified compliant. We develop a Bayesian Gaussian process (GP) classifier integrating covariates with multilayer network information through a Product of Experts (PoE) construction and incorporating a nondetection probability for missed positives. We derive identification bounds for the nondetection rate under an anchor condition and show that, for a fixed latent risk function, a common nondetection rate affects calibration but not ranking. Simulations show strong ranking performance relative to PU and network-based alternatives, while illustrating the difficulty of estimating the nondetection rate. In real audit data, the method identifies high risk firms with posterior uncertainty, estimates the number of undetected fraudulent cases, and identifies whether risk is driven by covariates, networks, or both. Of the six highest ranked firms, the two that had already been reinspected independently were both confirmed as noncompliant, providing an independent validation of these cases. The remaining four were selected by the tax authority for follow-up inspection.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Konstantinos Bourazas, Angelos Alexopoulos, Petros Dellaportas, Konstantinos Kalogeropoulos. 2026-10-01. Fraud Detection via Bayesian Positive-Unlabeled Learning with Gaussian Processes and Multilayer Networks. https://arxiv.org/abs/2609.38586

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Projection-based significance test for longitudinally collected functional data under a crossover design

Wearable devices for continuous electronic health monitoring often capture data at frequent intervals under a dense functional design. The focal point is the analysis of longitudinal functional data, wherein functional trajectories are observed repeatedly over time. This work is motivated by the interest in assessing the efficacy of a noninflammatory medication, meloxicam, on the daily activity levels of household cats with a pre-existing condition of osteoarthritis under a crossover design. An accelerometer records these activity profiles at a minute level over the entire study period. To this aspect, we propose an orthogonal projection-based pseudo-generalized F test to determine the significance of the functional treatment effect under a functional additive crossover model while accounting for carryover effects and baseline covariates. Under mild conditions, we derive the asymptotic null distribution of the test statistic and the theoretical power function when the projection function is estimated from the data. Numerical studies demonstrate that the proposed test maintains size, proves powerful in detecting the smooth effect of meloxicam, and exhibits high efficiency compared to bootstrap-based alternatives. Application of the test on the accelerometric activity profiles reveals a strong positive effect of meloxicam on the joint pain score of the cats after adjusting for other baseline covariates.

stat.ME↗

Statistical Methods for Crossover Trials: A Review of Classical and Recent Developments

A comprehensive review of the literature on crossover design is needed to highlight its evolution, applications, and methodological advancements across various fields. Given its widespread use in clinical trials and other research domains, understanding this design's challenges, assumptions, and innovations is essential for optimizing its implementation and ensuring accurate, unbiased results. This article extensively reviews the history and statistical inference methods for crossover designs. A primary focus is given to the AB-BA design as it is the most widely used design in literature. Extension from two periods to higher-order designs is discussed, and a general inference procedure for continuous response is studied. Analysis of multivariate and categorical responses is also reviewed in this context. Recent developments, including causal (potential outcomes) formulations of carryover, efficient use of period-specific baselines, generalized estimating equations for repeated measurements within periods, and estimands for incomplete crossover trials, are also discussed. Several open problems in this area are shortlisted.

stat.ME↗

A robust regression approach to synthetic control with interference

Synthetic control methods are widely used for policy evaluation, but most existing approaches rule out interference among units, compromising validity when such effects are present. We develop a framework that accommodates contaminated donor pools and unknown interference patterns through two stages: factor-model adjustment for unobserved confounding, followed by robust regression in which direct and interference effects appear as a sparse outlier component. We study two asymptotic regimes. When the number of units is fixed and at least half are unaffected by interference, high-breakdown robust regression yields consistent identification of valid controls and asymptotically normal inference. When the number of units diverges, we allow for sparse large and dense weak interference, with robust M-estimation remaining valid even when the post-intervention period is short. Unlike existing approaches requiring prespecification of valid controls or parametric modeling of interference, our framework relies only on coarse sparsity information and enables formal inference on both direct and interference effects. We assess the proposed methods through simulations and two empirical applications. An analysis of the US embassy relocation to Jerusalem reveals significant interference effects on conflict outcomes in Jordan, and an analysis of Beijing's air pollution policy uncovers spatial interference patterns consistent with prevailing wind directions.

stat.ME↗