arXiv ScienceSearch

arXiv subjects

Lukas Klein

Publications and source records attributed to Lukas Klein.

18 recordsLinked to original sources

Measuring spectral functions of doped magnets with Rydberg tweezer arrays

Spectroscopic measurements of single-particle spectral functions provide crucial insight into strongly correlated quantum matter by resolving the energy and spatial structure of elementary excitations. Here we introduce a spectroscopic protocol for single-charge injection with simultaneous spatial and energy resolution in a Rydberg tweezer array, effectively emulating scanning tunneling microscopy. By combining this protocol with single-atom-resolved imaging, we go beyond conventional spectroscopy by not only measuring the single-particle spectral function but also directly imaging the microscopic structure of the excitations underlying spectral resonances in frustrated $tJ$ Hamiltonians. We reveal resonances associated with the formation of bound magnetic polarons -- composite quasiparticles consisting of a mobile hole bound to a magnon -- and directly extract their binding energy, spatial extent, and spin character. Finally, by exploiting the spatial tunability of our platform, we measure the local density of states across different lattice geometries. Our work establishes Rydberg tweezer arrays as a powerful platform for spectroscopic studies of strongly correlated models, offering microscopic control and direct real-space access to emergent quasiparticles in engineered quantum matter.

cond-mat.quant-gas

Finally Outshining the Random Baseline: A Simple and Effective Solution for Active Learning in 3D Biomedical Imaging

Active learning (AL) has the potential to drastically reduce annotation costs in 3D biomedical image segmentation, where expert labeling of volumetric data is both time-consuming and expensive. Yet, existing AL methods are unable to consistently outperform improved random sampling baselines adapted to 3D data, leaving the field without a reliable solution. We introduce Class-stratified Scheduled Power Predictive Entropy (ClaSP PE), a simple and effective query strategy that addresses two key limitations of standard uncertainty-based AL methods: class imbalance and redundancy in early selections. ClaSP PE combines class-stratified querying to ensure coverage of underrepresented structures and log-scale power noising with a decaying schedule to enforce query diversity in early-stage AL and encourage exploitation later. In our evaluation on 24 experimental settings using four 3D biomedical datasets within the comprehensive nnActive benchmark, ClaSP PE is the only method that generally outperforms improved random baselines in terms of both segmentation quality with statistically significant gains, whilst remaining annotation efficient. Furthermore, we explicitly simulate the real-world application by testing our method on four previously unseen datasets without manual adaptation, where all experiment parameters are set according to predefined guidelines. The results confirm that ClaSP PE robustly generalizes to novel tasks without requiring dataset-specific tuning. Within the nnActive framework, we present compelling evidence that an AL method can consistently outperform random baselines adapted to 3D segmentation, in terms of both performance and annotation efficiency in a realistic, close-to-production scenario. Our open-source implementation and clear deployment guidelines make it readily applicable in practice. Code is at https://github.com/MIC-DKFZ/nnActive.

cs.CV

Initial data analysis of the national German transplantation registry with a focus on kidney transplantation

This study presents an Initial Data Analysis (IDA) of the German Transplantation Registry (TxReg) data for a better data understanding and to inform future data analyses. The IDA is focusing on data on first-time kidney-only transplantations in adult recipients from deceased donors between 2006 and 2016 and refers to data from 14,954 recipients and 9,964 donors across 25 tables. Investigated aspects include missing data patterns and structure, data consistency, and availability of event time data. Results show that missing data proportions vary widely, with some tables nearly complete while others have over 50% missing values. Missing data patterns are identified using a decision tree approach. An influx and outflux analysis demonstrates that some variables have high potential for imputing missing data, while others were less suitable for imputation. We identified 168 multi-sourced variables that are reported by multiple data providers in parallel leading to discrepancies for some variables but also providing opportunities for missing data imputation. Our findings on event time data demonstrate the importance of carefully selecting the variables used for event time analyses as results will strongly depend on this selection. In summary, our findings highlight the challenges when utilizing the TxReg data for research and provide recommendations for data preprocessing and analysis in future analyses.

stat.AP

nnActive: A Framework for Evaluation of Active Learning in 3D Biomedical Segmentation

Semantic segmentation is crucial for various biomedical applications, yet its reliance on large annotated datasets presents a bottleneck due to the high cost and specialized expertise required for manual labeling. Active Learning (AL) aims to mitigate this challenge by querying only the most informative samples, thereby reducing annotation effort. However, in the domain of 3D biomedical imaging, there is no consensus on whether AL consistently outperforms Random sampling. Four evaluation pitfalls hinder the current methodological assessment. These are (1) restriction to too few datasets and annotation budgets, (2) using 2D models on 3D images without partial annotations, (3) Random baseline not being adapted to the task, and (4) measuring annotation cost only in voxels. In this work, we introduce nnActive, an open-source AL framework that overcomes these pitfalls by (1) means of a large scale study spanning four biomedical imaging datasets and three label regimes, (2) extending nnU-Net by using partial annotations for training with 3D patch-based query selection, (3) proposing Foreground Aware Random sampling strategies tackling the foreground-background class imbalance of medical images and (4) propose the foreground efficiency metric, which captures the low annotation cost of background-regions. We reveal the following findings: (A) while all AL methods outperform standard Random sampling, none reliably surpasses an improved Foreground Aware Random sampling; (B) benefits of AL depend on task specific parameters; (C) Predictive Entropy is overall the best performing AL method, but likely requires the most annotation effort; (D) AL performance can be improved with more compute intensive design choices. As a holistic, open-source framework, nnActive can serve as a catalyst for research and application of AL in 3D biomedical imaging. Code is at: https://github.com/MIC-DKFZ/nnActive

cs.CV

Kinetically-induced bound states in a frustrated Rydberg tweezer array

Understanding how particles bind into composite objects is a ubiquitous theme in physics, from the formation of molecules to hadrons in quantum chromodynamics and the pairing of charge carriers in superconductors. The formation of bound states usually originates from attractive interactions between particles. However, the binding can also arise purely from the motion of dopants due to kinetic frustration, which is potentially related to unconventional pairing in moir\'e materials. Here, we report the first direct observation of kinetically-induced bound states between holes and magnons using a Rydberg atom array quantum simulator of the bosonic $t$-$J$ model in frustrated ladders and 2D lattices. First, we demonstrate the formation of mobile one-hole-one-magnon bound states. We then construct three-particle one-hole-two-magnon bound states and reveal the underlying binding mechanism by observing kinetically-induced singlet correlations. Finally, we investigate how mobile dopants structure their magnetic environment in a spin-balanced 2D triangular lattice, showing that a hole induces $120^\circ$ antiferromagnetic order, while a doublon dopant generates in-plane ferromagnetic correlations. Our results demonstrates compelling evidence of kinetically-induced binding, opening a new avenue to understand novel pairing mechanisms in correlated quantum materials like superconductors in moir\'e superlattices.

quant-ph

Probing spin-motion coupling of two Rydberg atoms by a Stern-Gerlach-like experiment

We propose and implement a protocol to measure the state-dependent motion of Rydberg atoms induced by dipole-dipole interactions. Our setup enables simultaneous readout of both the atomic internal state and position on a one-dimensional array of optical tweezers. We benchmark the protocol using two atoms in the same Rydberg state, which experience van der Waals repulsion, and measure velocities in agreement with theoretical predictions. When preparing the atoms in a different pair state, we observe an oscillatory dynamics that we attribute to the proximity of a macrodimer bound state. Finally, we perform a Stern-Gerlach-like experiment in which a superposition of the two previous pair states results in the separation of the atomic wavepacket into two macroscopically distinct trajectories, thereby demonstrating spin-motion coupling mediated by the interactions.

physics.atom-ph

Benchmarking direct and indirect dipolar spin-exchange interactions between two Rydberg atoms

We report on the experimental characterization of various types of spin-exchange interactions between two individual atoms, where pseudo-spin degrees of freedom are encoded in different Rydberg states. For the case of the direct dipole-dipole interaction between states of opposite parity, such as between $nS$ and $nP$, we investigate the effects of positional disorder arising from the residual atomic motion, on the coherence of spin-exchange oscillations. We then characterize an indirect dipolar spin exchange, i.e., the off-diagonal part of the van der Waals effective Hamiltonian that couples the states $nS$ and $(n+1)S$. Finally, we report on the observation of a new type of dipolar coupling, made resonant using addressable light-shifts and involving four different Rydberg levels: this exchange process is akin to electrically induced F\"orster resonance, but featuring local control. It exhibits an angular dependence distinct from the usual $1-3\cos^2(\theta)$ form of the resonant dipolar spin-exchange.

physics.atom-ph

Realization of a doped quantum antiferromagnet with dipolar tunnelings in a Rydberg tweezer array

Doping an antiferromagnetic Mott insulator is central to our understanding of a variety of phenomena in strongly-correlated electrons, including high-temperature superconductors. To describe the competition between tunneling $t$ of hole dopants and antiferromagnetic (AFM) spin interactions $J$, theoretical and numerical studies often focus on the paradigmatic $t$-$J$ model, and the direct analog quantum simulation of this model in the relevant regime of high-particle density has long been sought. Here, we realize a doped quantum antiferromagnet with next-nearest neighbour (NNN) tunnelings $t'$ and hard-core bosonic holes using a Rydberg tweezer platform. We utilize coherent dynamics between three Rydberg levels, encoding spins and holes, to implement a tunable bosonic $t$-$J$-$V$ model allowing us to study previously inaccessible parameter regimes. We observe dynamical phase separation between hole and spin domains for $|t/J|\ll 1$, and demonstrate the formation of repulsively bound hole pairs in a variety of spin backgrounds. The interference between NNN tunnelings $t'$ and perturbative pair tunneling gives rise to light and heavy pairs depending on the sign of $t$. Using the single-site control allows us to study the dynamics of a single hole in 2D square lattice (anti)ferromagnets. The model we implement extends the toolbox of Rydberg tweezer experiments beyond spin-1/2 models to a larger class of $t$-$J$ and spin-$1$ models.

quant-ph

Tomonaga-Luttinger Liquid Behavior in a Rydberg-encoded Spin Chain

Quantum fluctuations can disrupt long-range order in one-dimensional systems, and replace it with the universal paradigm of the Tomonaga-Luttinger liquid (TLL), a critical phase of matter characterized by power-law decaying correlations and linearly dispersing excitations. Using a Rydberg quantum simulator, we study how TLL physics manifests in the low-energy properties of a spin chain, interacting under either the ferromagnetic or the antiferromagnetic dipolar XY Hamiltonian. Following quasi-adiabatic preparation, we directly observe the power-law decay of spin-spin correlations in real-space, allowing us to extract the Luttinger parameter. In the presence of an impurity, the chain exhibits tunable Friedel oscillations of the local magnetization. Moreover, by utilizing a quantum quench, we directly probe the propagation of correlations, which exhibit a light-cone structure related to the linear sound mode of the underlying TLL. Our measurements demonstrate the influence of the long-range dipolar interactions, renormalizing the parameters of TLL with respect to the case of nearest-neighbor interactions. Finally, comparison to numerical simulations exposes the high sensitivity of TLLs to doping and finite-size effects.

quant-ph

Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities

The various limitations of Generative AI, such as hallucinations and model failures, have made it crucial to understand the role of different modalities in Visual Language Model (VLM) predictions. Our work investigates how the integration of information from image and text modalities influences the performance and behavior of VLMs in visual question answering (VQA) and reasoning tasks. We measure this effect through answer accuracy, reasoning quality, model uncertainty, and modality relevance. We study the interplay between text and image modalities in different configurations where visual content is essential for solving the VQA task. Our contributions include (1) the Semantic Interventions (SI)-VQA dataset, (2) a benchmark study of various VLM architectures under different modality configurations, and (3) the Interactive Semantic Interventions (ISI) tool. The SI-VQA dataset serves as the foundation for the benchmark, while the ISI tool provides an interface to test and apply semantic interventions in image and text inputs, enabling more fine-grained analysis. Our results show that complementary information between modalities improves answer and reasoning quality, while contradictory information harms model performance and confidence. Image text annotations have minimal impact on accuracy and uncertainty, slightly increasing image relevance. Attention analysis confirms the dominant role of image inputs over text in VQA tasks. In this study, we evaluate state-of-the-art VLMs that allow us to extract attention coefficients for each modality. A key finding is PaliGemma's harmful overconfidence, which poses a higher risk of silent failures compared to the LLaVA models. This work sets the foundation for rigorous analysis of modality integration, supported by datasets specifically designed for this purpose.

cs.AI

Navigating the Maze of Explainable AI: A Systematic Approach to Evaluating Methods and Metrics

Explainable AI (XAI) is a rapidly growing domain with a myriad of proposed methods as well as metrics aiming to evaluate their efficacy. However, current studies are often of limited scope, examining only a handful of XAI methods and ignoring underlying design parameters for performance, such as the model architecture or the nature of input data. Moreover, they often rely on one or a few metrics and neglect thorough validation, increasing the risk of selection bias and ignoring discrepancies among metrics. These shortcomings leave practitioners confused about which method to choose for their problem. In response, we introduce LATEC, a large-scale benchmark that critically evaluates 17 prominent XAI methods using 20 distinct metrics. We systematically incorporate vital design parameters like varied architectures and diverse input modalities, resulting in 7,560 examined combinations. Through LATEC, we showcase the high risk of conflicting metrics leading to unreliable rankings and consequently propose a more robust evaluation scheme. Further, we comprehensively evaluate various XAI methods to assist practitioners in selecting appropriate methods aligning with their needs. Curiously, the emerging top-performing method, Expected Gradients, is not examined in any relevant related study. LATEC reinforces its role in future XAI research by publicly releasing all 326k saliency maps and 378k metric scores as a (meta-)evaluation dataset. The benchmark is hosted at: https://github.com/IML-DKFZ/latec.

cs.CV

Enhancing predictive imaging biomarker discovery through treatment effect analysis

Identifying predictive covariates, which forecast individual treatment effectiveness, is crucial for decision-making across different disciplines such as personalized medicine. These covariates, referred to as biomarkers, are extracted from pre-treatment data, often within randomized controlled trials, and should be distinguished from prognostic biomarkers, which are independent of treatment assignment. Our study focuses on discovering predictive imaging biomarkers, specific image features, by leveraging pre-treatment images to uncover new causal relationships. Unlike labor-intensive approaches relying on handcrafted features prone to bias, we present a novel task of directly learning predictive features from images. We propose an evaluation protocol to assess a model's ability to identify predictive imaging biomarkers and differentiate them from purely prognostic ones by employing statistical testing and a comprehensive analysis of image feature attribution. We explore the suitability of deep learning models originally developed for estimating the conditional average treatment effect (CATE) for this task, which have been assessed primarily for their precision of CATE estimation while overlooking the evaluation of imaging biomarker discovery. Our proof-of-concept analysis demonstrates the feasibility and potential of our approach in discovering and validating predictive imaging biomarkers from synthetic outcomes and real-world image datasets. Our code is available at \url{https://github.com/MIC-DKFZ/predictive_image_biomarker_analysis}.

eess.IV

Improving Explainability of Disentangled Representations using Multipath-Attribution Mappings

Explainable AI aims to render model behavior understandable by humans, which can be seen as an intermediate step in extracting causal relations from correlative patterns. Due to the high risk of possible fatal decisions in image-based clinical diagnostics, it is necessary to integrate explainable AI into these safety-critical systems. Current explanatory methods typically assign attribution scores to pixel regions in the input image, indicating their importance for a model's decision. However, they fall short when explaining why a visual feature is used. We propose a framework that utilizes interpretable disentangled representations for downstream-task prediction. Through visualizing the disentangled representations, we enable experts to investigate possible causation effects by leveraging their domain knowledge. Additionally, we deploy a multi-path attribution mapping for enriching and validating explanations. We demonstrate the effectiveness of our approach on a synthetic benchmark suite and two medical datasets. We show that the framework not only acts as a catalyst for causal relation extraction but also enhances model robustness by enabling shortcut detection without the need for testing under distribution shifts.

cs.CV

Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance Assessment

Active Learning (AL) aims to reduce the labeling burden by interactively selecting the most informative samples from a pool of unlabeled data. While there has been extensive research on improving AL query methods in recent years, some studies have questioned the effectiveness of AL compared to emerging paradigms such as semi-supervised (Semi-SL) and self-supervised learning (Self-SL), or a simple optimization of classifier configurations. Thus, today's AL literature presents an inconsistent and contradictory landscape, leaving practitioners uncertain about whether and how to use AL in their tasks. In this work, we make the case that this inconsistency arises from a lack of systematic and realistic evaluation of AL methods. Specifically, we identify five key pitfalls in the current literature that reflect the delicate considerations required for AL evaluation. Further, we present an evaluation framework that overcomes these pitfalls and thus enables meaningful statements about the performance of AL methods. To demonstrate the relevance of our protocol, we present a large-scale empirical study and benchmark for image classification spanning various data sets, query methods, AL settings, and training paradigms. Our findings clarify the inconsistent picture in the literature and enable us to give hands-on recommendations for practitioners. The benchmark is hosted at https://github.com/IML-DKFZ/realistic-al .

cs.CV

A Call to Reflect on Evaluation Practices for Failure Detection in Image Classification

Reliable application of machine learning-based decision systems in the wild is one of the major challenges currently investigated by the field. A large portion of established approaches aims to detect erroneous predictions by means of assigning confidence scores. This confidence may be obtained by either quantifying the model's predictive uncertainty, learning explicit scoring functions, or assessing whether the input is in line with the training distribution. Curiously, while these approaches all state to address the same eventual goal of detecting failures of a classifier upon real-life application, they currently constitute largely separated research fields with individual evaluation protocols, which either exclude a substantial part of relevant methods or ignore large parts of relevant failure sources. In this work, we systematically reveal current pitfalls caused by these inconsistencies and derive requirements for a holistic and realistic evaluation of failure detection. To demonstrate the relevance of this unified perspective, we present a large-scale empirical study for the first time enabling benchmarking confidence scoring functions w.r.t all relevant methods and failure sources. The revelation of a simple softmax response baseline as the overall best performing method underlines the drastic shortcomings of current evaluation in the abundance of publicized research on confidence scoring. Code and trained models are at https://github.com/IML-DKFZ/fd-shifts.

cs.CV

From Correlation to Causation: Formalizing Interpretable Machine Learning as a Statistical Process

Explainable AI (XAI) is a necessity in safety-critical systems such as in clinical diagnostics due to a high risk for fatal decisions. Currently, however, XAI resembles a loose collection of methods rather than a well-defined process. In this work, we elaborate on conceptual similarities between the largest subgroup of XAI, interpretable machine learning (IML), and classical statistics. Based on these similarities, we present a formalization of IML along the lines of a statistical process. Adopting this statistical view allows us to interpret machine learning models and IML methods as sophisticated statistical tools. Based on this interpretation, we infer three key questions, which we identify as crucial for the success and adoption of IML in safety-critical settings. By formulating these questions, we further aim to spark a discussion about what distinguishes IML from classical statistics and what our perspective implies for the future of the field.

cs.CV

Robust Polarization Gradient Cooling of Trapped Ions

We implement three-dimensional polarization gradient cooling of trapped ions. Counter-propagating laser beams near $393\,$nm impinge in lin$\,\perp\,$lin configuration, at a frequency below the S$_{1/2}$ to P$_{3/2}$ resonance in $^{40}$Ca$^+$. We demonstrate mean phonon numbers of $5.4(4)$ at a trap frequency of $2\pi \times 285\,$kHz and $3.3(4)$ at $2\pi\times480\,$kHz, in the axial and radial directions, respectively. Our measurements demonstrate that cooling with laser beams detuned to lower frequencies from the resonance is robust against an elevated phonon occupation number, and thus works well for an initial ion motion far out of the Lamb-Dicke regime, for up to four ions, and for a micromotion modulation index $\beta\leq 0.1$. Still, we find that the spectral impurity of the laser field influences both, cooling rates and cooling limits. Thus, a Fabry-P\'{e}rot cavity filter is employed for efficiently suppressing amplified spontaneous emission of the diode laser.

physics.atom-ph

Collection of micromirror-modulated light in the single-pixel broadband hyperspectral microscope

Digital micromirror device (DMD) serves in a significant part of computational optical setups as a means of encoding an image by the desired pattern. The most prominent is its usage in the so-called single-pixel camera experiment. This experiment often requires an efficient and homogenous collection of light from a relatively large chip on a small area of an optical fiber or spectrometer slit. Moreover, this effort is complicated by the fact that the DMD acts as a diffractive element, which causes severe spectral inhomogeneities in the light collection. We studied the effect of light diffraction via a whiskbroom hyperspectral camera in a broad spectral range. Based on the knowledge, we designed a variety of different approaches to light collection. We mapped the efficiency and spectral homogeneity of each of the configuration - namely its ability to couple the light into commercially available fiber spectrometers working in the visible and IR range (up to 2500 nm). We found the integrating spheres to provide homogeneous light collection, which, however, suffers from very low efficiency. The best compromise between the performance parameters was provided by a combination of an engineered diffuser with an off-axis parabolic mirror. We used this configuration to create a computational microscope able to carry out hyperspectral imaging of a sample in a broad spectral range (400-2500 nm). We see such a setup as an ideal tool to carry out spectrally-resolved transmission microscopy in a broad spectral range.

physics.ins-det