arXiv ScienceSearch

arXiv subjects

David Yang

Publications and source records attributed to David Yang.

At least 19 recordsLinked to original sources

Supra Cognitive Modes: A Routed Architecture for Agent Memory

Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories. We describe Supra Cognitive Modes (SCM), an architecture that maps explicit or automatically selected per-query modes to retrieval and synthesis payloads over one shared ingest substrate. A frozen semantic classifier and runtime gates dispatch queries among fused lexical and dense lookup, graph or iterative multi-hop handling, and stratified long-form synthesis. The substrate combines multi-granularity embeddings, extracted triples, fact-version metadata, and optional asynchronous enrichments. We characterize the deployed configuration on three benchmarks: Long-term Conversational Memory (LoCoMo; n = 1,986), MemoryAgentBench (MAB; n = 3,671), and LongMemEval (n = 500). The reference run records 84.87% on LoCoMo factoid categories and 68.61% on adversarial abstention, 61.49% on MAB across two repetitions, and 86.00% on LongMemEval. A repository-backed reproduction produces similar aggregate scores and supports task- and mode-conditioned failure analysis. Raw baseline outputs, aligned end-to-end timing for LoCoMo and LongMemEval, and complete token ledgers are unavailable; stored rows also omit some final runtime decisions. The results characterize one implemented routed configuration and its diagnostic failure patterns, while source inspection verifies the per-query control interface and shared-substrate design. Causal routing effects, efficiency gains, and statistical significance remain outside the available evidence.

cs.AI

Assessing global drivers of forest transpiration using clustered machine learning models

Understanding the environmental drivers of forest transpiration is critical for improving global predictions of water availability and ecosystem health. Due to many competing controls on plant water stress and ecosystem transpiration, however, these drivers may vary widely across tree species which have adapted hydraulically to local climate conditions. Here, clustered machine learning models were used to analyze global drivers of forest transpiration rates using the SAPFLUXNET database. Sap flux data from a total of ninety-five sites spanning seven biomes were grouped using two clustering strategies: by biome and by plant functional type. Two supervised machine learning algorithms, a random forest algorithm and a neural network algorithm, were used to predict rates of sap flux for each cluster. The performance and feature importance in each model were analyzed and compared to evaluate the environmental variables that control each cluster's performance. By defining site clusters, these models are able to predict transpiration and its environmental drivers across a wide variety of geographical sites and tree species. Unlike models trained on the entire dataset, high-performing clustered models achieved R$^2$ values to measurement data in the range of 0.74 to 0.90, with the highest performance being achieved in mid-sized clusters of up to thirty-six sites. There was high variance in feature importance between clusters, indicating that key predictors of transpiration varied strongly across both plant functional type and biome. Overall, water-limited climates tended to be more controlled by soil moisture, whereas climates with high mean annual temperature tended to be more controlled by solar radiation and less dependent on air temperature. These findings provide insights into how forest transpiration responds to environmental factors across a wide range of climate types and tree species.

q-bio.QM

On Moy-Prasad quotients over Laurent series fields

Let $k$ be an algebraically closed field and $G$ a connected reductive group over $k((t))$ satisfying some conditions. We define a stratification by conjugacy classes of twisted Levi subgroups of $G$ on each Moy-Prasad quotient $\mathfrak{k}_{x,r}/\mathfrak{k}_{x,r+}$ of $G$. We then calculate the strata in terms of the associated twisted Levi subgroups. This calculation is necessary for several followup papers on the local geometric Langlands program.

math.RT

Quantitative 3D imaging of highly distorted micro-crystals using Bragg ptychography

Bragg coherent diffraction imaging (BCDI) fails to reliably retrieve phases in micro-crystals exhibiting strong strain inhomogeneities, which restricts its applicability. Here we show that three-dimensional Bragg ptychography (3DBP) overcomes this limitation by enabling stable inversion for large lattice distortions. Using a combination of experimental measurements and numerical tests, we compare the performance limits of the two approaches and demonstrate that 3DBP tolerates lattice distortions more than six times larger than BCDI. We also establish the sensitivity of both methods on a weakly distorted crystal, for which 3DBP yields smoother amplitude and phase fields with reduced short-length-scale artifacts. 3DBP thus provides a reliable route for imaging micro-crystals with large lattice distortions, expanding the scope of coherent X-ray Bragg microscopy to strongly deformed systems.

physics.optics

OSF: On Pre-training and Scaling of Sleep Foundation Models

Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from substantial heterogeneity across recording devices and cohorts. There have been growing efforts to build general-purpose foundation models (FMs) for sleep physiology, but lack an in-depth understanding of the pre-training process and scaling patterns that lead to more generalizable sleep FMs. To fill this gap, we curate a massive corpus of 166,500 hours of sleep recordings from nine public sources and establish SleepBench, a comprehensive, fully open-source benchmark. Leveraging SleepBench, we systematically evaluate four families of self-supervised pre-training objectives and uncover three critical findings: (1) existing FMs fail to generalize to missing channels at inference; (2) channel-invariant feature learning is essential for pre-training; and (3) scaling sample size, model capacity, and multi-source data mixture consistently improves downstream performance.With an enhanced pre-training and scaling recipe, we introduce OSF, a family of sleep FMs that achieves state-of-the-art performance across nine datasets on diverse sleep and disease prediction tasks. Further analysis of OSF also reveals intriguing properties in sample efficiency, hierarchical aggregation, and cross-dataset scaling.

cs.LG

Vision Transformer for Multi-Domain Phase Retrieval in Coherent Diffraction Imaging

Bragg coherent diffraction imaging (BCDI) phase retrieval becomes rapidly difficult in the strong-phase regime, where a crystal contains distortions beyond half a lattice spacing. An important special case is the phase domain problem, where blocks of a crystal are displaced with sharp jumps at domain walls. The strong-phase, here defined as beyond $\pm \pi/2$, generates split Bragg peaks and dense fringe structure for which classical iterative solvers often stagnate or return different solutions from different initialisations. Here, we introduce an unsupervised Fourier Vision Transformer (Fourier ViT) to solve this block-phase, multi-domain phase-retrieval problem directly from measured 2D Bragg diffraction intensities. Fourier ViT couples reciprocal-space information globally through multiscale Fourier token mixing, while shallow convolutional front and back-ends provide local filtering and reconstruction. We validate the approach on large-scale synthetic datasets of Voronoi multi-domain crystals with strong-phase contrast under realistic noise corruptions, and on experimental diffraction from a $\mathrm{La}_{2-x}\mathrm{Ca}_x\mathrm{MnO}_4$ nanocrystal. Across the regimes considered, Fourier ViT achieves the lowest reciprocal-space mismatch ($\chi^2$) among the compared methods and preserves domain-resolved phase reconstructions for increasing numbers of domains. On experimental data, with the same real-space support, Fourier ViT matches the iterative benchmark $\chi^2$ while improving robustness to random initialisations, yielding a higher success rate of low-$\chi^2$ reconstructions than the complex convolutional neural network baseline.

physics.optics

A Direct Second-Order Method for Solving Two-Player Zero-Sum Games

We introduce, to our knowledge, the first direct second-order method for computing Nash equilibria in two-player zero-sum games. To do so, we construct a Douglas-Rachford-style splitting formulation, which we then solve with a semi-smooth Newton (SSN) method. We show that our algorithm enjoys local superlinear convergence. In order to augment the fast local behavior of our SSN method with global efficiency guarantees, we develop a hybrid method that combines our SSN method with the state-of-the-art first-order method for game solving, Predictive Regret Matching$^+$ (PRM$^+$). Our hybrid algorithm leverages the global progress provided by PRM$^+$, while achieving a local superlinear convergence rate once it switches to SSN near a Nash equilibrium. Numerical experiments on matrix games demonstrate order-of-magnitude speedups over PRM$^+$ for high-precision solutions.

cs.GT

DualProtoSeg: Simple and Efficient Design with Text- and Image-Guided Prototype Learning for Weakly Supervised Histopathology Image Segmentation

Weakly supervised semantic segmentation (WSSS) in histopathology seeks to reduce annotation cost by learning from image-level labels, yet it remains limited by inter-class homogeneity, intra-class heterogeneity, and the region-shrinkage effect of CAM-based supervision. We propose a simple and effective prototype-driven framework that leverages vision-language alignment to improve region discovery under weak supervision. Our method integrates CoOp-style learnable prompt tuning to generate text-based prototypes and combines them with learnable image prototypes, forming a dual-modal prototype bank that captures both semantic and appearance cues. To address oversmoothing in ViT representations, we incorporate a multi-scale pyramid module that enhances spatial precision and improves localization quality. Experiments on the BCSS-WSSS benchmark show that our approach surpasses existing state-of-the-art methods, and detailed analyses demonstrate the benefits of text description diversity, context length, and the complementary behavior of text and image prototypes. These results highlight the effectiveness of jointly leveraging textual semantics and visual prototype learning for WSSS in digital pathology.

cs.CV

ConStruct: Structural Distillation of Foundation Models for Prototype-Based Weakly Supervised Histopathology Segmentation

Weakly supervised semantic segmentation (WSSS) in histopathology relies heavily on classification backbones, yet these models often localize only the most discriminative regions and struggle to capture the full spatial extent of tissue structures. Vision-language models such as CONCH offer rich semantic alignment and morphology-aware representations, while modern segmentation backbones like SegFormer preserve fine-grained spatial cues. However, combining these complementary strengths remains challenging, especially under weak supervision and without dense annotations. We propose a prototype learning framework for WSSS in histopathological images that integrates morphology-aware representations from CONCH, multi-scale structural cues from SegFormer, and text-guided semantic alignment to produce prototypes that are simultaneously semantically discriminative and spatially coherent. To effectively leverage these heterogeneous sources, we introduce text-guided prototype initialization that incorporates pathology descriptions to generate more complete and semantically accurate pseudo-masks. A structural distillation mechanism transfers spatial knowledge from SegFormer to preserve fine-grained morphological patterns and local tissue boundaries during prototype learning. Our approach produces high-quality pseudo masks without pixel-level annotations, improves localization completeness, and enhances semantic consistency across tissue types. Experiments on BCSS-WSSS datasets demonstrate that our prototype learning framework outperforms existing WSSS methods while remaining computationally efficient through frozen foundation model backbones and lightweight trainable adapters.

cs.CV

IMACT-CXR: An Interactive Multi-Agent Conversational Tutoring System for Chest X-Ray Interpretation

IMACT-CXR is an interactive multi-agent conversational tutor that helps trainees interpret chest X-rays by unifying spatial annotation, gaze analysis, knowledge retrieval, and image-grounded reasoning in a single AutoGen-based workflow. The tutor simultaneously ingests learner bounding boxes, gaze samples, and free-text observations. Specialized agents evaluate localization quality, generate Socratic coaching, retrieve PubMed evidence, suggest similar cases from REFLACX, and trigger NV-Reason-CXR-3B for vision-language reasoning when mastery remains low or the learner explicitly asks. Bayesian Knowledge Tracing (BKT) maintains skill-specific mastery estimates that drive both knowledge reinforcement and case similarity retrieval. A lung-lobe segmentation module derived from a TensorFlow U-Net enables anatomically aware gaze feedback, and safety prompts prevent premature disclosure of ground-truth labels. We describe the system architecture, implementation highlights, and integration with the REFLACX dataset for real DICOM cases. IMACT-CXR demonstrates responsive tutoring flows with bounded latency, precise control over answer leakage, and extensibility toward live residency deployment. Preliminary evaluation shows improved localization and diagnostic reasoning compared to baselines.

cs.AI

Spatio-temporal migration of antiferromagnetic domain walls in Sr2IrO4

By laser pump-probe time-resolved coherent magnetic X-ray diffraction imaging, we have measured the migration velocity of antiferromagnetic domain walls in the Mott insulator Sr2IrO4 at 100 K. During the laser-induced demagnetization, we observe domain walls moving at 3x10^6 m/s, significantly faster than acoustic velocities. This is understood to arise from a purely electronic spin contribution to the magnetic structure without any role for coupling to the crystal lattice.

cond-mat.str-el

Effect of nearby Metals on Electro-Quasistatic Human Body Communication

In recent decades Human Body Communication has emerged as a promising alternative to traditional radio wave communication, utilizing the body's conductive properties for low-power connectivity among wearables. This method harnesses the human body as an energy-efficient channel for data transmission within the electro-quasistatic frequency range, enabling advancements in human-machine interaction. While prior work has noted the role of parasitic return paths in such capacitively coupled systems, the influence of surrounding metallic objects on these paths, which are critical for EQS wireless signaling, has not been fully explored. This paper fills that gap with a structured study of how various conducting objects, from non-grounded (floating) metals and grounded metals to enclosed metallic environments such as elevators and cars, affect the body-communication channel. We present a theoretical framework supported by finite element method simulations and experiments with wearable devices. Results show that metallic objects within 20 cm of devices can reduce transmission loss by about 10 dB. When a device ground connects to a grounded metallic object, channel gain can increase by at least 20 dB. Contact area during touch-based interactions with grounded metals produces contact-impedance dependent high-pass channel characteristics. Proximity to metallic objects introduces variability within a critical distance, with grounded metals producing a larger overall effect than floating metals. These findings improve understanding of body-centric communication links and inform design for healthcare, consumer electronics, defense, and industrial applications.

eess.SP

Ultrafast heat transfer in single palladium nanocrystals seen with an X-ray free-electron laser

We report transient highly strained structural states in individual palladium (Pd) nanocrystals, electronically heated using an optical laser, which precede their uniform thermal expansion. Using an X-ray free-electron laser probe, the evolution of individual 111 Bragg peaks is measured as a function of delay time at various laser fluences. Above a laser fluence threshold at a sufficient pump-probe delay, the Bragg peak splits into multiple peaks, indicating heterogeneous strain, before returning to a single peak, corresponding to even heat distribution throughout the lattice expanded crystal. Our findings are supported by a lattice displacement and strain model of a single nanocrystal at different delay times, which agrees with the experimental data. Our observations have implications for understanding femtosecond laser interactions with metals and the potential photo-catalytic performance of Pd.

cond-mat.mtrl-sci

Unveiling Nano-scale Crystal Deformation using Coherent X-ray Dynamical Diffraction

Visualization of internal deformation fields in crystalline materials helps bridge the gap between theoretical models and practical applications. Applying Bragg coherent diffraction imaging under X-ray dynamical diffraction conditions provides a promising approach to the longstanding challenge of investigating the deformation fields in micron-sized crystals. Here, we present an automatic differentiation-based reconstruction method that integrates dynamical scattering theory to accurately reconstruct deformation fields in large crystals. Using this forward model, our simulated and experimental results demonstrate that three-dimensional local strain information inside a large crystal can be accurately reconstructed under coherent X-ray dynamical diffraction conditions with Bragg coherent X-ray diffraction imaging. These findings open an avenue for extending the investigation of local deformation fields to microscale crystals while maintaining nanoscale resolution, leveraging the enhanced coherence and brightness of advanced X-ray sources.

physics.app-ph

Beyond the First Read: AI-Assisted Perceptual Error Detection in Chest Radiography Accounting for Interobserver Variability

Chest radiography is widely used in diagnostic imaging. However, perceptual errors -- especially overlooked but visible abnormalities -- remain common and clinically significant. Current workflows and AI systems provide limited support for detecting such errors after interpretation and often lack meaningful human--AI collaboration. We introduce RADAR (Radiologist--AI Diagnostic Assistance and Review), a post-interpretation companion system. RADAR ingests finalized radiologist annotations and CXR images, then performs regional-level analysis to detect and refer potentially missed abnormal regions. The system supports a "second-look" workflow and offers suggested regions of interest (ROIs) rather than fixed labels to accommodate inter-observer variation. We evaluated RADAR on a simulated perceptual-error dataset derived from de-identified CXR cases, using F1 score and Intersection over Union (IoU) as primary metrics. RADAR achieved a recall of 0.78, precision of 0.44, and an F1 score of 0.56 in detecting missed abnormalities in the simulated perceptual-error dataset. Although precision is moderate, this reduces over-reliance on AI by encouraging radiologist oversight in human--AI collaboration. The median IoU was 0.78, with more than 90% of referrals exceeding 0.5 IoU, indicating accurate regional localization. RADAR effectively complements radiologist judgment, providing valuable post-read support for perceptual-error detection in CXR interpretation. Its flexible ROI suggestions and non-intrusive integration position it as a promising tool in real-world radiology workflows. To facilitate reproducibility and further evaluation, we release a fully open-source web implementation alongside a simulated error dataset. All code, data, demonstration videos, and the application are publicly available at https://github.com/avutukuri01/RADAR.

cs.CV

Edge-boosted graph learning for functional brain connectivity analysis

Predicting disease states from functional brain connectivity is critical for the early diagnosis of severe neurodegenerative diseases such as Alzheimer's Disease and Parkinson's Disease. Existing studies commonly employ Graph Neural Networks (GNNs) to infer clinical diagnoses from node-based brain connectivity matrices generated through node-to-node similarities of regionally averaged fMRI signals. However, recent neuroscience studies found that such node-based connectivity does not accurately capture ``functional connections" within the brain. This paper proposes a novel approach to brain network analysis that emphasizes edge functional connectivity (eFC), shifting the focus to inter-edge relationships. Additionally, we introduce a co-embedding technique to integrate edge functional connections effectively. Experimental results on the ADNI and PPMI datasets demonstrate that our method significantly outperforms state-of-the-art GNN methods in classifying functional brain networks.

cs.LG