arXiv ScienceSearch

arXiv subjects

Dawei Wang

Publications and source records attributed to Dawei Wang.

At least 19 recordsLinked to original sources

Dual IMU System for Accurate Gastrointestinal Motility Tracking

Gastrointestinal (GI) motility disorders impair the coordinated movement of food through the digestive tract. The Wireless Motility Capsule was developed to indirectly assess GI motility by recording pressure, pH, and temperature along the GI tract. Although an inertial measurement unit (IMU) enables precise and direct motion tracking, a single-IMU system is highly susceptible to artifacts caused by body movement and respiration. To address this limitation, we propose a dual-IMU system designed to distinguish intrinsic GI motility from extrinsic body motion. One IMU is integrated into an ingestible capsule to monitor internal movement, while a second IMU is worn externally to capture whole-body dynamics. Experimental validation was conducted using a turntable to simulate random body movements and a Stewart platform oscillating at 0.1 Hz to emulate GI motility. Three sensor fusion algorithms, including the Extended Kalman Filter, Madgwick Filter, and Mahony Filter, were benchmarked to calculate the real-time relative orientation between the two IMUs using quaternion analysis. The results show that the dual-IMU system effectively differentiates platform motion from turntable motion, demonstrating its potential to isolate GI-specific motility. Pearson correlation coefficients greater than 0.8 between the baseline and noise-cancelled motility signals indicate strong agreement and effective noise suppression.

eess.SP

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.

cs.AI

Deterministic Minimum-Leakage Continuous-Variable Quantum Key Distribution with Phase-Conjugated Twin Beams

Minimum-leakage continuous-variable quantum key distribution suppresses Eve's Holevo information by engineering the signal ensemble during state preparation. The existing symmetric realization relies on Alice-side heralding. Here we propose an alternative two-mode symmetric realization that is deterministic without heralding. The minimum-leakage condition of the new protocol requires transmitting two phase-conjugated twin beams (PCTB). We show that the PCTB protocol is related to the heralding protocol through a common entanglement-based source but corresponds to a different prepare-and-measure decomposition. The two protocols attain the same secret key rate per transmitted optical mode in the very-large-squeezing limit. For finite squeezing, however, the PCTB protocol requires approximately 3 dB less squeezing to achieve the same performance. We further optimize mode-symmetric Gaussian two-mode entangling-cloner attacks and find that correlated ancillary modes provide only a limited advantage over independent attacks under the minimum-leakage condition. These results establish phase-conjugated twin beams as a deterministic and experimentally appealing route to symmetric minimum-leakage CV-QKD.

quant-ph

Entropy engineering of BF-BT-based high-entropy ceramics for ultra-high energy storage performance

Dielectric capacitors are promising for pulsed power applications, but the energy storage performance of lead-free bulk ceramics is often limited by low breakdown strength and large ferroelectric hysteresis. Herein, a high entropy perovskite oxide BaTi0.2Zr0.2Sn0.2Hf0.2Nb0.1Sc0.1O3 was introduced into the BF-BT matrix to develop lead free high entropy ferroelectric ceramics. Multicomponent B-site substitution induces lattice distortion, enhanced pseudocubic characteristics, relaxor behavior, and grain refinement. These effects suppress polarization hysteresis and electrical conduction, resulting in a significant increase in breakdown strength. A maximum breakdown strength of 840 kV/cm1 and a recoverable energy density of 10.55 J/cm3 were achieved. Entropy-induced microstructural heterogeneity promotes a more uniform electric-field distribution and delays dielectric breakdown. This work demonstrates entropy engineering as an effective route to achieving high breakdown strength and superior energy-storage performance in lead-free ferroelectric ceramics.

cond-mat.mtrl-sci

MinSurf: resolving the atomic-scale stability landscape of mineral surfaces

Mineral surfaces govern interfacial reactivity in carbon mineralization, geo-energy storage, contaminant immobilization, heterogeneous catalysis and electrochemical interface engineering. Yet atomistic simulations often rely on commonly used facets or facet-level stability criteria, while distinct atomic terminations of the same crystallographic orientation are rarely resolved systematically because experimental characterization and density functional theory (DFT) calculations remain costly across large surface spaces. Here we present MinSurf, a high-throughput framework that resolves mineral surface selection as a surface-energy and morphology problem. MinSurf integrates surface enumeration, DFT labelling, machine-learning interatomic potentials and Wulff construction to predict stable terminations, surface-energy landscapes and equilibrium crystal morphologies. Applied to ten representative minerals, MinSurfSet comprises 764 surface slabs, with 90 corresponding oriented unit cells constructed as bulk references for surface-energy evaluation. The resulting MinNEP model predicts DFT surface energies with a mean absolute error of 0.0119 eV per Angstrom squared and achieves an overall acceleration of 1.14 x 10^4 relative to DFT. MinNEP preserves the DFT-derived morphology-determining surface-energy hierarchy and reproduces the dominant Wulff-exposed facets, while X-ray diffraction provides an independent crystallographic consistency check for alpha-quartz benchmark. By linking atomic terminations, surface energies and equilibrium morphologies, MinSurf provides reproducible and physically representative surface models for high-throughput simulations of mineral interfaces across energy, environmental and advanced inorganic materials.

cond-mat.mtrl-sci

Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classification, their performance degrades significantly as the label space expands a phenomenon we define as Performance Collapse in Long Sequence Recognition. Through an information theoretic analysis, we reveal that this collapse stems from a fundamental conflict between the escalating information entropy and the prominent attention dilution and decay within attention mechanisms, which impairs the model's ability to maintain a sufficient signal-to-noise ratio when processing extremely long prompts. To mitigate this, we propose Divide-and-Conquer Inference (DCI), a novel test-time scaling strategy for visual recognition with MLLMs. DCI recursively decomposes complex global classification tasks into multiple simpler, localized subproblems and employs a dynamic pruning mechanism to compress the search space. This method effectively improves the local signal to noise ratio and model accuracy by mitigating the inherent weight dilution issues in long-sequence inference. Moreover, while traditional self-attention incurs a prohibitive quadratic computational complexity, DCI achieves more favorable scaling behavior and substantially accelerates inference in large scale classification scenarios. Extensive experiments on benchmarks such as ImageNet-1K and ImageNet-21K demonstrate that DCI consistently improves classification accuracy. This enables lightweight open-source models to rival or even surpass frontier closed-source giants without any additional training or fine-tuning. As a model-agnostic, plug-and-play paradigm, DCI offers an efficient approach for scaling the inferential precision of MLLMs in large-scale scenarios.

cs.CV

Kirigami-Structured Electronic Capsule for Long-Term Continuous Gastric Monitoring

Ingestible electronic systems enable non-invasive, in situ sensing within the gastrointestinal (GI) tract, yet clinical translation has been limited by uncontrolled transit, short operational lifetimes, and unreliable wireless communication that prevent continuous monitoring. Here, we present a gastric-resident ingestible robotic platform that achieves week-long operation through integration of a bioinspired, electrically triggered release mechanism with a kirigami-enabled electronic architecture. A kirigami-patterned flexible printed circuit board spans the capsule body and deployable superelastic arms, enabling high-density integration of sensing, power management, and wireless modules within a constrained volume while tolerating large mechanical deformation during gastric residence. Stable retention and on-demand disassembly are achieved using thermally responsive polycaprolactone joints that transition from rigid to compliant states under electrical activation, avoiding dependence on variable chemical triggers. Reliable telemetry in the highly attenuating gastric environment is maintained using a dual-band Bluetooth Low Energy and sub-gigahertz module with RSSI- and throughput-aware adaptive transmission, balancing link robustness and energy consumption. We demonstrate long-term, continuous monitoring of gastric radiation exposure, enabling early detection of dose accumulation and providing a promising in vivo alternative to wearable or handheld dosimeters. Swine studies confirm stable gastric residence, sustained real-time telemetry, and safe gastrointestinal passage following triggered disassembly. This work establishes kirigami-enabled integration as a scalable strategy for long-term gastric-resident robotic systems.

eess.SY

Step-level Denoising-time Diffusion Alignment with Multiple Objectives

Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single reward function under a KL regularization constraint. In practice, however, human preferences are inherently pluralistic, and aligned models must balance multiple downstream objectives, such as aesthetic quality and text-image consistency. Existing multi-objective approaches either rely on costly multi-objective RL fine-tuning or on fusing separately aligned models at denoising time, but they generally require access to reward values (or their gradients) and/or introduce approximation error in the resulting denoising objectives. In this paper, we revisit the problem of RL fine-tuning for diffusion models and address the intractability of identifying the optimal policy by introducing a step-level RL formulation. Building on this, we further propose Multi-objective Step-level Denoising-time Diffusion Alignment (MSDDA), a retraining-free framework for aligning diffusion models with multiple objectives, obtaining the optimal reverse denoising distribution in closed form, with mean and variance expressed directly in terms of single-objective base models. We prove that this denoising-time objective is exactly equivalent to the step-level RL fine-tuning, introducing no approximation error. Moreover, we provide numerical results, which indicate our method outperforms existing denoising-time approaches.

cs.LG

A Capsule-Sized Multi-Wavelength Wireless Optical System for Edge-AI-Based Classification of Gastrointestinal Bleeding Flow Rate

Post-endoscopic gastrointestinal (GI) rebleeding frequently occurs within the first 72 hours after therapeutic hemostasis and remains a major cause of early morbidity and mortality. Existing non-invasive monitoring approaches primarily provide binary blood detection and lack quantitative assessment of bleeding severity or flow dynamic, limiting their ability to support timely clinical decision-making during this high-risk period. In this work, we developed a capsule-sized, multi-wavelength optical sensing wireless platform for order-of-magnitude-level classification of GI bleeding flow rate, leveraging transmission spectroscopy and low-power edge artificial intelligence. The system performs time-resolved, multi-spectral measurements and employs a lightweight two-dimensional convolutional neural network for on-device flow-rate classification, with physics-based validation confirming consistency with wavelength-dependent hemoglobin absorption behavior. In controlled in vitro experiments under simulated gastric conditions, the proposed approach achieved an overall classification accuracy of 98.75% across multiple bleeding flow-rate levels while robustly distinguishing diverse non-blood gastrointestinal interference. By performing embedded inference directly on the capsule electronics, the system reduced overall energy consumption by approximately 88% compared with continuous wireless transmission of raw data, making prolonged, battery-powered operation feasible. Extending capsule-based diagnostics beyond binary blood detection toward continuous, site-specific assessment of bleeding severity, this platform has the potential to support earlier identification of clinically significant rebleeding and inform timely re-intervention during post-endoscopic surveillance.

eess.IV

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-making emerging as core general-purpose capabilities for adapting to diverse scenarios and tasks. Real-time strategy (RTS) games serve as an ideal testbed for evaluating these two capabilities, as their inherent gameplay requires both macro-level strategic planning and micro-level tactical adaptation and action execution. Existing RTS game-based environments either suffer from relatively high computational demands or lack support for textual observations, which has constrained the use of RTS games for LLM evaluation. Motivated by this, we present TowerMind, a novel environment grounded in the tower defense (TD) subgenre of RTS games. TowerMind preserves the key evaluation strengths of RTS games for assessing LLMs, while featuring low computational demands and a multimodal observation space, including pixel-based, textual, and structured game-state representations. In addition, TowerMind supports the evaluation of model hallucination and provides a high degree of customizability. We design five benchmark levels to evaluate several widely used LLMs under different multimodal input settings. The results reveal a clear performance gap between LLMs and human experts across both capability and hallucination dimensions. The experiments further highlight key limitations in LLM behavior, such as inadequate planning validation, a lack of multifinality in decision-making, and inefficient action use. We also evaluate two classic reinforcement learning algorithms: Ape-X DQN and PPO. By offering a lightweight and multimodal design, TowerMind complements the existing RTS game-based environment landscape and introduces a new benchmark for the AI agent field. The source code is publicly available on GitHub(https://github.com/tb6147877/TowerMind).

cs.AI

Intelligent recognition of GPR road hidden defect images based on feature fusion and attention mechanism

Ground Penetrating Radar (GPR) has emerged as a pivotal tool for non-destructive evaluation of subsurface road defects. However, conventional GPR image interpretation remains heavily reliant on subjective expertise, introducing inefficiencies and inaccuracies. This study introduces a comprehensive framework to address these limitations: (1) A DCGAN-based data augmentation strategy synthesizes high-fidelity GPR images to mitigate data scarcity while preserving defect morphology under complex backgrounds; (2) A novel Multi-modal Chain and Global Attention Network (MCGA-Net) is proposed, integrating Multi-modal Chain Feature Fusion (MCFF) for hierarchical multi-scale defect representation and Global Attention Mechanism (GAM) for context-aware feature enhancement; (3) MS COCO transfer learning fine-tunes the backbone network, accelerating convergence and improving generalization. Ablation and comparison experiments validate the framework's efficacy. MCGA-Net achieves Precision (92.8%), Recall (92.5%), and mAP@50 (95.9%). In the detection of Gaussian noise, weak signals and small targets, MCGA-Net maintains robustness and outperforms other models. This work establishes a new paradigm for automated GPR-based defect detection, balancing computational efficiency with high accuracy in complex subsurface environments.

cs.CV

Lightweight framework for underground pipeline recognition and spatial localization based on multi-view 2D GPR images

To address the issues of weak correlation between multi-view features, low recognition accuracy of small-scale targets, and insufficient robustness in complex scenarios in underground pipeline detection using 3D GPR, this paper proposes a 3D pipeline intelligent detection framework. First, based on a B/C/D-Scan three-view joint analysis strategy, a three-dimensional pipeline three-view feature evaluation method is established by cross-validating forward simulation results obtained using FDTD methods with actual measurement data. Second, the DCO-YOLO framework is proposed, which integrates DySample, CGLU, and OutlookAttention cross-dimensional correlation mechanisms into the original YOLOv11 algorithm, significantly improving the small-scale pipeline edge feature extraction capability. Furthermore, a 3D-DIoU spatial feature matching algorithm is proposed, which integrates three-dimensional geometric constraints and center distance penalty terms to achieve automated association of multi-view annotations. The three-view fusion strategy resolves inherent ambiguities in single-view detection. Experiments based on real urban underground pipeline data show that the proposed method achieves accuracy, recall, and mean average precision of 96.2%, 93.3%, and 96.7%, respectively, in complex multi-pipeline scenarios, which are 2.0%, 2.1%, and 0.9% higher than the baseline model. Ablation experiments validated the synergistic optimization effect of the dynamic feature enhancement module and Grad-CAM++ heatmap visualization demonstrated that the improved model significantly enhanced its ability to focus on pipeline geometric features. This study integrates deep learning optimization strategies with the physical characteristics of 3D GPR, offering an efficient and reliable novel technical framework for the intelligent recognition and localization of underground pipelines.

cs.CV

DMA: Online RAG Alignment with Human Feedback

Retrieval-augmented generation (RAG) systems often rely on static retrieval, limiting adaptation to evolving intent and content drift. We introduce Dynamic Memory Alignment (DMA), an online learning framework that systematically incorporates multi-granularity human feedback to align ranking in interactive settings. DMA organizes document-, list-, and response-level signals into a coherent learning pipeline: supervised training for pointwise and listwise rankers, policy optimization driven by response-level preferences, and knowledge distillation into a lightweight scorer for low-latency serving. Throughout this paper, memory refers to the model's working memory, which is the entire context visible to the LLM for In-Context Learning. We adopt a dual-track evaluation protocol mirroring deployment: (i) large-scale online A/B ablations to isolate the utility of each feedback source, and (ii) few-shot offline tests on knowledge-intensive benchmarks. Online, a multi-month industrial deployment further shows substantial improvements in human engagement. Offline, DMA preserves competitive foundational retrieval while yielding notable gains on conversational QA (TriviaQA, HotpotQA). Taken together, these results position DMA as a principled approach to feedback-driven, real-time adaptation in RAG without sacrificing baseline capability.

cs.AI

Towards a deeper fundamental understanding of (Al,Sc)N ferroelectric nitrides

Density Functional Theory (DFT) calculations, within the virtual crystal alloy approximation, are performed, along with the development of a Landau-type model employing a symmetry-allowed analytical expression of the internal energy and having parameters being determined from first principles, to investigate properties and energetics of Al1-xScxN ferroelectric nitrides in their hexagonal forms. These DFT computations and this model predict the existence of two different types of minima, namely the 4-fold-coordinated wurtzite (WZ) polar structure and a 5-times paraelectric hexagonal phase (to be denoted as H5), for any Sc composition up to 40%. The H5 minimum progressively becomes the lowest energy state within hexagonal symmetry as the Sc concentration increases from 0 to 40%. Furthermore, the model points out to several key findings. Examples include the crucial role of the coupling between polarization and strains to create the WZ minimum, in addition to polar and elastic energies, and that the origin of the H5 state overcoming the WZ phase as the global minimum within hexagonal symmetry when increasing the Sc composition mostly lies in the compositional dependency of only two parameters, one linked to the polarization and another one being purely elastic in nature. Other examples are that forcing Al1-xScxN systems to have no or a weak change in lattice parameters when heating them allows to reproduce well their finite-temperature polar properties, and that a value of the axial ratio close to that of the ideal WZ structure does imply a large polarization at low temperatures but not necessarily at high temperatures because of the ordered-disordered character of the temperature-induced formation of the WZ state. Such findings should allow for a better fundamental understanding of (Al,Sc)N ferroelectric nitrides, which may be used to design efficient devices operating at low voltages.

cond-mat.mtrl-sci

Domain-Wall Mediated Polarization Switching in Ferroelectric AlScN: Strain Relief and Field-Dependent Dynamics

While scandium-doped aluminum nitride (AlScN) exhibits robust ferroelectricity and excellent thermal stability, its utility is limited by an exceptionally high coercive field ($E_c$) for polarization switching. Unraveling the atomistic switching dynamics is therefore critical for tailoring $E_c$. Here, we combine density functional theory and machine-learning molecular dynamics to elucidate the polarization switching mechanisms in AlScN over various Sc concentrations and applied electric fields. We find that excessive lattice strain strictly prohibits collective polarization switching, but the pre-existing domain walls relieve strain and lead to a distinct switching dynamics -- dictating a field-dependent switching mechanism. At low electric fields, switching occurs via gradual domain-wall propagation consistent with the Kolmogorov-Avrami-Ishibashi model. In contrast, high fields stimulate additional nucleation, driving a rapid, homogeneous reversal process described by the simultaneous non-linear nucleation and growth model. These findings highlight the critical role of domain-wall dynamics and suggest domain engineering as a viable strategy to tailor coercive fields in AlScN and related ferroelectrics.

cond-mat.mtrl-sci

AdaptJobRec: Enhancing Conversational Career Recommendation through an LLM-Powered Agentic System

In recent years, recommendation systems have evolved from providing a single list of recommendations to offering a comprehensive suite of topic focused services. To better accomplish this task, conversational recommendation systems (CRS) have progressed from basic retrieval augmented LLM generation to agentic systems with advanced reasoning and self correction capabilities. However, agentic systems come with notable response latency, a longstanding challenge for conversational recommendation systems. To balance the trade off between handling complex queries and minimizing latency, we propose AdaptJobRec, the first conversational job recommendation system that leverages autonomous agent to integrate personalized recommendation algorithm tools. The system employs a user query complexity identification mechanism to minimize response latency. For straightforward queries, the agent directly selects the appropriate tool for rapid responses. For complex queries, the agent uses the memory processing module to filter chat history for relevant content, then passes the results to the intelligent task decomposition planner, and finally executes the tasks using personalized recommendation tools. Evaluation on Walmart's real world career recommendation scenarios demonstrates that AdaptJobRec reduces average response latency by up to 53.3% compared to competitive baselines, while significantly improving recommendation accuracy.

cs.IR

Thermal noise induced probability switching in magnetic tunnel junction based on spin-circuit simulation

The probability switching characteristics in spin transfer torque magnetic tunnel junctions (STT-MTJs) are simulated by considering thermal noise using a spin-circuit module. Thermal noise significantly affects the probability switching for pulse durations exceeding 10 ns, while no probability switching properties are observed for pulses shorter than 1 ns due to the precessional switching. For pulse durations between 1 ns and 10 ns, the occurrence of mixed probability and abrupt switching suggests that thermal noise partially influences the switching properties. These results demonstrate the effectiveness of our simulation model in capturing the MTJ properties under the influence of thermal noise. The spin-circuit module used in this study lays the groundwork for future circuit system designs utilizing MTJ devices, such as true random number generators and neural network computing.

physics.app-ph

The Digital Cybersecurity Expert: How Far Have We Come?

The increasing deployment of large language models (LLMs) in the cybersecurity domain underscores the need for effective model selection and evaluation. However, traditional evaluation methods often overlook specific cybersecurity knowledge gaps that contribute to performance limitations. To address this, we develop CSEBenchmark, a fine-grained cybersecurity evaluation framework based on 345 knowledge points expected of cybersecurity experts. Drawing from cognitive science, these points are categorized into factual, conceptual, and procedural types, enabling the design of 11,050 tailored multiple-choice questions. We evaluate 12 popular LLMs on CSEBenchmark and find that even the best-performing model achieves only 85.42% overall accuracy, with particular knowledge gaps in the use of specialized tools and uncommon commands. Different LLMs have unique knowledge gaps. Even large models from the same family may perform poorly on knowledge points where smaller models excel. By identifying and addressing specific knowledge gaps in each LLM, we achieve up to an 84% improvement in correcting previously incorrect predictions across three existing benchmarks for two cybersecurity tasks. Furthermore, our assessment of each LLM's knowledge alignment with specific cybersecurity roles reveals that different models align better with different roles, such as GPT-4o for the Google Senior Intelligence Analyst and Deepseek-V3 for the Amazon Privacy Engineer. These findings underscore the importance of aligning LLM selection with the specific knowledge requirements of different cybersecurity roles for optimal performance.

cs.CR