arXiv ScienceSearch

arXiv subjects

Yunshan Wang

Publications and source records attributed to Yunshan Wang.

4 recordsLinked to original sources

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning

Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-Language Action (VLA) models. However, whether this verbalized Chain-of-Thought truthfully reflects the policy's underlying decision process remains poorly understood. We distinguish between functional reasoning, in which reasoning improves task performance, and faithful reasoning, in which reasoning truly reflects the policy's internal decision process. We argue that SoTA alignment strategies offer a necessary but insufficient notion of faithfulness, admitting reasoning whose intermediate steps can mask the causal links in action prediction through confounding factors (e.g., reasoning that is ungrounded in the environment and internally disconnected or inconsistent), restricting policy generalization. We study this gap through a human evaluation of a SoTA reasoning model for autonomous driving, revealing an inconsistent coupling between reasoning quality and downstream trajectory improvement. We then operationalize a behavioral surrogate for embodied faithfulness through a learned critic, Pinocchio, scoring observation grounding and stepwise coherence, and use this critic as a dense reward signal in post-training an embodied policy with reinforcement learning. Across withheld driving benchmarks, our post-trained planner improves faithfulness by 4% and 18% over SoTA alignment and trajectory error post-training baselines, respectively, while maintaining competitive downstream task performance. Finally, on a synthetic out-of-distribution test set, post-training for faithfulness improves policy responsiveness to rare counterfactual scenarios by 1.6x that of a SoTA policy, suggesting that faithful reasoning traces contribute to more robust, generalizable, and interpretable embodied intelligence. Project page: https://mjf-su.github.io/pinocchio/

cs.RO

DeferredSeg:A Multi-Expert Deferral Framework for Medical Image Segmentation

Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit overconfidence or underconfidence, leading to unreliable confidence scores for segmentation masks, especially in ambiguous regions. This undermines the trustworthiness required for clinical deployment. Motivated by the learning-to-defer (L2D) paradigm, we introduce DeferredSeg, a deferral-aware segmentation framework, i.e., a Human--AI collaboration system that determines whether to defer predictions to human experts in specific regions. DeferredSeg extends the base segmentor with an aggregated deferral predictor and additional routing channels that dynamically route each pixel to either the base segmentor or a human expert. To train this routing efficiently, we introduce a pixel-wise surrogate collaboration loss that supervises deferral decisions. In addition, to preserve spatial coherence within deferral regions, we propose a spatial-coherence loss that enforces smooth deferral masks, thereby enhancing reliability. Beyond single-expert deferral, we further extend the framework to a multi-expert setting by introducing multiple discrepancy experts for collaborative decision-making. To prevent overloading or underutilizing individual experts, we further design a load-balancing penalty that evenly distributes workload across expert branches. We evaluate DeferredSeg on three challenging medical datasets using MedSAM and CENet as the base segmentor for fair comparison. Experimental results show that DeferredSeg consistently outperforms the baseline, demonstrating its effectiveness for trustworthy dense medical segmentation. Moreover, the proposed framework is model-agnostic and can be readily applied to other segmentation architectures.

cs.CV

A wafer-scale ultrasensitive programmable chiroptical sensor

Chiroptical enantioselective sensing is gaining traction across various applications. However, intrinsic molecular chiroptical responses are weak, and existing amplification approaches add synthesis, manufacturing, or operational complexity that limits sensitivity, scalability, and dynamic control. Here, we present a fundamentally new sensing paradigm merging adsorption-driven chirality induction with wafer-scale optical transduction in a programmable heterostructure containing twisted aligned carbon nanotubes (CNTs) and phase change materials (PCMs). Chiral molecules adsorb onto CNTs to form chiroptically active composites that are macroscopically assembled by alignment and rotational stacking, yielding large ultraviolet circular dichroism (CD). We resolve molecule concentration and handedness in a single device without lithography, hotspot delivery, or differential protocols, achieving sub-$\mu$M sensitivity for CD-silent glucose and chiral amino acids enabled by $>10^5\,\mathrm{M^{-1}}$ adsorption constants. We validate adsorption using molecular dynamics simulations, reproduce experimental results using chiral transfer matrix simulations, and realize sensor programmability by tuning the PCM layer. This platform enables cost-effective in-situ enantiomer monitoring in aqueous environments.

physics.optics

A novel approach for classifying Monoamine Neurotransmitters by applying Machine Learning on UV plasmonic-engineered Auto Fluorescence Time Decay Series (AFTDS)

This study introduces a hybrid approach integrating advanced plasmonic nanomaterials and machine learning (ML) for high-precision biomolecule detection. We leverage aluminum concave nanocubes (AlCNCs) as an innovative plasmonic substrate to enhance the native fluorescence of neurotransmitters, including dopamine (DA), norepinephrine (NE), and 3,4-Dihydroxyphenylacetic acid (DOPAC). AlCNCs amplify weak fluorescence signals, enabling probe-free, label-free detection and differentiation of these molecules with great sensitivity and specificity. To further improve classification accuracy, we employ ML algorithms, with Long Short-Term Memory (LSTM) networks playing a central role in analyzing time-dependent fluorescence data. Comparative evaluations with k-Nearest Neighbors (KNN) and Random Forest (RF) demonstrate the superior performance of LSTM in distinguishing neurotransmitters. The results reveal that AlCNC substrates provide up to a 12-fold enhancement in fluorescence intensity for DA, 9-fold for NE, and 7-fold for DOPAC compared to silicon substrates. At the same time, ML algorithms achieve classification accuracy exceeding 89%. This interdisciplinary methodology bridges the gap between nanotechnology and ML, showcasing the synergistic potential of AlCNC-enhanced native fluorescence and ML in biosensing. The framework paves the way for probe-free, label-free biomolecule profiling, offering transformative implications for biomedical diagnostics and neuroscience research.

q-bio.BM