arXiv ScienceSearch

arXiv subjects

Kexin Zhang

Publications and source records attributed to Kexin Zhang.

At least 19 recordsLinked to original sources

Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection

We present our solution to the LUMPI track of the UCF UrbanTwin Sim2Real LiDAR Challenge at the 6th DriveX Workshop, ECCV 2026. The detector must be trained only on synthetic data and is evaluated on 50 held-out real LiDAR frames; a separate 50-frame synthetic submission is evaluated for point-cloud realism. Our method addresses the Sim2Real gap at three levels. First, we align synthetic scans to the 50k-point test density and build a 30k-record training pool using UT-LUMPI geometry, RangeLDM-based sampling diversification, rare-class copy-paste, and pedestrian-oriented augmentation. Second, complementary DSVT detectors and Car/Bus PointPillars specialists are trained under the same synthetic-only constraint. Third, predictions are integrated by class-aware routing, asymmetric agreement fusion, constrained residual-recall supplementation, class-coverage auditing, and selective box-size calibration. The realism branch is optimized independently with radial-density matching, weak affine calibration, and calibrated set mixing. The final submission obtains a Combined Score of 0.4692, a Detection Score of 0.1797, a Realism Score of 0.9035, and 3D mAP@0.5 of 0.1258.

cs.CV

Solution for UCF UrbanTwin V2X-Real Track: Sim-to-Real Urban LiDAR 3D Object Detection

Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method organizes digital-twin scans, diffusion-redrawn scans, density-stabilized scans, and pedestrian morphology-aligned samples into a unified training pool with complementary roles. Within a common DSVT detection formulation, source-specialized expert branches preserve those roles while optimizing for the same detection objective. At inference, a predefined class-aware fusion pathway integrates geometry-stable and calibration-aware branches for vehicles, sampling-complementary branches for trucks, and morphology-consistent evidence for pedestrians. A label-free point-cloud center blend then refines geometric localization. On the UrbanTwin V2X-Real hidden test set, the unified system achieves a combined score of 0.7421, with 3D mAP@0.5 of 0.4518 and a realism score of 0.8871. The results indicate that a stable, interpretable collaboration among data sources is more valuable than unconstrained aggregation of model outputs.

cs.CV

Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation

This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.

cs.CV

ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing

Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transformer backbones often brings substantial computational cost without consistently improving mutation-sensitive prediction. We introduce ProtLingo, an efficient PLM framework that augments a pretrained single-sequence backbone with conditional local memory and sparse expert routing. ProtLingo maps contextual residue representations into route-specific discrete codes, composes centered local windows into latent $N$-gram addresses, and retrieves reusable residual signals associated with recurring local sequence contexts. In parallel, selected feed-forward blocks are upcycled into sparse Mixture-of-Experts layers with shared and routed experts, enabling residue-dependent computation while activating only a subset of parameters. Experiments on protein fitness prediction, FLIP benchmarks, and supervised contact prediction show that ProtLingo achieves competitive performance with a 150M-scale backbone, including strong parameter efficiency on mutation-effect prediction and preserved long-range structural representations.

cs.AI

CompEvo: Competition-Induced Evolution for Multi-Agent in News-Driven Time Series Forecasting

News-driven time series forecasting uses evolving textual events together with historical observations to predict future values, supporting applications such as market risk monitoring and resource scheduling. In multi-agent settings, two challenges still remain. The first is degeneration of thought, where agents converge to similar evidence-seeking behaviors. The second is insufficient theoretical grounding, where strategy updates are often heuristic and lack a principled formulation. To address the above challenges, we propose CompEvo, a competition-induced evolution framework for multi-agent news-driven time series forecasting. For theoretical grounding, we introduce an evolutionary game formulation to guarantee equilibrium existence and optimization convergence. Building on this formulation, we construct a trainable multi-agent evolution framework that integrates strategy execution, fitness-based differentiable selection, and competition-induced strategy evolution. CompEvo enables heterogeneous agents to explore diverse news evidence, converts forecasting feedback into differentiable influence weights, and evolves agent strategies under competitive pressure to preserve effective logic while maintaining diversity. Experiments on four real-world datasets show that CompEvo reduces RMSE by 27.3% and MAPE by 26.2% on average over strong baselines. Further analysis indicates that CompEvo successfully maintains diverse and specialized agent behaviors.

cs.NE

Beyond Harassment: Exploring the Harm Experienced by People with Disabilities in Social Virtual Reality

People with disabilities (PWD) are increasingly engaging in social virtual reality (VR) platforms, where immersive and embodied interactions can intensify negative experiences. While prior work has examined harassment in VR, little is known about the harms experienced by PWD and the perceived severity associated with different harassment and disability types. Unlike harassment, which represents behaviors, harm is more critical to designing effective protections, as it reflects the consequences and impact; the realism of VR and the vulnerability resulting from disability identity can further amplify such impact. To characterize and model harms for PWD, we conducted a literature review, followed by an online survey with 67 PWD to understand participants' harassment experiences and resulting harms in social VR. We identified 19 types of harm in 5 categories, and reported the severity perception of each type of harm. Finally, we analyzed our results from the critical disability theory perspective, summarized the uniqueness of harm in social VR, and discussed design implications for specialized safety mechanisms that mitigate harm for PWD.

cs.HC

Asymptotics of the principal eigenvalue of an elliptic operator on closed and orientable Riemannian manifolds: small diffusion

This paper is concerned with the asymptotic behavior of the principal eigenvalue $\lambda(D)$ of the elliptic eigenvalue problem \[ -D\Delta_{M}u - a\langle \nabla_M f, \nabla_M u\rangle_g + c u = \lambda(D)u, \] posed on a closed orientable Riemannian manifold $(M,g)$, in the small-diffusion limit $D \to 0^+$. Under the assumption that $f$ is a Morse function on $M$, we establish that the limiting value $\lim_{D\to 0}\lambda(D)$ is completely characterized by the critical points of $f$ and the associated Riemannian Hessian, specifically through the values of $c$ and the Riemannian Hessian eigenvalues at those points.

math.AP

NavSight in the Wild: Understanding Real-World Use of a Mobile Augmented Reality Application for People with Low Vision in Outdoor Navigation

The ability to navigate outdoors safely and independently is crucial yet challenging for people with low vision (PLV). While various augmented reality (AR) systems for low vision have been designed and evaluated in ideal lab environments, no research has investigated their real-world feasibility and challenges. We present NavSight, a mobile AR application that assists PLV in outdoor navigation by recognizing important outdoor objects (e.g., curb, vehicle) and rendering real-time visual augmentations. Through a seven-day diary study with 12 PLV in real-world settings, we characterize the impact of NavSight on scene perception, users' configuration strategies on what objects to augment and how to augment them across scenarios, how users made sense of and responded to recognition errors, and the social acceptability of using NavSight in public. We further identify environmental factors affecting recognition, such as weather conditions, lighting and shadows, and nonstandard road markings and textures, as well as usability issues in daily use. We discuss these real-world challenges and derive design implications for future AI-powered assistive AR systems for outdoor use.

cs.HC

Safety vs. Social Image: Co-Designing Protection Mechanisms Against Ableist Harassment with People with Disabilities in Social Virtual Reality

People with disabilities (PWD) increasingly use avatars to express disability identities in social virtual reality (VR), but greater visibility also invites targeted harassment. Existing safety features are often insufficient, overlooking PWD's experiences and needs. To address this gap, we co-designed protection mechanisms with 11 PWD to reveal their values and needs. Our research employed a social lens to interpret harassment behaviors and protection mechanisms. Inspired by Hall's Proxemics Theory that interpersonal distances indicate social intent and boundaries, we divided social VR spaces into four proxemic zones (Intimate, Personal, Social, and Public) and used them to structure our protection mechanism co-design. We also provided different protection mechanism probes (Inform, Educate, Consent, and Combat) to elicit participant preferences. Our study highlighted the role of social proximity in shaping PWD's harassment perception and protection preferences and revealing PWD's unique social values and needs (e.g., managing harassment with optimism and resilience, prioritizing social image over safety). We proposed design recommendations for protection mechanisms that protect PWD while maintaining their desired social images.

cs.HC

Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge

The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J &F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.

cs.CV

Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

Multiple Sequence Alignments (MSAs) provide protein language models with explicit evolutionary context, but their large depth makes subsampling unavoidable under limited token budgets. Existing strategies, including random selection, identity-based filtering, and diversity-driven sampling, are effective heuristics, yet provide limited control over the evolutionary signals retained in the subset. In this work, we recast MSA subsampling as an explicit optimization problem, where key evolutionary measures, including query identity and diversity, are treated as controllable objectives. Building on this view, we introduce AP-REASONER, an Affinity-Propagation-based factor-graph approach. With evolution-aware unary factors, exemplar-consistency factors, and two control knobs, AP-REASONER performs factor-graph reasoning through message passing to infer a fixed-budget MSA subset. Experiments on long-range contact prediction and conformational ensemble prediction show that AP-REASONER outperforms baseline subsamplers on structure-sensitive downstream tasks and enables controllable recovery of alternative protein conformations. These results highlight the value of modeling MSA subsampling as a controllable optimization problem, where factor-graph reasoning offers an effective alternative to heuristic selection.

cs.LG

What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision

People with low vision (PLV) struggle to perceive complex scenes like busy kitchens and crowded streets, which contain many objects, visual clutter, and dynamic elements. Prior AR systems for low vision either enhance low-level visual features or augment task-relevant objects for single tasks in simple settings, leaving multi-object augmentation in complex scenes underexplored. Informed by a formative study characterizing important objects and their perceived importance for PLV, we built SceneGlance, a wearable AR system that recognizes important objects and visually distinguishes them by importance level. Through a controlled lab study with 12 PLV in a mock-up kitchen scene and a free-form think-aloud study with 13 PLV navigating an outdoor route, we found that AR distinction on object importance shifted PLV's attention toward objects of higher importance, and supported perception strategies such as building mental snapshots from the augmentation distribution and hierarchical scanning by importance. However, this attention shift came with a tradeoff of reduced overall scene recall. The studies also surfaced challenges posed by AR augmentations in complex scenes, such as adjacent augmentations blending or interfering with each other, yielding design implications for more practical AR vision enhancement systems in the complex real world.

cs.HC

Intrinsically low thermal conductivity of stoichiometric lithium niobate:Experimental measurement and microscopic origin

With the rapid development of integrated electro-optic and nonlinear optical devices based on lithium niobate (LiNbO$_3$, LN), thermal management is becoming a critical area of focus. However, experimental measurement of thermal transport in stoichiometric LiNbO$_3$ (sLN) remains scarce, and the intrinsic microscopic mechanisms remain to be established. Here, we combine the laser pump-probe technique of frequency-domain thermoreflectance (FDTR) with state-of-the-art machine-learned atomistic simulations to comprehensively investigate thermal transport in sLN. The measured and simulated room-temperature thermal conductivity ($\kappa$) values of sLN agree well, which are orders-of-magnitude lower than that of many classic and emerging semiconductors such as silicon. Furthermore, the temperature-dependent $\kappa$ exhibits a $T^{-\alpha}$ scaling with $\alpha$ near unity, suggesting that thermal transport is dominated by intrinsic phonon-phonon scattering. By comparing sLN with cubic boron arsenide (cBAs) which serves as an ultrahigh-$\kappa$ benchmark, we reveal that harmonic properties are not responsible for the low $\kappa$ of sLN, which feature phonon heat capacity and group velocities that are either higher than or comparable to those in cBAs. Instead, the low $\kappa$ originates from substantially stronger anharmonicity and larger scattering phase space. These two factors collectively suppress phonon lifetimes by 1-2 orders of magnitude, leading to a maximum phonon mean free path of approximately 140 nm. As a result, notable size effects emerge in thin-film sLN below 1 $\mu$m, with $\kappa$ dropping to half the bulk value at 10 nm. Altogether, our findings establish a fundamental understanding of thermal transport in sLN and provide atomistic insights for thermal management in advanced lithium niobate technologies.

cond-mat.mtrl-sci

On Privacy-Preserving Image Transmission in Low-Altitude Networks: A Swin Transformer-Based Framework with Federated Learning

The rapid development of low-altitude economy has driven the proliferation of Unmanned Aerial Vehicle (UAV) applications, including logistics, inspection, and emergency response. However, transmitting high-volume image data from UAVs to ground stations faces significant challenges due to limited bandwidth and stringent privacy requirements. To address these issues, a Semantic Communication (SC) framework based on Federated Learning (FL) is proposed for efficient and privacy-preserving image transmission. A Swin Transformer-based Semantic Communication (STSC) architecture is designed to extract multi-scale semantic features under constrained bandwidth conditions. Dedicated communication and computing nodes are deployed on UAVs to enhance real-time coverage and flexibility. Meanwhile, a FL mechanism enables global model training across distributed devices without sharing raw data, thus preserving user privacy. Simulation experiments conducted on the CIFAR-10 dataset demonstrate that the proposed STSC framework achieves at least 5.7 dB improvement in Peak Signal-to-Noise Ratio (PSNR) compared to DeepJSCC baselines, while also showing superior convergence and generalization performance. The framework effectively integrates UAV-assisted deployment with SC and privacy protection, offering a practical solution for bandwidth-constrained image transmission in low-altitude networks.

eess.IV

A Variable-Spot-Size and Multi-Frequency Square-Pulsed Source (SPS) Approach for Comprehensive Characterization of Anisotropic Thermal Transport Properties in Multilayered Thin Films

Multilayered thin-film structures are frequently encountered in industrial applications, where accurate thermal property characterization is essential for performance optimization. These films, typically ranging from nanometers to micrometers in thickness, often exhibit anisotropic thermal conductivity and non-bulk heat capacity, which are challenging to measure. In this study, we introduce a variable-spot-size and multi-frequency square-pulsed source (SPS) method for the simultaneous determination of anisotropic thermal conductivities, heat capacities, and interfacial thermal conductance in multilayered systems. By leveraging a broad modulation frequency range (1 Hz to 10 MHz) and tunable laser spot sizes, the SPS method enhances sensitivity to different thermal parameters across layers. We validate this approach on a silicon-on-insulator (SOI) sample comprising a 1.59 um Si layer, 1.03 um SiO2 layer, and a silicon substrate with a 122 nm aluminum (Al) transducer. The SPS method successfully extracts seven key thermal parameters, including the in-plane and cross-plane thermal conductivities and heat capacity of the Si film, the thermal conductivity and heat capacity of the SiO2 layer, the thermal conductivity of the substrate, and the interfacial thermal conductance between Al and Si. Temperature-dependent measurements from 80 to 500 K showed excellent agreement with literature values and first-principles predictions, confirming the method's accuracy and reliability. These results demonstrate the SPS method as a powerful tool for comprehensive thermal characterization of complex multilayered structures, with implications for both fundamental research and practical applications.

physics.app-ph

Depth-Resolved Thermal Conductivity of HFCVD Diamond Films via Square-Pulsed Thermometry

The integration of high-thermal-conductivity diamond films onto silicon carbide (SiC) substrates offers a promising pathway for thermal management in high-power electronic devices. Here, we investigate the depth-dependent thermal conductivity of a ~5 {\mu}m-thick diamond film grown on SiC by hot-filament chemical vapor deposition (HFCVD) using square-pulsed source (SPS) thermometry. Electron backscatter diffraction (EBSD) and transmission electron microscopy (TEM) reveal pronounced grain coarsening from the nucleation interface to the film surface. By combining frequency-dependent thermal penetration with a depth-resolved thermal transport model, we quantitatively reconstruct the thermal conductivity profile. The thermal conductivity increases sharply from ~60 W m^(-1) K^(-1) near the nucleation region to ~200 W m^(-1) K^(-1) at the surface, directly reflecting the underlying microstructural evolution. These results provide a physically grounded understanding of graded heat transport in HFCVD diamond and offer practical guidance for engineering diamond-based thermal management layers for next-generation power devices.

cond-mat.mtrl-sci

Towards Training-Free Scene Text Editing

Scene text editing seeks to modify textual content in natural images while maintaining visual realism and semantic consistency. Existing methods often require task-specific training or paired data, limiting their scalability and adaptability. In this paper, we propose TextFlow, a training-free scene text editing framework that integrates the strengths of Attention Boost (AttnBoost) and Flow Manifold Steering (FMS) to enable flexible, high-fidelity text manipulation without additional training. Specifically, FMS preserves the structural and style consistency by modeling the visual flow of characters and background regions, while AttnBoost enhances the rendering of textual content through attention-based guidance. By jointly leveraging these complementary modules, our approach performs end-to-end text editing through semantic alignment and spatial refinement in a plug-and-play manner. Extensive experiments demonstrate that our framework achieves visual quality and text accuracy comparable to or superior to those of training-based counterparts, generalizing well across diverse scenes and languages. This study advances scene text editing toward a more efficient, generalizable, and training-free paradigm. Code is available at https://github.com/lyb18758/TextFlow

cs.CV

Asymptotics of the principal eigenvalue of an elliptic operator on closed and orientable Riemannian manifolds

This paper investigates the asymptotic behavior of the principal eigenvalue $\lambda(s)$, as $s\to+\infty$, for the following elliptic eigenvalue problem \begin{equation*}\label{E} -\Delta_{M}u-s\langle \nabla_M f, \nabla_M u\rangle_g +c u=\lambda(s)u, \end{equation*} defined on an orientable and closed Riemannian manifold $(M,g)$. Assuming $f$ is a Morse function defined on $M$, we find that the limit $\lim\limits_{s\to+\infty} \lambda(s)$ is determined by the minimum value of the function $c$ over the set of the maximum points of $f$, a result that is independent of the curvature of manifold.

math.AP