arXiv ScienceSearch

arXiv subjects

Jian Song

Publications and source records attributed to Jian Song.

At least 19 recordsLinked to original sources

Rough backward stochastic differential equations

We develop an intrinsic well-posedness theory for nonlinear backward stochastic differential equations (BSDEs) driven simultaneously by Brownian motion and a (level-$2$) rough path of finite $p$-variation. Unlike earlier approaches based on smooth approximation or transformation methods, we formulate the equation directly by viewing its solution as a rough semimartingale (in the sense of \cite{friz2023rough}). This framework is particularly suited to BSDEs, whose Brownian martingale component is only implicitly defined and lacks the \emph{a priori} time regularity required by stochastic-sewing-based controlled rough path methods \cite{fhl21,allan2024rough}. We establish comparison, existence, uniqueness, and stability under the natural regularity condition $H\in C_b^γ$, $γ>p$. The main analytical difficulty, namely, the loss of integrability arising from nonlinear composition, is overcome through conditional $p$-variation norms with BMO-type properties. Finally, by randomizing the rough driver as a Brownian rough path, we establish a direct correspondence between rough BSDEs and backward doubly stochastic differential equations (BDSDEs).

math.PR

Scaling limit for the pinning model in correlated Gaussian environment beyond the $L^2$-regime

In this paper, we study the scaling limit of the pinning model in correlated Gaussian environment. The tail probability of the underlying renewal process of the model has a polynomial decay with exponent $α>0$. The covariance of the Gaussian environment $\{ω_n\}_{n\in\mathbb N}$ is given by $\text{Cov}_{\mathbb P}(ω_n,ω_m)\sim |n-m|^{2H-2}$ with $H\in(0,1)$. Assuming $α\in(0,\frac12]$, $H\in(\frac12,1)$ and $α+2H>2$, we show that the partition function of the disordered pinning model, under the appropriate scaling, converges in distribution to the $L^1$-solution of the fractional stochastic heat equation driven by Gaussian noise correlated in time and localized at the origin. In particular, it is known that the solution is not $L^2$-integrable when $α<\frac12$.

math.PR

Finite-Time Singularities of the Kähler--Ricci Flow on Fano Bundles

We study collapsing finite-time singularities of the unnormalized Kähler--Ricci flow on Fano bundles arising in the analytic minimal model program. For a Fano bundle $X^n\rightarrow Y^m$, we prove a maximal splitting theorem by the Kähler-Ricci flow that every tangent space at a fixed limiting point splits globally as \((\C^m,g_{\rm E},J_0)\times(Z',d',J')\). If the fibre has complex dimension one, we prove a global Type-I bound for the full curvature tensor and show that the ambient and intrinsic diameters of every fibre are uniformly comparable to \(\sqrt{T-t}\). Furthermore, every tangent flow at any fixed limiting point is the round shrinking cylinder \(\C^m\times\PP^1\).

math.DG

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

On-policy distillation (OPD), which leverages a pre-trained, specialized teacher model to provide dense supervisory signals, has achieved significant success in Large Language Models (LLMs) and has recently been adapted to flow matching models. However, this paradigm suffers from two major issues: First, training a separate, task-specific teacher for every new objective incurs high computational costs. Second, the discrepancy between teacher and student distributions often leads to compounding errors along the generation trajectory. In this paper, we introduce \textbf{Self-OPD}, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision. At each timestep, Self-OPD branches the deterministic next-state prediction into $K$ stochastic SDE candidates, rolls them out with the ODE sampler, and compares their rewards against a deterministic self-reference baseline to obtain normalized advantages. The velocity field is optimized with an all-branch pull-push objective, where high-advantage branches attract the student and low-advantage branches repel it under direction-aware attenuation and SDE-variance normalization. For multi-objective alignment, Self-OPD fuses normalized scores at the reward level, avoiding direct gradient conflict. Experiments on single and mixed reward benchmarks show that Self-OPD outperforms prior RL and OPD methods without task-specific teachers.

cs.CV

Homojunction-induced thermopower enhancement in polymer films

It has been more than twenty years since conductive polymers began to receive attention as an emerging thermoelectric material. However, the trade-off between electrical conductivity (σ) and thermopower (S) has proven to be a major challenge that has obstructed their use in actual devices. Here we report the discovery that the thermopower of the p- and n-type legs of organic thermogenerators can be substantially enhanced, without significant deterioration of σ, by constructing an in-plane segmented structure consisting of a homojunction with different doping levels on either side. In such segmented layers, the S is abnormally higher than the average value of the constituent parts when applying a forward temperature gradient (heating the heavily doped counterpart), while it is lower upon a reverse temperature gradient. Typically, for a two-stage segmented film of p-type PDPP-Se, an abnormally large S of 210 uV K-1 and σ of 2.5*10^4 S m-1 are obtained, resulting in a large power factor (PF) of 1100 uW m-1 K-2 and a record ZT of 1.36 at room temperature. The enhanced thermopower is attributed to an additional voltage developed at the homojunction under heating as explained by kinetic Monte Carlo simulations. This finding provides a breakthrough approach to the modulation of thermoelectric transport properties of conductive polymers.

cond-mat.mtrl-sci

Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed. Post-event optical imagery, however, may be unavailable because of cloud, smoke, or darkness. The Bright Challenge evaluated all-weather building damage mapping from a submeter-resolution pre-event optical image and a post-event SAR image. Participants were required to detect and delineate each building and assign exactly one of three mutually exclusive damage labels. The challenge extended the globally distributed \textsc{Bright} dataset with instance-level annotations for about 291,000 buildings across 16 disaster events spanning seven disaster types. The final phase was evaluated exclusively on two 2025 events absent from training: a wildfire event in California and a hurricane in Jamaica. A total of 157 participants made 1,289 submissions, and 46 teams entered the final phase. The two winning solutions achieved test mAPs of 0.182 and 0.181, approximately 8.7 times the public baseline of 0.021, but remained far below the best in-domain holdout score of 0.513. Across teams ranked in both phases, performance declined sharply and the rank order changed substantially. The two leading solutions independently favored modality-specific encoding, staged or late optical--SAR fusion, and an optical-dominant separation of building localization from damage recognition. The winning method additionally used scene-aware threshold adjustment and pseudo-label adaptation. These results identify cross-event generalization and stable severity discrimination as the principal remaining challenges. All data, annotations, baseline code, and winning solutions are publicly available at https://github.com/ChenHongruixuan/BRIGHT.

cs.CV

Towards Ultra-High Reliability in Wi-Fi 8: IEEE 802.11bn Core Mechanisms, mmWave Integration, and Performance Verification

As the demand for wireless connectivity expands from high-speed data transmission to high-reliability applications, such as the Industrial Internet of Things and immersive communications, traditional Wi-Fi technologies optimized primarily for peak throughput face new challenges in reliability and latency. Consequently, Wi-Fi 8 aims to achieve ultra-high reliability (UHR), improve communication performance in complex environments, and drive the transition from high-speed connectivity to highly reliable intelligent connectivity. This article provides a comprehensive review of the core mechanisms of Wi-Fi 8 and conducts system-level performance verification. We focus on the key enhancement mechanisms at the physical (PHY) and medium access control (MAC) layers in IEEE 802.11bn, elaborating on their theoretical principles and key application scenarios. Additionally, this paper explores the potential role of integrated millimeter-wave (IMMW) technology as a complementary solution for spectrum expansion in the Wi-Fi 8 era, analyzing its basic architecture and implementation. Finally, system-level simulations are performed to verify the effectiveness of the key technologies in IEEE 802.11bn in achieving their performance targets, while further validating the robust performance of the IMMW scheme under practical hardware impairments.

cs.NI

UXBench: Benchmarking User Experience in AI Assistants

As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present UXBench, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through comprehensive analysis of model behavior and performance gaps, we show that user feedback prediction is a learnable capability, where a reward model trained from in-the-wild feedback signals can achieve well-calibrated accuracy. We further document the systematic biases of LLM-as-a-judge evaluation protocols and compare typical response strategies that directly affect user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing to a user-centric scaling law that shapes the success of AI assistants.

cs.CL

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 15 percentage point gain in end-to-end task success (from 75% in a multimodal baseline system to 90% in an improved configuration identified through a prior ablation study) is sufficient to produce consistent and measurable differences in user perception. The baseline system combines Whisper for speech recognition, Florence-2 for open-vocabulary object detection, LLaMA 3.1 for action extraction, and an interval Type-2 fuzzy logic controller for motion execution. The improved configuration replaces the perception and language modules with Grounding DINO + SAM and Qwen 3.5 9B, respectively, while retaining the same controller. A within-subject user study with 24 participants compared both systems on the same tabletop object-grasping task. After interacting with each configuration, participants rated perceived speed, reliability, and overall competence and fluency on a 7-point Likert scale. Results show that 17 out of 24 participants (70.83%) preferred the improved system (exact binomial test, p = 0.043, h = 0.43), and all three perceptual constructs were rated significantly higher for the improved configuration after Holm correction, with large to very large effect sizes (p < 0.001). These findings confirm that the identified technical improvements are perceptible to users in direct interaction and underscore the importance of complementing benchmark evaluation with user-centred evidence when assessing robotic manipulation pipelines.

cs.RO

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery

LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and an execution environment, they can propose, validate, and iterate scientific solutions, and have produced results that outperform human-designed approaches. As model capabilities continue to improve, we argue that the bottleneck for autonomous scientific discovery is shifting from prescribing agent workflows to designing agent environments: the resources, constraints, and interfaces that shape agent behavior. We frame this as environment engineering: building environments that amplify productive behaviors, such as open-ended exploration, systematic artifact management, and inter-agent collaboration, while suppressing harmful behaviors, such as reward hacking and high-friction human oversight. We present EurekAgent, an environment-engineered agent system for metric-driven autonomous scientific discovery. EurekAgent engineers the environment along four dimensions: permissions engineering for bounded agent execution and isolated evaluation; artifact engineering for filesystem and Git-based collaboration; budget engineering for budget-aware exploration; and human-in-the-loop engineering for easy human supervision and intervention. EurekAgent sets new state-of-the-art results on multiple mathematics, kernel engineering, and machine learning tasks, including new state-of-the-art 26-circle packing results discovered with less than $11 in total API cost. We open-source our code and results, and call for environment engineering as a core research direction for developing reliable autonomous research agents.

cs.AI

Stochastic partial differential equations associated with Feller processes

For the stochastic partial differential equation $\frac{\partial u}{\partial t}=\mathcal L u +u\dot W$ where $\dot W$ is Gaussian noise colored in time and $\mathcal L$ is the infinitesimal generator of a Feller process $X$, we obtain Feynman-Kac type of representations for the Stratonovich and Skorohod solutions as well as for their moments. The regularity of the law and the Hölder continuity of the solutions are also studied.

math.PR

RAM-H1200: A Unified Evaluation and Dataset on Hand Radiographs for Rheumatoid Arthritis

Rheumatoid arthritis (RA) assessment from hand radiographs requires multi-level analysis and modeling of anatomical structures and fine-grained local pathological changes. However, existing public resources do not support such unified multi-level analysis, often lacking full-hand coverage, fine-grained annotations, and consistent integration with clinical scoring systems. In particular, annotations that enable quantitative analysis of bone erosion (BE) remain scarce. RAM-H1200 contains 1,200 hand radiographs collected from six medical centers, with multi-level annotations including (i) whole-hand bone structure instance segmentation, (ii) pixel-level BE masks, (iii) SvdH-defined joint regions of interest, and (iv) joint-level SvdH scores for both BE and joint space narrowing (JSN). It is designed to evaluate whether models can jointly capture anatomical structure, localized erosive pathology, and clinically standardized RA severity from hand radiographs. The proposed BE masks enable, for the first time, quantitative BE analysis beyond coarse categorical grading by providing explicit spatial supervision for lesion extent and morphology. To our knowledge, RAM-H1200 is the first public large-scale benchmark that jointly supports whole-hand bone structure instance segmentation, pixel-level BE delineation, and clinically grounded joint-level SvdH scoring for both BE and JSN. Results across benchmark tasks show that anatomical modeling is substantially more mature than quantitative BE analysis: whole-hand bone segmentation achieves strong performance, whereas BE segmentation remains a major open challenge. By unifying anatomical structure modeling, quantitative lesion analysis, and clinically grounded SvdH scoring, RAM-H1200 provides a single benchmark for comprehensive RA analysis on hand radiographs.

cs.CV

Controlled fields, rough stochastic calculus, and Itô-Wentzell-Alekseev-Gröbner identities

We develop a calculus of space-time controlled fields for rough stochastic systems. This approach provides a unified composition rule for evaluating random fields along rough semimartingales and yields a rough stochastic Itô-Wentzell formula under natural and verifiable regularity assumptions. Our motivation comes from works of Hudde et al. (2024) and, independently, Del Moral and Singh (2022) where the authors established, respectively, Itô-Alekseev-Gröbner, backward Itô-Wentzell, and diffusion interpolation formulas.

math.PR

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we introduce RPC-Bench, a large-scale question-answering benchmark built from review-rebuttal exchanges of high-quality computer science papers, containing 15K human-verified QA pairs. We design a fine-grained taxonomy aligned with the scientific research flow to assess models' ability to understand and answer why, what, and how questions in scholarly contexts. We also define an elaborate LLM-human interaction annotation framework to support large-scale labeling and quality control. Following the LLM-as-a-Judge paradigm, we develop a scalable framework that evaluates models on correctness-completeness and conciseness, with high agreement to human judgment. Experiments reveal that even the strongest models (GPT-5) achieve only 68.2% correctness-completeness, dropping to 37.46% after conciseness adjustment, highlighting substantial gaps in precise academic paper understanding. Our code and data are available at https://rpc-bench.github.io/.

cs.CL

Early Preconfiguration Failure: A Novel Predictor of the Repetitive Subconcussion

Early diagnosis and assessment of repetitive subconcussive (rSC) brain injuries are crucial for early clinical intervention. Conventional methods, largely relying on slow fMRI, fail to capture millisecond-level early cortical dynamics, particularly spatiotemporal features associated with pre-configuration dynamics. This study introduces a novel approach integrating dynamic hierarchical spatial features and cortical early behavioral time-domain sensitivity, utilizing EEG and visual attention tasks. We analyzed cortical early behaviors in 24 healthy controls (HC), 21 rSC patients,and a validation cohort of 25 cTBI patients from public datasets. Results reveal distinct temporal patterns in HC: elevated integration at 0-100 ms, rebound dynamics at 100-200ms, and visual perception integration peaks at 200-600 ms. In contrast, rSC patients exhibited significantly impaired dynamic features, with reduced integration levels indicating a decline in pre-configuration dynamics. Signed center distance (SCD) analysis of separation-integration trajectories showed significantly lower early SCD values in rSC patients compared to HC, while cTBI patients displayed negative SCD values, reflecting irreversible damage. Machine learning classification achieved optimal performance in distinguishing between HC, rSC, and cTBI groups using early cortical features, highlighting the critical role of millisecond-level cortical dynamics in rSC diagnosis.

q-bio.NC

PLAS-Net: Pixel-Level Area Segmentation for UAV-Based Beach Litter Monitoring

Accurate quantification of the physical exposure area of beach litter, rather than simple item counts, is essential for credible ecological risk assessment of marine debris. However, automated UAV-based monitoring predominantly relies on bounding-box detection, which systematically overestimates the planar area of irregular litter objects. To address this geometric limitation, we develop PLAS-Net (Pixel-level Litter Area Segmentor), an instance segmentation framework that extracts pixel-accurate physical footprints of coastal debris. Evaluated on UAV imagery from a monsoon-driven pocket beach in Koh Tao, Thailand, PLAS-Net achieves a mAP_50 of 58.7% with higher precision than eleven baseline models, demonstrating improved mask fidelity under complex coastal conditions. To illustrate how the accuracy of the masking affects the conclusions of environmental analysis, we conducted three downstream demonstrations: (i) power-law fitting of normalized plastic density (NPD) to characterize fragmentation dynamics; (ii) area-weighted ecological risk index (ERI) to map spatial pollution hotspots; and (iii) source composition analysis revealing the abundance-area paradox: fishing gear constitutes a small proportion of the total number of items, but has the largest physical area per unit item. Pixel-level area extraction can provide more valuable information for coastal monitoring compared to methods based solely on counting.

cs.CV

Virtual Polarization Modulation: Enabling CSI-Free DCO-OFDM over Dynamic OWC Channels

In dynamically varying optical wireless communication (OWC) links, conventional quadrature amplitude modulation (QAM) in optical orthogonal frequency-division multiplexing (OFDM) requires frequent channel estimation and equalization, incurring pilot overhead and processing latency. This paper proposes a virtual polarization modulation (VPM)-based direct-current-biased optical OFDM (DCO-OFDM) scheme that maps each data symbol onto the three-dimensional Stokes space and places its corresponding Jones vector across two adjacent OFDM subcarriers. Using a rotation-based analytical framework, closed-form symbol error rate (SER) expressions are derived for arbitrary spherical constellations, along with upper and lower bounds and high signal-to-noise ratio (SNR) approximations. The framework is further extended to practical OWC scenarios with frequency-selective channels and atmospheric turbulence. Monte Carlo (MC) simulations validate the theoretical results. The results show that under practical OWC impairments, VPM outperforms QAM with least-squares (LS) channel estimation and minimum mean square error (MMSE) equalization. At a target SER of $10^{-5}$, 16-VPM achieves SNR gains of approximately 7.5 dB and 4 dB over equalized 16-QAM and 8-QAM, respectively, in frequency-selective channels, and a 6 dB advantage over equalized 16-QAM under atmospheric turbulence. By eliminating the need for channel state information, the proposed VPM-based DCO-OFDM provides a robust and low-latency solution for dynamic OWC links.

physics.optics