arXiv ScienceSearch

arXiv subjects

Tong Wang

Publications and source records attributed to Tong Wang.

At least 19 recordsLinked to original sources

Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation

Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at https://github.com/wangtong627/LAMP.

cs.CV

Exact solutions of nonreciprocal Su-Schrieffer-Heeger model with domain walls

We investigate domain-wall physics in a non-Hermitian Su-Schrieffer-Heeger model featuring the non-Hermitian skin effect and its higher-dimensional extensions, focusing on analytical eigenstate solutions under open boundary conditions. By gluing two Su-Schrieffer-Heeger chains with inverted coupling ratios together, we identify two distinct interface geometries, and derive closed-form quantization conditions for the complex wave number and wave functions using symmetry-based ansatzes. Besides conventional skin modes and topological zero modes, the domain walls generate additional localized states at finite energy that detach from the bulk open-boundary-condition continuum. We validate the analytical results by exact diagonalization, and further generalize the interface construction to a two-dimensional Lieb lattice, where competing nonreciprocities produce tunable funneling regimes towards codimension-one and codimension-two interfaces.

cond-mat.mes-hall

Integrated yoctosecond-precision timing detector

Precise timing detection is essential for exploring ultrafast phenomena in fields ranging from free-electron lasers to ultra-high-power laser facilities. However, achieving simultaneous high resolution and large dynamic range remains a fundamental challenge, and state of the art systems are often constrained by their physical size, power requirements, and limited scalability. Here we introduce an integrated dual electro optic sub cycle timing detector (DEST) that overcomes these limitations. In a proof of principle measurement, the device resolves timing jitter as small as 11 yoctoseconds (ys, $10^{-24}$ s) at 1 MHz--equivalent to the transit time of light across two protons--while maintaining an unambiguous measurement range of 6.15 ps and a dynamic range exceeding 155 dB. The core detection unit is miniaturized to chip scale dimensions of 18 mm x 2 mm x 1 mm on a thin film lithium niobate platform, ensuring inherent stability and immunity to environmental disturbances. Moreover, the architecture naturally lends itself to massive parallelization through array integration, with the potential to push timing precision to the sub 10 rontosecond (rs, $10^{-27}$ s) level within a 1 $m^2$ footprint. This combination of extreme sensitivity, wide dynamic range, compact size, and scalability opens new avenues for detecting previously inaccessible weak signals, including those from gravitational waves, quantum vacuum fluctuations, and beyond.

physics.optics

Probing Triton's Space Environment and Internal Structure: An Integrated Detection-and-Interpretation Framework

Triton, Neptune's largest moon, is a prime ocean-world target. Constraining ocean thickness, composition, and conductivity is essential for habitability assessment, but magnetic induction alone cannot resolve the thickness-conductivity degeneracy, and magnetic perturbations from Triton's space currents can obscure the internal induction signal. We present an integrated detection-and-interpretation concept linking four physically consistent calculations. Using `PlanetProfile', we construct a common radial interior structure (temperature, density, conductivity, seismic-wave speed). We then use `MoonMag' to compute the degree-one magnetic-induction response from that conductivity profile at the synodic, rotational, and orbital periods. We perform a multi-fluid `SWMF' simulation with the induced dipole as the inner-boundary condition and develop a Coulomb-gauge Poisson reconstruction to isolate space-current magnetic fields. Finally, we develop the `TritonSeis' workflow, three-dimensional seismic forward modeling plus hierarchical travel-time inversion, to constrain the ice-ocean and ocean-rock interface depths. We find that induction is substantially more sensitive to ocean conductivity than to layer thickness, and that space-current fields are comparable in amplitude to the internal induction signal. A five-station synthetic recovery test resolves both interfaces to first order, with errors of +8.4% for the ice shell and -12.5% for the ocean. Under a conservative noise assumption, the minimum detectable magnitudes are approximately 3.8-4.6 at epicentral distances of 100-1000 km. The Poisson reconstruction and end-to-end seismic recovery are, to our knowledge, the first such quantitative demonstrations for Triton. Coordinated magnetic, plasma, and seismic measurements are complementary and can break the conductivity-thickness degeneracy, providing a framework for future Triton exploration.

astro-ph.EP

From Generation to Simulation: How Far Are World Models from Being True Simulators?

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments. Yet the remaining distance from generation to simulation lacks a systematic assessment. We present a capability-based study using an external yardstick: eight capabilities of a traditional simulator, namely asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, and evaluation metrics. We trace three main technical routes--latent dynamics, video generation, and joint-embedding prediction--and map exactly 200 representative works published from 2018 to June 2026 onto these capabilities. Our analysis shows that world models have achieved functional substitution in interaction and controllability for specific scenarios, but remain short of traditional simulators in formal guarantees of physical laws, structured state feedback, and reproducible long-horizon evolution. State feedback is the most neglected cross-route shortcoming: only 6 of 163 implementation papers expose a runtime interface for querying entity states or physical parameters. We identify six research directions: formalized physics, a unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization. Project page: https://github.com/AtongWang/world-model-simulators

cs.AI

OVIBench: Benchmarking Online Video Question Answering under Interruption

Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting. OVIBench categorizes interruptions into three types: Cancellation, False Trigger, Correction and supports both open-ended and multiple-choice evaluations. To enable large-scale and reproducible testing, we develop an offline simulation protocol that reproduces interruption during generation under a unified temporal setup, together with a multi-dimensional metric suite for assessing interruption understanding and response generation. Experiments demonstrate that OVIBench effectively distinguishes models' interruption-handling abilities, especially in following correction requests. Finally, we construct a train set OVI-Train for interruption-aware fine-tuning. Models fine-tuned on this dataset achieve significant gains on OVIBench, validating the effectiveness of our benchmark and data design. OVIBench, OVI-Train, and the evaluation code will be released.

cs.CV

Observation of Subharmonic Charge-Density-Wave Correlations in La-Based Cuprates

Pair-density-wave (PDW) correlations have been proposed as an important ingredient in the complex phase diagram of high-$T_{\rm c}$ cuprates, yet bulk-sensitive experimental signatures remain scarce. Here we report resonant x-ray scattering measurements revealing a subharmonic charge-density-wave (CDW) scattering response in Sr-doped $1/8$-LBCO. The subharmonic response appears at approximately half the primary CDW ordering wave vector and emerges within a physically relevant temperature regime associated with the development of in-plane superconducting correlations. Comparable subharmonic behavior is also observed in a chemically distinct La-based cuprate, LSCO, within the same stripe-ordered, layer-decoupled regime. Together, these observations identify a bulk-sensitive scattering signature that is consistent with PDW correlations in La-based superconducting cuprates.

cond-mat.supr-con

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled training sets, and transfer poorly across analytical objectives and experimental datasets. Here we introduce UltraIR, a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. UltraIR is pretrained on approximately 60 million simulated IR spectra using spectral reconstruction, molecular fingerprint similarity alignment, and functional-group prediction, then adapted to downstream objectives with task-specific labels or targets. Across functional-group prediction, molecular structure elucidation, physicochemical property prediction, mixture-component identification and quantification, bacterial classification, medicinal-herb geographic origin traceability and constituent quantification, microplastics classification, and soil property prediction, UltraIR outperforms conventional machine-learning and task-specific deep-learning baselines. It performs strongly with limited labeled experimental spectra and in zero-shot inference for the same analytical task across Fourier-transform infrared spectrometers and laboratories, providing a route to adaptable, data-efficient chemical sensing from complex real-world samples.

cs.LG

The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior

As large language models continue to improve, agentic systems are becoming increasingly important, and tools are a key design dimension because they determine how agents access information and take action in their environments. Prior work on agent tooling has primarily focused on expanding what agents can do, but has paid less systematic attention to how those capabilities are organized and exposed to the model. We refer to this latter design dimension as tool architecture. We study tool architecture in coding agents through controlled experiments on repository-level issue fixing, comparing six tool architectures that hold the underlying information and actions similar while varying how they are organized and exposed to the model, across three actors and a total of 11,700 trajectories. Our experiments show that, even when tools provide similar capabilities, tool architecture changes agent behavior: Compared to a basic architecture where the agent has only the bash tool, more structured low-level interfaces improve consistency across repeated attempts by up to 4.7 $\times$; natural-language search broadens repository exploration and increases access to relevant files by more than 11%; and Python CodeAct-style interfaces achieve similar task performance with 41.6% fewer steps and 56.3% lower token usage. By contrast, lightweight text-based cognitive-scaffolding tools, such as tools that let the agent record intermediate reasoning, have limited effect on actor behavior.

cs.SE

Measurement-induced generation of Schrödinger cat states in cavity QED

Schrödinger cat states, representing coherent superpositions of macroscopically distinguishable states, are indispensable nonclassical resources for continuous-variable quantum information processing. Existing generation protocols typically rely on strong nonlinear interactions, complicated control techniques, or engineered dissipation, posing challenges for experimental implementation. Here, we propose a simple measurement-based protocol for generating Schrödinger cat states in a cavity-QED system by combining coherent driving, dispersive atom--cavity interactions, and atomic postselection. The atom--cavity interaction establishes coherent correlations between the atomic and photonic degrees of freedom, while the subsequent atomic postselection projects the cavity field onto a non-Gaussian superposition state with pronounced Wigner negativity. Numerical simulations based on the Lindblad master equation show that the generated Schrödinger cat states remain robust against moderate cavity dissipation. Our results demonstrate that conditional atomic measurements provide an effective and experimentally accessible approach for preparing nonclassical cavity states without relying on strong optical nonlinearities or engineered dissipation.

quant-ph

Simulational and theoretical studies of the Anderson transition in the chiral symmetry classes with weak topology

Combining lattice model simulations with a field theory study of effective theories, we investigate the nature of the Anderson transition in chiral symmetry classes with one-dimensional (1D) weak topology. In the simulation study, we extend previous transfer matrix analyses to the chiral symplectic class, and study numerical Lyapunov exponents via a finite-size scaling (FSS) analysis that assumes spatially isotropic scaling. The analysis shows that, as in the other two chiral symmetry classes, the weak topology induces an intermediate quasi-localized (QL) phase between metal and Anderson insulator phases. In this QL phase, the localization length of wave functions diverges exclusively along the direction of the 1D weak topology. In the field theory study, we revisit and extend our previous two-dimensional (2D) renormalization group (RG) analysis to all three chiral classes, now newly incorporating a one-loop renormalization of the weak topological term in the analysis. The revised analysis reveals that a quasi-localized strong-coupling fixed point previously reported in the chiral unitary class is unstable under this new inclusion; instead, the strong-coupling phase is entirely governed by a stable fixed point with conventional localized character. Nevertheless, in the chiral unitary and chiral symplectic classes, the RG analysis still yields the hallmark of the 1D weak topology through the spatially anisotropic scaling of the Anderson transition criticality. These theoretical findings suggest that the quasi-localized phase observed numerically in 2D models may be an artifact of the spatially isotropic scaling assumption in the FSS analysis. A conclusive numerical identification of this phase therefore requires a finite-size scaling approach that accommodates generic (anisotropic) spatial scaling.

cond-mat.dis-nn

Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval

Composed Image Retrieval (CIR) is the task of retrieving a target image from a database using a multimodal query, which consists of a reference image and a modification text. The text specifies how to alter the reference image to form a ''mental image'', based on which CIR should find the target image in the database. The fundamental challenge of CIR is that this ''mental image'' is not physically available and is only implicitly defined by the query. The contemporary literature pursues zero-shot methods and uses a Large Multimodal Model (LMM) to generate a textual description for a given multimodal query, and then employs a Vision-Language Model (VLM) for textual-visual matching to search for the target image. In contrast, we address CIR from first principles by directly generating the ''mental image'' for more accurate matching. Particularly, we prompt an LMM to generate a ''mental image'' for a given multimodal query and propose to use this ''mental image'' to search for the target image. As the ''mental image'' has a synthetic-to-real domain gap with real images, we also generate a synthetic counterpart for each real image in the database to facilitate matching. In this sense, our method uses LMM to construct a ``paracosm'', where it matches the multimodal query and database images. Hence, we call this method Paracosm. Notably, Paracosm is a training-free zero-shot CIR method. It significantly outperforms existing zero-shot methods on challenging benchmarks, achieving state-of-the-art performance for zero-shot CIR.

cs.CV

Deep Multitask Learning for Mixed-Type Outcomes with Shared Sparsity

Most existing multitask learning approaches are limited by their reliance on task-specific loss functions tailored to the scale and type of each outcome. When outcomes differ across tasks, these losses are generally not directly comparable, which makes it difficult to formulate a unified objective and may limit information sharing across tasks. We propose a multitask transformation framework in which task-specific responses may differ through unknown monotone transformations. Motivated by high-dimensional biological applications in which the predictor dimension may diverge with the sample size while only a common subset of predictors is informative, we consider shared sparsity across tasks. Under this framework, we estimate the target functions and identify important predictors by optimizing a smoothed rank-based criterion with a group-Lasso penalty, implemented through a multitask deep neural network with a shared first layer. We establish the nonasymptotic excess-risk bounds, and variable-selection consistency for the proposed estimator. Simulation studies show that the proposed method achieves competitive prediction and variable-selection performance compared with competing approaches. Analyses of gene-expression studies with continuous, binary, and mixed outcomes further illustrate that the proposed method improves prediction and identifies biologically meaningful shared predictors.

stat.ML

Composed Object Retrieval: Object-level Retrieval via Composed Expressions

Retrieving fine-grained visual content based on user intent remains a challenge in multimodal systems. Although current Composed Image Retrieval (CIR) methods combine reference images with retrieval texts, they are constrained to image-level matching and cannot localize specific objects. To this end, we propose Composed Object Retrieval (COR), a new object-level retrieval task that retrieves target object(s) from candidate objects in a target image and grounds the retrieved result with pixel-level masks. Given a reference object, its mask, a target image, and a retrieval text describing the desired modification, COR requires models to perform composed visual-textual reasoning rather than relying on explicit category names. This setting introduces several challenges, including fine-grained compositional matching, negative-object filtering under visually similar distractors, and flexible single- or multi-object retrieval. We construct COR125K, the first large-scale COR benchmark, containing 125,541 retrieval triplets across 408 categories with base/novel splits for evaluating category-level generalization. We also present CORE, a unified end-to-end model that integrates reference region encoding, adaptive vision-text interaction, and region-level contrastive learning to align composed representations with target objects while suppressing background and distractors. Extensive experiments demonstrate that CORE significantly outperforms existing CIR-based pipelines and strong baselines in both base and novel categories, establishing a simple and effective foundation for fine-grained object-level multimodal retrieval. Code will be released publicly at https://github.com/wangtong627/COR.

cs.CV

ARTEMIS: Agent-guided Reliability-aware Temporal Mask Evolution for Imperfectly Supervised Video Polyp Segmentation

Imperfectly supervised video polyp segmentation (VPS) aims to learn dense, temporally consistent masks from inexpensive supervision, including weak annotations (points, scribbles) and semi-supervision with few densely labeled frames. This setting is clinically valuable but challenging due to weak contrast, ambiguous boundaries, motion blur, and specular highlights, compounded by sparse pixel-level guidance. While SAM2 can generate dense masks from sparse inputs, direct pseudo-labeling often yields geometry-degraded masks with boundary leakage, underutilizes temporal consistency, and ignores reliability. To address these issues, we propose ARTEMIS, a unified framework for imperfectly supervised VPS driven by agent-guided reliability-aware temporal mask evolution. ARTEMIS initializes coarse masks from available supervision: SAM2 converts points/scribbles, while dense labels serve as reliable anchors. A debate-and-judge vision-language agent selects reliable temporal anchors under weak supervision, which are propagated bidirectionally with SAM2 to refine unreliable or unlabeled frames. Finally, ARTEMIS trains the segmenter using temporal reliability-aware robust learning, incorporating reliability-guided reference selection, a Reference Prototype Transport Module, and reliability-aware robust loss. These components assess mask reliability, evolve anchors over time, transport target identity across frames, and down-weight noisy supervision instead of discarding difficult samples. Experiments on SUN-SEG and CVC-ClinicDB-612 under scribble, point, and limited-label settings demonstrate that ARTEMIS achieves state-of-the-art performance. Code will be released at https://github.com/wangtong627/ARTEMIS.

cs.CV

Instability Caused by Integration of IBRs under Strong Grid Connections -- A Practical Case Study on Large-scale Energy Storage Systems

It has been well known that inverter-based resources (IBRs) can lead to converter-driven stability issues under weak grid connections. However, as the number of IBRs increases, instabilities can also occur even under strong grid connections. A practical case is presented to demonstrate this conclusion, using large-scale energy storage systems (ESSs) as an example. In this study, the ESSs induce oscillations with a frequency of 150 Hz in the d-q coordinates while providing both capacitive and inductive reactive power support (achieved by ESS functional control loops) to the connected power system. Theoretical analysis reveals that under strong grid connections, the dynamic interactions among power conversion systems (PCSs) of ESSs can be superimposed and intensified as the ESS scale extends, which reduces oscillation damping and leads to system instability. This indicates that ESS functional control loops also have potential instability risks when providing supports to power systems, which should be carefully examined. Finally, major impact factors are identified to mitigate the oscillations, and the conclusions are validated based on the SIMULINK platform. This paper provides valuable practical insights into system instabilities even under strong grid conditions, emphasizing the importance of functional control design and careful planning of the scale for IBR-dominated systems.

eess.SY

Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios

Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, most existing detectors rely on static inference with globally shared parameters, limiting their ability to adapt to diverse and unstructured traffic scenarios. As a result, a single static model often struggles to simultaneously handle both clear near-range samples and challenging conditions such as distant small targets or adverse weather environments. To address this limitation, we propose CBDES MoE TSR, a hierarchically decoupled heterogeneous mixture-of-experts(MoE) framework for traffic sign recognition. The proposed framework departs from the conventional globally shared parameter paradigm by introducing a heterogeneous You Only Look Once (YOLO) expert pool together with a lightweight gating network, enabling an image-level dynamic routing mechanism. Based on the semantic characteristics of the input image, the gating module selectively activates the most suitable expert model from the expert pool, enabling a shift from fixed parameter fitting to on-demand dynamic representation. This design enhances feature extraction capability for specific scenarios while maintaining controlled inference overhead. Experimental results demonstrate that the proposed method achieves a remarkable balance between detection accuracy and efficiency on the composite traffic sign dataset. Specifically, our method attains an mAP50-95 of 76.8%, yielding a 2.3% improvement over the baseline method (74.5%) while simultaneously reducing computational overhead by approximately 39.4%. These findings robustly validate the effectiveness of the proposed approach.

cs.CV

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and essentially degrades the task to text rendering. Moreover, the conditions imposed by pre-trained glyph encoder limit the scope of editable text. To address these issues, this paper proposes a self-prompting scene text editing method that constructs style and glyph prompts directly from the original image, without introducing additional style or glyph encoders. We employ a two-stage training strategy: the diffusion transformer is first trained on large-scale self-supervised data and then refined using a small set of paired images. By leveraging the in-context learning capability of the Multi-Modal Diffusion Transformer (MM-DiT), it achieves open-vocabulary and style-consistent text editing. Experimental results on various languages demonstrate that our method achieves the state-of-the-art performance in both text accuracy and style consistency. Our project page: hongxiii.github.io/mstedit.

cs.CV