arXiv ScienceSearch

arXiv subjects

Cheng Fan

Publications and source records attributed to Cheng Fan.

13 recordsLinked to original sources

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.

cs.AI

RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs

Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods select tokens by importance, diversity, or spatial coverage, but treat retained tokens as interchangeable and do not explicitly track which object-related regions are already covered. We present RoRA, a training-free framework that casts visual token pruning as role-oriented regional evidence allocation. Given a fixed budget, RoRA partitions tokens into a protected semantic core, complementary context, and fine-grained detail. It first calibrates text-conditioned attention with a positional prior and a prompt-calibrated object prior, then builds Attention-Anchored Regions (AARs) from high-confidence anchors as lightweight proxies for covered object support. Context is explored mainly outside AARs, while a small AAR-guided budget restores local detail; pairwise similarity is used only for context-stage redundancy filtering. Under matched budgets, RoRA consistently outperforms strong training-free baselines across LLaVA and Qwen-VL families, retaining most of the unpruned accuracy even at aggressive pruning ratios, e.g., 96.5% of full performance at 88.9% pruning on LLaVA-1.5, and improving over D2Pruner by about 5% on Qwen3-VL at 75-90% pruning. At a 66.7% pruning ratio, RoRA requires only 0.7 ms for token selection and reduces end-to-end inference time by 24.6%, corresponding to a 1.33x speedup over unpruned inference on an NVIDIA H800.

cs.CV

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation

Referring Remote Sensing Image Segmentation (RRSIS) has achieved significant progress through the integration of VLMs and the Segment Anything Model (SAM). However, this progress largely relies on strong pre-trained capabilities, while leaving two fundamental limitations insufficiently addressed: (1) Architectural Weak-Coupling, where the unidirectional flow forces reliance on coarse VLM prompts and wastes SAM's pixel-level structural guidance, causing localization drift; and (2) Object-Centric Semantic Bias, where models overemphasize dominant object semantics while remaining insensitive to spatial reasoning crucial for RRSIS. Motivated by these observations, we propose CROSS, a tightly integrated paradigm for RRSIS. First, we introduce Linguistic-Guided Cascaded Distillation (LGCD) to bridge the architectural gap, which distills SAM's geometric affinities as soft regularizers into VLM intermediate layers, injecting dense structural priors to refine localization. Second, Perspective-Spatial Contrastive Learning (PSCL) imposes cross-anchored constraints by mining mask-filtered deceptive distractors and spatial-linguistic counterfactuals as hard negatives, explicitly shattering semantic shortcuts to enforce genuine logical consistency. Extensive experiments on RRSIS benchmarks demonstrate that CROSS achieves state-of-the-art performance and maintains precise localization even under severe spatial description perturbations, standing as a robust new paradigm for RRSIS.

cs.CV

From Question Answering to Task Completion: A Survey on Agent System and Harness Design

LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. As agent systems have evolved from prompt engineering to workflows and context engineering, harness engineering, and agent-native training with co-evolution, a central question has become increasingly important: where does the bottleneck in agent performance reside, in the foundation model, in the execution harness, or in the coupling between them? This survey examines LLM-based agents through a model-harness lens. We first clarify the functional definition of agents and the implementation view of an LLM-based agent as a foundation model coupled with an execution harness. We then analyze the limits of model-centric scaling, trace four paradigms of agent engineering, and decompose the execution harness into six coupled runtime responsibilities: observation, context, control, action, state, and verification. Using this decomposition, we map task properties and domain pressures to harness configurations, review benchmark and evaluation practices, and synthesize model-harness evidence on how runtime design affects long-horizon task completion, efficiency, and reliability. Finally, we identify open challenges in value-aware evaluation, safety, harness generalization, and model-harness co-evolution. Rather than treating agents as models with auxiliary tools, this survey argues that agent quality -- including success, efficiency, safety, and generalization -- emerges from the interaction between model capability, runtime infrastructure, task structure, and evaluation design. A collection of papers discussed in this survey is provided in https://github.com/ggjy/Awesome-Agent-Engineering.

cs.AI

TDDFT Gradients and Nonadiabatic Couplings with Minimal Auxiliary Basis Set Approximation for Fewest-Switches Surface Hopping Dynamics

The electronic structure calculations remain a major bottleneck in ab initio nonadiabatic molecular dynamics. We develop an efficient TDDFT-based FSSH implementation in the GPU4PySCF package for medium-sized molecular systems. Our approach combines density fitting, TDDFT with minimal auxiliary basis sets (TDDFT-ris), and an approximate Z-vector solver to reduce the computational cost of TDDFT excited states and derivative coupling calculations. These approximations introduce negligible errors in realistic FSSH workloads while maintaining high computational efficiency. Benchmark results show that, for 73-atom systems with a triple-$\zeta$ basis set, individual electronic structure calculations are completed within one minute on a single NVIDIA A100 GPU.

physics.chem-ph

GPU Accelerated Minimal Auxiliary Basis Approach TDDFT for Large Organic Molecules

We introduce a GPU-accelerated implementation of time-dependent density functional theory with the minimal auxiliary basis approach (TDDFT-risp) in GPU4PySCF, together with large system demonstrations carried out using the Tamm--Dancoff approximation (TDA-risp). The method combines GPU-accelerated three-center integral evaluation, tensor contractions, exchange-space truncation, omission of hydrogen atoms from the auxiliary basis, and a host memory assisted Davidson solver. On the EXTEST42 benchmark set, a conservative 40 eV exchange cutoff yields excitation-energy errors relative to standard TDA of about 0.03--0.05 eV for low-lying states. For systems of 300 to 3000 atoms, we demonstrate that TDA-risp calculations of 15 low-lying excited states with $\omega$B97XD/def2-SVP complete on a single A100 GPU with wall times ranging from minutes to hours. These results position GPU-TDDFT-risp as a practical route toward excited-state calculations for large organic and biomolecular systems with thousands of atoms.

physics.chem-ph

Analytical Excited-State Gradients and Derivative Couplings in TDDFT with Minimal Auxiliary Basis Set Approximation and GPU Acceleration

Calculating excited-state gradients and derivative couplings using time-dependent density functional theory (TDDFT) remains a computationally demanding task. An efficient variant, TDDFT with resolution of the identity and a minimal auxiliary basis (TDDFT-ris), has been developed to accelerate excitation energy calculations. However, the formulation and implementation of analytical derivatives for this method have not yet been reported. In this work, we present an implementation of analytical excited-state gradients and derivative couplings within the TDDFT-ris framework. Benchmark calculations on medium-sized organic molecules demonstrate a two- to three-fold speedup for both gradients and derivative couplings compared to standard TDDFT. The accuracy of the TDDFT-ris approach is assessed for gradient-dependent applications, including geometry optimizations, emission energy calculations, and the localization of minimum-energy crossing points. Overall, the TDDFT-ris method provides reliable approximations for most cases, with noticeable errors mainly occurring in derivative couplings between nearly degenerate states.

physics.chem-ph

Performing Path Integral Molecular Dynamics Using Artificial Intelligence Enhanced Molecular Simulation Framework

This study employed an artificial intelligence-enhanced molecular simulation framework to enable efficient Path Integral Molecular Dynamics (PIMD) simulations. Owing to its modular architecture and high-throughput capabilities, the framework effectively mitigates the computational complexity and resource-intensive limitations associated with conventional PIMD approaches. By integrating machine learning force fields (MLFFs) into the framework, we rigorously tested its performance through two representative cases: a small-molecule reaction system (double proton transfer in formic acid dimer) and a bulk-phase transition system (water-ice phase transformation). Computational results demonstrate that the proposed framework achieves accelerated PIMD simulations while preserving quantum mechanical accuracy. These findings show that nuclear quantum effects can be captured for complex molecular systems, using relatively low computational cost.

physics.chem-ph

Spatial Control of Charge Doping in n-Type Topological Insulators

Spatially controlling the Fermi level of topological insulators and keeping its electronic states stable are indispensable processes to put this material into practical use for semiconductor spintronics devices. So far, however, such a method has not been established yet. Here we show a novel method for doping hole into n-type topological insulators Bi$_2$X$_3$ (X= Se, Te) that overcomes the shortcomings of the previous reported methods. The key of this doping is to adsorb H$_2$O on Bi$_2$X$_3$ decorated with a small amount of carbon, and its trigger is the irradiation of photon with sufficient energy to excite core-electrons of the outermost layer atoms. This method allows controlling the doping amount by the irradiation time, and acts as photolithography. Such a tunable doping makes it possible to design the electronic states at the nanometer scale, and thus paves a promising avenue toward the realization of novel spintronics devices based on topological insulators.

cond-mat.mes-hall

Generating High-Precision Force Fields for Molecular Dynamics Simulations to Study Chemical Reaction Mechanisms using Molecular Configuration Transformer

Theoretical studies on chemical reaction mechanisms have been crucial in organic chemistry. Traditionally, calculating the manually constructed molecular conformations of transition states for chemical reactions using quantum chemical calculations is the most commonly used method. However, this way is heavily dependent on individual experience and chemical intuition. In our previous study, we proposed a research paradigm that uses enhanced sampling in molecular dynamics simulations to study chemical reactions. This approach can directly simulate the entire process of a chemical reaction. However, the computational speed limits the use of high-precision potential energy functions for simulations. To address this issue, we present a scheme for training high-precision force fields for molecular modeling using a previously developed graph-neural-network-based molecular model, molecular configuration transformer. This potential energy function allows for highly accurate simulations at a low computational cost, leading to more precise calculations of the mechanism of chemical reactions. We applied this approach to study a Claisen rearrangement reaction and a Carbonyl insertion reaction catalyzed by Manganese.

physics.chem-ph

Carrier doping of Bi$_2$Se$_3$ surface by chemical adsorption -- a DFT study

Bi$_2$Se$_3$ is one of the most promising topological insulators, but it suffers from intrinsic n-doping due to Se-vacancies, which shifts the Fermi level into the bulk conduction band, leading to topologically trivial carriers. Recently it was shown that this Fermi-level shift can be compensated by a locally controlled surface p-doping process, through water adsorption and XUV irradiation. Here, the microscopic mechanism of this surface doping is studied by means of density functional theory (DFT) focusing on the adsorption of H$_2$O, OH, O, C and CH on Bi$_2$Se$_3$. We find that water adsorption has a negligible doping effect while hydroxyl groups lead to n-doping. Carbon adsorption on Se vacancies gives rise to p-doping but it also strongly modifies the electronic band structure around the Dirac point. Only if the Se vacancies are filled with atomic oxygen, the experimentally observed p-doping without change of the topological surface bands is reproduced. Based on the DFT results, we propose a reaction path where photon absorption gives rise to water splitting and the produced O atoms fill the Se vacancies. Adsorbed OH groups appear as intermediate states and carbon impurities may have a catalytic effect in agreement with experimental observations.

cond-mat.mtrl-sci

A Generalized Nucleation Theory for Ice Crystallization

Despite the simplicity of the water molecule, the kinetics of ice nucleation under natural conditions can be complex. We investigated spontaneously grown ice nuclei using all-atom molecular dynamics simulations and found significant differences between the kinetics of ice formation through spontaneously formed and ideal nuclei. Since classical nucleation theory can only provide a good description of ice nucleation in ideal conditions, we propose a generalized nucleation theory that can better characterize the kinetics of ice crystal nucleation in general conditions. This study provides an explanation on why previous experimental and computational studies have yielded widely varying critical nucleation sizes.

cond-mat.soft

Receiver Operating Characteristic Curves and Confidence Bands for Support Vector Machines

Many problems that appear in biomedical decision making, such as diagnosing disease and predicting response to treatment, can be expressed as binary classification problems. The costs of false positives and false negatives vary across application domains and receiver operating characteristic (ROC) curves provide a visual representation of this trade-off. Nonparametric estimators for the ROC curve, such as a weighted support vector machine (SVM), are desirable because they are robust to model misspecification. While weighted SVMs have great potential for estimating ROC curves, their theoretical properties were heretofore underdeveloped. We propose a method for constructing confidence bands for the SVM ROC curve and provide the theoretical justification for the SVM ROC curve by showing that the risk function of the estimated decision rule is uniformly consistent across the weight parameter. We demonstrate the proposed confidence band method and the superior sensitivity and specificity of the weighted SVM compared to commonly used methods in diagnostic medicine using simulation studies. We present two illustrative examples: diagnosis of hepatitis C and a predictive model for treatment response in breast cancer.

stat.ML