arXiv ScienceSearch

arXiv subjects

Jianwei Yin

Publications and source records attributed to Jianwei Yin.

At least 19 recordsLinked to original sources

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are aggregated across a batch. We formulate these dimen- sions as Intra-Trajectory Feedback Attribution and Inter-Trajectory Objec- tive Aggregation, and introduce BATON (Bayesian Attribution and Trajectory Objective Normalization), a dual-axis policy optimization framework. BATON instantiates the first axis with Bayesian Feedback Attribution, which constructs a feedback-conditioned posterior over sampled actions, and the second with Trajec- tory Mass Normalization (TMN), which assigns equal optimization mass to com- plete trajectories. Experiments with GRPO and GiGPO on ALFWorld, WebShop, and SearchQA show that both axes provide independent gains and that their combi- nation consistently achieves the strongest overall performance across model scales.

cs.AI

Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications

Learning-augmented algorithms use fallible predictions while retaining formal performance guarantees. This survey synthesizes prediction interfaces, error measures, consistency--robustness trade-offs, and five representative construction mechanisms across online optimization, caching, learned data structures, graph problems, and mechanism design. An orthogonal theorem-level axis distinguishes achieved upper bounds from matched asymptotic dependence. Formal guarantees are separated from empirical systems evidence, with explicit treatment of prediction cost, feedback, and composition. The resulting synthesis states sufficient conditions for limited end-to-end reasoning and delineates open problems in cost-aware prediction, endogenous error, semantic predictors, and benchmarking.

cs.LG

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/ASK/REFUSE/CONFIRM) and introduces a self-revealing upper bound Acc <= GAR (Gold Action Recall); the two modes show up as bound violation (Acc > GAR, exposing state-grader masking of miscalibration) and large bound slack (GAR >> Acc, localizing execution failure within TOOL_CALL). We validate it on a panel of tool-calling models across multiple multi-turn benchmarks. Across our panel, the diagnostic reveals action-class miscalibration as a substantial failure mode the state grader cannot see. This gap inflates standing for heavily tool-trained families, which our diagnostic separates from families with context-appropriate action choice. Calibration is reshapable through context-only perturbations, but the reshape is heterogeneous: a single perturbation moves accuracy in opposite directions across families (up to +11.5 vs -21.0 pp on the same scenario), and its effect further depends on the perturbation mechanism. We argue that multi-turn tool-calling evaluations should supplement aggregate accuracy with action-class diagnostics that expose what the model actually does in each scenario.

cs.CL

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).

cs.RO

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.

cs.AI

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols

With the rapid advancement of generative models, privacy, copyright, safety, and reliability risks have attracted growing attention. To mitigate these risks, machine unlearning has been increasingly adapted from traditional classification models to generative settings. Despite notable progress, existing studies remain fragmented in their target definitions, unlearning mechanisms, and evaluation protocols, making objective comparison difficult across models, modalities, and applications. Moreover, modality-specific surveys often overlook the shared structure of Generative Model Unlearning (GenMU). To address this gap, we provide a comprehensive review of GenMU and formulate it as target-constrained distributional projection: given a target event, an unlearning operator transforms the generative distribution to suppress target-related outputs while preserving useful behavior and controlling operator cost. Under this framework, an unlearning request is specified by a target event, implemented through an unlearning operator, and assessed by empirical evidence over target suppression, distribution preservation, and operator cost. We further reorganize existing GenMU studies under this view. From this perspective, we provide the first explicit and unified account of how GenMU connects to mainstream applications, including privacy protection, copyright and style protection, safety alignment, hallucination mitigation, and deployment defense. Finally, we identify key open problems and future directions toward reliable, scalable, robust, and auditable GenMU. We consistently maintain the related open-source materials at https://github.com/caxLee/GenMU-Survey.

cs.LG

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Current approaches generally involve either expensive task-specific training or adapting pre-trained Text-to-Speech (TTS) models. However, both paradigms face challenges: task-specific methods often degrade fidelity in unedited regions, whereas TTS adaptations struggle with a trade-off between editing naturalness and temporal fidelity. To address these issues, we propose AST, an Adaptive, Seamless, and Training-free speech editing framework. Built upon pre-trained TTS, AST leverages Latent Recomposition to stitch preserved source segments with synthesized targets, guaranteeing fidelity in unedited regions. To break the quality-controllability trade-off, we introduce Adaptive Weak Fact Guidance (AWFG), which modulates a mel-space signal to ensure seamless boundary transitions without disrupting the generative manifold. Furthermore, to address evaluation gaps in temporal fidelity, we propose a new benchmark suite: the LibriSpeech-Edit dataset and a novel metric, Word-level Dynamic Time Warping (WDTW). Extensive experiments demonstrate that AST consistently outperforms existing task-specific and fine-tuned speech editing methods across content accuracy, perceptual quality, speaker preservation, and temporal fidelity. Remarkably, AST achieves state-of-the-art zero-shot speech editing performance without any task-specific training or paired editing data, validating the effectiveness of latent recomposition and AWFG in bridging the quality-controllability trade-off.

cs.SD

Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion

Repository-level code completion remains challenging for large language models (LLMs) due to cross-file dependencies and limited context windows. Prior work addresses this challenge using Retrieval-Augmented Generation (RAG) frameworks based on semantic indexing or structure-aware graph analysis, but these approaches incur substantial computational overhead for index construction and maintenance. Motivated by common developer workflows that rely on lightweight search utilities (e.g., ripgrep), we revisit a fundamental yet underexplored question: how far can simple, index-free lexical retrieval support repository-level code completion before more complex retrieval mechanisms become necessary? To answer this question, we systematically investigate lightweight, index-free, intent-aware lexical retrieval through extensive empirical analysis. We first introduce Naive GrepRAG, a baseline framework in which LLMs autonomously generate ripgrep commands to retrieve relevant context. Despite its simplicity, Naive GrepRAG achieves performance comparable to sophisticated graph-based baselines. Further analysis shows that its effectiveness stems from retrieving lexically precise code fragments that are spatially closer to the completion site. We also identify key limitations of lexical retrieval, including sensitivity to noisy matches from high-frequency ambiguous keywords and context fragmentation caused by rigid truncation boundaries. To address these issues, we propose GrepRAG, which augments lexical retrieval with a lightweight post-processing pipeline featuring identifier-weighted re-ranking and structure-aware deduplication. Extensive evaluation on CrossCodeEval and RepoEval-Updated demonstrates that GrepRAG consistently outperforms state-of-the-art (SOTA) methods, achieving 7.04-15.58 percent relative improvement in code exact match (EM) over the best baseline on CrossCodeEval.

cs.SE

Service Ecosystem Evolution: A Comprehensive Survey from Complex Network Perspectives

Digital society increasingly relies on complex service ecosystems, formed by interconnected services from technology giants. However, the growing scale and intricate dependencies of these ecosystems pose significant challenges to their evolution, frequently leading to systemic failures during upgrades or restructuring. To address these challenges, current research is shifting from the perspective of single services to the ecosystem. Drawing upon the synthesis and induction of current research, we present a comprehensive survey focused on the study of service ecosystem evolution, employing a novel three-stage analytical framework. This framework structures the evolutionary lifecycle and provides a systematic way to organize and review existing research, filling the gap caused by the lack of surveys specifically focused on service ecosystem evolution. Additionally, we pioneer the application of complex network theory to analyze service ecosystem evolution, providing a novel perspective to capture the inherent connectivity, dynamics, and emergent properties often missed by traditional approaches. Building on our analysis, we point out critical research gaps and propose three specific future directions. We provide a robust theoretical foundation and methodological guidance for understanding and guiding the sustainable evolution of modern service ecosystems.

physics.soc-ph

Continual Video-MLLM Adaptation over Evolving Domains

Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remains underexplored. In real-world deployments, video data often arrives continuously from heterogeneous domains, requiring the model to acquire new domain-specific knowledge without overwriting previously learned capabilities. Existing continual learning methods typically rely on shared adaptation spaces, which can induce severe cross-domain interference and catastrophic forgetting. We propose Distribution-Aware Expert Routing, a parameter-efficient framework for continual Video-MLLM adaptation over evolving domains. DAER maintains domain-isolated lightweight experts while keeping the pretrained Video-MLLM backbone frozen, thereby decoupling domain-specific adaptation from the general multimodal knowledge of the pretrained model. To enable fine-grained specialization, we introduce an intra-domain distribution-aware routing mechanism that matches each input to expert-level prototype reservoirs using MMD. To address the absence of task identities at inference time, we further propose an inter-domain routing mechanism that performs prototype matching in a discriminative subspace for robust domain identification. In addition, we introduce adaptive domain merging to improve parameter scalability and adopt a two-stage optimization strategy to stabilize expert specialization during continual learning. We evaluate DAER by curating a domain-incremental benchmark built from ten VidQA datasets covering diverse visual environments and reasoning demands. Experiments on two strong Video-MLLM backbones show that DAER consistently outperforms prior methods.

cs.CV

Agentic Services Computing

Services computing has evolved from Web services and microservices to cloud-native and serverless paradigms. These approaches established mature principles for describing, composing, deploying, operating, and governing reusable software functions. LLM-based agents now introduce a fundamentally different service form. Service value in this paradigm emerges not only from invoking predefined functions but also from delegating goals to autonomous entities. These entities understand context, use tools, collaborate with peers, and act across open environments. This shift raises a core question for services computing. How can goal-driven, stateful, tool-mediated, and accountable autonomous behavior be engineered and managed as a service? Recent studies on LLM agents and multi-agent systems provide important foundations. A clear service-centered research roadmap for this emerging paradigm nevertheless remains absent. This work introduces Agentic Services Computing (ASC) to address this gap. ASC extends services computing from managing reusable functional endpoints to engineering and governing autonomous service entities. It defines agentic services as service-oriented autonomous agents. Related research is organized through a lifecycle view that connects service objects, system structures, enabling infrastructure, evaluation metrics, application evidence, and open challenges. This service-centered perspective establishes a foundation for future service ecosystems. Autonomous agents can thus be systematically described, composed, delivered, monitored, audited, and evolved as first-class services.

cs.SE

Industrial Data-Service-Knowledge Governance: Toward Integrated and Trusted Intelligence

The convergence of artificial intelligence, cyber-physical systems, and distributed networking has accelerated the evolution of industrial intelligence across edge, cloud, and cross-organizational communication environments. However, existing governance mechanisms remain fragmented across data management, service orchestration, and knowledge-based decision-making, making it difficult to ensure reliability, accountability, compliance, and explainability throughout the industrial intelligence stack. To address this gap, we present TRISK (TRusted Industrial Data-Service-Knowledge governance), a conceptual and taxonomic framework for trustworthy industrial intelligence. TRISK is grounded in a five-dimensional trust model covering quality, security, privacy, fairness, and explainability, and formalizes how trust is constructed, propagated, aggregated, and fed back across data, service, and knowledge layers in networked industrial systems. Through a structured synthesis of more than 100 representative studies, standards, and technical reports, we examine data governance as the foundation of trust construction, service governance as the mediation layer for trustworthy execution, and knowledge governance as the semantic anchor for reasoning, validation, and feedback adaptation. We further discuss industrial implementation patterns, cross-industry implications, and the role of emerging communication and computing technologies. Finally, we outline a future research roadmap toward adaptive, verifiable, and human-aligned industrial governance for Industry 5.0.

cs.CE

Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction

Visual token reduction has emerged as an effective strategy for accelerating Multimodal Large Language Models (MLLMs). Many existing methods prune tokens by ranking text-visual attention scores. However, we show that attention is often dominated by a model-induced prior: even without textual instruction, MLLMs tend to focus on certain task-agnostic regions. Consequently, the attention scores of instruction-conditioned tokens are suppressed, increasing the risk that these tokens are discarded during pruning. To address this issue, we propose Prior-Corrected Token Reduction (PriorTR), a training-free token reduction method that explicitly separates task-conditioned attention from the model-induced prior. PriorTR estimates the attention map of the prior, and contrasts it with the task-conditioned attention distribution to measure the additional usable information contributed by each visual token. Importantly, PriorTR computes both the model-induced prior and the task-conditioned posterior within a single forward pass by introducing a null token that serves as an instruction-agnostic probe in the attention block. This design avoids duplicated propagation. Extensive experiments across multiple multimodal benchmarks and MLLMs demonstrate that PriorTR consistently improves the trade-off between accuracy and efficiency over strong training-free baselines, particularly under aggressive token budgets.

cs.CV

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performance. Existing token pruning methods typically rely on single-layer signals, such as attention scores or token similarities, which overlook the cross-layer transformation of visual representations and may exhibit positional bias in multimodal token sequences. To address this limitation, we propose a training-free token pruning framework based on Cross-Layer Spectral Evolution (CLSE). Instead of measuring token importance from single-layer feature magnitudes, CLSE quantifies how token representations evolve across Transformer layers in the frequency domain. This evolution reflects the transition from high-frequency structural details to low-frequency semantic abstractions. We observe that tokens with stronger spectral redistribution across layers are more likely to be semantically active and should therefore be preserved. By modeling cross-layer token dynamics, CLSE provides a stable importance criterion that mitigates positional bias. Extensive experiments on both image and video benchmarks demonstrate that CLSE achieves a superior trade-off between efficiency and accuracy under aggressive token reduction. Across multiple MLLMs, CLSE reduces FLOPs, KV cache memory, and latency while maintaining competitive or improved performance.

cs.CV

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus obscuring real progress. Rebuilding high-quality benchmarks such as V*Bench requires substantial human annotation, yet each static release can quickly become another leaked artifact. We propose ReKey, a live benchmark protocol that randomly regenerates the answer-bearing local detail, or visual key, in real images at evaluation time. Using human-validated edit slots, ReKey samples fresh instances with new answers, construction-grounded labels, and controlled visual-search difficulty. On V*Bench, the ReKey regenerated benchmark reveals a sharp score jump across eight frontier vision-language models (VLMs): The original items score 9.5--18.8 percentage points higher than the regenerated variants. By making the visual key renewable, ReKey keeps evaluation fresh as models and training data evolve.

cs.CV

Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions

Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones. However, current approaches to fine-tuned SVs suffer from two limitations. First, they require careful selection of steering factors on a per-SV basis to balance steering effectiveness and generation quality at inference time. Second, they operate as full-sequence SVs (FSSVs), which can sacrifice generation quality regardless of factor selection due to excessive intervention on the model generation process. To address the first limitation, we propose joint training of steering factors and directions, such that post-hoc factor selection is no longer required. Using neural network scaling theory, we find that moderately large initialization sizes and learning rates for steering factors are essential for stability and efficiency of joint training. To tackle the second limitation, we draw inspiration from representation fine-tuning and introduce Prompt-only SV (PrOSV), an SV that intervenes only on a few prompt tokens. Our empirical results show that PrOSV outperforms traditional FSSVs on AxBench when using our joint training scheme. We also find that PrOSV achieves a better tradeoff between general model utility and adversarial robustness than FSSV.

cs.LG

Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness

Breakthroughs in large language models and multimodal generation technologies have propelled the digital reconstruction of human mental traits, emotional patterns, and long-term memory from science fiction toward engineering practice. Yet current research and industry practices at the intersection of AI and digital humans remain hampered by fundamental conceptual ambiguities: the essential differences between next-generation intelligent agents and traditional virtual humans, the construction pathways for digital entities possessing self-identity, and the core technical and ethical challenges confronting this domain all demand urgent clarification. This paper systematically examines the transformative logic underlying the transition from traditional virtual humans to the ``Soul Computing'' paradigm, driven by frontier AI technologies. We first analyze the evolutionary patterns of human consciousness and memory mechanisms, reassessing the core value of massive multimodal digital fragments in the reverse reconstruction of individual mental worlds. On this basis, we formally delineate the academic connotations of narrow and broad Soul Computing for the first time, clarifying its academic boundaries and essential distinctions from Affective Computing, Historical Reconstruction, and Mortal Computation. We argue that Soul Computing systems must architecturally construct an ``Intensional'' core rather than serving as purely ``Extensional'' functional carriers, thereby enabling the fundamental transition of AI from toolhood to living agency.

cs.AI

Mitigating Manifold Departure: Uncertainty-Aware Subspace Rectification for Trustworthy MLLM Decoding

MLLMs frequently hallucinate objects inconsistent with visual inputs. This issue is typically attributed to the over-reliance on language priors, which can override the visual context. Recent training-free decoding strategies address this by penalizing language priors. However, these methods overlook the dual nature of language priors, where they can be both helpful and harmful depending on the alignment with visual evidence. In particular, blindly suppressing language priors often disrupts the model's semantic manifold, leading to performance degradation, a phenomenon we term Manifold Departure. To address this, we propose Manifold-Guided Adaptive Projection (MGAP), a geometry-aware, training-free decoding method that mitigates hallucinations while preserving representation structure. MGAP first constructs a language-prior subspace from blind hidden states via SVD. During decoding, MGAP projects each multimodal hidden state onto this subspace and applies a consistency-aware gate to adaptively attenuate only the projected prior component, yielding a subspace-selective update that largely preserves the orthogonal semantic components. Extensive experiments on POPE and CHAIR show that MGAP outperforms prior decoding baselines, achieving stronger hallucination suppression without sacrificing coherence.

cs.LG