arXiv ScienceSearch

arXiv subjects

Ling-Yue Ge

Publications and source records attributed to Ling-Yue Ge.

3 recordsLinked to original sources

SkillAA: Attribution-Guided Skill-Graph Updating with Targeted Validation and Rollback

External skills provide domain procedures without parameter updates, but existing methods often edit skills directly from failed rollouts without structured routing from an observed failure to an editable location; existing skill graphs also underuse semantic boundaries, object addresses, and topological dependencies for skill retrieval, targeted updating, and scoped validation. We introduce SkillAA (Skill Abductive Attribution), a structured skill-optimization framework for frozen language models. It represents skill applicability, execution, and composition in a unified graph, allowing the same structure to support skill selection, attribution-guided repair, and update validation. SkillAA contrasts successful and failed executions to route candidate repairs to specific graph objects, updates only the selected local structure, and uses Local and Big Gates to screen candidate changes before commitment. With gpt-5.6-sol, SkillAA reaches 81.5%, 66.7%, and 91.2% on SearchQA, LiveMath, and DocVQA, respectively, and attains the highest observed mean in every main setting. These results support the utility of attribution-guided graph editing and graph-scoped validation.

cs.AI

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capabilities under multi-criteria constraints. To address this gap, we introduce MapTab, a multimodal benchmark designed to assess holistic multi-criteria reasoning in MLLMs through route-planning tasks. MapTab requires models to perceive and ground visual information from map images while integrating route attributes, such as Time and Price, from structured tables. It covers two scenarios: Metromap, spanning metro networks in 160 cities across 52 countries, and Travelmap, featuring 168 representative tourist attractions from 19 countries. Overall, MapTab includes 328 images, 196,800 route-planning queries, and 3,936 QA queries, incorporating four key criteria: Time, Price, Comfort, and Reliability. Extensive evaluations of 21 representative MLLMs show that current models still struggle with multicriteria multimodal reasoning. Notably, when visual perception is unreliable, multimodal reasoning can even underperform unimodal approaches. MapTab therefore offers a challenging and realistic testbed for systematically evaluating and advancing MLLMs across core perception, integration, numerical comparison, and route planning capabilities.

cs.LG

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry structural obligations, including capability coverage, message compatibility, validation, final-answer aggregation, and parser-compatible output protocols. Existing systems either fix the role inventory and lose adaptivity, or allow unconstrained generation to induce role drift, removing structurally necessary roles and breaking answer contracts. We formulate this as contract-preserving role evolution, requiring every committed edit to preserve five structural contracts (capability, communication, validation, aggregation, output protocol). We instantiate this formulation in SERO, a Self-Evolving Role Orchestration framework that evolves a typed role-card pool through credit-guided retrieval, a credit-ranked communication DAG with a protected terminal aggregator and conditional validator repair, and a contextual-bandit controller whose LLM-proposed edits are committed only when they preserve the contracts and improve task score. Experiments on real-world reasoning benchmarks across three LLM backbones confirm the value of contract-preserving role evolution.

cs.CL