arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,675 records · Page 93Linked to original sources

When History Helps and Hurts: Selective History Use across Multimodal Turns

Reliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical evidence despite conflicting new observations. Existing multi-turn evaluations rarely separate these history-use demands from underlying question difficulty. To address this gap, we introduce ReTurn, a benchmark of 7,000 base tasks spanning visual and audio evidence for evaluating selective history use. For task-carrying history, Reconfirm/Reground require applying a historical question to current media while varying historical agreement; for evidence-carrying history, Retrieve/Rebind require answering a current question using historical media while varying current-media competition. Each pair preserves the target question, media, and answer. Tasks support open-ended and multiple-choice evaluation, with matched single-turn counterparts serving as answerability references. Across 13 omni-modal, vision-language, and audio-language models, median model-level open-ended accuracy falls from 93.7% with direct input to 72.3% in conversation. Behavioral probes show that high question recall can coexist with weaker task application, while competing media can redirect answers away from historical targets. Supervised adaptation yields only partial gains. ReTurn provides a controlled framework for assessing whether multimodal models select and use the historical information required by each request.

cs.CL↗

Computing Lower Bounds on the Non-negative Rank via SAT Solvers

Finding the extension complexity (xc) of a polytope is equivalent to finding the non-negative rank of its slack matrix. We provide a formulation for the rectangle and refined rectangle covering, bounding the non-negative rank of diverse non-negative matrices. While the bounds are known, our formulation enables us to use boolean satisfiability (SAT) solvers, which proves to be a strong tool. We obtain improved values for lower bounds on the non-negative rank of multiple matrices. In particular, we determined the xc of some regular polygons.

math.OC↗

"Hot-Blooded" vs "Cold-Blooded": Simulating the Behavioral Phenotypes of Childhood Aggression via Generative Agents

This study examines the construct validity of LLM-based generative agents in simulating reactive, proactive, and co-occurring aggression in children. Four distinct agents were instantiated using a theory-driven parameterization grounded in the social information processing model. A total of 1,920 simulation runs were conducted across eight social scenarios, employing a hybrid blind-coding pipeline to extract 32 quantitative behavioral indicators. Results demonstrate robust discriminant validity relative to a non-aggressive baseline, with large effect sizes. High cross-seed reliability confirms that behavioral differentiation is driven by underlying psychological parameters rather than model stochasticity. Qualitative narrative analyses further converged with established empirical literature. Overall, these findings indicate that theory-parameterized LLM agents can accurately reproduce distinct aggression subtypes, offering a scalable, highly controllable framework for hypothesis generation, intervention piloting, and the refinement of psychological measurement tools.

cs.HC↗

Tell Robot What Not to Do: A Negation Understanding Perspective

Instruction following enables robots to perform diverse tasks specified in natural language, making it a fundamental capability for human-robot interaction. Beyond communicating desired outcomes, users also need to specify constraints on what not to do. We investigate how to enable vision-language-action models (VLAs) to follow negated instructions, where robots must accomplish task goals while respecting explicit exclusions. To this end, we propose NegaAlign, a parameter-efficient, plug-and-play framework that extends pretrained VLAs to follow negated instructions through image-language supervision alone. Specifically, we introduce Negation Transformation Layers into selected layers of the vision-language backbone to reshape intermediate instruction representations. Meanwhile, a teacher-guided alignment mechanism is designed to align instruction-relevant visual tokens, transferring action-relevant grounding from instructions that satisfy the negated constraint. The training phase uses supervision constructed from existing demonstrations and updates only the inserted layers, keeping all pretrained parameters frozen, including the action generator. We further introduce NegaBench, a simulation benchmark spanning 10 scenarios across five domains for systematically evaluating manipulation under negated constraints. Experiments across GR00T, $π_0$, and $π_{0.5}$ demonstrate consistent improvements in negated instruction following. With 11.6M trainable parameters, NegaAlign increases the negated-instruction success rate of $π_{0.5}$ from 2.60% to 88.45% on NegaBench and from 12.4% to 88.8% on real-world tasks, while retaining performance on affirmative instructions.

cs.RO↗

Interval-valued SHAP in Tree-Based Models

Shapley values are among the most popular feature-attribution explanations. Efficient approaches for computing/estimating Shapley values for tree-based models, which are state-of-the-art for tabular data sets, have been developed. However, it is known that Shapley values can be (highly) unrobust due to small and realistic changes. In this paper, we propose an imprecise Dirichlet model (IDM) based method to analyze the robustness of Shapley values in decision trees and random forests. Technically, it is done by quantifying and analyzing the interval-valued Shapley values when a few unannotated instances are randomly introduced to the leaves of the trees. The interval-valued Shapley values can be defined following common principles in handling incomplete data: the pessimistic and averaging principles. We derive various theoretical results that lead to efficient computation of the interval-valued Shapley values. We also show that the proposed method can be straightforwardly generalized to the case of Banzhaf values. We then present various case studies and experiments to illustrate the behaviour of the proposed interval-valued Shapley values and their applications in debiasing uninformative features.

cs.LG↗

Look Where You Can: Active View Selection for CAD Reconstruction under Occlusion

CAD reconstruction methods assume a luxury reality rarely grants: unrestricted visual access to the object, photographed from any desired angle. Real objects, however, are scene-embedded, bolted against walls, wedged into corners, resting on floors, where the scene renders much of the view sphere unreachable and the remaining views unequally informative. We introduce \textbf{SightCAD}, a framework for parametric CAD reconstruction that treats view feasibility as a first-class constraint. In this work we consider objects from standard CAD benchmarks embedded in realistic indoor scenes with physically derived visibility constraints over a discrete view sphere. A learned view selector must choose $K$ feasible views for a vision--language model (VLM) that generates executable CadQuery code, scored by geometric fidelity of the executed solid. Because reward arrives only after discrete view selection, autoregressive generation, and CAD-kernel execution, we propose a joint training paradigm in which the view selector and the CAD-generation VLM are trained together against this reward. The learned selection policy departs sharply from random, uniform, and coverage-greedy alternatives, outperforming surface-area maximization (SA-max) by up to $6.4$ Intersection-over-Union (IoU) points across budgets $K\in\{1,\dots,5\}$. The full system surpasses strong external baselines on scene-embedded, occluded multi-view renders of DeepCAD and Fusion360 objects ($+21$ and $+17$ effective-mIoU points over the best baseline, respectively), as well as on test-time domain-canonicalized real images from the industrial T-LESS benchmark and on both synthetic and real images from the MP6D industrial metal-parts benchmark, while producing the highest rate of executable programs of any method compared (invalid-code rate ${\leq}1.5\%$).

cs.CV↗

Capturing a Moving Target on Star Graphs by Two Communicating Mobile Agents

We study a problem of searching for a mobile target in an $m$-ray star graph, a natural generalization of linear search to multiple directions. The target is placed adversarially on one of the rays and may move with constant speed. We investigate a two-robot setting, where cooperation and communication play a central role. We study two communication models: the Face-to-Face (F2F) model, where robots communicate only upon meeting, and the Sender-Receiver (S/R) model, where communication is asymmetric. We focus on the {\em away model}, in which the target moves {\em away} from the origin with speed $v<1$. We design search strategies that minimize the competitive ratio and analyze the problem under three knowledge assumptions: \emph{NoDistance}, \emph{NoSpeed}, and \emph{NoKnowledge}. For each setting, we derive upper bounds on the competitive ratio and for some cases, we derive the lower bound. Our results highlight how the number of rays and the communication model influence the competitive ratio.

cs.DC↗

Reliability-Aware Future Conditioning for Temporally Robust Robot Manipulation

A generated video of a task the robot is about to perform is useful guidance only if it depicts the phase the robot is actually in. We show that temporal misalignment can turn a task-consistent generated future into actively harmful guidance. On CALVIN, a five-frame early shift nearly erases the benefit of generated futures, reducing success from 81.3% to 54.8% against 54.0% without futures; imposed timing shifts reduce it even further to 34.2%, 19.8 points below the future-free policy. We introduce Reliability-Aware Future Conditioning (RAFC), which treats this as a control problem rather than a generation problem. At every step, RAFC estimates how far to trust the received clip and which nearby temporal hypothesis to prefer, falling back toward a static branch when neither fits, and it learns both from task reward alone without shift labels or alignment supervision. RAFC sits on top of Future-Experience Conditioning (FEC), which builds the clip once from task grounding, a robot-free digital-twin rollout, and mask-free video diffusion. Under deliberately off-grid phase shifts and rate mismatch, RAFC substantially improves success under temporal mismatch. Candidate ensembling accounts for most of the recovery near alignment, while learned reliability adds a further 7.0 percentage points over uniform averaging of the identical candidate bank under off-grid shifts. The gain holds on the evaluated task sets and survives on a Franka under natural timing mismatch nobody imposed, where aggregate success rises from 26.7% to 56.7%. All resources will be made publicly available. https://future-condition.github.io/.

cs.RO↗

Stochastic Grouping Conformal Prediction for Effective Subgroup Reliability

Conformal prediction offers a distribution-free coverage guarantee, making it especially attractive for clinical applications. Standard conformal prediction, however, provides such guarantees only at the population level, and its prediction sets can exhibit coverage disparities across clinically important subgroups. A natural remedy is to calibrate within predefined groups. However, this can require access to sensitive subgroup attributes and is prone to a worst-group bottleneck: protecting the most difficult subgroup can inflate prediction sets for all, increasing cognitive burden on decision makers. To this end, we propose Stochastic Grouping Conformal Prediction (SGCP), a conformal framework for subgroup-reliable uncertainty quantification. It learns a stochastic grouping map that allows each sample to draw calibration information from others with similar calibration behavior, yielding a local score law that boosts reliability across subpopulations. We prove that SGCP retains the standard coverage guarantee. Experiments on synthetic and real-world benchmarks show that it consistently reduces subgroup coverage gaps while achieving smaller or comparable prediction set sizes relative to existing baselines.

cs.LG↗

From k+1 to k heads the descriptive trade-off is non-recursive

We prove that no recursive function can upper bound the increase in the size of description when a two-way deterministic finite automaton with k+1 heads is replaced by an equivalent two-way deterministic finite automaton with k heads. This is true for all k, and remains true if the automata are unary and/or nondeterministic.

cs.FL↗

MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement

Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of model intelligence by scaling RL compute. Prior to RL, we conduct mid-training on a broad multimodal corpus to provide ample exploration space, and build a solid infrastructure on the pretrained hybrid-SWA architecture to support subsequent scale-up. We scale RL compute along three dimensions: (1) larger batches and higher throughput, with an asynchronous training that consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M; (2) more diverse and complex environments, spanning code, general, visual, and cyber domains under a mixture of agent harnesses; and (3) more grader compute, via groupwise agentic grading that yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions. To keep training stable at scale, we freeze the MoE router and establish a multi-layer defense against reward hacking. We further build infrastructure for mixed-task agentic RL, including a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency. We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.

cs.CL↗

A Bregman proximal linearized ADMM for fractional programming with nonlinear coupling constraints

This paper considers a nonconvex fractional optimization problem with composite structure and nonlinear equality constraints. We reformulate the problem as a non-fractional min-max problem with corresponding optimal solutions. By linearizing the nonlinear constraint terms in the augmented Lagrangian and incorporating Bregman proximal linearization and a relaxation factor in $(0, 2)$ into the alternating direction method of multipliers (ADMM), we develop a Bregman proximal linearized ADMM with guaranteed subsequential convergence to a lifted critical point. Convergence of the entire sequence is established under the Kurdyka--Łojasiewicz (KL) property. We further derive convergence rates under either a Hölderian value proximity error bound condition or the KL property, with the corresponding exponent in $[0, 1)$. In particular, we establish superlinear convergence for exponents in $(0, 1/2)$, together with finite, linear, and sublinear convergence in the remaining cases. Numerical experiments demonstrate the effectiveness of the proposed algorithm.

math.OC↗

Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

A no-code fix resolves an invalid bug report by directing the user to change a setting, update to a version where the problem is already fixed, or adjust their workflow. Manually verifying whether a proposed no-code fix resolves the reported bug takes considerable developer time. This study proposes an automated, execution-based pipeline for evaluating the capability of large language models (LLMs) to generate no-code fixes in a real browser environment. We evaluate 322 no-code fixes generated by the 12 configurations released with the benchmark of a previous study, covering bug reports categorized as Faulty Configuration, Wrong Version, or External System & Dependency. An executor agent applies each fix by following its natural-language instructions, and an issue-specific checker determines whether the reported bug persists. We repeat the pipeline with three executors: two Computer-Use Agents, OpenCUA-72B and Claude Sonnet 5, and one multimodal agentic LLM, Meta's Muse Glimmer. Only 17.6% of the candidate issues could be set up and passed both sanity gates. Across the 322 fixes, 14.6% to 49.7% resolved the bug depending on the executor, and the strongest configuration, Claude Opus 4.6 in the Vanilla pipeline, resolved up to 74.1% of its fixes under Claude Sonnet 5. Changing only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors reached the same verdict on only 46.9% of the fixes. Compared with human execution, the executors matched the human consensus for 66.1% to 88.1% of the sampled fixes. Even under the best executor, fewer than half of the LLM-generated no-code fixes resolve the reported bug, so such fixes need verification before they reach users. Execution-based verification can provide this, but the measured capability depends strongly on the executor, which evaluations must report and control.

cs.SE↗

From Asymptotic to Designer-Assigned-Time Control: A Review of Stability Notions, Design Mechanisms, and Controller Architectures

Many control tasks require a target to be reached not only eventually but on time. Finite-, fixed-, predefined-, and prescribed-time control address this need, yet the labels are used loosely: a settling time that grows with the initial condition, a bound that holds for all initial conditions, a deadline chosen by the designer, and a limit attained only at the terminal instant often share one name. This review aims to make such claims comparable. It separates three questions that are often conflated: what temporal property is promised, which feature of the Lyapunov analysis produces it, and which controller architecture carries it into the closed loop. A recurring question is whether a guarantee proven for an idealized loop survives once observers, adaptation, disturbances, actuator limits, and digital implementation are included. The examined studies are therefore audited one by one, recording what each claims, which variable is actually certified, and how the result is validated. The audit shows that estimation and approximation errors often reduce an exact guarantee to a practical one, and that experimental evidence comes mostly from fast electromechanical systems. A scalar benchmark with two complementary tunings shows how much of the apparent difference between methods stems from conservative bounds, initial-condition dependence, and numerical tolerance, and how saturation and sampling can delay or remove a deadline. The review closes with open problems, among them deciding which deadlines a given plant can meet, combining certificates across interconnected subsystems, and preserving a guarantee through implementation.

eess.SY↗

Pythagorean pairs in boxes of prime powers

Let $p_i$ be the $i$th prime, and let $f(N)$ be the largest cardinality of a subset of \[ \left\{\prod_{i=1}^N p_i^{a_i}:0\le a_i<N,\ a_i\in\mathbb{Z}\right\} \] containing no two distinct legs of an integer right triangle. We prove \[ f(N) \ll N^N (\log\log N)^{-1/1296}. \]

math.NT↗

MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation

Research idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-objective nature, where ideas should balance novelty, plausibility and feasibility. While recent LLM-based approaches have made progress through carefully designed prompts or agent pipelines, they are constrained by predefined, static ideation workflows. To address this limitation, we propose MindFlow, a framework that explicitly formulates ideation as a graph-structured Flow in Mind, which is composed of modular thinking operators and modeled by a probabilistic mind supernet. Given a research topic, a controller dynamically samples thinking flows to generate candidate ideas. This open-ended problem is optimized using a tournament-based relative ranking, enabling the controller to progressively favor higher-quality thinking flows. We further introduce an evaluation protocol that jointly assesses problem finding and problem solving, going beyond title- or abstractonly judgments. Across diverse topics, MindFlow shows its superiority as an explicit, controllable and optimizable research idea innovator.

cs.AI↗

Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps

Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not explicitly model the reliability or pose informativeness, i.e., how strongly a correspondence constrains the camera pose, of individual correspondences. We show that the natural way to learn it -- using the per-correspondence error to constrain the learning of confidence -- suffers from a depth-dependent bias: small pixel errors reside predominantly at large depths and do not lead to high pose informativeness. Instead, in our method (CELL), we learn a per-correspondence confidence end-to-end through the pose, using a differentiable probabilistic PnP whose log-partition term encourages weight configurations that yield a better-constrained pose distribution. The learned confidence is used in three ways: (i) it reweights the flow supervision in a decoupled training scheme that keeps pose gradients out of the flow/edge backbone; (ii) it drives a probabilistic correspondence selection at test time; and (iii) together with the network's edge-probability it weights a final edge-matching refinement. We further design a partial-completion depth representation that adds signal without hallucinating across large gaps. On M3ED and DSEC our full system improves over the LEAR baseline on the majority of the evaluated sequences: it reduces the median translation error by up to 26.9% and the median rotation error by up to 15.8%.

cs.CV↗

Nonlinear Schrödinger equation in an exterior domain

We study the exact controllability of the cubic nonlinear Schrödinger equation in the exterior $Ω=\mathbb{R}^n\setminusΘ$ of a non-trapping obstacle $Θ$, with internal controls supported outside a large ball, and the corresponding boundary control problem on $Ω_0=B_{R_0}\setminusΘ$ with Dirichlet controls acting only on the outer sphere $\partial B_{R_0}$. Using the local smoothing effect of Burq, Gérard and Tzvetkov [9], we prove observability inequalities for the linear Schrödinger equation in the whole scale of Sobolev spaces $H^σ_D(Ω)$, $σ\in[-2,2]$, associated with the Dirichlet Laplacian; they require only the non-trapping assumption (no star-shapedness of $Θ$), and their proofs use no normal boundary traces. Combined with Strichartz-type estimates and a perturbation argument, this yields local exact controllability in $H^1_0(Ω)$ and in $H^2(Ω)\cap H^1_0(Ω)$, in dimensions two and three, around the zero solution and, for internal controls and under a unique continuation assumption, around nontrivial trajectories. The same method applies to the quintic equation, which is energy-critical in dimension three. Several open problems are discussed.

math.AP↗