arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Acceleration of Newton's method for nonlinear algebraic systems by preflattening and elimination techniques

In this paper, we propose a new approach called {\em preflattening} to speed up the resolution of nonlinear algebraic systems by Newton-type methods. Preflattening is meant to be for nonlinear systems what preconditioning is for linear ones. The idea is to transform the system into an equivalent one that is more favorable to Newton's method. To this end, we introduce the notion of {\em flatness}, which is intended to play for nonlinear systems a role similar to that of the condition number for linear systems. For a physical model of interest in the geosciences, namely the chemical equilibrium problem, we combine this strategy with the nonlinear elimination technique and demonstrate the value of the resulting approach through numerical simulations.

math.NA↗

Formally Certifying the Vertex Set of a Polyhedron Faster than Informal Enumeration

The computation of the vertices of a polyhedron described by a system of linear inequalities is a central problem in polyhedral computation. It is a fundamental step in the conversion between H-representations, by linear inequalities, and V-representations, by vertices and extreme rays. This operation plays an important role both in the study of polyhedra and their combinatorics in mathematics and in applications to software and system verification. We present a certificate-based approach for formally verifying the computation of the vertices of a polyhedron. Given an informally computed list of vertices, our method allows to certify in the proof assistant Rocq that the list is complete, or even exact. The cornerstone of the method is a new completeness criterion based on an abstract simplicial complex that generalizes a triangulation of the normal fan of the polyhedron. A significant advantage over previous approaches is that the usually expensive numerical computations are essentially reduced to membership tests to the polyhedron, while the other steps are cheap combinatorial tests. We implement the certification method and prove its correctness in the proof assistant Rocq. We experiment with it on a variety of polyhedra, including Birkhoff polytopes, cross-polytopes, cubes, permutahedra, hypersimplices, and high-dimensional polytopes involved in the disproof of the Hirsch conjecture. Our experiments show that certification with the Rocq-to-OCaml extracted checker is typically 1.5x to over 5x faster than vertex enumeration by the state-of-the-art informal C implementation lrslib of the reverse search method.

cs.LO↗

Puffin: Probabilistic Learning of Spatial Detail From Coarse Observations

High-resolution socioeconomic variables are important for applications such as urban planning, public health, disaster response, and resource allocation. In practice, however, these variables are often observed only at a coarse spatial resolution. We introduce Puffin, a probabilistic framework for statistical disaggregation that raises the resolution of coarse totals using high-resolution satellite embeddings as covariates. Instead of predicting a single value for each fine-resolution subregion, Puffin learns a probability distribution and is trained through an aggregation-aware likelihood. At inference, Puffin conditions these predictions on the observed regional total and splits it among the subregions. The resulting fine-scale estimates are consistent with the observed aggregate and come with calibrated uncertainty, without requiring fine-resolution labels for training. We evaluate Puffin on German and US census, employment, and election data across population, jobs, and other count variables, and study when statistical disaggregation succeeds or fails across regions, countries, and targets.

cs.LG↗

Not Every Change Is Necessary: Recoverable Drift in Large Language Model Unlearning

Machine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achieving the desired forgetting level can still leave collateral changes that impair non-target behavior. Our recovery comparisons suggest that some of these changes can be reversed while preserving observed forgetting performance. In this work, we present Propose-Then-Project Unlearning (PTP-U), a framework that combines targeted forgetting with the recovery of non-target capabilities. PTP-U first applies local analytic edits to weaken target knowledge associations, then aligns non-target output distributions with those of the original model to recover capabilities while maintaining fixed forgetting constraints. Both stages serve a common goal: satisfying the forgetting requirements while preserving fluent generation and performance on non-target tasks. Across three benchmarks, PTP-U achieves the strongest forgetting-retention trade-off among evaluated methods, reaching 81.22%-91.03% forgetting while preserving 94.20% non-target utility on average. At matched forgetting, PTP-U consistently retains higher non-target utility.

cs.CL↗

Bounds on the maximum number of limit cycles of piecewise linear Lienard systems II. The discontinuous case

This paper concerns the planar Liénard system \(\dot x=y-F(x),\ \dot y=-x\), where \(F(x)\) is a piecewise linear function with exactly \(m\) jump points and no fold points. Tonnelier [SIAM J. Appl. Math. 63 (2002)] conjectured that the maximum number of limit cycles of the system is \(2m\). The conjecture was confirmed for \(m=1\) in [J. Lond. Math. Soc. 113 (2026)], and a lower bound of \(2m\) for arbitrary \(m\) was established by Chen et al. [arXiv:2608.19542]. In this paper, we construct systems with at least \(4m-2\) hyperbolic crossing limit cycles for every positive integer \(m\), thereby disproving Tonnelier's conjecture for \(m\geq2\). We also establish the uniform upper bound \(2^{224(m+1)^2}\) for the total number of crossing, grazing, sliding, and composite limit cycles.

math.DS↗

Multi-period Mean-Expectile Portfolio Optimization under Wasserstein Ambiguity: Reformulation, Degeneracy and the Role of the Ground Metric

Expectiles are the only law-invariant risk measures that are both coherent and elicitable. Unlike Conditional Value-at-Risk (CVaR), however, they do not admit a Rockafellar--Uryasev representation that admits tractable Wasserstein reformulations. We address this difficulty by developing an envelope theorem for worst-case expectiles that characterizes the worst-case expectile over a Wasserstein ambiguity set as the unique root of a worst-case expectation with a two-piece affine integrand. This representation permits direct application of standard Wasserstein duality. Using this result, we reformulate a multi-period tri-level mean--expectile portfolio problem as four parametric linear programs with constraints. We establish four structural properties of the proposed model: an endogenously damped price of robustness, a decision-dependent critical radius beyond which the expectile tail component becomes inactive, exact recovery of the nominal model at zero ambiguity, and a characterization of how the Wasserstein ground metric determines whether the limiting portfolio becomes more concentrated or more diversified. Numerical experiments on 90 FTSE constituents over 3,341 out-of-sample trading days show that the expectile model outperforms a CVaR model matched on ambiguity set, radius, ground metric, and trade-off weight in all nine parameter cells---significantly so whenever the radius is non-trivial. The experiments further confirm the predicted degeneracy under the ground metric.

q-fin.CP↗

From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding

Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.

cs.MM↗

Random walk in a regenerating random environment

In this paper we consider random walks moving in a dynamic random environment that is completely resampled with probability $p$ after each step of the random walker. This resampling yields a straightforward regeneration structure and limit theorems, such as a law of large numbers with a limiting speed $v(p)$ that depends on $p$. We investigate several properties of $v(p)$ as a function of $p$, such as differentiability and monotonicity. Similarily, we study the limiting diffusivity $σ^2(p)$ in a central limit theorem.

math.PR↗

Event-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational Agents

For persistent and personalized conversational agents, memory systems can enable them to remember, update, and reason over long histories by storing past interactions and retrieving relevant information. Existing memory systems typically follow two paradigms: flat-structured memory and graph-based memory. The former is lightweight but leaves event relations and state updates implicit, while the latter explicitly models memory structure but incurs additional construction cost and introduces irrelevant relations over long histories. To address these limitations, we propose QGMem, a novel memory construction and activation framework motivated by human memory, in which experience is organized into events and query-relevant events are modeled by graph as working memory. QGMem converts long dialogue histories into event-indexed atomic memory units that preserve individual experiences and consolidates related units into dynamic memory traces that retain state trajectories and current states. When a query arrives, hybrid memory retrieval gathers complementary candidate memories, and query-aware reranking activates the most relevant units as a compact working memory. To expose relational dependencies in the working memory and support conflict-aware reasoning, QGMem organizes the working memory as a local graph, which is then encoded as a graph token and provided to the LLM together with the textual working memory to improve evidence utilization during answer generation. Experiments across six benchmarks validate the framework and show consistent gains in retrieval, multi-hop evidence composition, conflict resolution, and ultra-long dialogue reasoning with compact contexts and moderate inference cost.

cs.CL↗

Multipartite Entanglement Geometry in AdS/BCFT

Multipartite entanglement measures probe how quantum correlations are organized beyond what can be inferred from pairwise entropies. We study how this structure is reshaped by a physical boundary in a holographic BCFT. The end-of-the-world brane opens new relative-homology channels for multipartite entanglement, allowing connected bulk webs to reorganize and factorize at scales that need not coincide with any ordinary Ryu--Takayanagi transition. In static AdS$_3$/BCFT$_2$, this produces a characteristic hierarchy of multipartite phases and a nontrivial dependence of genuine multi-entropy on the boundary. In the thermofield-double state, genuinely connected webs can instead be supported through the Einstein--Rosen bridge, leading to growth, reconnection and saturation scales distinct from the usual Hartman--Maldacena transitions. These results show that multipartite entanglement is sensitive not only to extremal-surface areas but also to how bulk entanglement structures are allowed to join, split and terminate in the presence of a boundary.

hep-th↗

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Project Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone, we describe how to build a competitive pointwise decoder-only reranker using a simple two-step recipe comprising pre-training from scratch followed by supervised fine-tuning, starting only from commonly available datasets. Contrary to the dominant approach in the literature, we do not rely on existing open-weight backbones from third parties, and thus we are fully in control of model training, from end to end. We were able to accomplish the bulk of our experiments using no more than a handful of GPUs. This report articulates the importance and benefits of our approach, and we share artifacts that enable transparent, independent reproduction of all aspects of model training. Beyond data, code, and configurations that capture our efforts, we also release checkpoints for our family of Gaggle models, demonstrating the feasibility of our approach and providing a first step toward validating our broader thesis.

cs.IR↗

Neural Decoding as Cognitive Inference

The brain maintains stable cognition despite continuously changing neural activity. How to extract stable cognitive states from variable neural observations remains a central problem in neural decoding. Existing neural decoding methods map neural observations to predefined external labels based on the stimulus-response principle, often capturing recording-specific spurious correlations. Inspired by how the brain infers the world, and specifically by Bayesian brain theory, we recast neural decoding as cognitive inference constrained by brain-intrinsic priors, yielding high-level meta-neural semantic representations. In decoding experiments spanning five neural recording modalities and three cognitive domains (motor, perception and internal mentation), our cognitive inference method reorganized the geometry of neural observation representations, yielding meta-neural semantic representations that exhibited consistent geometric relationships across cognitive tasks and enabled the recovery of stable cognitive states from variable neural observations. Our work provides an account of how the brain maintains relatively stable cognition despite continual changes in the external environment. Cognitive stability is sustained through cognitive inference from changing neural activity, without requiring fixed neural activity patterns.

q-bio.NC↗

Revisiting Identity and Spectra Dispersion in Media-Bridged Time Series Forecasting: Linking Multivariate Signals and Narrative Flows

Media-bridged time series forecasting is expanding to encompass traditional "multivariate" and emerging "multimodal" (e.g., through textual assistance). Existing Time Series Forecasting (TSF) models still rely on paradigm-specific relation, fusion, and temporal modules, hindering a common forecasting backbone across numerical and pre-aligned narrative-flow settings. To explore this, we propose the Multimedia Identity-Aware Prism Network (MIDAPN), a unified spatiotemporal forecasting backbone based on media-general graph adaptation and automatic temporal learning: (1) Following media pre-alignment, our Multimedia Identity-Aware Graph (MIDAG) revisits identity through static essence, dynamic behavior, and latent commonality, inducing affinities that extend variable-specific dependencies across media. Contextual Identity Modulation (CIM) further refines discriminative aggregation. (2) We develop Spectral Prism Convolution (SPConv) to automatically perform hierarchical temporal analysis, balancing coarse trends and fine-grained details. Meanwhile, its Adaptive Search Guidance configures a scale-efficient architecture for temporal-dimension reconstruction. These decoupled yet synergistic components jointly address media identity disentanglement and temporal-scale mismatch. Comprehensive evaluations involving 16 SOTA TSF models across 13 "multivariate" and 12 "multimodal" datasets, alongside targeted long-context comparisons against 14 time series foundation models and fused pretrained language models, demonstrate MIDAPN's consistent superiority and broad shared backbone compatibility. The code is available at \href{https://github.com/leijieruilq/MIDAPN/tree/main}{https://github.com/MIDAPN}.

cs.CV↗

On Calabi-Yau varieties that are linear sections of Grassmannians of lines

In this paper we consider the $n-3$ dimensional linear section $X_Π$ of the Grassmannian $\mathbb G(1,n)$ of lines in $\mathbb P^n$ with a general linear space $Π$ of codimension $n+1$, with $n=2k+1$ odd and $n\geq 5$, that is a Calabi-Yau variety. Related to $X_Π$ there is a degree $k+1$ Pfaffian hypersurface $\mathcal C_Π$ that is the intersection of the dual of $\mathbb G(1,n)$ with $Π^\perp$. We prove that $\mathcal C_Π$ is rational and we provide a birational parametrization of it. Moreover we study in some detail the congruence of lines of $\mathbb P^n$ given by the centers of the linear complexes parametrised by the points of $\mathcal C_Π$.

math.AG↗

Covering radius of rank-metric codes via covering lifts and clubs

We introduce a generator-matrix-based geometric approach to the covering radius of $\mathbb F_{q^m}$-linear rank-metric codes. Starting from the $q$-system associated with the code, we define the notion of covering lift and show that the covering radius can be recovered from the hyperplane weights of such lifts. This yields an exact criterion, in terms of the length and effective length of the code, for the covering radius to attain its largest possible value, together with the upper bound $ρ(\mathcal{C})\le \min\{m,n\}-1$ in all remaining cases. We then specialize our approach to $1$-dimensional codes, for which the effective length coincides with the minimum rank distance. In this setting, codes attaining $ρ(\mathcal{C})= \min\{m,n\}-1$ are related to covering lifts that are clubs. More generally, we characterize this case through linear maps into suitable quotient spaces, obtaining equivalent interpretations in terms of generalised evasive subspaces and auxiliary matrix rank-metric codes. These results determine the covering radius for several boundary values of the minimum distance and for all $1$-dimensional codes with extension degree $m\in\{3,4,5\}$. For $m=6$ and every $q$ we provide constructions and, for $q\in\{2,3\}$, computational results showing that $1$-dimensional rank-metric codes with the same length and minimum distance can have different covering radii.

math.CO↗

Uniqueness of pressureless flow

We consider the pressureless Euler equations in one spatial dimension. These equations model the dynamics of particles on a line that interact only through perfectly inelastic collisions. The flow was first shown to be unique for given initial conditions by Huang and Wang, assuming the entropy and the strong initial continuity of energy conditions hold. In this paper, we prove an analogous uniqueness theorem using a Lagrangian interpretation. The key ingredients are showing that each solution satisfying the entropy condition and an initial kinetic energy bound has this particular type of Lagrangian interpretation and a monotonicity inequality involving conditional expectation.

math.AP↗

Box dimension profiles and intermediate dimensions

For a non-empty bounded set $E\subset\R^d$, the upper and lower box dimensions of its orthogonal projections are constant onto almost all $m$-dimensional subspaces. These almost-sure values are called the upper and lower $m$-box dimension profiles of $E$, respectively. We study their relationship with intermediate dimensions, which interpolate between Hausdorff and box dimensions by restricting the relative sizes of covering sets. We establish several equivalent characterizations of both quantities and obtain upper and lower bounds for box dimension profiles in terms of intermediate dimensions, with the lower bounds also involving the upper Assouad spectrum. We further derive explicit formulas for the box dimension profiles in terms of the box dimension and the initial growth rate of the intermediate dimensions, under a suitable condition on the quasi-Assouad dimension.

math.CA↗

A Gauge-Invariant Clustering Coefficient for Complex-Weighted Bipartite Networks

Structural measures such as the clustering coefficient and the average shortest-path length characterise how a network is organised. These measures are typically formulated for real, non-negative edge weights. A class of quantum and photonic architectures has complex edge weights instead, whose phases determine whether alternative routes interfere constructively or destructively. These architectures are also bipartite, so triangles are absent and the smallest closed cycle is a four-node square. Existing measures address complex weights and bipartite structure separately: bipartite clustering coefficients quantify clustering through four-node squares but do not contain phase information, while interferometric coefficients retain phase but are defined on triangles. In this work, we define a clustering coefficient for complex-weighted bipartite networks which reduces to the classical bipartite coefficient when the phases vanish, becomes negative when alternative routes cancel, and can be obtained for all nodes from a single sparse matrix product. We show that the phase accumulated around a square is the smallest gauge-invariant carrier of structural phase information in a bipartite network. The clustering coefficient factorises into a topological contribution and a phase contribution, making a bipartite Watts--Strogatz ensemble analytically tractable. For phases uniformly distributed on $[-Δ,Δ]$, the mean phase contribution is $(\sinΔ/Δ)^4$, independent of node, degree, and topology. We also derive closed-form expressions for the clustering coefficient, open-path visibility, and phase variance. We numerically verify these predictions and further show that phase disorder increases coherent distance.

physics.soc-ph↗