arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

Secure Aggregate Encryption with Identity-Based Authentication for Multi-Vendor FPGA Cloud Deployment

Secure deployment of FPGA bitstreams in heterogeneous multi-vendor cloud infrastructures requires scalable authorization, authenticated device access, and bitstream protection throughout the deployment lifecycle. Existing approaches typically address these functions through separate mechanisms, increasing coordination and key management requirements. This paper presents SAEID, a secure FPGA deployment framework that integrates aggregate authorization, certificate-free identity-based device authentication, and identity bound bitstream verification within a pairing based cryptographic framework, with symmetric cryptography used for session binding and AES 256 GCM based bitstream protection. Building on AgEID, an aggregate encryption scheme that enables individual decryption for authorized FPGA devices, SAEID extends the underlying framework to heterogeneous multi vendor deployments by supporting multiple FPGA vendors and IP providers, capability-aware authorization, and dynamic device membership. SAEID provides protection against unauthorized access to future deployments after device revocation and to prior deployments by newly enrolled devices, while retaining constant-size aggregate ciphertexts within each vendor domain and individual decryption. Experimental results demonstrate the practical performance of SAEID, with identity-based device authentication completing in approximately 4.35 ms and device membership updates requiring 4.95 to 5.09 micro seconds for device addition. The aggregate-encryption component exhibits scalable behavior with increasing device-set size. The complete SAEID software decryption path was also validated on a physical ZC702 Cortex A9 platform, requiring 191.126 ms. These results demonstrate scalable aggregate authorization with bounded authentication and dynamic-membership overhead for secure multi-vendor FPGA deployment.

cs.CR↗

Sharing of Gaussian tripartite steering in de-Sitter space

We investigate the redistribution and directional properties of Gaussian tripartite quantum steering in de-Sitter space within the continuous-variable framework. By expressing the Bunch-Davies vacuum in terms of the open-chart modes through a Bogoliubov transformation, we derive the covariance matrix of the resulting multipartite Gaussian state and analyze both one-to-two and two-to-one steering configurations. We find that de-Sitter curvature generally suppresses the steering initially shared among the accessible modes, but its influence depends strongly on the steering direction. In particular, while some steering configurations undergo sudden death at finite curvature, two-to-one steering toward a nongravitational mode remains finite throughout the parameter range considered, demonstrating a robustness of tripartite steering absent in its bipartite counterpart. Remarkably, the maximum steering asymmetry occurs at the critical point at which steering in one direction undergoes sudden death, marking the transition from two-way to one-way steering. We further show that the de-Sitter gravitational field can generate tripartite steering involving modes separated by the cosmological horizon. Depending on the steering partition, this curvature-induced steering exhibits persistent one-way or asymmetric two-way behavior, as well as a sudden-birth transition from one-way to two-way steering. These results demonstrate that de-Sitter curvature not only degrades preexisting quantum correlations but also redistributes and generates directional multipartite steering across causally disconnected regions.

gr-qc↗

Relative Patch Response Learning for Generalizable AI-Generated Image Detection

Generative models can now synthesize highly realistic images, simultaneously increasing the risks of misinformation and visual forgery. Therefore, detecting AI-generated images becomes more essential, and a reliable detector must generalize to unseen generators and stay robust to unseen perturbations in the wild. Existing detectors are typically trained on either independently collected real and generated images or aligned real-generated pairs designed to mitigate content bias. Building on aligned pairs, recent methods form a mixed view by replacing some patches of the real image with their generated counterparts. However, we find that self-attention lets real and generated patches interact, so the feature of each patch no longer reflects its own source alone. This contextual shift makes a per-patch source label an imprecise target. To this end, we propose Relative Patch Response Learning (PRL). Instead of labeling each patch, PRL compares the same patch across two mixed views of an aligned pair and learns from its patch response, the change of its score between the views. (i) To give precise supervision under the contextual shift, a relative response objective measures the responses of source-changed patches against those of source-unchanged patches, which respond to the shift alone. (ii) To provide a reliable reference for the shift, a reference coherence objective keeps each group of source-unchanged patches moving as a whole. (iii) Since the two views contain different amounts of generated content, an area ranking objective asks the view with the larger generated area to have a higher mean patch score. Extensive experiments demonstrate the superior performance of PRL, which surpasses the best prior methods by 4.3% and 5.9% in average balanced accuracy across eight standard and three in-the-wild benchmarks, respectively.

cs.CV↗

How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices

Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, evidence sources, evaluators, and scoring procedures to explain the capabilities assessed by different designs. Our synthesis highlights three recurring lessons: output checks, process checks, and human studies provide complementary information; evaluator calibration is specific to the property being assessed; and resource budgets and attempt selection are integral to interpreting performance comparisons. We identify diagnostic evaluation designs and documented gaps in supporting evidence, and translate these comparisons into reporting and audit recommendations for specific evaluation settings. The survey helps readers navigate existing evaluations, select appropriate benchmarks, and design subsequent studies.

cs.AI↗

Benchmarking Universal Machine-Learning Interatomic Potentials for Temperature-Dependent Elasticity of Binary and High-Entropy Refractory Carbides

Accurate prediction of temperature-dependent elasticity is important for assessing refractory carbides under high-temperature conditions, but the computational cost of ab initio molecular dynamics (AIMD) limits systematic investigations across compositions and temperatures. Universal machine-learning interatomic potentials (uMLIPs) offer an efficient alternative, yet their accuracy for this task remains insufficiently established. Here, we benchmark nine uMLIPs against consistent AIMD reference data for five binary and two high-entropy (HE) carbides between 300 and 1200 K. Elastic constants and the corresponding bulk, shear, and Young's moduli are obtained using stress-strain molecular dynamics, explicitly sampling thermal atomic motion and anharmonic effects beyond thermal expansion alone. We assess absolute elastic properties and normalized thermal softening separately. MACE-MH-1 achieves the lowest overall mean absolute percentage error (5.7%). MACE-MH-1 and DPA4-Mini reproduce thermal softening most accurately, with mean deviations of 2.5 and 2.3 percentage points, respectively. All models generally underestimate stiffness, with larger equilibrium volumes relative to the AIMD reference likely contributing to this trend for most models. C12 exhibits the largest model-dependent errors. The HE carbides are described with accuracy comparable to that of the binary carbides, indicating no apparent accuracy penalty from chemical complexity. Accuracy instead varies with transition-metal composition, with group-V carbides, particularly TaC, presenting the greatest challenge. These results identify promising pretrained models for finite-temperature elasticity in carbides and show why accurate absolute stiffness and thermal softening must be assessed independently.

cond-mat.mtrl-sci↗

Coarse ellipticity and De Giorgi-Nash-Moser theory in the optimal range

We extend the theory of De Giorgi-Nash-Moser to elliptic equations with symmetric coefficients $\mathbf{a}(x)$ which are possibly degenerate and unbounded but satisfy a coarse ellipticity condition. We prove local upper bounds for weak subsolutions, a weak Harnack inequality for nonnegative supersolutions, and a Harnack inequality for nonnegative solutions. The coarse ellipticity hypothesis requires spatial moments of coarse-grained matrices, suitably discounted and summed across scales, to be finite. In particular, it holds if, for some $α,β\geq0$ and $1 0$. For $α=β=0$, it corresponds to the results of Bella and Schäffner [8].

math.AP↗

From GeV to PeV: Orbital Modulation and Wind-Regulated Ultrahigh-energy Emission in Cygnus~X-3

Recent LHAASO observations reveal ultrahigh-energy (UHE) gamma-ray emission from Cygnus~X-3 with evidence for orbital modulation, providing a rare opportunity to study particle acceleration and transport on binary scales. We develop a common binary--jet model that jointly explains the contemporaneous GeV--PeV spectra and orbital light curves. The GeV emission is produced by anisotropic inverse-Compton scattering of stellar photons by relativistic electrons, while multi-PeV protons escaping from the inner region produce TeV--PeV gamma rays through photohadronic and hadronuclear interactions. The GeV orbital modulation constrains the binary--jet geometry, allowing the UHE observations to probe the escaping proton population and its interaction with the circum-binary environment. We find that the UHE emission is strongly regulated by the accretion-driven wind structure. In the present state, a fast disk wind displaces much of the denser Wolf--Rayet (WR) wind from the escaping-proton trajectories, suppressing hadronuclear emission. Changes in the strength, collimation, or orientation of the disk wind can expose the escaping protons to a much denser WR wind and enhance the PeV gamma-ray luminosity by orders of magnitude. This provides a possible connection to the much brighter, though unconfirmed, PeV activity reported in the 1980s, and links UHE gamma-ray activity to the accretion and outflow state of the binary.

astro-ph.HE↗

Beyond special quasirandom structures: free energies from energy cumulants

Predicting the finite-temperature stability of a disordered phase requires its free energy. It is traditionally approximated by combining the energy of a special quasirandom structure (a small cell mimicking a random alloy) with the ideal configurational entropy. This approximation neglects the short-range order that develops on cooling. In this study, we show that the conventional approximation is the first term of an infinite expansion of the free energy in energy cumulants. Higher-order cumulants are estimated from energies of random arrangements, without Monte Carlo or molecular dynamics. In a nine-element refractory system, adding the second cumulant reduced free energy errors by about an order of magnitude. The expansion gives phase diagrams with order--disorder transitions, defect concentrations, and short-range order. To demonstrate its utility, we used a foundation interatomic potential to screen lithium-excess rocksalt oxides across 26 elements for synthesizable disordered phases rich in the Li$_4$ clusters needed for lithium percolation. The approach applies to any lattice and chemistry, and brings disordered phases within reach of routine screening.

cond-mat.mtrl-sci↗

legoESM: a modular, differentiable, multiscale, AI-ready Earth system model built with AI agents

Earth system models (ESMs) have grown tremendously in realism, yet key uncertainties persist in the climate response to greenhouse-gas forcing, particularly due to cloud radiative feedbacks. In addition, their software architecture was not designed for accelerator hardware or modern artificial intelligence (AI). Here we present legoESM, a composable, differentiable, multiscale ESM written in JAX. It builds on decades of community-developed parameterizations and numerical methods, recast in a unified framework by AI coding agents under a human-specified scientific contract and verified through benchmarking. Dynamical cores, physics schemes, grids, complexity levels and components are swappable like building blocks, and can use conventional physics or machine-learned emulators. A single code base spans metre-scale large-eddy simulation to global simulations and weather to climate. End-to-end differentiability enables gradient-based calibration, variational data assimilation and online training. legoESM modular architecture enables systematic evaluation of diverse model variants to explore structural uncertainty and test hypotheses. legoESM produces realistic simulations across scales, reduces land-surface temperature bias through gradient-based calibration, and scales efficiently on GPUs to kilometer-scale simulations. It offers an open, community infrastructure for hypothesis testing, research and teaching in Earth sciences and a template for multiscale physical systems.

physics.ao-ph↗

STEMMA: Song-to-Stem Multi-Audio Reasoning for Large Audio Language Models

Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.

cs.SD↗

On ruled normal surfaces associated with rectifying curves

In this paper we study the geometrical entities as Gauss and mean curvatures attached to the ruled normal surfaces associated with rectifying curves. A particular and special example of rectifying curves is also investigated, via its attached ruled normal surface. Our main result is the proof of the non-existence of rectifying curves that are simultaneously Bertrand curves.

math.DG↗

Breaking linearity in the sum-rank metric: MSRD codes from switched $σ$-rational normal curves

We construct a family of scalar-closed, non-additive maximum sum-rank distance codes (MSRD) by switching norm components of a $σ$-rational normal curve and lifting the resulting point set through a linearized Reed--Solomon syndrome map. We characterize the admissible switches by a multiplicative stability condition on the norm classes. In particular, for $1 \leq n_i \leq m$, $N=n_1+\ldots+n_{\ell}$ and $2\leqδ\leq N-1$, every partition of $\mathbb{F}_q^*$ satisfying a specific condition, referred to as Condition $(\diamond)$, yields a code in $\mathbb{F}_{q^m}^{N}$ of size $q^{m(N-δ+1)}$ and minimum sum-rank distance $δ$; the code is non-additive whenever a curve component is retained. For $\ell\geq2$ these appear to be the first non-additive MSRD codes in the literature, all previously known families with more than one block being $\mathbb{F}_q$-linear. Moreover, every code of the family has the same sum-rank weight distribution as an $\mathbb{F}_{q^m}$-linear MSRD code with the same parameters, although the geometry of the switched set distinguishes it from linearized Reed--Solomon codes. In the single-block case, an explicit rank isometry identifies the construction with the cone codes of Durante, Grimaldi and Longobardi.

cs.IT↗

Universal weighted sampling of positive and normal operator orbits

We study common weighted sampling designs for operator systems that are exactly observable at nonnegative integer times. For positive operators with original frame condition number at most $κ$, we construct a common integer design with arbitrarily small relative Gramian error and $O(\log R)$ distinct times up to $R$. The optimal logarithmic counting coefficient for preserving the frame property satisfies $C^*(κ)\simπ^{-2}\logκ$ as $κ\to\infty$. Real and integer times give the same infimum, and we determine its asymptotics as $κ\downarrow1$. A finite bound on the sampled condition number strictly increases this coefficient. For tight frames, we determine the exact optimum under this constraint and construct an integer design attaining it. Without this constraint, the infimum is never attained. For positive operators, a design valid for every finite $κ$ requires counting faster than $\log R$, with arbitrarily slow divergence of the ratio. For unrestricted normal operators and fixed $κ>1$, at least linear counting is necessary. Under the logarithmic sector condition $|\argλ|\le M(-\log|λ|)$ with fixed $M\ge0$, logarithmic sampling is possible with a sharp leading coefficient. For fixed Borel spectral sets satisfying a radial density condition, we characterize logarithmic sampling and determine the minimum upper counting exponent, attained by an integer design. Counterexamples show that the radial condition cannot be omitted.

math.FA↗

Evaluating Exact Output and Checkpoint-State Prediction in Real Programs

We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when missing responses count as wrong. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Source-clustered analyses preserve this accuracy gap, while adjusted Python models give no evidence of a positive incremental association between cumulative state load and error. Changed inputs and checkpoint tasks alter several factors together, so the gaps do not isolate trace length or an internal state-tracking mechanism. The benchmark exposes errors hidden by short-output scores alone.

cs.SE↗

Compressional heating above 1 keV on the LM26 magnetized target fusion machine

The LM26 device has achieved a peak electron temperature of $T_e = 1180 \pm 65$ eV as measured by filtered X-ray diodes, supported by adjacent Thomson scattering measurements of $T_e = 1090 \pm 40$ eV, with a deuterium ion temperature of $T_i = 459 \pm 33$ eV as inferred from neutron yield. LM26 compresses a spherical tokamak plasma inside an initially 1.76 m diameter solid lithium liner imploded by theta-pinch coils. Radial compression by a factor of 2.75 increases $T_e$ over fivefold, while $T_i$ increases as much as twofold. The integrated observation of magnetic flux compression, electron heating, ion heating, and neutron production in a magnetized plasma compressed by a large lithium liner is an important milestone for magnetized target fusion

physics.plasm-ph↗

Tracking High-Impact Researchers Through Highly Cited Tails

It is well recognised that field weighted citations are approximately lognormally distributed. This skewness reflects competition for attention, whereby most publications are gradually forgotten and only a minority continue to be recalled years later. Exploiting this observation, we construct a researcher-level indicator based on authorship of top 10% cited papers in All Science Journal Code (ASJC) bins and apply it to publications affiliated with Irish institutions over a typical funding cycle. Restricting attention largely to original research papers, we find that researchers authoring strictly more than one ASJC top-decile paper in a given year exhibit persistent participation in the highly cited tail of the literature in subsequent years, thus the simple track record indicator possesses predictive value. In contrast, the median researcher in cohorts of Research Ireland and European Research Council awardees selected through qualitative peer review either authors no top-decile paper or less than the median researcher in our ranking. These results demonstrate that research excellence as judged by qualitative peer review and influence on the research literature identify substantially different cohorts of researchers. In the absence of evidence, funding bodies and universities may conflate these two. While research reputation may be enhanced through many channels, highly cited papers constitute the most widely observable manifestation of international scholarly influence. Our analysis indicates that literature influence may be improved through greater use of data in track record pre-screening prior to qualitative peer review.

physics.soc-ph↗

Femtosecond Three-Dimensional Imaging of Single-Protein with Hard X-ray Laser

The extremely intense pulses of X-ray free-electron lasers (XFELs) have enabled imaging of radiation-sensitive samples, such as macromolecular microcrystals, beyond radiation damage limits. These sources have the potential to deliver biomolecular single-particle imaging, similar to cryo-electron microscopy but without the need for cryo-fixation and with temporal resolution from femtoseconds to milliseconds. While this possibility was recognized before XFELs were built, the biological single-particle imaging work-flow has previously only been demonstrated on large virus particles. Based on decades of improvements in X-ray beam focusing, particle delivery, diffraction detection, and advanced analysis, here we demonstrate imaging of a single molecular complex, the giant-hemoglobin erythrocruorin (Ery) with X-ray laser pulses. Two-dimensional classes of diffraction patterns could be reconstructed to 15 Angstrom resolution, and 3D images to approximately 20 Angstrom, while the 3D merged intensity in reciprocal space extended beyond 20 Angstrom. The resolution discrepancy is likely due to heterogeneity caused by gas-phase compaction of the complexes. With increased throughput, this approach could be used to reveal in-situ structural details during mass spectrometry studies of biomolecules, while improvements in sample delivery may provide ultrafast snapshot imaging of biological single-particles in their native-state beyond the limitations of radiation damage.

physics.bio-ph↗

A Security Meta-Model for Retrieval-Augmented Generation Systems

Retrieval-Augmented Generation (RAG) systems extend large language models (LLMs) with external knowledge through a multi-stage pipeline. While this architecture can improve the factual grounding of generated answers, it introduces structural attack surfaces that extend beyond those of standalone LLMs. In this paper, we introduce a security meta-model that captures explicit causal relationships between RAG surfaces, attacks, weaknesses, risks, and CIA impact (Confidentiality, Integrity, Availability). Its purpose is to provide security engineers with a structured and user-friendly framework for gathering and assessing the risks, weaknesses, and mitigations relevant to their RAG deployment. We designed the meta-model through an iterative, structured analysis of 43~publications (2023--2026) and instantiated it as a catalog populated with the security threats and remediations reported in the literature. Filtering the catalog according to a deployment configuration produces a risk profile containing the risks applicable to that deployment. An interactive web visualizer lets users navigate the catalog as a graph, follow causal chains, and explore stakeholder-specific views. Analysis of the catalog revealed a persistent imbalance between attack-focused and defense-focused research, a concentration of threats at ingestion, and coverage gaps affecting output integrity. Coverage is assessed against the OWASP LLM Top~10, and operational applicability is illustrated across textual, graph-based, and multimodal RAG configurations.

cs.CR↗