arXiv ScienceSearch

subject

cs.CE

cs.CE: explore 113 source-linked works published from 2021 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-15. Counts describe this index, not the complete source archives.

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.

cs.CL

Physics-Informed Neural Networks for Depth-Averaged Granular Avalanche Dynamics on Curved Topography

Physics-informed neural networks (PINNs) provide a mesh-free framework for solving governing equations, but their application to granular avalanche dynamics over curved terrain remains largely unexplored. This study extends a depth-averaged PINN formulation based on the Savage-Hutter equations to an exponentially curved chute with spatially varying inclination and a strain-rate-dependent Mohr-Coulomb earth-pressure closure. The model is validated against measured front- and rear-edge trajectories from a laboratory granular-avalanche experiment, with selected observations withheld from training. A staged temporal curriculum proved essential for accurate prediction, reducing the held-out trajectory error by approximately two orders of magnitude compared with training over the full time domain from the outset. Sparse-data experiments further showed that observation placement was more influential than observation number within the configurations tested. Four observations bracketing the transition from acceleration to deceleration achieved nearly the same accuracy as the eight-observation reference configuration, whereas observations clustered at early or late times performed poorly. The results demonstrate the importance of both training strategy and informative data placement when applying PINNs to granular flows over curved topography.

cond-mat.soft

Performance evaluation of variational quantum eigensolver and quantum dynamics algorithms on the advection-diffusion equation

Near-term quantum algorithms are a promising route to solving partial differential equations, but gauging their true potential requires separating algorithmic performance from sampling and hardware noise. We benchmark a ground-state variational quantum eigensolver (VQE), cast as a variational quantum linear solver, against the Trotterization, variational quantum imaginary time evolution, and adaptive variational quantum dynamics simulation methods applied to the one-dimensional advection-diffusion equation in the recent quantum-dynamics study by Alipanah et al. [Phys. Rev. Res. 7, 043318 (2025)] at matched grid and problem size. On a noiseless state-vector simulator the $N=4$ VQE drives the final-time infidelity to a numerical floor ($\sim\!10^{-14}$) once the depth reaches $L\approx5$, an \emph{algorithmic ceiling} set by exact expectation values. Evaluating the same solver with a finite number $S$ of measurement shots, still without hardware noise, makes the infidelity sampling limited, following $1-f\approx c/S$ (a best-case readout-sampling estimate, with the solution's signs assumed known), providing a regime-matched comparison with the shot-based emulator of Alipanah \emph{et al.}\ and explaining the gap to their noisy hardware runs ($>10^{-1}$). The benchmark thus decomposes the near-term error budget into algorithmic, sampling, and hardware contributions, with a matched-depth resource comparison. The formulation applies without modification across $N=4,5,6$ qubits and to a two-dimensional (eight-qubit, $16\times16$) problem evolved to $t=1$, where the state-vector VQE holds a $\sim\!10^{-7}$ algorithmic-ceiling infidelity against the sampling-limited $\sim\!10^{-5}$ of the corresponding shot-based simulation, a difference of measurement regime rather than algorithmic superiority.

quant-ph

On the Application of Hybrid Mixed Domain Decomposition Methods to Permanent Magnet Synchronous Machines

In this work, we study the application of a hybrid mixed domain decomposition(HMDD) method for the rotor-stator coupling of a permanent magnet synchronous machine. For this, we derive a variational formulation on the electric machine inspired by hybridized discontinuous Galerkin methods using a mixed magnetostatics problem, an affine material law and boundary conditions respecting the symmetry of the motor. We are then able to locate the resulting finite element method within the HMDD framework. This enables us naturally to transfer the well-posedness results and error estimates for the HMDD method to the finite element method considered in this work. Lastly, as a proof of concept, we consider an academic example and compare the resulting magnetic flux density and potential lines to their counterparts obtained by a well-established in-house code using iso-geometric analysis.

cs.CE

Do simulated agents move like real people?

Human mobility is increasingly represented using synthetic populations that offer scalable alternatives when individual-level observations are unavailable or sensitive. Yet validation typically emphasizes aggregate statistics, which can obscure whether simulated agents traverse transportation networks in ways that resemble real travelers. Here, we develop a path-centric framework that combines direct path-level comparisons with higher-order network models to compare observed and simulated mobility on a shared metropolitan road network. Observed and simulated paths share broad statistical regularities and short-range memory. Beyond these similarities, however, simulated mobility underrepresents long paths, exhibits greater redundancy among long route sequences, covers a smaller and partly different portion of the network, and is more predictable overall. These discrepancies show that agreement in aggregate mobility patterns does not imply fidelity in how travelers move through the underlying infrastructure. Higher-order path analysis therefore offers a framework for validating synthetic mobility at the spatial and sequential scales relevant to scientific inference, urban planning, and policy.

cs.CE

Disciplined Bilevel Programming

Bilevel optimization provides a natural modeling language for hierarchical decision problems. However, applying existing numerical solvers usually requires substantial manual analysis and reformulation. In this paper, we introduce disciplined bilevel programming (DBLP), a symbolic framework that allows users to specify and solve optimistic bilevel problems in a high-level, human-readable way that is close to the mathematical formulation. For problems with a disciplined nonlinear upper problem and a convex lower problem satisfying the disciplined parameterized programming rules, DBLP automatically canonicalizes the lower problem into conic form and constructs an equivalent single-level reformulation using the conic Karush-Kuhn-Tucker conditions. We relax the resulting complementarity constraint and use a gap continuation procedure to approximately solve a sequence of smooth nonlinear problems. We implement DBLP in the open-source Python package BLVPY, an extension of CVXPY for bilevel programming. We demonstrate the modeling and solution capabilities of BLVPY on a range of bilevel optimization problems from several application domains. The proposed framework and implementation allow users to specify and solve bilevel optimization problems within a few lines of code, without prior expertise in bilevel modeling and numerical optimization.

math.OC

HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields

Reconstructing oscillatory wave fields from scattered sensors is a severely underdetermined inverse problem. Beyond the challenges of general physical-field reconstruction, wave responses are complex-valued, frequency-sensitive, and highly oscillatory, while costly simulation and sensing often leave only extreme-sparse observations. Existing low-rank, operator, and diffusion approaches are largely designed for real-valued, smoother fields; dense pixel-space diffusion is particularly inefficient for oscillatory complex fields and difficult to scale to 3D. We propose HarmoCore, which places a generative prior in a compact, continuous, and structured wave-field latent. HarmoCore represents joint real--imaginary channels with Functional Tucker cores over shared continuous spatial bases, learns a frequency-conditioned core diffusion prior, and performs Diffusion Posterior Sampling directly in core space. At fixed sensor coordinates, the multilinear decoder induces an explicit likelihood guidance operator, avoiding dense pixel-space correction. Optional target-equation residual guidance further promotes physical consistency. Experiments on 2D Helmholtz, 2D synthetic wave fields, and 3D Helmholtz show substantial gains under 1%--2% sensing while remaining practical in three dimensions.

cs.LG

Aerodynamic Shape Design Space Exploration with Deep Latent Diffusion Model

We propose DiffGeo, a latent space diffusion-based generative framework for aerodynamic design space exploration under extreme data scarcity. DiffGeo combines a learned latent space model for automatic shape parameterization, with a diffusion sampler to directly generate novel, geometry-valid and controllable designs. We validate the approach on a series of case studies: (i) a 2D airfoil generation benchmark, where DiffGeo's latent diffusion model is compared against GAN- and VAE-based baselines in terms of sample quality, diversity and constraint adherence under limited data; (ii) integration into a surrogate-based optimization pipeline, where DiffGeo's conditional sampling produces task-informed airfoil data that improve both surrogate modeling and optimization performance; and (iii) extension to 3D turbomachinery blade prototyping, where DiffGeo generates realistic and high-performance blade geometries from a small set of reference designs. Throughout these investigations, DiffGeo achieves high-quality and diverse shape generation with at least an order of magnitude less data than alternatives, decouples geometry representation from design targets for flexible reuse, and seamlessly incorporates complex design constraints via energy-based conditioning. These capabilities demonstrate DiffGeo's potential to enhance early-stage design by automating design space exploration--improving efficiency, expanding design diversity and embedding engineering knowledge through controllable guidance.

cs.CE

VPID: An Integrated Framework for Vulnerability Prioritization and Intrusion Detection in Enterprise Networks

Small enterprises face increasingly serious threats to their internal networks but often lack the financial resources, computing capacity, and specialist staff required to deploy resource intensive security platforms. This paper designs and implements VPID, a lightweight framework for vulnerability prioritization and intrusion detection that consists of two principal modules: controlled vulnerability validation and intelligent intrusion defense. The first module uses OpenVAS for asset mapping and vulnerability identification, applies a decision tree to prioritize vulnerabilities, and employs a rule engine to generate targeted validation payloads. The second module captures network traffic using Scapy, analyzes it through a detection pipeline that combines a decision tree with multinomial Naive Bayes, verifies traffic assessed as high risk using Snort rules, and performs blocking and alerting through iptables. The evaluation uses 550,000 network flow samples containing normal and attack traffic for detector training, together with 15,000 labeled vulnerability records. On the vulnerability ranking test set, the decision tree achieves a precision of 91.8%, a recall of 89.5%, and an F1 score of 90.6%. On an independent test set containing 55,000 traffic samples, the combined detection pipeline achieves a precision of 94.5%, a recall of 88.3%, and an F1 score of 91.3%, while maintaining a false positive rate below 1.5%.

cs.CE

Don't You Know, Pump it Up! Investigating Cryptocurrency Manipulation in Telegram-Driven Activity

Telegram plays a pivotal role in cryptocurrency communication and has been repeatedly associated with coordinated schemes, such as pump-and-dump manipulation. However, existing studies typically focus on known manipulation chats or a limited set of cryptocurrencies, leaving open the question of how Telegram is leveraged for mass promotional activity (shilling) at scale. Moving beyond these limitations, this work analyzes the interplay between information flows and market activity across public Telegram channels. To this end, we propose a scalable framework that (i) classifies crypto-related messages using a fine-tuned encoder model to filter semantic noise, (ii) detects anomalous spikes in cryptocurrency mentions via adaptive thresholding, and (iii) validates temporal associations between social bursts and market movements using quasi-experimental econometric methods (RDD and DiD). We apply this framework to one year of public Telegram data (14,499 channels and over 20 million messages) aligned with transaction data for more than 17,000 cryptocurrencies. Our analysis identifies 47 events consistent with potential pump-and-dump activity and 73 sustained market reactions, showing that manipulative signals are characterized by extreme temporal synchronization and precede price movements by seconds. Notably, psycholinguistic analysis reveals that pump-and-dump messages are linguistically indistinguishable from organic discussions, highlighting the limits of text-based detection alone. Finally, we estimate the cumulative financial volume of detected pump-and-dump events to exceed $200 million and release a public cryptocurrency dictionary and a fine-tuned classifier to support future research.

cs.SI

Connectome-Based Modelling Reveals Orientation Maps in the Drosophila Optic Lobe

The ability to extract oriented edges from visual input is a core computation across animal vision systems. Orientation maps, long associated with the layered architecture of the mammalian visual cortex, systematically organise neurons by their preferred edge orientation. Despite lacking cortical structures, the Drosophila melanogaster brain contains feature-selective neurons and exhibits complex visual detection capacity, raising the question of whether map-like vision representations can emerge without cortical infrastructure. We integrate a complete fruit fly brain connectome with biologically grounded spiking neuron models to simulate neuroprocessing in the fly visual system. By driving the network with oriented stimuli and analysing downstream responses, we show that coherent orientation maps can emerge from purely connectome-constrained dynamics. These results suggest that species of independent origin could evolve similar visual structures.

cs.CE

Risk-Sensitive Reward Composition for Conditional GFlowNets

Generative Flow Networks (GFlowNets) for structure-based drug design condition on one rigid protein structure. A flexible target holds several distinct structural shapes, its conformations, each occupied for a fraction of the simulation time. Scoring a candidate against all of them raises an open question: how do K scores become one reward? The designer cannot choose arbitrarily. Populations carry simulation error, and biology dictates which conformations are deal-breakers, so a candidate that fails one is disqualified, not merely ranked lower. No standard rule captures this. We compose the reward from a conditional value-at-risk (CVaR), a worst-case score rule, and an ambiguity radius expressing distrust in the stated weights. Together, these define a family of targets, amortised by a single conditional GFlowNet. We answer whether such a sampler can be trained on fully enumerable synthetic worlds, where every error is exact rather than estimated. Pricing the tail rather than averaging moves 2-10 times more mass to candidates that pass every conformation. One network covers the family to within 0.37-2.7x the error of a perfect sampler. An exact-KL oracle, a copy trained on the true target, shows if a shortfall is the optimiser's or the architecture's. When good candidates are rare, exploration decides: injecting unseen states finds 0.987-1.000 of good regions, while reweighting visited finds 0.35-0.76.

cs.CE

TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding

Understanding the economic intent of Ethereum transactions is critical for user safety, yet current tools expose only raw on-chain data or surface-level intent, leading to widespread ``blind signing'' (approving transactions without understanding them). Through interviews with 16 Web3 users, we find that effective explanations should be structured, risk-aware, and grounded at the token-flow level. Motivated by these findings, we formulate TxSum, a new domain-grounded NLP task for DeFi transaction explanation, and construct a dataset of 187 complex Ethereum transactions with 2,375 token-flow annotations and transaction-level summaries. We further introduce MATEX, a grounded multi-agent framework for high-stakes transaction explanation. It selectively retrieves external knowledge under uncertainty and audits explanations against raw traces to improve token-flow-level factual consistency. MATEX achieves the strongest overall explanation quality, especially on micro-level factuality and intent quality. It improves user comprehension on complex transactions from 52.9% to 76.5% over the strongest baseline and raises malicious-transaction rejection from 36.0% to 88.0%, while maintaining a low false-rejection rate on benign transactions.

cs.CE

M-Tensor Formalism: A Non-iterative High Dimensional Least Squares Regression for Nonlinear Models with Scarce Data

We present a multilinear regression framework based on tensor algebra tailored to high-dimensional contexts where data is scarce. We exploit algebraic properties of a partial tensor product, namely the m-tensor product, to leverage structured equations with separated variables. The proposed method combines kernel properties along with tensor algebra to prevent explicit construction of the exponentially large feature space and tackle approximations up to hundreds of parameters while avoiding the fixed-point strategy. This is achieved by only ever employing the regression operator in a factorized form. We present this formalism along with different regularization techniques suited for low amount of data with a high number of parameters while preserving well-known matrix-based properties. We demonstrate complexity scaling on a general benchmark to show robustness for engineering problems and ease of implementation.

cs.CE

A non-intrusive approach to index-aware learning

We present a non-intrusive version of the index-aware learning framework introduced in arXiv:2309.00958. Index-aware learning itself is an approach for learning the time and parameter dependent solutions of differential-algebraic equations (DAEs), in particular those describing electric circuits. A central feature of the approach is that it ensures the learned solutions to fulfill the inherent constraints of the DAE, such as e.g. Kirchhoff's laws in the case of electric circuits. This is achieved by leveraging a decoupling of the DAE into its differential and algebraic parts, with the non-intrusive version of the approach additionally relying on results from arXiv:2604.20475 and arXiv:2107.07755. We illustrate the approach using a filtered buck converter as an example and compare both the intrusive and non-intrusive versions. The code for the example is openly available.

cs.CE

A Methodology for Integrating Life Cycle Assessment into a Multidisciplinary Design Analysis and Optimization Framework for Sustainable Launcher Development

The increasing number of orbital and sub-orbital launches makes it necessary to investigate the environmental impacts of launch vehicles and incorporate eco-design considerations into their development. In response, the European Space Agency has promoted Life Cycle Assessment (LCA) as a standardization methodology to mitigate environmental impacts of present and future space missions. This need is further amplified in the NewSpace, where numerous configurations and innovative technologies are explored, reinforcing the importance of integrating environmental considerations. At early design stages, launch vehicle architecture can be formalized through a multi-physics optimization problem based on Multidisciplinary Design Analysis and Optimization (MDAO) methods, where disciplines such as propulsion, aerodynamics, structure, and trajectory are coupled to obtain trade-offs among candidate configurations. This paper proposes a methodology to integrate an LCA discipline within an MDAO framework for launch vehicle design. The approach relies on parametric life-cycle inventories depending on design and coupling variables, covering component and propellant production as well as transport to the launch site. Launch emissions are evaluated from optimized trajectory profiles and characterized in terms of climate change impact. The methodology is illustrated on a representative expendable launch vehicle, where multi-objective optimizations assess trade-offs between performance and environmental indicators. Results highlight antagonistic behaviors among environmental impact categories, emphasizing the importance of carefully defining environmental objectives in eco-design studies. The generic nature of the methodology lays the foundation for integrating LCA into early-stage launch vehicle design, enabling exploration of trade-offs between performance, cost, and environmental considerations.

math.OC

The PUR-1 Cyber-Physical Digital Twin

Digital twin technologies have the potential to improve operational flexibility and responsiveness capabilities of nuclear systems. To provide decision support, cyber event characterization, state estimation, predictive control, and real-time dynamic processing of operational data, however, an efficient digital twin needs to integrate multiple models (data-driven as well as physics-based) with explainability while at the same time maintain two-way synchronization with the physical facility at a time constant less than its operational cycle. In this work, we present the Purdue University Reactor One Digital Twin (PUR-1 DT), a cyber-physical digital twin with a complete high-fidelity physics-based and AI-driven virtual model stack (neutronics, thermal-hydraulics, point kinetics) which provides closed-loop explainable diagnostics, forecasting, predictive control, and action recommendation back to the reactor via two-way communications and a cyber-physical testbed. We demonstrate real-time synchronized state estimation and short-term forecasting over a full reactor operational cycle and conduct a series of benchmarking experiments to validate accuracy and latency. Our results show good agreement with experimental results and lay the groundwork for further development and experimental demonstration of DT-enabled functionalities in real-world facilities.

cs.CE

FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation

Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent systems have automated this process, scaling factor mining far beyond manual effort. However, these automated approaches optimize directly for returns and rarely check whether a generated factor still expresses the economic hypothesis that motivated it. We identify this inconsistency between mathematical form and economic meaning as a structural failure mode of return-oriented automation. The resulting factors blur the line between real signals and spurious correlations and break down across regime shifts. We propose FaVOR (Factor Validation through Observable Reasoning), an agentic framework that restructures factor mining around hypothesis-level evidence rather than return outcomes. In place of the standard hypothesis-to-formula leap, FaVOR enforces a three-stage consistency loop tying mathematical form to economic rationale throughout. (1) Decomposition splits a broad economic hypothesis into independent observable conditions. (2) Validation checks whether each factor reflects its intended condition. (3) Integration merges them into a composite whose structure remains interpretable. On the CSI 500 and S&P 500 in 2025, FaVOR outperforms existing baselines while remaining effective across regimes. FaVOR shows that hypothesis-grounded factor discovery produces signals that are interpretable by construction, regime-robust, and economically faithful. The code is available at https://github.com/damilab/FaVOR.

cs.AI
Compare source metadata on this page
WorkPublishedSource identifierSource
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis2026-09-022609.03218arxiv
Physics-Informed Neural Networks for Depth-Averaged Granular Avalanche Dynamics on Curved Topography2026-09-022609.05542arxiv
Performance evaluation of variational quantum eigensolver and quantum dynamics algorithms on the advection-diffusion equation2025-03-31Phys. Rev. E 114, 025306 (2026)arxiv
On the Application of Hybrid Mixed Domain Decomposition Methods to Permanent Magnet Synchronous Machines2026-05-292605.31032arxiv
Do simulated agents move like real people?2026-05-302606.00733arxiv
Disciplined Bilevel Programming2026-09-012609.00644arxiv
HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields2026-09-012609.00679arxiv
Aerodynamic Shape Design Space Exploration with Deep Latent Diffusion Model2026-09-01AIAA Journal, 2026arxiv
VPID: An Integrated Framework for Vulnerability Prioritization and Intrusion Detection in Enterprise Networks2026-09-012609.00819arxiv
Don't You Know, Pump it Up! Investigating Cryptocurrency Manipulation in Telegram-Driven Activity2026-09-012609.01176arxiv
Connectome-Based Modelling Reveals Orientation Maps in the Drosophila Optic Lobe2026-09-012609.01330arxiv
Risk-Sensitive Reward Composition for Conditional GFlowNets2026-09-012609.01929arxiv
TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding2025-12-072512.06933arxiv
M-Tensor Formalism: A Non-iterative High Dimensional Least Squares Regression for Nonlinear Models with Scarce Data2026-02-092602.08509arxiv
A non-intrusive approach to index-aware learning2026-05-292605.30955arxiv
A Methodology for Integrating Life Cycle Assessment into a Multidisciplinary Design Analysis and Optimization Framework for Sustainable Launcher Development2026-06-242606.25945arxiv
The PUR-1 Cyber-Physical Digital Twin2026-08-312608.30186arxiv
FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation2026-08-312608.30192arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.