arXiv ScienceSearch

arXiv subjects

Simon Martin

Publications and source records attributed to Simon Martin.

10 recordsLinked to original sources

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation

Enterprise software systems commonly expose business functionality through both relational databases and REST APIs. Accessing these interfaces requires specialized technical knowledge, as users must determine whether a request requires a database query or an API operation and understand the corresponding schemas, endpoints, and parameters. This creates demand for natural language interfaces that translate user requests into SQL queries and REST API calls. While large language models (LLMs) show promise for structured code generation, they typically lack reliable knowledge of enterprise-specific schemas, endpoints, and documentation. Retrieval-augmented generation (RAG) addresses this limitation by grounding generation in external documentation. However, prior work largely studies SQL query generation and REST API call generation separately, despite enterprise documentation environments often containing both database schemas and API specifications. We systematically evaluate standard RAG, Self-RAG, and CoRAG across SQL query generation, REST API call generation, and a combined task requiring routing between both operation types. Using SAP Transactional Banking as a realistic enterprise use case, we constructed an execution-validated dataset and compared retrieval strategies under database-only, API-only, and mixed-documentation settings. Retrieval augmentation proved essential for reliable enterprise structured generation, substantially improving performance over a no-retrieval baseline. CoRAG achieved the best results in the combined SQL query and REST API call setting, with statistically significant improvements in exact-match accuracy over standard RAG, primarily driven by stronger SQL query generation under mixed-documentation retrieval conditions. Overall, findings show that retrieval strategy substantially affects structured generation performance under mixed-documentation settings.

cs.SE

High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks

We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher-student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with the input dimension, and the sample size grows quadratically. This scaling aims to describe overparameterized neural networks in which feature learning still plays a central role. In the high-dimensional limit, we derive a dynamical characterization of the gradient flow, in the spirit of dynamical mean-field theory (DMFT). Under l2-regularization, we analyze these equations at long times and characterize the performance and spectral properties of the resulting estimator. This result provides a quantitative understanding of the effect of overparameterization on learning and generalization, and reveals a double descent phenomenon in the presence of label noise, where generalization improves beyond interpolation. In the small regularization limit, we obtain an exact expression for the perfect recovery threshold as a function of the network widths, providing a precise characterization of how overparameterization influences recovery.

math.OC

Dimensionality-Changing Transition from a Non-Fermi Liquid to a Spin-Solid in a Multichannel Kondo Lattice

A multichannel Kondo system, where a single quantum spin couples to multiple channels of an electronic bath, provides one of the simplest examples of a zero-dimensional non-Fermi liquid. It is natural to ask: what happens when an extensive number of such systems are coupled together? A simple renormalization group argument implies that in a chain of SU(N) multichannel quantum systems, where each spin is coupled to its own bath of K channels, the individual spins dynamically decouple at low energy when N>K, resulting in a 'sliding' non-Fermi liquid. Using Quantum Monte Carlo (QMC) simulations, we find evidences of a continuous, 'dimensionality-changing' phase transition out of this non-Fermi liquid into a valence-bond solid phase as the intersite coupling is increased. Remarkably, at the critical point, correlations exhibit a power-law behavior even along the direction in which the spins are coupled, indicating the breakdown of dynamical decoupling at the transition. We also develop an RG scheme to understand the universal aspects of this transition.

cond-mat.str-el

A Perturbative Approach to Symmetric Mass Generation

The Landau paradigm has been a powerful framework for understanding phase transitions involving spontaneous symmetry breaking. In contrast, phase transitions between two symmetric phases, where neither phase breaks any symmetry, remain less explored. One intriguing class of such transitions involves "symmetric mass generation" (SMG), where interactions drive a transition from a gapless symmetric phase to a gapped symmetric phase. In this work, we develop a controlled perturbative approach to study a class of such transitions, based on an $\epsilon$-expansion around the critical dimension where the SMG-inducing-interaction becomes marginal. Applying this method to two distinct models, we identify a single-parameter-tuned transition in each case, which we conjecture captures the universal critical behavior of the SMG transition in these models. We compute universal quantities associated with these transitions.

cond-mat.str-el

The Spoils of Algorithmic Collusion: Profit Allocation Among Asymmetric Firms

We study the propensity of independent algorithms to collude in repeated Cournot duopoly games. Specifically, we investigate the predictive power of different oligopoly and bargaining solutions regarding the effect of asymmetry between firms. We find that both consumers and firms can benefit from asymmetry. Algorithms produce more competitive outcomes when firms are symmetric, but less when they are very asymmetric. Although the static Nash equilibrium underestimates the effect on total quantity and overestimates the effect on profits, it delivers surprisingly accurate predictions in terms of total welfare. The best description of our results is provided by the equal relative gains solution. In particular, we find algorithms to agree on profits that are on or close to the Pareto frontier for all degrees of asymmetry. Our results suggest that the common belief that symmetric industries are more prone to collusion may no longer hold when algorithms increasingly drive managerial decisions.

econ.GN

Bayes-optimal learning of an extensive-width neural network from quadratically many samples

We consider the problem of learning a target function corresponding to a single hidden layer neural network, with a quadratic activation function after the first layer, and random weights. We consider the asymptotic limit where the input dimension and the network width are proportionally large. Recent work [Cui & al '23] established that linear regression provides Bayes-optimal test error to learn such a function when the number of available samples is only linear in the dimension. That work stressed the open challenge of theoretically analyzing the optimal test error in the more interesting regime where the number of samples is quadratic in the dimension. In this paper, we solve this challenge for quadratic activations and derive a closed-form expression for the Bayes-optimal test error. We also provide an algorithm, that we call GAMP-RIE, which combines approximate message passing with rotationally invariant matrix denoising, and that asymptotically achieves the optimal performance. Technically, our result is enabled by establishing a link with recent works on optimal denoising of extensive-rank matrices and on the ellipsoid fitting problem. We further show empirically that, in the absence of noise, randomly-initialized gradient descent seems to sample the space of weights, leading to zero training loss, and averaging over initialization leads to a test error equal to the Bayes-optimal one.

stat.ML

Multi-Modal Dataset Creation for Federated Learning with DICOM Structured Reports

Purpose: Federated training is often hindered by heterogeneous datasets due to divergent data storage options, inconsistent naming schemes, varied annotation procedures, and disparities in label quality. This is particularly evident in the emerging multi-modal learning paradigms, where dataset harmonization including a uniform data representation and filtering options are of paramount importance. Methods: DICOM structured reports enable the standardized linkage of arbitrary information beyond the imaging domain and can be used within Python deep learning pipelines with highdicom. Building on this, we developed an open platform for data integration and interactive filtering capabilities that simplifies the process of assembling multi-modal datasets. Results: In this study, we extend our prior work by showing its applicability to more and divergent data types, as well as streamlining datasets for federated training within an established consortium of eight university hospitals in Germany. We prove its concurrent filtering ability by creating harmonized multi-modal datasets across all locations for predicting the outcome after minimally invasive heart valve replacement. The data includes DICOM data (i.e. computed tomography images, electrocardiography scans) as well as annotations (i.e. calcification segmentations, pointsets and pacemaker dependency), and metadata (i.e. prosthesis and diagnoses). Conclusion: Structured reports bridge the traditional gap between imaging systems and information systems. Utilizing the inherent DICOM reference system arbitrary data types can be queried concurrently to create meaningful cohorts for clinical studies. The graphical interface as well as example structured report templates will be made publicly available.

cs.IR

Real World Federated Learning with a Knowledge Distilled Transformer for Cardiac CT Imaging

Federated learning is a renowned technique for utilizing decentralized data while preserving privacy. However, real-world applications often face challenges like partially labeled datasets, where only a few locations have certain expert annotations, leaving large portions of unlabeled data unused. Leveraging these could enhance transformer architectures ability in regimes with small and diversely annotated sets. We conduct the largest federated cardiac CT analysis to date (n=8,104) in a real-world setting across eight hospitals. Our two-step semi-supervised strategy distills knowledge from task-specific CNNs into a transformer. First, CNNs predict on unlabeled data per label type and then the transformer learns from these predictions with label-specific heads. This improves predictive accuracy and enables simultaneous learning of all partial labels across the federation, and outperforms UNet-based models in generalizability on downstream tasks. Code and model weights are made openly available for leveraging future cardiac CT analysis.

eess.IV

On the Impact of Overparameterization on the Training of a Shallow Neural Network in High Dimensions

We study the training dynamics of a shallow neural network with quadratic activation functions and quadratic cost in a teacher-student setup. In line with previous works on the same neural architecture, the optimization is performed following the gradient flow on the population risk, where the average over data points is replaced by the expectation over their distribution, assumed to be Gaussian.We first derive convergence properties for the gradient flow and quantify the overparameterization that is necessary to achieve a strong signal recovery. Then, assuming that the teachers and the students at initialization form independent orthonormal families, we derive a high-dimensional limit for the flow and show that the minimal overparameterization is sufficient for strong recovery. We verify by numerical experiments that these results hold for more general initializations.

math.OC

Critical phase induced by Berry phase and dissipation in a spin chain

Motivated by experiments on spin chains embedded in a metallic bath, as well as closed quantum systems described by long-range interacting Hamiltonians, we study a critical SU(N) spin chain perturbed by dissipation, or equivalently, after space-time rotation, long-range spatial interactions. The interplay of dissipation and the Wess-Zumino (Berry phase) term results in a rich phase diagram with multiple renormalization-group fixed points. For a range of the exponent that characterizes the dissipative bath, we find a second-order phase transition between the fixed point that describes an isolated critical spin chain and a dissipation-induced-ordered phase. More interestingly, for a different range of the exponent, we find a stable, gapless, nonrelativistic phase of matter whose existence necessarily requires coupling to the dissipative bath. Upon tuning the exponent, we find that the fixed point corresponding to this gapless, stable phase "annihilates" the fixed point that describes the transition out of this phase to the ordered phase. We also study a relativistic version of our model, and we identify a new critical point. We discuss the implications of our work for Kondo lattice systems and engineered long-range interacting quantum systems.

cond-mat.str-el