arXiv ScienceSearch

arXiv subjects

Wei Dai

Publications and source records attributed to Wei Dai.

At least 19 recordsLinked to original sources

this-that-model-1.0: A typed decision model that decides in 30 ms, for a millionth of a cent

Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears without a person. What the program needs back is not prose. It is one of n declared options and a number it can threshold. Today that costs a round trip to a frontier model -- hundreds of milliseconds, a per-token bill, and a parser -- for a question that is usually a conjunction of three clauses. this-that-model-1.0 is a 2B-parameter typed decision model. Its answer is read directly from the hidden state at a designated position and restricted to the option set the caller declared, so no text is generated, nothing can be malformed, and every question in a request is answered in the same forward pass. It decides in 30.9 ms on one laptop GPU and generates zero output tokens doing it, where a frontier API call costs 8758 ms and the hosted systems that answer these questions well spend between 21 and 212 generated tokens per question thinking first, billed for every one. It sustains 32 decisions per second on one consumer GPU and never lets the state leave the machine. On a third party's recorded cohort of 68 decision questions, on their inputs and their wording, it scores 0.941 with a Brier score of 0.042, against 0.765 and 0.133 for the hosted service Jev on the same items. One pass of our 42-family internal suite takes 32 seconds and 0.000217 USD of electricity; the most accurate hosted model we measured needs 155.2 minutes and 10.636 USD. We also report where it loses. On multi-step arithmetic, which a single forward pass cannot carry intermediate results through, it scores 0.560 against their 0.98 to 1.00, and a targeted second training round improved the five task families it was written for and transferred to none of the other 13. The model is open-sourced in https://huggingface.co/flock-io/this-that-model-1.0

cs.CL

Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics

Reasoning language models are trained to produce solutions, not to refuse them, and this bias persists when the problem they are handed is false. Asked to prove a corrupted theorem, a strong model will typically comply and produce a confident derivation of something untrue. We present Euston, an 8B mathematical claim-verification model trained to resist exactly this. Training data were generated with GraphSynth, a probabilistic factor-graph generator that couples attribute-level diversity to decode-time structural masking and span-synchronized verification, yielding 3{,}026 matched true/corrupted statement pairs (6,052 statements) drawn from arXiv papers spanning 2010--2025. We fine-tuned DeepSeek-R1-8B with GRPO under a rule-based, zero-API reward for 189 steps on four H100 GPUs. On a balanced 200-true/200-false held-out split, balanced accuracy rises from 29.50% to 63.75% and the discrimination gap---the difference between the rate of calling false statements false and the rate of calling true statements false moves from -0.5% (z=-0.1) to +27.5% (z=+6.0). Critically, the gain is not purchased with general mathematical ability: AIME 2026 accuracy under official semantics is 65.00% against a 69.17% base, a difference of -4.17% that is not statistically significant, whereas an earlier run of the same recipe on a smaller GraphSynth corpus collapsed to 40.00%. Median response length also falls from 19,217 to 18,296 tokens and the truncation rate from 25.8% to 8.3%, so the improvement does not come from thinking longer. We report the result together with the confounds that bound its interpretation, principally the all-false composition of the official evaluation sets and the low precision implied at realistic error prevalence.

cs.CL

Joint Structure Identification and Newton Acceleration via Proximal Line Search for Nonconvex Optimization

We consider composite optimization with a smooth, possibly nonconvex function and a separable convex polyhedral term. Such problems have a simple nonsmooth structure but arise widely in applications. Existing Newton-type methods may fail to identify nonsmooth coordinates during direction computation, move coordinates away from those reached by the full step during line search, or rely on separate first-order steps for identification. In this work, we propose a unified inexact proximal Newton method that couples structure identification with second-order acceleration through a proximal line search designed to retain the nonsmooth structure revealed at the local model. In particular, on a predicted working set, we solve a quadratic model inexactly but enforce its constraints exactly, allowing coordinates to reach new kinks without leaving the selected affine intervals. This set also includes small-margin kink coordinates to exploit the Hessian coupling effect. To globalize the direction, we construct a proximal line search using an isotropic proximal model whose linear term is chosen so that the initial trial exactly recovers the full step. Backtracking increases the proximal curvature until sufficient decrease holds, allowing coordinates to remain at kinks reached by the full step while damping other components. Generating each trial point requires only proximal operator evaluations. We prove finite backtracking and finite active-set identification under strict complementarity and a non-singular reduced Hessian, together with local Q-superlinear convergence under a directional Dennis-More condition and vanishing relative inexactness. Numerical results verify the effectiveness of the proposed method.

eess.SP

Hunting for Axions in REactor neutrino COherent scattering Detection Experiment

Nuclear power plants are not only vital sources of clean energy but also powerful facilities for probing new physics beyond the Standard Model. Due to the intense gamma-ray flux and an appropriate energy conditions, they are particularly well-suited for searches of light hypothetical particles such as sub-MeV axions and axion-like particles (ALPs). In this work, we propose to search for the ALPs in the REactor Neutrino COherent scattering Detection Experiment (RECODE), where two low-threshold, high-purity germanium detectors are placed at 11 m (near point) and 22 m (far point) from a 3.4 GW nuclear reactor at Sanmen nuclear power plant. With a 10 kg$\cdot$year exposure, we demonstrate that the expected sensitivities to the ALP couplings to the electrons and photons are competitive with or surpass the available results from the beam-dump experiments. A planned upgrade to 100 kg$\cdot$year will fully cover the so-called {$\it$ cosmological triangle} region, probing unexplored parameter space relevant to axions.

hep-ph

On the sharp quantitative stability of critical points of the Hardy-Littlewood-Sobolev inequality in $\mathbb{R}^{n}$ with $n\geq3$

Assume $n\geq3$ and $u\in \dot{H}^1(\mathbb{R}^n)$. Recently, Piccione, Yang and Zhao \cite{Piccione-Yang-Zhao} established a nonlocal version of Struwe's decomposition in \cite{Struwe-1984}, i.e., if $Γ(u):=\left\|Δu+D_{n,α}\int_{\mathbb{R}^{n}}\frac{|u|^{p_α}(y) }{|x-y|^α}\mathrm{d}y |u|^{p_α-2} u\right\|_{\dot{H}^{-1}} \rightarrow 0$ and $u\geq 0$, then $dist(u,\mathcal{T})\to 0$, where $dist(u,\mathcal{T})$ denotes the $\dot{H}^1(\mathbb{R}^n)$-distance of $u$ from the manifold of sums of Talenti bubbles. In this paper, we establish the nonlocal version of the quantitative estimates of Struwe's decomposition in Ciraolo, Figalli and Maggi \cite{CFM} for one bubble and $n\geq3$, Figalli and Glaudo \cite{Figalli-Glaudo2020} for $3\leq n\leq5$ and Deng, Sun and Wei \cite{DSW} for $n\geq6$ and two or more bubbles. We prove that for $n\geq 3$ and $0<α<n$, \[dist (u,\mathcal{T})\leq C\begin{cases} Γ(u)\left|\log Γ(u)\right|^{\frac{1}{2}}\quad&\text{if } \,\, n\geq 6, \,\, ν\geq2 \,\, \text{and} \,\, α=\frac{n+2}{2}, \\ Γ(u) \quad&\text{for any other cases,}\end{cases}\] where $ν$ denotes the number of bubbles. Furthermore, we show that this inequality is sharp for $n\geq 6$ and $α=\frac{n+2}{2}$. It should be emphasized that, in our paper, we have developed new techniques to deal with the strong singular case $4<α<n$, which can not be handled by reduction methods in previous works. We believe that our method can also be applied to other problems related to the physically interesting Hartree equation.

math.AP

Nanocavity Confinement by Orthogonal Valley- and SSH- Topological Interfaces In Glide-Symmetric Photonic Crystal Structures

Valley photonic crystals enable valley-dependent transport and chirality-selective emission, but incorporating wavelength-scale localization remains challenging. Existing valley-photonic-crystal cavities rely on finite defects or local lattice modifications that require structure-specific optimization and offer limited continuous control. Here, we theoretically and experimentally demonstrate two-dimensional nanocavity confinement using two orthogonal domain walls in a glide-symmetric valley photonic crystal. A valley domain wall confines the guided interface mode transversely, while an SSH-like domain wall localizes it longitudinally. Starting from a glide-symmetry-protected Dirac point in a bearded-interface waveguide, controlled displacements of adjacent triangular holes open a topological gap in the continuous guided-mode dispersion. The displacement amplitude $ΔR$ tunes the gap, mode volume, and intrinsic radiative $Q$ factor. Implemented in a silicon photonic-crystal slab, the structure exhibits localized resonances within the topological mode gap and systematic spectral tuning with $ΔR$. The maximum measured loaded $Q$ factor is $1.2\times10^{4}$. This approach enables continuously tunable, high-$Q$ nanocavities integrated into topological waveguide networks for compact resonant devices and enhanced light--matter interactions.

physics.optics

Optimal Rigidity Results for the $k$-Hessian Equation of Lane--Emden Type

In this paper, we establish optimal Liouville theorems and classification results for the \(k\)-Hessian Lane--Emden equation \[ σ_k(-D^2u)=u^p\quad\text{in }\R^n,\qquad -D^2u\in\overline{Γ_k},\qquad u\geq 0, \] where \(2\leq k<\frac{n}{2}\) and $p>0$. Let $p_- = \frac{nk}{n-2k}$ and the critical Hessian--Sobolev exponent $p_* = \frac{(n+2)k}{n-2k}$. Phuc and Verbitsky proved nonexistence of positive solutions for \(k 2k\), without any additional assumption. For the limiting case \(n=2k\), we classify finite-mass solutions to the $\frac{n}{2}$-Hessian Liouville equation under a proper asymptotic condition $u(x)\rightarrow-\infty$ as $|x|\rightarrow\infty$. In particular, we provide the fully nonlinear counterparts of the classical Liouville and classification theorems of Gidas--Spruck, Gidas--Ni--Nirenberg, and Caffarelli--Gidas--Spruck.

math.AP

Wideband Large-Array Processing and Sparse Design for Angle Imaging

This paper shows that wideband large-array processing can recover a large number of angle pixels with far fewer antenna elements. The key advantage of wideband signaling is that different frequencies induce different virtual arrays, whose union forms a virtual array with a substantially increased number of effective virtual elements. Thus, a sparse physical array can support far more spatial samples than physical antennas. Motivated by this capability, we study the recovery of angular responses across the full field of view $[-90^\circ, 90^\circ)$, discretized according to the improved angular resolution, and refer to this sensing regime as angle imaging. However, the resulting virtual array is inherently irregular, clustered, and does not automatically guarantee stable recovery. To address this challenge, we introduce a coverage criterion that estimates the number of stably recoverable angle pixels, without computationally intensive singular-value-based conditioning tests over candidate image dimensions. For systems satisfying this criterion, we theoretically establish deterministic condition-number bounds that characterize stable angle imaging. Building on this criterion, we derive non-uniform sparse array designs that minimize the number of physical antennas while maintaining recovery over the full field of view. Simulation results show that the proposed criterion provides practical guidance for stable system design, and that the resulting sparse arrays can recover substantially more angle pixels than the number of physical antennas, with representative designs supporting over ten times as many angle pixels as physical antennas.

eess.SP

Non-radial solutions for the quasi-linear Hénon type $N$-Laplacian Liouville equation

In this paper, we investigate the following quasi-linear weighted $N$-Laplacian Liouville equation \begin{equation*}\label{0} -Δ_N u=|x|^{Nα}e^{u}, \qquad x\in \R^N, \end{equation*} where $N \geq 2$. For $\al>0$, by carefully studying the linearized problem and applying the approximation method and bifurcation theory, we prove that, when the parameter $α$ equals to the critical values $α(k):=\frac{\sqrt{k(N-1)(k+N-2)}}{N-1}-1$ for $k \geq 2$, there exist non-radial solutions $u$ (bifurcating from $U_{α(k)}$) to the above quasi-linear Hénon type Liouville equation such that $u\sim \ln|x|$, $|\nabla u|= O(|x|^{-1})$ at $\infty$ and $\int_{\R^N}|x|^{Nα}e^{u}\md x=N\left(\frac{N^2}{N-1}\right)^{N-1}(α+1)^{N-1}ω_N$. One should note that, $α(k)=k-1$ for $k\geq2$ when $N=2$. Our results successfully extend the existence result of J. Prajapat and G. Tarantello in \cite{PT} concerning the $2$-dimension and Laplacian case (i.e., $N=2$) to the more general $N$-dimension and $N$-Laplacian cases ($N\geq 2$), and extend the results of F. Gladiali, M. Grossi, and S. L. N. Neves in \cite{GGN} and the authors in \cite{DDGL} from $1<p<N$ to the much more complicated limiting case $p=N$. We introduced some new ideas and overcame a series of crucial difficulties, including the nonlinearity nature of the $N$-Laplacian $Δ_N$, the lack of Green integral representation formula and critical weighted Sobolev embedding inequality, the absence of Kelvin type transforms for linearized/difference equations, the invariance of the total mass under scalings of $u$, and the signs-changing and divergence (to $-\infty$) at $\infty$ of the solutions, which makes the suitable choices of the approximate problems, the (normalized) approximate function sequences and the working space to be quite difficult.

math.AP

SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens

Foundation models such as Segment Anything Model 2 (SAM2) have transformed natural-image and video segmentation, and recent work has begun adapting them to medical imaging. These adaptations, however, are largely general-purpose models that treat MRI as one modality among many; large-scale, MRI-specific modelling and benchmarking remain limited, even though MRI's low soft-tissue contrast leaves many boundaries effectively invisible on individual slices. We present SAMRI-3D, a benchmark and method for 3D MRI segmentation with SAM2. The SAMRI-3D benchmark is the largest MRI-only evaluation to date - 10,392 volumes from 34 datasets (27 public, 7 in-house) spanning 12 anatomical domains and 10+ sequences, with explicit seen/unseen splits. Freezing the image encoder and fine-tuning only the lightweight decoder and memory modules raises mean Dice from 0.58 (zero-shot SAM2) to 0.76, surpassing recent SAM-based medical models (SAMed-2 0.69, Medical-SAM2 0.49, SAM-Med3D 0.37) with strong statistical significance. To target invisible boundaries, we introduce Global Volume Tokens (GVT): persistent memory tokens trained with a Truncated Signed Distance Field (TSDF) reconstruction objective that is discarded at inference (zero added cost). This full model, SAMRI-3D, attains the best accuracy (0.78) and lowest variance across all 34 datasets and, uniquely, shows no drop on 8 held-out datasets (0.79 unseen vs. 0.78 seen); per-sequence analysis confirms the TSDF objective helps most where per-slice contrast is weakest. We will release the benchmark, code, and models in this paper.

cs.CV

Constraint-Anchored Reasoning Traces

Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thought (CoT) trace corrupts all downstream reasoning. We find that in state-of-the-art open-source MLLMs, once the first error occurs, the reasoning cascades into failure across all remaining steps in 65% of such cases (a metric we term the snowball rate). Existing mitigations-sampling multiple chains, post-hoc self-verification, or full program synthesis-either lack symbolic grounding, catch errors too late, or sacrifice the flexibility of natural language reasoning. We propose Constraint-Anchored Reasoning Traces (CART), a neuro-symbolic framework that trains MLLMs to interleave natural language reasoning steps with symbolic constraint assertions: lightweight, machine-checkable statements about visual content (e.g., count(red_objects) = 3). A dual-pronged Constraint Propagation Module-combining a learned neural grounding head with Boolean Constraint Propagation-continuously verifies these anchors against extracted visual features and checks their mutual logical consistency. When a contradiction is detected, a backtrack controller halts generation and reverts to the last consistent checkpoint, preventing error propagation. A variable-frequency emission mechanism allows the model to adaptively control anchor density, avoiding trace bloat. We construct 218K training instances by augmenting GQA, CLEVR-CoGenT, and VCR with ground-truth constraint annotations derived from scene graphs, and fine-tune open-source MLLMs (LLaVA-NeXT, Qwen2-VL) via LoRA. On five benchmarks, CART reduces the snowball rate from 0.65 to 0.14, improves GQA accuracy by +4.6 percentage points over trainingonly baselines, and achieves 89.1 F1 on POPE-all with at most 18% inference overhead.

cs.AI

Global existence for quasilinear wave equations on hyperbolic space

The main purpose of this paper is to study the global solvability for a general class of quasilinear shifted wave equations on hyperbolic spaces, for smooth initial data with small amplitude. In contrast to the case of Euclidean spaces, when the space dimension is three, we do not need to assume structural conditions like the null conditions to ensure global existence. To achieve this, we establish the energy and local energy estimates for perturbed wave operators on $\mathbb{R}\times \mathbb{H}^n$. These estimates allow time-dependent metric perturbations and require only suitable smallness together with polynomial decay in the radial variable $r$. As a byproduct, for semilinear problems with power-type nonlinearities and radial data, we also obtain global solutions with low-regularity. In particular, we prove an analog of the radial Glassey conjecture on hyperbolic space.

math.AP

Readout-Induced Leakage in Superconducting Circuits with Nonlinear Couplings

In superconducting circuits, drive-induced unwanted transitions limit the readout power, thereby constraining readout speed and fidelity. When such transitions excite the qubit into leakage states, they produce correlated errors that are particularly harmful for quantum error correction. Native nonlinear qubit-readout resonator coupling is a promising alternative to conventional linear hybridization because it provides intrinsic Purcell protection and stricter selection rules for multiphoton processes. In realistic devices, however, we show that such a coupling alone neither eliminates nor necessarily suppresses drive-induced transitions. Instead, if not appropriately engineered, these couplings often worsen the situation by introducing additional parasitic processes. Moreover, the rates of these unwanted transitions remain sensitive to the choice of readout frequency, regardless of the coupling mechanism. We demonstrate that readout-induced leakage can thus vary by orders of magnitude even when readout frequencies differ by less than ~7%. Our results establish that the benefits of native nonlinear couplings are realized only through informed device design, including the spectral placement of relevant auxiliary modes and elimination of parasitic ones.

quant-ph

Dean of LLM Tutors: A Framework for Automated Quality Review of AI-generated Feedback

Large language model (LLM) tutors are increasingly used to generate educational feedback, but existing research has focused mainly on feedback generation rather than feedback evaluation. As a result, LLM-generated feedback may offer limited pedagogical value and carry risks of hallucination. The current study introduces DeanLLM, an automated review framework for comprehensively evaluating feedback generated by LLM tutors before it is shared with students. We developed a 16-dimension evaluation framework covering feedback content, educational effectiveness, and hallucination risks, and validated it using using human-expert annotations of LLM-generated tutor feedback on synthetic computer science assignment submissions derived from real coursework. We then examined whether LLMs could serve as automated LLM-generated tutor feedback reviewers, and used the best-performing reviewer to benchmark tutor feedback generated by 10 commercial LLMs. Psychometric analyses supported the reliability of the proposed framework and showed that human reviewers tended to evaluate feedback holistically, whereas the LLM reviewer separated rubric dimensions more mechanically. Standard zero-shot and few-shot prompting showed limited agreement with human experts for content-quality judgments. Supervised fine-tuning of GPT-4.1 with human-labelled examples containing scores only, without explanatory rationales, achieved the strongest alignment with expert judgments. Reasoning LLMs were particularly effective at hallucination detection and produced automated tutor feedback with stronger educational effectiveness and factuality than lightweight models. The findings indicate that DeanLLM offers a scalable way for automatically improving the reliability and safety of LLM tutor feedback, while also demonstrating that reviewer calibration and model choice remain critical for educational deployment.

cs.CY

SCALEFeedback: A Large-Scale Dataset of Synthetic Computer Science Assignments for LLM-generated Educational Feedback Research

Using Large Language Models (LLMs) to give educational feedback to students for their assignments has attracted much attention in the AI in Education (AIED) field. Yet, there is currently no large-scale open-source dataset of student assignments that includes detailed assignment descriptions, rubrics, and student submissions across various courses. As a result, research on generalisable methodology for automatic generation of effective and responsible educational feedback remains limited. In this paper, we introduce a synthetic computer science university assignment dataset for LLM-based educational feedback research, called SCALEFeedback (Synthetic Computer science Assignments for LLM Educational Feedback Research). The dataset is generated via Sophisticated Assignment Mimicry (SAM) framework specifically designed to synthesise this dataset and that utilizes one-to-one LLM-based imitation from real assignment descriptions, rubrics, and student submissions. Our open-source dataset contains 10,000 synthetic student submissions spanning 155 assignments across 59 university-level computer science courses. Technical validation confirmed that the synthetic dataset closely resembles real data while successfully eliminating personally identifiable information present in the source material. The creation of this dataset is a valuable contribution to researchers who aim to develop LLM-based generalisable methods for offering high-quality, automated educational feedback in a scalable way.

cs.CY

Optimizing Pump Conditions of Parametric Amplifiers for Fast Multiplexed Readout of Superconducting Qubits

Low-noise parametric amplifiers are widely used as the first-stage amplifier in qubit readout chains. The performance of parametric amplifiers depends sensitively on the choice of the pump condition. We propose a strategy for determining the pump condition that is tailored for fast multiplexed readout. Choosing the amplifier pump to maximize the signal-to-noise ratio (SNR) improvement at the readout frequency of the limiting qubit--the qubit that requires the longest readout time to reach a target SNR--minimizes the total multiplexed readout time. We demonstrate our pump calibration strategy experimentally on a five-qubit multiplexed readout chain with a traveling-wave parametric amplifier. Using our strategy, we reduce the multiplexed readout time by 320 ns compared to optimizing the average SNR improvement on all qubits, without degrading the target SNR for any qubit.

quant-ph

CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure visual quality but offer no mechanism for assessing cultural faithfulness. Consequently, a model that replaces a Namaste with a handshake receives the same score as one that generates the gesture correctly. We propose CultureScore, a compositional evaluation framework that decomposes cultural faithfulness into three granular dimensions: Identity (who is represented), Context (culturally localized background), and Behavior (normative gestures and interactions). We operationalize this framework through an evaluation suite spanning 10 countries, yielding 6,174 generated videos across three state-of-the-art models. Our evaluation reveals that no current model achieves culturally faithful video generation: the best-performing model reaches only 56.8\% overall CultureScore, with Behavior the most challenging dimension, which remains below 52.1\% across all models. Furthermore, human preference rankings align directionally with CultureScore but are inverted relative to VideoScore; the highest-scoring model on visual quality was ranked last by annotators, underscoring that cultural faithfulness is an essential criterion for equitable video generation. Data and code are available at https://huggingface.co/datasets/ankurani/CultureScore.

cs.CV

On $τ$-tilting graphs for quasi-silted algebras

We prove that the $τ$-tilting graph of any quasi-silted algebra is connected and has the reachable-in-face property. Our approach utilizes $τ$-reduction and wall and chamber structures. In particular, we observe a sufficient condition on the wall and chamber structure under which the connectivity of $τ$-tilting graphs is preserved under taking quotients of algebras. As an immediate consequence, the connectivity of $τ$-tilting graphs is also established for several new classes of algebras.

math.RT