arXiv ScienceSearch

arXiv subjects

Jackson Lee

Publications and source records attributed to Jackson Lee.

3 recordsLinked to original sources

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR

Reinforcement learning with verifiable rewards has made post-training highly effective when correctness can be checked automatically. However, many important model behaviors require satisfying several qualitative criteria at once. Rubric-based rewards address this setting by grading prompt-specific criteria and aggregating them into a scalar reward. Yet standard static aggregations conflate a criterion's human-assigned importance with its current usefulness as an optimization signal. We show that this assumption breaks down in rubric RL: many important criteria are already saturated or currently unreachable, while criteria that distinguish rollouts are not necessarily those with the largest human weights. We introduce POW3R, a policy-aware rubric reward framework that preserves human weights and category balance as the rubric objective while adapting criterion-level reward weights during training. POW3R uses rollout-level contrast to emphasize criteria that currently separate the policy's outputs, making the GRPO reward more informative without changing the underlying evaluation target. Across three base policies on two datasets spanning multimodal and text-only settings, POW3R wins $24$ of $30$ base-policy/metric comparisons, improving both mean rubric reward and strict completion (the fraction of prompts whose response satisfies every required rubric criterion) over vanilla GRPO with rubric rewards, and reaches the same plateau in $2.5$--$4\times$ fewer training steps. Rubric rewards should therefore distinguish what should matter in the final answer from what can teach the current policy.

cs.AI

Machine Learning Phase Field Reconstruction in a Bose-Einstein Condensate

A basic challenge in experimental physics is the extraction of information related to variables that are not directly measured. The challenge is particularly severe in quantum systems where one may be interested in correlations of operators that are not diagonal in the measurement basis. In this paper we take a step towards addressing this issue in the context of Boson superfluids, where standard in-situ imaging yields only the spatially resolved density, leaving the phase field - crucial for identifying topological defects such as vortices and confirming superfluidity - indirectly encoded. Previous work has shown that the location of vortices in the phase field may be detected, but has not solved the problems of fully reconstructing the phase or identifying the charge (vortex vs. antivortex). This paper shows that a combination of a deep machine learning (ML) model and classical computer vision post-processing steps can address this gap. We use realistic snapshots of the thermal state of a two-dimensional BEC in a harmonic trap using synthetic data obtained from projected Gross-Pitaevskii equation simulations to train a U-Net-based architecture to infer the absolute values of the phase field gradients from an observed density field, and then employ a separate ML model to locate the positions of the vortex cores and a post-processing graphical analysis to determine with high accuracy the phase field, including the quantized charge of each vortex.

cond-mat.quant-gas

Machine-learning the spectral function of a hole in a quantum antiferromagnet

Understanding charge motion in a background of interacting quantum spins is a fundamental problem in quantum many-body physics. The most extensively studied model for this problem is the so-called $t$-$t'$-$t''$-$J$ model, where the determination of the parameter $t'$ in the context of cuprate superconductors is challenging. Here we present a theoretical study of the spectral functions of a mobile hole in the $t$-$t'$-$t''$-$J$ model using two machine learning techniques: K-nearest Neighbors regression (KNN) and a feed-forward neural network (FFNN). We employ the self-consistent Born approximation to generate a dataset of about $1.3 \times 10^5$ spectral functions. We show that for the forward problem, both methods allow for the accurate and efficient prediction of spectral functions, allowing for e.g. rapid searches through parameter space. Furthermore, we find that for the inverse problem (inferring Hamiltonian parameters from spectra), the FFNN can, but the KNN cannot, accurately predict the model parameters using merely the density-of-state. Our results suggest that it may be possible to use deep learning methods to predict materials parameters from experimentally measured spectral functions.

cond-mat.str-el