arXiv ScienceSearch

arXiv subjects

Patrick Wilhelm

Publications and source records attributed to Patrick Wilhelm.

11 recordsLinked to original sources

Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback

Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget. We study the fixed-budget decision problem behind this practice: under the same post-training budget, should one use a larger policy, train a smaller policy longer, generate more rollout search, or spend compute on stronger reward feedback? We introduce a FLOP-accounting framework for GRPO post-training that decomposes compute into rollout/search, policy-update/learning, and reward- or feedback-model evaluation. Across LoRA-adapted Qwen2.5 policies, we find conditional allocation frontiers: the best observed allocation changes with model size, compute budget, reward system, and evaluation target. Same-FLOP model-size comparisons show that model choice and training allocation are coupled because larger policies consume more per-token compute and therefore buy fewer updates or rollouts under the same budget. Reward systems also change the accounting: rule-based rewards spend nearly all non-update compute on policy rollouts, while PRM-style feedback allocates a visible part of the budget to reward-model inference. We present RACE as a diagnostic pilot-grid protocol, not a guarantee of held-out improvement, for identifying allocation regimes before expensive validation runs; our results suggest that RL post-training papers should report total FLOPs together with how compute is divided among model size, search, learning, and feedback.

cs.LG

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting in Gameable ALFWorld and WebShop. Agents are instrumented with activation-based reward-hack scores, token-level entropy, and decision-context features. We find that adapters fine-tuned on \textit{School-of-Reward-Hacks} dataset can transfer reward-hack tendencies into agentic action selection, especially when the environment exposes proxy-reward affordances. However, mitigating such behavior cannot rely on activation dynamics alone. High reward-hack activation identifies a latent policy state, but does not necessarily imply an immediate exploit action. Across next-step prediction tasks, entropy and context-calibrated internal features improve risk estimation over reward-hack activation alone. Activation-direction steering further reduces proxy-exploit behavior in selected mixed-adapter regimes. Overall, our results support context-calibrated internal monitoring for agents: reward-hack activation identifies a latent policy state, while entropy and decision context help determine when that state becomes risky action.

cs.AI

Revisiting Gradient Staleness: Evaluating Distance Metrics for Asynchronous Federated Learning Aggregation

In asynchronous federated learning (FL), client devices send updates to a central server at varying times based on their computational speed, often using stale versions of the global model. This staleness can degrade the convergence and accuracy of the global model. Previous work, such as AsyncFedED, proposed an adaptive aggregation method using Euclidean distance to measure staleness. In this paper, we extend this approach by exploring alternative distance metrics to more accurately capture the effect of gradient staleness. We integrate these metrics into the aggregation process and evaluate their impact on convergence speed, model performance, and training stability under heterogeneous clients and non-IID data settings. Our results demonstrate that certain metrics lead to more robust and efficient asynchronous FL training, offering a stronger foundation for practical deployment.

cs.LG

Monitoring Emergent Reward Hacking During Generation via Internal Activations

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of completed responses, it remains unclear whether such behavior can be identified during generation. We propose an activation-based monitoring approach that detects reward-hacking signals from internal representations as a model generates its response. Our method trains sparse autoencoders on residual stream activations and applies lightweight linear classifiers to produce token-level estimates of reward-hacking activity. Across multiple model families and fine-tuning mixtures, we find that internal activation patterns reliably distinguish reward-hacking from benign behavior, generalize to unseen mixed-policy adapters, and exhibit model-dependent temporal structure during chain-of-thought reasoning. Notably, reward-hacking signals often emerge early, persist throughout reasoning, and can be amplified by increased test-time compute in the form of chain-of-thought prompting under weakly specified reward objectives. These results suggest that internal activation monitoring provides a complementary and earlier signal of emergent misalignment than output-based evaluation, supporting more robust post-deployment safety monitoring for fine-tuned language models.

cs.CL

Noise-aware Client Selection for carbon-efficient Federated Learning via Gradient Norm Thresholding

Training large-scale Neural Networks requires substantial computational power and energy. Federated Learning enables distributed model training across geospatially distributed data centers, leveraging renewable energy sources to reduce the carbon footprint of AI training. Various client selection strategies have been developed to align the volatility of renewable energy with stable and fair model training in a federated system. However, due to the privacy-preserving nature of Federated Learning, the quality of data on client devices remains unknown, posing challenges for effective model training. In this paper, we introduce a modular approach on top to state-of-the-art client selection strategies for carbon-efficient Federated Learning. Our method enhances robustness by incorporating a noisy client data filtering, improving both model performance and sustainability in scenarios with unknown data quality. Additionally, we explore the impact of carbon budgets on model convergence, balancing efficiency and sustainability. Through extensive evaluations, we demonstrate that modern client selection strategies based on local client loss tend to select clients with noisy data, ultimately degrading model performance. To address this, we propose a gradient norm thresholding mechanism using probing rounds for more effective client selection and noise detection, contributing to the practical deployment of carbon-efficient Federated Learning.

cs.LG

Beyond Test-Time Compute Strategies: Advocating Energy-per-Token in LLM Inference

Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks but come with substantial energy and computational costs, particularly in request-heavy scenarios. In many real-world applications, the full scale and capabilities of LLMs are often unnecessary, as Small Language Models (SLMs) can provide accurate responses for simpler text generation tasks. When enhanced with advanced reasoning strategies, such as Chain-of-Thought (CoT) prompting or Majority Voting, SLMs can approach the performance of larger models while reducing overall computational requirements. However, these strategies can also introduce additional energy costs, creating an energy-accuracy trade-off. Our analysis examines these trade-offs in test-time compute strategies for smaller models compared to larger ones, using the MMLU benchmark. Additionally, we explore the input-output token dynamics of transformer architectures, which result in nonlinear hardware energy operation curves for LLMs. To bridge AI research with its physical impact, we propose \textit{energy efficiency metrics}, including Energy-per-Token, as complements to traditional accuracy benchmarks. Beyond model selection, we propose controlled reasoning in CoT token generation, using operating curves to regulate reasoning depth dynamically. This vision integrates a energy-aware routing mechanism, ensuring that model selection and inference strategies balance accuracy for sustainable AI deployment.

cs.CL

Carbon-Aware Quality Adaptation for Energy-Intensive Services

The energy demand of modern cloud services, particularly those related to generative AI, is increasing at an unprecedented pace. To date, carbon-aware computing strategies have primarily focused on batch process scheduling or geo-distributed load balancing. However, such approaches are not applicable to services that require constant availability at specific locations due to latency, privacy, data, or infrastructure constraints. In this paper, we explore how the carbon footprint of energy-intensive services can be reduced by adjusting the fraction of requests served by different service quality tiers. We show that adapting this quality of responses with respect to grid carbon intensity can lead to additional carbon savings beyond resource and energy efficiency. Building on this, we introduce a forecast-based multi-horizon optimization that reaches close-to-optimal carbon savings and is able to automatically adapt service quality for best-effort users to stay within an annual carbon budget. Our approach can reduce the emissions of large-scale LLM services, which we estimate at multiple 10,000 tons of CO2 annually, by up to 10%.

cs.DC

Experimental determination of the dissociative recombination rate coefficient for rotationally-cold CH$^{+}$ and its implications for the diffuse cloud chemistry

Observations of CH$^+$ are used to trace the physical properties of diffuse clouds, but this requires an accurate understanding of the underlying CH$^+$ chemistry. Until this work, the most uncertain reaction in that chemistry was dissociative recombination (DR) of CH$^+$. Using an electron-ion merged-beams experiment at the Cryogenic Storage Ring, we have determined the DR rate coefficient of the CH$^+$ electronic, vibrational, and rotational ground state applicable for different diffuse cloud conditions. Our results reduce the previously unrecognized order-of-magnitude uncertainty in the CH$^+$ DR rate coefficient to $\sim \pm 20\%$ and are applicable at all temperatures relevant to diffuse clouds, ranging from quiescent gas to gas locally heated by processes such as shocks and turbulence. Based on a simple chemical network, we find that DR can be an important destruction mechanism at temperatures relevant to quiescent gas. As the temperature increases locally, DR can continue to be important up to temperatures of $ \sim 600\,\mathrm{K} $ if there is also a corresponding increase in the electron fraction of the gas. Our new CH$^+$ DR rate coefficient data will increase the reliability of future studies of diffuse cloud physical properties via CH$^+$ abundance observations.

astro-ph.GA

Laser-probing the rotational cooling of molecular ions by electron collisions

We present state-selected measurements of rotational cooling and excitation rates of CH$^+$ molecular ions by inelastic electron collisions. The experiments are carried out at the Cryogenic Storage Ring, making use of a monoenergetic electron beam at matched velocity in combination with state-sensitive laser-dissociation of the CH$^+$ ions for simultaneous monitoring of the rotational level populations. Employing storage times of up to 600 s, we create conditions where electron-induced cooling to the $J = 0$ ground state dominates over radiative relaxation, allowing for the experimental determination of inelastic electron collision rates to benchmark state-of-the-art theoretical calculations. On a broader scale, our experiments pave the way to probe inelastic electron collisions for a variety of molecular ions relevant in various plasma environments.

physics.atom-ph

Interplay of Fractional Chern Insulator and Charge-Density-Wave Phases in Twisted Bilayer Graphene

We perform an extensive exact diagonalization study of interaction driven insulators in spin- and valley-polarized moir\'{e} flat bands of twisted bilayer graphene aligned with its hexagonal boron nitride substrate. In addition to previously reported fractional Chern insulator phases, we provide compelling evidence for competing charge-density-wave phases at multiple fractional fillings of a realistic single-band model. A thorough analysis at different interlayer hopping parameters, motivated by experimental variability, and the role of kinetic energy at various Coulomb interaction strengths highlight the competition between these phases. The interplay of the single-particle and the interaction induced hole dispersion with the inherent Berry curvature of the Chern bands is intuitively understood to be the driving mechanism for the ground-state selection. The resulting phase diagram features remarkable agreement with experimental findings in a related moir\'{e} heterostructure and affirms the relevance of our results beyond the scope of graphene based materials.

cond-mat.str-el

The Cryogenic Storage Ring CSR

An electrostatic cryogenic storage ring, CSR, for beams of anions and cations with up to 300 keV kinetic energy per unit charge has been designed, constructed and put into operation. With a circumference of 35 m, the ion-beam vacuum chambers and all beam optics are in a cryostat and cooled by a closed-cycle liquid helium system. At temperatures as low as (5.5 $\pm$ 1) K inside the ring, storage time constants of several minutes up to almost an hour were observed for atomic and molecular, anion and cation beams at an energy of 60 keV. The ion-beam intensity, energy-dependent closed-orbit shifts (dispersion) and the focusing properties of the machine were studied by a system of capacitive pickups. The Schottky-noise spectrum of the stored ions revealed a broadening of the momentum distribution on a time scale of 1000 s. Photodetachment of stored anions was used in the beam lifetime measurements. The detachment rate by anion collisions with residual-gas molecules was found to be extremely low. A residual-gas density below 140 cm$^{-3}$ is derived, equivalent to a room-temperature pressure below 10$^{-14}$ mbar. Fast atomic, molecular and cluster ion beams stored for long periods of time in a cryogenic environment will allow experiments on collision- and radiation-induced fragmentation processes of ions in known internal quantum states with merged and crossed photon and particle beams.

physics.atom-ph