arXiv ScienceSearch

arXiv subjects

Shiyang Chen

Publications and source records attributed to Shiyang Chen.

At least 19 recordsLinked to original sources

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show that this context-management layer is a safety-critical failure surface: in-context governance constraints that agents reliably obey while visible can be silently removed by compaction, causing the same agent to perform prohibited tool actions later in the session. We call this failure mode Governance Decay. We introduce ConstraintRot, a benchmark of long-horizon agent scenarios with deterministic tool-call grading, and measure compaction-induced violations across seven model families. Across 1,323 episodes, violation rises from 0% with the policy in full context to 30% after compaction, reaching 59% for some models; when the constraint survives the summary, violation remains 0%, but when it is dropped, violation reaches 38%. We further study a Compaction-Eviction Attack, in which adversarial in-context content biases the summarizer to omit a legitimate policy, and show that optimized injections defeat every evaluated model. Finally, we propose Constraint Pinning, a simple training-free mitigation that quarantines governance constraints from lossy compaction and restores violation to 0% in our benchmark. These results identify context management as a first-class governance surface for deployed LLM agents.

cs.AI

Electromagnetic Shower Reconstruction and Identification in FASER's Emulsion Detector for LHC Forward Neutrino Measurements

We present methods for electromagnetic shower reconstruction and identification in the FASERnu emulsion detector using 100 GeV and 200 GeV electron test-beam data from the CERN SPS H4 beamline. The reconstruction employs a clustering-based algorithm without energy-dependent tuning to determine shower axes. A multi-level identification chain comprising track pre-selection, a cut-based selection, and a BDT classifier achieves combined background rejection rates of 99.99% (100 GeV) and 99.94% (200 GeV). The method reaches total reconstruction and identification efficiencies of 58.9% (100 GeV) and 70.8% (200 GeV) evaluated from simulated samples. Energy reconstruction using the total number of reconstructed segments as the calorimetric estimator yields relative biases of +0.6% (100 GeV) and -0.8% (200 GeV), with resolutions of 25.4% and 22.6%, respectively. Systematic uncertainties on the energy reconstruction are dominated by variations in emulsion film detection efficiency, with totals of (+10.9%/-8.2%) at 100 GeV and (+10.3%/-6.9%) at 200 GeV. The methodology provides a validated framework for electron neutrino identification with the FASERnu detector at the LHC.

hep-ex

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a lens concurrent work sets aside -- the model's attention to labeled tool-definition segments. On real BFCL failures, by per-candidate attention argmax the model attends most to the correct tool 80% of the time (vs. 21% chance), and the gold is the under-attended segment on only 10%: it looks at the right tool and still picks wrong. This directly refutes the intuitive "crowded-harness / lost-in-the-middle" explanation: the failure is at the decision readout, not the harness, and we pin it there three ways. (1) Input vs. readout: repairing the prompt (reordering or duplicating the gold tool) recovers <=23% of failures, while readout-side interventions recover 59-91%. (2) Representation-invariance: two gold-pointed interventions in different representations -- an additive attention-logit bias and a residual-stream steering vector -- recover largely the same failures (per-task Jaccard 0.865 pooled, 0.79-0.91 per model), so the bottleneck is localized to the readout independent of which representation is poked. (3) A training-free, gold-free selector: per-segment attention closes most of the gold-free-vs-oracle gap on BFCL (+11.9 pts pooled function-name selection vs. +17.9-pt oracle headroom) and adds +14.9 pts on Seal-Tools; every model positive (exact McNemar p<=8e-4 each). Scopes differ: the causal attention-bias dose-response is bidirectional and monotonic on 10 mask-honoring models (3-32B), the full 0.5-32B span carrying only the correlational diagnostic; the deployable selector is evaluated on 5 single-turn models and does not yet transfer to a multi-turn loop.

cs.AI

Stochastic Path Sampler For Lattice Field Theory

In lattice field theory, target distributions are known only up to normalization, (\tilde{\pi}(\phi)\propto e^{-S(\phi)}), while the partition function is intractable. Markov chain Monte Carlo simulations often become inefficient near phase transitions or the continuum limit due to critical slowing down. In this work, we propose a novel sampler based on nonequilibrium thermodynamics, called Stochastic Path Sampler (SPS), which can generate configurations for the unnormalized target distribution without requiring training data. The central idea of SPS is to establish a trajectory-level balance for learnable forward and backward stochastic dynamics between two equilibrium states, namely the prior and target distributions. This is achieved by minimizing the path-space variational free energy, equivalently an entropy-production upper bound, defined by the log-ratio of forward and auxiliary backward trajectory measures, thereby enhancing the reversibility of the forward and backward processes. The learned forward process provides independent proposals, which are subsequently corrected by an extended-space Independence Metropolis--Hastings step. In two-dimensional (\phi^4) theory, we demonstrate that our neural sampler can achieve the same sampling quality as HMC but with a much shorter autocorrelation time in the critical region. This sampler offers a stochastic-quantization-inspired route to data-free proposal construction for lattice field theory by leveraging a variational free-energy principle derived from path-space irreversibility.

hep-lat

Stochastic first-passage modeling of single-event burnout in SiC power MOSFETs

Single-event burnout (SEB) in silicon carbide (SiC) power MOSFETs is often characterized by deterministic threshold quantities. Near the boundary between recovery and runaway, stochastic variability can make this threshold description probabilistic rather than sharp. This work introduces a first-passage perspective for stochastic threshold broadening in burnout. The process is described by a reduced electrothermal feedback-relaxation model with an absorbing boundary. The model combines carrier multiplication, avalanche feedback, localized heating, carrier loss, and thermal relaxation. Stochastic carrier and thermal terms represent unresolved event-level variability. The main finding is that finite fluctuations broaden the deterministic burnout threshold into a probabilistic transition band. Noise-induced subthreshold runaway also emerges, where nominally recoverable conditions can still fail through rare stochastic excursions. First-passage-time distributions resolve the time scale of burnout and survival probabilities further distinguish rapid feedback-dominated runaway from delayed stochastic failure. A feedback-relaxation phase diagram organizes recoverable, probabilistic, and rapidly unstable regimes. This framework provides a statistical-physics interpretation of threshold dispersion in single-event burnout of SiC power MOSFETs by linking coarse-grained electrothermal dynamics to probabilistic and time-resolved failure observables.

cond-mat.stat-mech

Momentum Measurement of Charged Particles in FASER's Emulsion Detector at the LHC

We present a momentum measurement method based on multiple Coulomb scattering (MCS) in the FASER$\nu$ emulsion detector. The measurement of charged-particle momenta is essential for studying neutrino interactions in the TeV energy range at the FASER experiment. This method exploits the sub-micron spatial resolution and long tracking length of the FASER$\nu$ detector, enabling momentum determination from a few GeV up to a few TeV. The performance was evaluated using Geant4-based Monte Carlo simulations and validated with muon test beam data in the momentum range 100-300 GeV. As a first probe of the method for higher momentum muons, background muons recorded by the FASER$\nu$ detector were examined, showing reconstructed momenta consistent with expectations from their angular spread.

hep-ex

Variational Autoregressive Networks Applied to $\phi^4$ Field Theory Systems

We combine reinforcement learning with variational autoregressive networks (VANs) to perform data-free training and sampling for the discrete Ising model and the continuous $\phi^4$ scalar field theory. We quantify the complexity of the target distribution via the KL divergence between the magnetization distribution and a reference Gaussian distribution, and observe that configurations with smaller KL divergence typically require fewer training steps. Motivated by this observation, we investigate transfer learning and show that fine-tuning models pretrained at a single value of $\kappa$ can reduce training time compared with training from a Gaussian field. In addition, inspired by single-site and cluster Monte Carlo updates, we introduce single-site and block Metropolis--Hastings (MH) updates on top of VAN proposals. These MH corrections systematically reduce the residual bias of pure VAN sampling in the parameter range we study, while maintaining high sampling efficiency in terms of the effective sample size (ESS). For both the Ising model and the $\phi^4$ theory, our results agree with standard Monte Carlo benchmarks within errors, and no clear critical slowing down is observed in the explored parameter ranges.

hep-lat

Incorporating GNSS Information with LIDAR-Inertial Odometry for Accurate Land-Vehicle Localization

Currently, visual odometry and LIDAR odometry are performing well in pose estimation in some typical environments, but they still cannot recover the localization state at high speed or reduce accumulated drifts. In order to solve these problems, we propose a novel LIDAR-based localization framework, which achieves high accuracy and provides robust localization in 3D pointcloud maps with information of multi-sensors. The system integrates global information with LIDAR-based odometry to optimize the localization state. To improve robustness and enable fast resumption of localization, this paper uses offline pointcloud maps for prior knowledge and presents a novel registration method to speed up the convergence rate. The algorithm is tested on various maps of different data sets and has higher robustness and accuracy than other localization algorithms.

cs.RO

Deal: Distributed End-to-End GNN Inference for All Nodes

Graph Neural Networks (GNNs) are a new research frontier with various applications and successes. The end-to-end inference for all nodes, is common for GNN embedding models, which are widely adopted in applications like recommendation and advertising. While sharing opportunities arise in GNN tasks (i.e., inference for a few nodes and training), the potential for sharing in full graph end-to-end inference is largely underutilized because traditional efforts fail to fully extract sharing benefits due to overwhelming overheads or excessive memory usage. This paper introduces Deal, a distributed GNN inference system that is dedicated to end-to-end inference for all nodes for graphs with multi-billion edges. First, we unveil and exploit an untapped sharing opportunity during sampling, and maximize the benefits from sharing during subsequent GNN computation. Second, we introduce memory-saving and communication-efficient distributed primitives for lightweight 1-D graph and feature tensor collaborative partitioning-based distributed inference. Third, we introduce partitioned, pipelined communication and fusing feature preparation with the first GNN primitive for end-to-end inference. With Deal, the end-to-end inference time on real-world benchmark datasets is reduced up to 7.70 x and the graph construction time is reduced up to 21.05 x, compared to the state-of-the-art.

cs.DC

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children or animals. Despite this, they attain high scores on many popular visual benchmarks, with headroom rapidly eroded by model progress. This creates a need for difficult benchmarks that remain relevant for longer. We introduce ZeroBench - a lightweight visual reasoning benchmark curated using adversarial filtering to be "impossible" for frontier LMMs at its original release, with initial SotA scores of 0% pass@1 and pass^5. We track progress on ZeroBench over the subsequent year, observing SotA reaching 6% pass^5 and 19% pass@5, indicating the potential longevity of the benchmark. We evaluate 46 LMMs on ZeroBench, compare performance to a human baseline, analyse strengths and weaknesses, chart a year of progress in visual capabilities, and publicly release ZeroBench at https://zerobench.github.io.

cs.CV

Exploring Generative Networks for Manifolds with Non-Trivial Topology

The expressive power of neural networks in modelling non-trivial distributions can in principle be exploited to bypass topological freezing and critical slowing down in simulations of lattice field theories. Some popular approaches are unable to sample correctly non-trivial topology, which may lead to some classes of configurations not being generated. In this contribution, we present a novel generative method inspired by a model previously introduced in the ML community (GFlowNets). We demonstrate its efficiency at exploring ergodically configuration manifolds with non-trivial topology through applications such as triple ring models and two-dimensional lattice scalar field theory.

hep-lat

KVDirect: Distributed Disaggregated LLM Inference

Large Language Models (LLMs) have become the new foundation for many applications, reshaping human society like a storm. Disaggregated inference, which separates prefill and decode stages, is a promising approach to improving hardware utilization and service quality. However, due to inefficient inter-node communication, existing systems restrict disaggregated inference to a single node, limiting resource allocation flexibility and reducing service capacity. This paper introduces KVDirect, which optimizes KV cache transfer to enable a distributed disaggregated LLM inference. KVDirect achieves this through the following contributions. First, we propose a novel tensor-centric communication mechanism that reduces the synchronization overhead in traditional distributed GPU systems. Second, we design a custom communication library to support dynamic GPU resource scheduling and efficient KV cache transfer. Third, we introduce a pull-based KV cache transfer strategy that reduces GPU resource idling and improves latency. Finally, we implement KVDirect as an open-source LLM inference framework. Our evaluation demonstrates that KVDirect reduces per-request latency by 55% compared to the baseline across diverse workloads under the same resource constraints.

cs.DC

PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips

We study a new vulnerability in commercial-scale safety-aligned large language models (LLMs): their refusal to generate harmful responses can be broken by flipping only a few bits in model parameters. Our attack jailbreaks billion-parameter language models with just 5 to 25 bit-flips, requiring up to 40$\times$ fewer bit flips than prior attacks on much smaller computer vision models. Unlike prompt-based jailbreaks, our method directly uncensors models in memory at runtime, enabling harmful outputs without requiring input-level modifications. Our key innovation is an efficient bit-selection algorithm that identifies critical bits for language model jailbreaks up to 20$\times$ faster than prior methods. We evaluate our attack on 10 open-source LLMs, achieving high attack success rates (ASRs) of 80-98% with minimal impact on model utility. We further demonstrate an end-to-end exploit via Rowhammer-based fault injection, reliably jailbreaking 5 models (69-91% ASR) on a GDDR6 GPU. Our analyses reveal that: (1) models with weaker post-training alignment require fewer bit-flips to jailbreak; (2) certain model components, e.g., value projection layers, are substantially more vulnerable; and (3) the attack is mechanistically different from existing jailbreak methods. We evaluate potential countermeasures and find that our attack remains effective against defenses at various stages of the LLM pipeline.

cs.CR

Learning phase transitions by siamese neural network

The wide application of machine learning (ML) techniques in statistics physics has presented new avenues for research in this field. In this paper, we introduce a semi-supervised learning method based on Siamese Neural Networks (SNN), trying to explore the potential of neural network (NN) in the study of critical behaviors beyond the approaches of supervised and unsupervised learning. By focusing on the (1+1) dimensional bond directed percolation (DP) model of nonequilibrium phase transition and the 2 dimensional Ising model of equilibrium phase transition, we use the SNN to predict the critical values and critical exponents of the systems. Different from traditional ML methods, the input of SNN is a set of configuration data pairs and the output prediction is similarity, which prompts to find an anchor point of data for pair comparison during the test. In our study, during test we set different bond probability $p$ or temperature $T$ as anchors, and discuss the impact of the configurations at this anchors on predictions. In addition, we use an iterative method to find the optimal training interval to make the algorithm more efficient, and the prediction results are comparable to other ML methods.

physics.comp-ph

FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design

Six-bit quantization (FP6) can effectively reduce the size of large language models (LLMs) and preserve the model quality consistently across varied applications. However, existing systems do not provide Tensor Core support for FP6 quantization and struggle to achieve practical performance improvements during LLM inference. It is challenging to support FP6 quantization on GPUs due to (1) unfriendly memory access of model weights with irregular bit-width and (2) high runtime overhead of weight de-quantization. To address these problems, we propose TC-FPx, the first full-stack GPU kernel design scheme with unified Tensor Core support of float-point weights for various quantization bit-width. We integrate TC-FPx kernel into an existing inference system, providing new end-to-end support (called FP6-LLM) for quantized LLM inference, where better trade-offs between inference cost and model quality are achieved. Experiments show that FP6-LLM enables the inference of LLaMA-70b using only a single GPU, achieving 1.69x-2.65x higher normalized inference throughput than the FP16 baseline. The source code is publicly available at https://github.com/usyd-fsalab/fp6_llm.

cs.LG

ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks

This study examines 4-bit quantization methods like GPTQ in large language models (LLMs), highlighting GPTQ's overfitting and limited enhancement in Zero-Shot tasks. While prior works merely focusing on zero-shot measurement, we extend task scope to more generative categories such as code generation and abstractive summarization, in which we found that INT4 quantization can significantly underperform. However, simply shifting to higher precision formats like FP6 has been particularly challenging, thus overlooked, due to poor performance caused by the lack of sophisticated integration and system acceleration strategies on current AI hardware. Our results show that FP6, even with a coarse-grain quantization scheme, performs robustly across various algorithms and tasks, demonstrating its superiority in accuracy and versatility. Notably, with the FP6 quantization, \codestar-15B model performs comparably to its FP16 counterpart in code generation, and for smaller models like the 406M it closely matches their baselines in summarization. Neither can be achieved by INT4. To better accommodate various AI hardware and achieve the best system performance, we propose a novel 4+2 design for FP6 to achieve similar latency to the state-of-the-art INT4 fine-grain quantization. With our design, FP6 can become a promising solution to the current 4-bit quantization methods used in LLMs.

cs.CL

Applications of Domain Adversarial Neural Network in phase transition of 3D Potts model

Machine learning techniques exhibit significant performance in discriminating different phases of matter and provide a new avenue for studying phase transitions. We investigate the phase transitions of three dimensional $q$-state Potts model on cubic lattice by using a transfer learning approach, Domain Adversarial Neural Network (DANN). With the unique neural network architecture, it could evaluate the high-temperature (disordered) and low-temperature (ordered) phases, and identify the first and second order phase transitions. Meanwhile, by training the DANN with a few labeled configurations, the critical points for $q=2,3,4$ and $5$ can be predicted with high accuracy, which are consistent with those of the Monte Carlo simulations. These findings would promote us to learn and explore the properties of phase transitions in high-dimensional systems.

physics.comp-ph

Exploring percolation phase transition in the three-dimensional Ising model with machine learning

The percolation study offers valuable insights into the characteristics of phase transition, shedding light on the underlying mechanisms that govern the formation of global connectivity within the system. We explore the percolation phase transition in the 3D cubic Ising model by employing two machine learning techniques. Our results demonstrate the capability of machine learning methods in distinguishing different phases during the percolation transition. Through the finite-size scaling analysis on the output of the neural networks, the percolation temperature and a correlation length exponent in the geometrical percolation transition are extracted and compared to those in the thermal magnetization phase transition within the 3D Ising model. These findings provide a valuable way essential for enhancing our understanding of the property of the QCD critical point, which belongs to the same universality class as the 3D Ising model.

nucl-th