arXiv ScienceSearch

arXiv subjects

G. Sun

Publications and source records attributed to G. Sun.

At least 19 recordsLinked to original sources

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.

cs.AI

Constraints on non-standard neutrino interactions from Borexino extended data-set

Neutrino non-standard interactions (NSI) constitute an active research field, as they are closely related to potential new physics associated with dark matter searches and exotic interactions arising from fundamental symmetry violations. The Borexino's unprecedented sensitivity to solar neutrinos, derived from its low background and precision spectral measurements, enables stringent constraints on potential deviations from the standard three-flavor neutrino oscillation paradigm. This work presents an update on the analysis of flavor-diagonal NSI using the full Borexino Phase-III data set, extending the study previously reported in Constraints on flavor-diagonal non-standard neutrino interactions from Borexino Phase-II (JHEP 2020, 38). The updated analysis incorporates the extended temporal and statistical coverage of Phase-III. The results indicate improved sensitivity to the diagonal NSI parameters, with constraints exceeding those obtained in Phase-II. Furthermore, a more general analysis that includes all possible off-diagonal NSI terms is presented for the first time, providing a comprehensive exploration of the NSI parameter space associated with the flavors of the incoming and outgoing neutrinos. This work once again underlines Borexino's critical role in probing new physics scenarios and reinforces its legacy in neutrino research. Detailed comparisons with Phase-II results are discussed, along with implications for theoretical models of NSI.

hep-ex

ELUTQ: Optimizing Quantization Accuracy under LUT-Based Computation for Edge LLMs

Weight quantization effectively reduces memory consumption and enable the deployment of Large Language Models on edge devices, yet existing hardware-friendly methods often rely on uniform quantization, which suffers from poor weight-distribution fitting and high dequantization overhead under low-bit settings. In this paper, we propose ELUTQ, an efficient quantization framework featuring a novel quantization format termed Hierarchical Linear Quantization (HLQ). HLQ is designed to better capture the statistical characteristics of weights and eliminate dequantization overhead using Bit-serial LUT-based GEMM operations. HLQ significantly improves model accuracy under low-bit settings and achieves performance comparable to QAT methods without any retraining of the weights. Moreover, an optimized quantization pipeline is integrated into ELUTQ, enabling it to complete the quantization of LLaMA 3.1-70B using only 64 GB of CPU memory and 48 GB of VRAM, reducing the hardware requirements for large-scale model quantization. To enable efficient deployment on edge devices, ELUTQ designs high-performance kernels to support end-to-end inference. Our 2-bit LLaMA3.1-8B achieves 1.5x speedup over AWQ on RTX 3090. Code is available at https://github.com/Nkniexin/ELUTQ.

cs.LG

Neutral pressure measurement in TCV tokamak using ASDEX-type pressure gauges

Probing the neutral gas distribution at the edge of magnetic confinement fusion devices is critical for plasma exhaust studies. In the TCV tokamak, a set of ASDEX-type hot ionization pressure gauges (APGs) has been installed for fast, in-situ measurements of the neutral pressure distribution in the TCV chamber. The APGs have been calibrated against baratron pressure gauges (BPGs) for pressures ranging from less than 1 mPa to several hundred mPa. A correction to account for the residual pressure in the pumped torus is proposed to improve the measurement accuracy. APG measurements in a series of plasma discharges with varied density ramp rates are analyzed and compared with the BPG pressure measurements. APG measurements feature a significantly faster time response and extend the BPG measurement range to lower pressures. Systematically higher neutral pressures measured with APGs compared to BPGs connected to the same TCV port, are attributed to the BPG's slower time response and a nonuniform neutral distribution in gauge ports during the plasma discharge. The initial APG operations in TCV have been proven successful, which validates the APG as an adequate pressure measurement technique for the upcoming TCV divertor upgrades.

physics.plasm-ph

SOLPS-ITER simulation of an X-point radiator in TCV

SOLPS-ITER simulation is performed to reproduce the X-point radiator recently observed in nitrogen-seeded TCV experiments, which is a scenario that may be favorable to solve the power exhaust problems in future fusion devices. The simulations reveal the transition from the detached regime without XPR to the XPR regime, when increasing the nitrogen seeding rate. A cold X-point core surrounded by ionizing and radiative mentals is formed inside the separatrix and slightly above the X-point, where more than 90% of the total input power is dissipated. The cold X-point core exhibits a temperature of approximately 1eV and features high recombination rate to host the convective fluxes from the ionizing mental. Increasing nitrogen seeding rate also moves the nitrogen ionization front away from the target faster than the nitrogen stagnation point, which enhances the divertor nitrogen leakage to the main chamber and benefits the XPR region cooling. Carbon radiation decreases as the nitrogen seeding increases, and carbon radiation contributes to above 5% of the core impurity radiation before entering the XPR, which decreases to 2.8% when reaching the XPR. Both baffled and unbaffled divertor geometries are simulated and compared, showing that baffles facilitate the access to XPR by increasing the X-point neutral density, but requires higher seeding rate to enter the XPR regime.

physics.plasm-ph

Investigating the influence of divertor baffles on nitrogen-seeded detachment in TCV with SOLPS-ITER simulations and TCV experiments

Plasma edge simulations with the SOLPS-ITER code are performed to study the influence of divertor baffles on nitrogen-seeded detachment in TCV single-null, L-mode discharges. Scans of nitrogen seeding rate are conducted in both baffled and unbaffled TCV divertors, where the nitrogen seeding with baffles is found to yield lower target temperatures and heat fluxes than with baffles-only and with seeding-only. The cumulative effects of baffles and seeding on target parameters are explained by the two-point model. The divertor neutral density and neutral compression increase with baffles, due to lower divertor to main chamber neutral conductance, as explained by a schematic neutral transport model with baffles. The nitrogen retention, defined as the ratio of average nitrogen nuclei density in divertor and main chamber, increases with the seeding rate if baffled, and remains constant if unbaffled. At the same outboard mid-plane separatrix plasma density, the nitrogen retention with baffles is lower than the unbaffled retention at low seeding levels and is higher at high seeding levels, which is explained by the changes of nitrogen ion and neutral transport with baffles and seeding. The baffled carbon retention is higher than the unbaffled retention due to lower divertor to main chamber carbon neutral conductance. Baffles increase the divertor radiation. The predicted trends of target parameters, the distribution of neutrals and radiations are well supported by TCV experiments, though discrepancies in the absolute values remain. The simulations yield an overall colder and denser divertor, consistent with previous SOLPS-ITER simulations of Ohmically heated L-modes in TCV. The successful comparison of simulation and experiment, together with the understanding gained from the neutral transport model, increases the confidence in the SOLPS simulations for the next TCV divertor upgrade.

physics.plasm-ph

MAD: Meta Adversarial Defense Benchmark

Adversarial training (AT) is a prominent technique employed by deep learning models to defend against adversarial attacks, and to some extent, enhance model robustness. However, there are three main drawbacks of the existing AT-based defense methods: expensive computational cost, low generalization ability, and the dilemma between the original model and the defense model. To this end, we propose a novel benchmark called meta adversarial defense (MAD). The MAD benchmark consists of two MAD datasets, along with a MAD evaluation protocol. The two large-scale MAD datasets were generated through experiments using 30 kinds of attacks on MNIST and CIFAR-10 datasets. In addition, we introduce a meta-learning based adversarial training (Meta-AT) algorithm as the baseline, which features high robustness to unseen adversarial attacks through few-shot learning. Experimental results demonstrate the effectiveness of our Meta-AT algorithm compared to the state-of-the-art methods. Furthermore, the model after Meta-AT maintains a relatively high clean-samples classification accuracy (CCA). It is worth noting that Meta-AT addresses all three aforementioned limitations, leading to substantial improvements. This benchmark ultimately achieved breakthroughs in investigating the transferability of adversarial defense methods to new attacks and the ability to learn from a limited number of adversarial examples. Our codes and attacked datasets address will be available at https://github.com/PXX1110/Meta_AT.

eess.IV

Parallel flows as a key component to interpret Super-X divertor experiments

The Super-X Divertor (SXD) is an alternative divertor configuration leveraging total flux expansion at the Outer Strike Point (OSP). While the extended 2-Point Model (2PM) predicts facilitated detachment access and control in the SXD configuration, these attractive features are not always retrieved experimentally. These discrepancies are at least partially explained by the effect of parallel flows which, when self-consistently included in the 2PM, reveal the role of total flux expansion on the pressure balance and weaken the total flux expansion effect on detachment access and control, compared to the original predictions. This new model can partially explain the discrepancies between the 2PM and experiments performed on tokamak \`a configuration variable (TCV), in ohmic L-mode scenarios, which are particularly apparent when scanning the OSP major radius Rt. In core density ramps in lower Single-Null (SN) configuration, the impact of Rt on the CIII emission front movement in the divertor outer leg - used as a proxy for the plasma temperature in the divertor - is substantially weaker than 2PM predictions. Furthermore, in OSP radial sweeps in lower and upper SN configurations, in ohmic L-mode scenarios with a constant core density, the peak parallel particle flux density at the OSP is almost independent of Rt, while the 2PM predicts a linear dependence. Finally, analytical and numerical modeling of parallel flows in the divertor is presented. It is shown that an increase in total flux expansion can favour supersonic flows at the OSP. Parallel flows are also shown to be relevant by analysing SOLPS-ITER simulations of TCV.

physics.plasm-ph

Performance assessment of a tightly baffled, long-legged divertor configuration in TCV with SOLPS-ITER

Numerical simulations explore the possibility to test the tightly baffled, long-legged divertor (TBLLD) concept in a future upgrade of the Tokamak \`a configuration variable (TCV). The SOLPS-ITER code package is used to compare the exhaust performance of several TBLLD configurations with existing unbaffled and baffled TCV configurations. The TBLLDs feature a range of radial gaps between the separatrix and the outer leg side walls. All considered TBLLDs are predicted to lead to a denser and colder plasma in front of the targets and improve the power handling by factors of 2-3 compared to the present, baffled divertor and by up to a factor of 12 compared to the original, unbaffled configuration. The improved TBLLD performance is mainly due to a better neutral confinement with improved plasma-neutral interactions in the divertor region. Both power handling capability and neutral confinement increases when reducing the radial gap. The core compatibility of TBLLDs with nitrogen seeding is also evaluated and the detachment window with acceptable core pollution for the proposed TBLLDs is explored, showing a reduction of required upstream impurity concentration up to 18% to achieve the detachment with thinner radial gap.

physics.plasm-ph

The Gamow Explorer: A gamma-ray burst observatory to study the high redshift universe and enable multi-messenger astrophysics

The Gamow Explorer will use Gamma Ray Bursts (GRBs) to: 1) probe the high redshift universe (z > 6) when the first stars were born, galaxies formed and Hydrogen was reionized; and 2) enable multi-messenger astrophysics by rapidly identifying Electro-Magnetic (IR/Optical/X-ray) counterparts to Gravitational Wave (GW) events. GRBs have been detected out to z ~ 9 and their afterglows are a bright beacon lasting a few days that can be used to observe the spectral fingerprints of the host galaxy and intergalactic medium to map the period of reionization and early metal enrichment. Gamow Explorer is optimized to quickly identify high-z events to trigger follow-up observations with JWST and large ground-based telescopes. A wide field of view Lobster Eye X-ray Telescope (LEXT) will search for GRBs and locate them with arc-minute precision. When a GRB is detected, the rapidly slewing spacecraft will point the 5 photometric channel Photo-z Infra-Red Telescope (PIRT) to identify high redshift (z > 6) long GRBs within 100s and send an alert within 1000s of the GRB trigger. An L2 orbit provides > 95% observing efficiency with pointing optimized for follow up by the James Webb Space Telescope (JWST) and ground observatories. The predicted Gamow Explorer high-z rate is >10 times that of the Neil Gehrels Swift Observatory. The instrument and mission capabilities also enable rapid identification of short GRBs and their afterglows associated with GW events. The Gamow Explorer will be proposed to the 2021 NASA MIDEX call and if approved, launched in 2028.

astro-ph.HE

Content-Aware Speaker Embeddings for Speaker Diarisation

Recent speaker diarisation systems often convert variable length speech segments into fixed-length vector representations for speaker clustering, which are known as speaker embeddings. In this paper, the content-aware speaker embeddings (CASE) approach is proposed, which extends the input of the speaker classifier to include not only acoustic features but also their corresponding speech content, via phone, character, and word embeddings. Compared to alternative methods that leverage similar information, such as multitask or adversarial training, CASE factorises automatic speech recognition (ASR) from speaker recognition to focus on modelling speaker characteristics and correlations with the corresponding content units to derive more expressive representations. CASE is evaluated for speaker re-clustering with a realistic speaker diarisation setup using the AMI meeting transcription dataset, where the content information is obtained by performing ASR based on an automatic segmentation. Experimental results showed that CASE achieved a 17.8% relative speaker error rate reduction over conventional methods.

cs.SD

Transformer Language Models with LSTM-based Cross-utterance Information Representation

The effective incorporation of cross-utterance information has the potential to improve language models (LMs) for automatic speech recognition (ASR). To extract more powerful and robust cross-utterance representations for the Transformer LM (TLM), this paper proposes the R-TLM which uses hidden states in a long short-term memory (LSTM) LM. To encode the cross-utterance information, the R-TLM incorporates an LSTM module together with a segment-wise recurrence in some of the Transformer blocks. In addition to the LSTM module output, a shortcut connection using a fusion layer that bypasses the LSTM module is also investigated. The proposed system was evaluated on the AMI meeting corpus, the Eval2000 and the RT03 telephone conversation evaluation sets. The best R-TLM achieved 0.9%, 0.6%, and 0.8% absolute WER reductions over the single-utterance TLM baseline, and 0.5%, 0.3%, 0.2% absolute WER reductions over a strong cross-utterance TLM baseline on the AMI evaluation set, Eval2000 and RT03 respectively. Improvements on Eval2000 and RT03 were further supported by significance tests. R-TLMs were found to have better LM scores on words where recognition errors are more likely to occur. The R-TLM WER can be further reduced by interpolation with an LSTM-LM.

cs.CL

Cross-Utterance Language Models with Acoustic Error Sampling

The effective exploitation of richer contextual information in language models (LMs) is a long-standing research problem for automatic speech recognition (ASR). A cross-utterance LM (CULM) is proposed in this paper, which augments the input to a standard long short-term memory (LSTM) LM with a context vector derived from past and future utterances using an extraction network. The extraction network uses another LSTM to encode surrounding utterances into vectors which are integrated into a context vector using either a projection of LSTM final hidden states, or a multi-head self-attentive layer. In addition, an acoustic error sampling technique is proposed to reduce the mismatch between training and test-time. This is achieved by considering possible ASR errors into the model training procedure, and can therefore improve the word error rate (WER). Experiments performed on both AMI and Switchboard datasets show that CULMs outperform the LSTM LM baseline WER. In particular, the CULM with a self-attentive layer-based extraction network and acoustic error sampling achieves 0.6% absolute WER reduction on AMI, 0.3% WER reduction on the Switchboard part and 0.9% WER reduction on the Callhome part of Eval2000 test set over the respective baselines.

cs.CL

Phase transitions in the $\mathbb{Z}_p$ and U(1) clock models

Quantum phase transitions are studied in the non-chiral $p$-clock chain, and a new explicitly U(1)-symmetric clock model, by monitoring the ground-state fidelity susceptibility. For $p\ge 5$, the self-dual $\mathbb{Z}_p$-symmetric chain displays a double-hump structure in the fidelity susceptibility with both peak positions and heights scaling logarithmically to their corresponding thermodynamic values. This scaling is precisely as expected for two Beresinskii-Kosterlitz-Thouless (BKT) transitions located symmetrically about the self-dual point, and so confirms numerically the theoretical scenario that sets $p=5$ as the lowest $p$ supporting BKT transitions in $\mathbb{Z}_p$-symmetric clock models. For our U(1)-symmetric, non-self-dual minimal modification of the $p$-clock model we find that the phase diagram depends strongly on the parity of $p$ and only one BKT transition survives for $p\geq 5$. Using asymptotic calculus we map the self-dual clock model exactly, in the large $p$ limit, to the quantum $O(2)$ rotor chain. Finally, using bond-algebraic dualities we estimate the critical BKT transition temperatures of the classical planar $p$-clock models defined on square lattices, in the limit of extreme spatial anisotropy. Our values agree remarkably well with those determined via classical Monte Carlo for isotropic lattices. This work highlights the power of the fidelity susceptibility as a tool for diagnosing the BKT transitions even when only discrete symmetries are present.

cond-mat.str-el

Observation of Multiphoton Frequency Conversion in Superconducting Circuits

Multiphoton up/down conversion in a transmon circuit, driven by a pair of microwaves tuned near and far off the qubit resonance, has been observed. The experimental realization of these high order non-linear processes is accomplished in the three-photon regime, when the transmon is coupled to weak bichromatic microwave fields with the same Rabi frequencies. A many-mode Floquet formalism, with longitudinal coupling, is used to simulate the quantum interferences in the absorption spectrum that manifest the multiphoton pumping processes in the transmon qubit. An intuitive graph theoretic approach is used to introduce effective Hamiltonians that elucidate main features of the Floquet results. The analytical solutions also illustrate how controllability is achievable for desired single- or multiphoton pumping processes in a wide frequency range.

quant-ph

The Global 21-cm Signal in the Context of the High-z Galaxy Luminosity Function

Motivated by recent progress in studies of the high-$z$ Universe, we build a new model for the global 21-cm signal that is explicitly calibrated to measurements of the galaxy luminosity function (LF) and further tuned to match the Thomson scattering optical depth of the cosmic microwave background, $\tau_e$. Assuming that the $z \lesssim 8$ galaxy population can be smoothly extrapolated to higher redshifts, the recent decline in best-fit values of $\tau_e$ and the inefficient heating induced by X-ray binaries (HMXBs; the presumptive sources of the X-ray background at high-$z$) imply that the entirety of cosmic reionization and reheating occurs at redshifts $z \lesssim 12$. In contrast to past global 21-cm models, whose $z \sim 20$ ($\nu \sim 70$ MHz) absorption features and strong $\sim 25$ mK emission features were driven largely by the assumption of efficient early star-formation and X-ray heating, our new fiducial model peaks in absorption at $\nu \sim 110$ MHz at a depth of $\sim -160$ mK and has a negligible emission component. As a result, a strong emission signal would provide convincing evidence that HMXBs are not the only drivers of cosmic reheating. Shallow absorption troughs should accompany strong heating scenarios, but could also be caused by a low escape fraction of Lyman-Werner photons. Generating signals with troughs at $\nu \lesssim 95$ MHz requires a floor in the star-formation efficiency in halos below $\sim 10^{9} M_{\odot}$, which is equivalent to steepening the faint-end of the galaxy LF. These findings demonstrate that the global 21-cm signal is a powerful complement to current and future galaxy surveys and efforts to better understand the interstellar medium in high-$z$ galaxies.

astro-ph.GA

Topological quasi-one-dimensional state of interacting spinless electrons

By decreasing the transversal confinement potential in interacting one-dimensional spinless electrons and populating the second energetically lowest sub-band, for not too strong interactions system transitions into a quasi-one-dimensional state with dominant superconducting correlations and one gapless mode. By combining effective field theory approach and numerical density matrix renormalization group simulations we show that this quasi-one-dimensional state is a topological state that hosts zero-energy edge modes. We also study the single-particle correlations across the interface between this quasi-one-dimensional and single-channel states.

cond-mat.str-el

Ground-state phases of rung-alternated spin-1/2 Heisenberg ladder

The ground-state phase diagram of Heisenberg spin-1/2 system on a two-leg ladder with rung alternation is studied by combining analytical approaches with numerical simulations. For the case of ferromagnetic leg exchanges a unique ferrimagnetic ground state emerges, whereas for the case of antiferromagnetic leg exchanges several different ground states are stabilized depending on the ratio between exchanges along legs and rungs. For the more general case of a honeycomb-ladder model for the case of ferromagnetic leg exchanges besides usual rung-singlet and saturated ferromagnetic states we obtain a ferrimagnetic Luttinger liquid phase with both linear and quadratic low energy dispersions and ground state magnetization continuously changing with system parameters. For the case of antiferromagnetic exchanges along legs, different dimerized states including states with additional topological order are suggested to be realized.

cond-mat.str-el