arXiv ScienceSearch

arXiv subjects

Kairui Zhang

Publications and source records attributed to Kairui Zhang.

At least 19 recordsLinked to original sources

VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring and assigns each feature an intrinsic token name: the token string whose embedding is nearest to that feature. Without reducing reconstruction quality compared with a standard SAE, VASAE produces dictionaries with vocabulary-aligned features. Using a 0.8 cutoff on the nearest-token alignment score, dictionaries trained on GPT-2-small post-residual streams align about 90% of features in layers 0--10. In Llama-3.1-8B, representative shallow and middle-layer dictionaries contain strongly aligned features, including 92.8% in the shallow layer, while the representative final-layer dictionary shows limited alignment. After subtracting the sentence-level mean sparse code, case studies show that many remaining intrinsic token names are relevant to nearby input tokens. These results suggest that vocabulary-aligned anchoring can connect learned features to intrinsic token names during training, complementing post hoc interpretation of learned dictionaries.

cs.CL

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory. These obstacles undermine real-time responsiveness and cause severe forgetting throughout prolonged interactions. In this work, we introduce LiveStarPro, a live streaming assistant that is designed for proactive video understanding over long-horizon streams. The design of LiveStarPro rests on three complementary components. The first component is Streaming Verification Decoding (SVeD), an inference framework that identifies the appropriate response timing through single-pass perplexity verification, thereby eliminating the dependency on explicit silence tokens. The second component is Streaming Causal Attention Masks (SCAM), a training strategy that enforces incremental video-language alignment over variable-length streams. The third component is Tree-Structured Hierarchical Memory (TSHM), a recursive memory architecture that organizes evicted historical information into event chains and consequently enables efficient retrieval from effectively unbounded video streams. To facilitate a comprehensive evaluation under realistic online conditions, we further present OmniStarPro, a large-scale benchmark that spans 15 diverse real-world scenarios and that extends to hour-scale streams for the assessment of long-term recall. Extensive experiments demonstrate that LiveStarPro consistently surpasses existing methods, attaining a 28.9% improvement in semantic correctness and an 18.2% reduction in timing error, while its streaming key-value cache further yields a 1.58x inference speedup over the same model without caching. The model and the code are publicly available at https://github.com/sotayang/LiveStarPro.

cs.CV

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding

Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding. However, a critical challenge remains in streaming scenarios: existing models typically pause video perception while generating responses, breaking real-time video-language synchrony and causing stutters. To address this, we introduce a novel paradigm for online video understanding: Streaming Video-Language Synchrony (SVLS), and present LyraV, a live streaming assistant built upon a hierarchical control framework with two core innovations. First, the Frame-Driven Transition Controller (FDTC), a training-free verification-based finite-state machine, makes high-level semantic decisions on when to continue speaking, start a new response, or stay silent. Second, the Streaming Token Pacer (SToP), a plug-and-play lightweight predictive module, dynamically adapts the language generation rate to match the pace of the visual content. Concretely, LyraV performs \emph{per-frame incremental, sub-budget decoding}: within each frame interval it emits only a small chunk of tokens that fits the real-time budget, so perception is never blocked for a full sentence. Together, these components enable LyraV to seamlessly interleave incoming video frames with generated word tokens, achieving a fine-grained synchrony. Extensive experiments conducted on five online and three offline benchmarks demonstrate that LyraV preserves the backbone's general understanding ability while substantially improving streaming synchrony and narrative fluency, delivering a 98.29\% synchrony with video playback and a real-time processing speed of 3.89 FPS. Interestingly, we observe an empirical capability in LyraV: dynamic reasoning over streaming tokens, enabling continuous interpretation and "thinking" alongside visual input.

cs.CV

Natural SUSY with mixed axion/axino dark matter

While supersymmetric models provide a solution to the big hierarchy problem, natural SUSY is also allowed by the little hierarchy problem. In supersymmetric models which include the Peccei-Quinn (PQ) solution to the strong CP problem, one expects the presence of an axion-axino-saxion supermultiplet with a micro-eV-scale axion and a saxion with mass of order the soft breaking scale. The axino mass is much more model-dependent, and may occur in the range of keV-TeV: over 9 orders of magnitude. This leads to the possibility of the axino as lightest SUSY particle (LSP) and the presence of mixed axion plus axino dark matter. The case of natural SUSY with higgsino-like WIMPs as LSP seems (nearly) excluded by multi-ton noble liquid WIMP detector limits, even in the case where the LSP has a depleted abundance compared to axions. We examine the case where the axino is LSP leading to mixed axion-axino dark matter in a natural SUSY context. We map out regions of PQ scale f_a vs. axino mass m_{\ta} parameter space where such a scenario remains viable in both the SUSY DFSZ and KSVZ axion models. For axino mass ~100 keV, we find solutions in accord with the measured dark matter abundance with mainly warm axino dark matter for f_a~ 10^{11} GeV and also solutions with mainly axion cold DM and a tiny axino contribution for higher f_a~ 3\times 10^{12} GeV.

hep-ph

Learning Associations in Reconfigurable Particle Packings via Local Cyclic Driving

We investigate associative-memory behavior in a reconfigurable particle packing programmed by purely local cyclic driving. The system is a two-dimensional bidisperse Lennard--Jones particle assembly with periodic boundaries evolved under athermal quasistatic relaxation. During training, a fixed set of input particles is driven cyclically while output particles are selected on-the-fly by a region-driving rule and driven according to a prescribed flow pattern; during retrieval, only the inputs are driven. Associative-memory performance is quantified by the cosine similarity between realized and target output displacement directions. Unlike physical learning systems with fixed architecture, learning here arises through emergent weight updates: localized rearrangements modify the contact network and reshape the effective mechanical couplings between inputs and outputs. Across task difficulty we identify three regimes. In an easy setting, the intrinsic mechanical response already produces coherent motion in the right-hand region under input-only driving, yielding high performance without training. In a hard setting, the desired mapping conflicts with the dominant collective drift, resulting in low baseline performance and only modest training gains; introducing intermittent relaxation cycles reduces train--retrieval mismatch and improves performance. In an intermediate quadrupolar task, repositioning the input--output geometry stabilizes the desired response and converts initially stochastic trajectories into reproducible learned motions. Together these results identify minimal physical ingredients for association-based functionality in athermally driven particulate media and motivate an association learning phase diagram for reconfigurable matter.

cond-mat.dis-nn

pMSSM versus complete models and the excellent prospects for top-squark discovery at HL-LHC

LHC sparticle search limits are usually performed within the context of simplified models and subsequently interpreted within the 19 parameter phenomenological MSSM (pMSSM) as to how many models avoid search limits for a particular sparticle mass, often including WIMP dark matter constraints. We provide a critical discussion of this procedure and how it can go wrong due to the introduction of new prejudices. By ameliorating these conditions, one is pushed into the more plausible four extra parameter non-universal Higgs model (NUHM4). Implementing a decoupling/quasi-degeneracy solution to the SUSY flavor and CP problems leads to first/second generation sfermions in the tens-of-TeV range. In this case, the natural solutions typically contain top-squarks in the 1-2 TeV range which are accessible to high-lumi LHC (HL-LHC) searches. This search channel, along with higgsino and wino pair production, may allow a nearly complete scan of natural/plausible parameter space by HL-LHC.

hep-ph

WIMP Dark Matter from a Natural Discrete Gauge Symmetry in the Standard Model

The internal structure of the Standard Model implies a natural $\mathbb{Z}_4 \times \mathbb{Z}_3$ discrete gauge symmetry. Cancellation of the corresponding Dai--Freed anomalies requires the introduction of three right-handed neutrinos and three additional Majorana fermions $\chi_i$. This gauge symmetry forbids the decay of the lightest fermion $\chi_1$ into Standard Model particles, rendering it automatically stable and providing a dark matter candidate without introducing an ad hoc stabilizing symmetry and domain-wall problem. The mass of $\chi_1$ is generated by the vacuum expectation value of a singlet scalar near the electroweak scale, naturally realizing a weakly interacting massive particle (WIMP) freeze-out scenario. Dark matter annihilation proceeds through scalar mediation, allowing the observed relic abundance to be reproduced while remaining consistent with current direct-detection constraints. It naturally realizes the secluded dark matter scenario and can be further tested in the next generation of experiments.

hep-ph

Supersymmetry at High Luminosity LHC

Weak-scale supersymmetry (SUSY) is well motivated as a technically natural solution to the gauge hierarchy problem. LHC limits on superpartners, however, have sharpened the Little Hierarchy problem, raising the question of why $m_{weak} \ll m_{soft}$. We review current collider and dark matter constraints and their implications for leading SUSY-breaking scenarios. Electroweak naturalness is revisited using a weak-scale measure that avoids ambiguities associated with high-scale fine-tuning. Also, within the string landscape, soft terms are statistically favored to be large, while anthropic selection enforces a weak scale near the observed value. This framework often termed \emph{stringy naturalness} -- naturally accommodates $m_h \simeq 125$ GeV while placing sparticles above current LHC limits. Updated HL-LHC projections for non-universal Higgs mass (NUHM) models -- the most plausible realization consistent with above picture -- show that searches for higgsinos, stops, and heavy Higgs bosons will soon begin to probe the core of the viable parameter space.

hep-ph

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and determine optimal response timing, often compromising real-time responsiveness and narrative coherence. To address these limitations, we introduce LiveStar, a pioneering live streaming assistant that achieves always-on proactive responses through adaptive streaming decoding. Specifically, LiveStar incorporates: (1) a training strategy enabling incremental video-language alignment for variable-length video streams, preserving temporal consistency across dynamically evolving frame sequences; (2) a response-silence decoding framework that determines optimal proactive response timing via a single forward pass verification; (3) memory-aware acceleration via peak-end memory compression for online inference on 10+ minute videos, combined with streaming key-value cache to achieve 1.53x faster inference. We also construct an OmniStar dataset, a comprehensive dataset for training and benchmarking that encompasses 15 diverse real-world scenarios and 5 evaluation tasks for online video understanding. Extensive experiments across three benchmarks demonstrate LiveStar's state-of-the-art performance, achieving an average 19.5% improvement in semantic correctness with 18.1% reduced timing difference compared to existing online Video-LLMs, while improving FPS by 12.0% across all five OmniStar tasks. Our model and dataset can be accessed at https://github.com/yzy-bupt/LiveStar.

cs.CV

Natural supersymmetry at a muon collider

There is great interest within the particle physics community for building a $\mu^+\mu^-$ collider with center-of-mass (CoM) energies ranging from $\sqrt{s}\sim$ 1-14 TeV. For Beyond-the-Standard-Model (BSM) physics, natural supersymmetry seems perhaps the most motivated, plausible extension of the Standard Model. Here, we examine what can be accomplished by a muon collider with regards to natural SUSY at various muon collider CoM energies. In natural SUSY -- especially in the guise that would emerge from the string landscape -- one expects sparticles to be spread over two orders of magnitude in mass values. A muon collider with highly variable beam energies would be most useful for targeting 2-body reaction thresholds and Higgs boson resonances.

hep-ph

Aspects of the WIMP quality problem and R-parity violation in natural supersymmetry with all axion dark matter

In supersymmetric models where the mu problem is solved via discrete R-symmetries, then both the global U(1)_{PQ} (Peccei-Quinn, needed to solve the strong CP problem) and R-parity conservation (RPC, needed for proton stability) are expected to arise as accidental, approximate symmetries. Then in some cases, SUSY dark matter is expected to be all axions since the relic lightest SUSY particles (LSPs) can decay away via small R-parity violating (RPV) couplings. We examine several aspects of this {\it all axion} SUSY dark matter scenario. 1. We catalogue the operator suppression which is gained from discrete R-symmetry breaking via four two-extra-field base models. 2. We present exact tree-level LSP decay rates including mixing and phase space effects and compare to results from simple, approximate formulae. 3. Natural SUSY models are characterized by light higgsinos with mass ~100-350 GeV so that the dominant sparticle production cross sections at LHC14 are expected to be higgsino pair production which occurs at the 10^2-10^4 fb level. Assuming nature is natural, the lack of an RPV signal from higgsino pair production in LHC data translates into rather strong upper bounds on nearly all trilinear RPV couplings in order to render the SUSY signal (nearly) invisible. Thus, in natural SUSY models with light higgsinos, the RPV-couplings must be small enough that the LSP has a rather high quality of RPC.

hep-ph

TSCAN: Context-Aware Uplift Modeling via Two-Stage Training for Online Merchant Business Diagnosis

A primary challenge in ITE estimation is sample selection bias. Traditional approaches utilize treatment regularization techniques such as the Integral Probability Metrics (IPM), re-weighting, and propensity score modeling to mitigate this bias. However, these regularizations may introduce undesirable information loss and limit the performance of the model. Furthermore, treatment effects vary across different external contexts, and the existing methods are insufficient in fully interacting with and utilizing these contextual features. To address these issues, we propose a Context-Aware uplift model based on the Two-Stage training approach (TSCAN), comprising CAN-U and CAN-D sub-models. In the first stage, we train an uplift model, called CAN-U, which includes the treatment regularizations of IPM and propensity score prediction, to generate a complete dataset with counterfactual uplift labels. In the second stage, we train a model named CAN-D, which utilizes an isotonic output layer to directly model uplift effects, thereby eliminating the reliance on the regularization components. CAN-D adaptively corrects the errors estimated by CAN-U through reinforcing the factual samples, while avoiding the negative impacts associated with the aforementioned regularizations. Additionally, we introduce a Context-Aware Attention Layer throughout the two-stage process to manage the interactions between treatment, merchant, and contextual features, thereby modeling the varying treatment effect in different contexts. We conduct extensive experiments on two real-world datasets to validate the effectiveness of TSCAN. Ultimately, the deployment of our model for real-world merchant diagnosis on one of China's largest online food ordering platforms validates its practical utility and impact.

cs.LG

Prospects for supersymmetry at high luminosity LHC

Weak scale supersymmetry (SUSY) is highly motivated in that it provides a 't Hooft technically natural solution to the gauge hierarchy problem. However, recent strong limits from superparticle searches at LHC Run 2 may exacerbate a so-called Little Hierarchy problem (LHP) which is a matter of practical naturalness: why is m_{weak}<< m_{soft}? We review recent LHC and WIMP dark matter search bounds as well as their impact on a variety of proposed SUSY models: gravity-, gauge-, anomaly-, mirage- and gaugino-mediation along with some dark matter proposals such as well-tempered neutralinos. We address the naturalness question. We also address the emergence of the string landscape at the beginning of the 21st century and its impact on expectations for SUSY. Rather generally, the string landscape statistically prefers large soft SUSY breaking terms but subject to the anthropic requirement that the derived value of the weak scale for each pocket universe (PU) within the greater multiverse lies with the ABDS window of values. This {\it stringy natural} (SN) approach implies m_h~ 125 GeV more often than not with sparticles beyond or well-beyond present LHC search limits. We review detailed reach calculations of the high-lumi LHC (HL-LHC) for non-universal Higgs mass models which present perhaps the most plausible realization of SUSY from the string landscape. In contrast to conventional wisdom, from a stringy naturalness point of view, the search for SUSY at LHC has only just begun to explore the interesting regimes of parameter space. We comment on how non-universal Higgs models could be differentiated from other expressions of natural SUSY such as natural anomaly-mediation and natural mirage mediation at HL-LHC.

hep-ph

All axion dark matter from supersymmetric models

Supersymmetric models accompanied by certain anomaly-free discrete R-symmetries Z_n^R are attractive in that 1. the R-symmetry (which can arise from compactified string theory as a remnant of the broken 10-d Lorentz symmetry) forbids unwanted superpotential terms while allowing for the generation of an accidental, approximate global U(1)_{PQ} symmetry needed to solve the strong CP problem and 2. they provide a raison d'etre for an otherwise ad-hoc R-parity conservation. We augment the minimal supersymmetric Standard Model (MSSM) by two additional Z_n^R- and PQ-charged fields X and Y wherein SUSY breaking at an intermediate scale m_{hidden} leads to PQ breaking at a scale f_a\sim 10^{11} GeV leading to a SUSY DFSZ axion. The same SUSY breaking can trigger R-parity breaking via higher-dimensional operators leading to tiny R-violating couplings of order (f_a/m_P)^N and a WIMP quality problem. For Z_4^R and Z_8^R, we find only an N=1 suppression. Then the lightest SUSY particle (LSP) of the MSSM becomes unstable with a lifetime of order ~ 10^{-3}-10 seconds so the LSPs all decay away before the present epoch. That leaves a universe with all axion cold dark matter and no WIMPs in accord with recent LZ-2024 WIMP search results.

hep-ph

Implications of Higgs mass for hidden sector SUSY breaking

Hidden sector SUSY breaking where charged hidden sector fields obtain SUSY breaking vevs once seemed common in dynamical SUSY breaking (DSB). In such a case, scalars can obtain large masses but gauginos and A-terms gain loop-suppressed anomaly-mediated contributions which may be smaller by factors of 1/16\pi^2 ~1/160. This situation leads to models such as PeV or mini-split supersymmetry with m(scalars)~ 160 m(gauginos). In order to generate a light Higgs mass m_h~ 125 GeV, the scalar mass terms are required in the 10-100 TeV range, leading to large, unnatural contributions to the weak scale. Alternatively, in gravity mediation with singlet hidden sector fields, then m(scalars)~ m(gauginos)~ A-terms and the large A-terms lift m_h ->125 GeV even for natural values of m(stop1)~ 1-3 TeV. Requiring naturalness, which is probabilistically preferred by the string landscape, then the measured Higgs mass seems to favor singlets in the hidden sector, which can be common in metastable and retrofitted DSB models.

hep-ph

Testable Flavored TeV-scale Resonant Leptogenesis with MeV-GeV Dark Matter in a Neutrinophilic 2HDM

We explore flavored resonant leptogenesis embedded in a neutrinophilic 2HDM. Successful leptogenesis is achieved by the very mildly degenerate two heavier right-handed neutrinos (RHNs), $N_2$ and $N_3$, with mass splitting of only $\Delta M_{32}/M_2 \sim \mathcal{O}(0.1\%-1\%)$. The lightest RHN, with MeV-GeV-scale mass, lies below the sphaleron freeze-out temperature and remains stable, serving as a dark matter candidate. The model enables TeV-scale leptogenesis while avoiding the extreme mass degeneracy plagued conventional resonant leptogenesis. Baryon asymmetry, neutrino masses, and potentially the dark matter relic density can be addressed within a unified and experimentally testable framework.

hep-ph

Living dangerously with decoupled first/second generation scalars: SUSY prospects at the LHC

The string landscape statistical draw to large scalar soft masses leads to a mixed quasi-degeneracy/decoupling solution to the SUSY flavor and CP problems where first/second generation matter scalars lie in the 20-40 TeV range. With increasing first/second generation scalars, SUSY models actually become more natural due to two-loop RG effects which suppress the corresponding third generation soft masses. This can also lead to substantial parameter space regions which are forbidden by the presence of charge and/or color breaking (CCB) minima of the scalar potential. We outline the allowed SUSY parameter space for the gravity-mediated three extra-parameter-non-universal Higgs model NUHM3. The natural regions with m_h~ 125 GeV, \Delta_{EW}<~ 30 and decoupled first/second generation scalar are characterized by rather heavy gluinos and EW gauginos, but with rather small \mu and top-squarks not far beyond LHC Run 2 limits. This scenario also explains why SUSY has so far eluded discovery at LHC in that the parameter space with small scalar and gaugino masses is all excluded by the presence of CCB minima.

hep-ph

Decoding the gaugino code, naturally, at high-lumi LHC

Natural supersymmetry with light higgsinos is most likely to emerge from the string landscape since the volume of scan parameter space shrinks to tiny volumes for electroweak unnatural models. Rather general arguments favor a landscape selection of soft SUSY breaking terms tilted to large values, but tempered by the atomic principle: that the derived value of the weak scale in each pocket universe lie not too far from its measured value in our universe. But that leaves (at least) three different paradigms for gaugino masses in natural SUSY models: unified (as in nonuniversal Higgs models), anomaly-mediation form (as in natural AMSB) and mirage mediation form (with comparable moduli- and anomaly-mediated contributions). We perform landscape scans for each of these, and show they populate different, but overlapping, positions in m(\ell\bar{\ell}) and m(wino) space. The first of these may be directly measurable at high-lumi LHC via the soft opposite-sign dilepton plus jets plus MET signature arising from higgsino pair production while the second of these could be extracted from direct wino pair production leading to same-sign diboson production.

hep-ph