arXiv ScienceSearch

arXiv subjects

Jin Chen

Publications and source records attributed to Jin Chen.

At least 19 recordsLinked to original sources

On general background of quantum non-invertible symmetry in 2D

We study the dual quantum symmetry $\mathrm{Rep}(G)$ in the two-dimensional theory $\widetilde{\mathfrak{T}}_{\mathrm{Rep}(G)}$ obtained by gauging a non-abelian symmetry $G$ of $\mathfrak{T}_G$, where for Lie groups $G$ the gauging is understood as flat gauging. We develop a general framework for computing partition functions of $\widetilde{\mathfrak{T}}_{\mathrm{Rep}(G)}$ in arbitrary non-invertible symmetry backgrounds, represented by topological defect networks of $\mathrm{Rep}(G)$, in terms of the partition functions of the original theory $\mathfrak{T}_G$. We also derive the inverse transformation, expressing partition functions of $\mathfrak{T}_G$ in $G$ backgrounds in terms of those of $\widetilde{\mathfrak{T}}_{\mathrm{Rep}(G)}$. We test the framework in several examples, including finite groups with multiplicity-free and higher-multiplicity fusion rules, and discuss a formal extension to compact Lie groups, focusing on $\mathrm{Rep}(SU(2))$.

hep-th

Aspire: Can Models Self-Evolve from Vague Goals?

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.

cs.CL

Momentum-resolved EELS study of collective charge excitations in 1$T$-TaS$_2$

We use momentum-resolved electron energy-loss spectroscopy (M-EELS) to study the low-energy charge excitations of 1$T$-TaS$_2$ across the nearly commensurate-to-commensurate charge-density-wave (CDW) transition. Single-crystal x-ray diffraction and elastic M-EELS measurements confirm the expected rotation of the CDW wave vector upon entering the commensurate phase. In the nearly commensurate phase, the low-energy M-EELS spectra reveal an acoustic phonon branch and two optical phonon features whose energies and dispersions are broadly consistent with previous calculations and inelastic x-ray measurements. Across the transition, the optical phonon energies remain nearly unchanged, while their spectral intensity develops a pronounced temperature dependence near the CDW ordering wave vector. At higher energies, the finite-momentum charge response undergoes a substantial redistribution of spectral weight below the transition, consistent with the opening of an energy gap. These results demonstrate that M-EELS provides simultaneous access to lattice dynamics and finite-momentum valence band charge excitations in 1$T$-TaS$_2$, revealing their evolution across the commensurate CDW transition.

cond-mat.str-el

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

cs.AI

Controlling the dynamics of an electric-field-driven droplet on a lubricant-infused micropillar surface

As a non-contact control approach, electric field (EF) can be utilised to drive droplet dynamics on a lubricant-infused surface (LIS), with numerous potential applications ranging from drug manufacturing to 3D printing. However, the resulting droplet dynamics remain poorly understood, especially as there are several possible droplet lubrication states on LIS. Here, we develop a lattice Boltzmann scheme that fully captures the interplay between the interfacial flows and electrohydrodynamics and harness it to investigate EF driven droplets on micropillar LIS. Combining simulations and analytical calculations, we establish quantitative expressions for the drag force and the electric force acting on a moving droplet. We demonstrate that the models can accurately capture droplet dynamics during programmable manipulation, including periodic motion and long-distance transport. Such reliable theoretical models can potentially transform precision control of droplet dynamics by removing the reliance on trial and error tests.

physics.flu-dyn

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell the difference: they concentrate rules in the user turn, while coding-agent benchmarks emphasize final task success. We introduce Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads. To separate compliance from coincidence we introduce Against-Prior Accuracy (AP-Acc), which scores only rules labeled as opposing unprompted defaults, observed by re-running tasks with the rule withheld across nine probe builds and curated otherwise. Across 12 frontier models, accuracy spans 72.1-85.9% and AP-Acc 66.1-78.6%; every model is worse on against-prior rules, by 3.6 to 7.4 points (mean 5.81), and the direction survives a common-support analysis with item-clustered intervals. Aggregate scores therefore overstate compliance by a model-specific margin: prior control leaves the top build unchanged and exchanges three adjacent rank pairs. A counterbalanced conflict pilot on nine separate builds adds a second result: pooled precedence does not follow prompt depth, with system prompts, project files, and user instructions ahead of tool and skill descriptions.

cs.AI

Freeform super-oscillatory optics for CMOS-integrated THz super-resolution imaging

The diffraction limit fundamentally constrains the spatial resolution of far-field imaging systems. While near-field techniques can circumvent this limit, their inherently short working distances (WD) severely restrict practical applications. Super-oscillatory lenses (SOLs) offer a far-field alternative; however, conventional SOLs are plagued by discrete operating wavelengths, low efficiencies (below 5%), and formidable trade-offs among numerical aperture, chromatic aberration, and depth of focus (DOF). Here, we introduce a nonlocal, nonlinear-curvature mechanism to design a freeform SOL that achieves ultrabroadband (0.3 to 1 THz), achromatic super-resolution focusing with an unprecedented efficiency of 44%. Operating at a 9 mm WD, the lens maintains a consistent sub-diffraction full-width at half-maximum (FWHM) of around 0.45 wavelength alongside an extended DOF of around 10 wavelengths. By integrating a compact 65-nm CMOS oscillator-radiator array, we establish an advanced imaging platform capable of resolving complex 2D and 3D sub-millimeter features (down to 0.15 mm). Readily scalable to the optical regime via two-photon lithography, this freeform SOL paradigm paves the way for next-generation, high-performance integrated photonics.

physics.optics

Zero-Mem: Zero-Token Memory Operations for LLM Agents

LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{https://github.com/TheMoon0815/Zero-mem}.

cs.CL

A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models

Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time. In this paper, we propose an infer-diagnose-refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.

cs.RO

An integrated super resolution THz 3D imaging system based on a linear nonlocal achromatic freeform Bessel beam lens and high power oscillator radiator array

High performance terahertz (THz) 3D imaging is critical for non-destructive evaluation. However, conventional architectures are fundamentally limited by severe chromatic aberrations, modest spatial resolution, restricted depths of focus (DOF), and the bulky nature of commercial transceivers. While metasurfaces offer a compact alternative, achieving broadband achromatic super-resolution with an extended DOF remains a formidable challenge. Here, we present a highly integrated 3D THz imaging platform that synergizes a 3D printed nonlocal freeform Bessel-beam lens with a high power, 65nm CMOS oscillator radiator array. Harnessing nonlocal interactions within the lens, we generate an achromatic super resolution Bessel beam (0.3 to 1 THz) with a subdiffraction full width at half maximum (FWHM) of 0.65{\lambda} and a robust 4.7-mm DOF. Crucially, the system overcomes conventional sidelobe limitations, enabling high-fidelity 2D imaging of intricate sub-millimeter targets (e.g., USAF 1951 charts and QR codes) alongside robust 3D volumetric imaging through highly scattering media, such as printed circuit boards. By converging standard CMOS technology with additive manufacturing, this work establishes a versatile, cost-effective paradigm for next-generation integrated THz photonics

physics.optics

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that language. We call this the Illusion of Cultural Alignment. To test this assumption directly, we introduce MSQA, a benchmark of 1,064 natively sourced questions across 11 language groups, five cultural dimensions, and three difficulty tiers. Unlike translated benchmarks, MSQA targets locally grounded knowledge and reduces shortcuts from English-centric cross-lingual transfer. Evaluating 18 LLMs, we find substantial cultural degradation and a pronounced Locality Effect: cultural competence tracks pre-training exposure more closely than general reasoning ability. We further show that common inference-time remedies do not dissolve the illusion. Models remain overconfident on unfamiliar cultural questions, repeated sampling yields unstable rather than reliable correctness, and retrieval augmentation helps unevenly on long-tail facts. These findings indicate that cultural alignment cannot be inferred from multilingual ability alone and requires deeper intervention than calibration, sampling, or retrieval at inference time

cs.CL

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical user interfaces to complete long-horizon, high-value professional workflows across diverse domains. Current GUI benchmarks still predominantly focus on general-purpose software, relatively simple applications, and short-horizon tasks, leaving it largely unknown whether modern agents can follow user instructions to autonomously operate domain-specific professional software and accomplish economically valuable work in an end-to-end manner. To bridge this gap, we introduce Workflow-GYM, a benchmark for long-horizon GUI tasks centered on professional domains and specialized software environments. Through extensive experiments on state-of-the-art models, we find that even the strongest models achieve only slightly above 30% success rates, highlighting that professional long-horizon GUI workflows remain highly challenging for current GUI agents. Further analysis reveals that current agents struggle to maintain long-horizon workflow consistency, frequently exhibiting workflow stage omission, error propagation, objective drift, and insufficient understanding of professional software environments. Our findings provide important insights into the limitations of current agent systems and suggest key directions for the next generation of GUI-agent research.

cs.AI

Defect Conformal Manifolds along RG Domain Walls between $\mathbb Z_N$-Parafermions and Minimal Models

We investigate the renormalization group (RG) domain walls interpolating between the $\mathbb{Z}_N$ parafermion theory (the critical $N$-state Potts model) and the Virasoro minimal model $\mathcal{M}_{N+1}$. These flows are genuinely non-perturbative and an explicit construction of Gaiotto type RG domain wall remains elusive. We bypass this limitation by employing a bottom-up approach centered on the emergence of ``phantom currents". By tracking the preserved non-invertible symmetries ($\mathfrak{so}(3)_N$) along the flow, we extract the exact spectrum of these currents localized on the defect. We demonstrate that the presence of a spin-1 phantom current allows the interface to be marginally deformed, dynamically generating a continuous defect conformal manifold. Furthermore, we show that an extra spin-2 operator, crucially as a $W^{(3)}$-algebra descendant of the spin-1 phantom current, rigidly constrains the UV-IR stress tensor mixing via the cluster decomposition principle. This algebraic framework enables the exact computation of the parameter-dependent transmission rate across the conformal manifold, which we observe strictly vanishes in the large-$N$ limit as a consequence of macroscopic target space collapse.

hep-th

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation

Scaling recommendation models is a central challenge in recommender systems. Recently, RankMixer has emerged as an effective solution, operating on a unified token representation and alternating between token mixing and per-token feedforward networks (P-FFNs) to achieve scalable performance. However, RankMixer suffers from \textit{embedding collapse}, where learned representations have low effective rank, limiting expressivity and underutilizing the expanded representation space. Through empirical analysis and theoretical insights, we identify rigid token mixing and P-FFN modules as the primary causes of this phenomenon, jointly inducing a \textbf{damped oscillatory trajectory} in effective-rank evolution across layers. To address it, we propose RankElastor, a novel architecture that produces spectrum-robust representations with provable collapse mitigation. RankElastor introduces two components: (i) \textbf{parameterized full mixing}, which enables expressive token mixing with improved spectral robustness; and (ii) \textbf{GLU-improved P-FFNs}, which stabilize representation spectra through GLU-style FFN modules. Extensive experiments on large-scale industrial datasets demonstrate that RankElastor consistently improves recommendation performance, mitigates embedding collapse, and exhibits robust scaling behavior. Code is available at this GitHub repository: https://github.com/vasile-paskardlgm/RankElastor

cs.LG

RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction

The use of CT imaging is important for screening, diagnosis, therapy planning, and prognosis of lung cancers. Unfortunately, due to differences in imaging protocols and scanner models, CT images acquired by different means may show large differences in noise statistics, contrast, and texture. In this study, we develop a novel conditional MeanFlow pipeline for CT image reconstruction. We introduce a conditional MeanFlow network that models the reconstruction trajectory by predicting image-conditioned flow fields given intermediate image states. The image reconstruction network is trained with a MeanFlow consistency loss along with the image reconstruction loss. In order to provide a spatially adaptive refinement process, we integrate a regional reinforcement learning-driven policy network into our approach. The policy network receives information about the MeanFlow rollouts and provides predictions in terms of tile-wise refinement budgets, stopping criteria, and total budget allocation of refinement processes. Our policy network is trained through reinforcement learning in a policy gradient framework, where the goal of the training reward is to maximize reconstruction quality while minimizing unnecessary computations and avoiding instabilities. In this way, our approach combines conditional flow-based reconstruction with reinforcement learning-based spatial reconstruction control. Our results show high accuracy in the tumor ROI, with the average radiomic feature CCC being $0.93 \pm 0.09$, an average PSNR of $31.94 \pm 2.64$, and average SSIM of $0.97 \pm 0.03$. Moreover, there is an improvement in the overall quality of images, with an average PSNR of $34.23 \pm 1.71$ and average SSIM of $0.95 \pm 0.01$.

cs.CV

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

cs.CL

RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems

The scaling laws for recommender systems have been increasingly validated, where MetaFormer-based architectures consistently benefit from increased model depth, hidden dimensionality, and user behavior sequence length. However, whether representation capacity scales proportionally with parameter growth remains unexplored. Prior studies on RankMixer reveal that the effective rank of token representations exhibits a damped oscillatory trajectory across layers, failing to increase consistently with depth and even degrading in deeper layers. Motivated by this observation, we propose RankUp, an architecture designed to mitigate representation collapse and enhance expressive capacity through randomized permutation splitting over sparse features, a multi-embedding paradigm, global token integration and crossed pretrained embedding tokens. RankUp has been fully deployed in large-scale production across Weixin Video Accounts, Official Accounts and Moments, yielding GMV improvements of 3.41%, 4.81% and 2.12%, respectively.

cs.IR

Classification of 2D Fermionic Systems with a $\mathbb Z_2$ Flavor Symmetry

We classify superfusion categories describing two-dimensional fermionic systems equipped with the universal fermion-parity symmetry, implemented by a topological defect line (TDL) $Z$, and an additional $\mathbb{Z}_2$ flavor symmetry generated by a $W$ TDL. Depending on whether $W$ is m-type or q-type, its fusion rules lead to three distinct classes, and solving the super-pentagon equations yields 16 consistent superfusion categories. These are labeled by invariants $(\nu_W,\nu_Z,\nu_{WZ})$, which determine the $\mathbb{Z}_8$ anomaly classes of the symmetries generated by $W$, $Z$, and $WZ$. We also provide explicit realizations using multiple Majorana fermions and comment on implications for fermionic CFTs and gapped phases.

hep-th