arXiv ScienceSearch

arXiv subjects

Yukun Zhang

Publications and source records attributed to Yukun Zhang.

At least 19 recordsLinked to original sources

Welfare-Opaque Income: Taxation under AI-Agent Delegation

We study income taxation when an AI agent implements economically relevant choices through a rule hidden from the government. Alongside unobserved productive ability, this hidden preference-to-execution mapping creates \emph{double unobservability}: the same observable tax-base response can carry different welfare consequences. We call the resulting income \emph{welfare-opaque}. Our constructions show that tax-base statistics can coincide while reform welfare effects differ, even when mechanical welfare weights are identical. We derive an optimal-tax condition that adds a response-weighted execution wedge to the familiar sufficient statistics. A higher marginal rate gains a corrective benefit under local over-execution and an additional cost under local under-execution. Observing the wedge identifies the welfare effect of a marginal reform at the prevailing schedule; bounds on it deliver bounds on that effect. A controlled laboratory compares 4,500 model runs across five AI engines. Faithful delegation selects the score maximizer in essentially all runs. Conflicted objectives produce heterogeneous responses: Claude largely preserves the score maximizer, GLM moves predominantly downward, and GPT-mini and Qwen show concentrated lower-tail increases. Qwen also makes substantial downward adjustments. Different engines locate their departures at different points and in different directions of the designed distribution. Explicit scores align model rankings; formula-based objective instructions yield more uneven agreement. Qwen shows a clear positive tax-by-objective interaction, but its direction does not generalize across engines and the pooled sign depends on its inclusion. The analysis identifies execution information as a complement to conventional tax-base statistics.

cs.CY

The Organization of Inference: Information, Resource Constraints, and AI Production

The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow's binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value

cs.AI

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

Agent harnesses supply planning guidance, organize execution, and check completion. We study how these components affect success, erroneous acceptance, and cost in two Retail experiments and an Airline pilot in $τ^2$-bench. The primary comparison pairs prewritten task-specific plans (Fixed) with shuffled policy text matched in word count (Sham), isolating the contribution of guidance content. Across 265 matched cells, Fixed improves oracle-verified success by 7.17 percentage points (90\% task-clustered bootstrap interval, 1.15--13.36 points), with gains concentrated in higher-complexity tasks. A read-only terminal verifier rejects 61\% of Retail oracle-invalid episodes while withholding 17\% of correct ones, at less than one cent of additional cost per episode. Which component matters more depends on the loss assigned to erroneous acceptance: at low liability the planning gain dominates; at high liability the verifier's avoided false passes dominate---and a standalone verifier captures nearly all the false-pass benefit of the full planning-plus-verification stack at a fraction of its cost.

cs.AI

One- and two-dimensional cluster states for topological phase simulation and measurement-based quantum computation

Quantum entanglement is a fundamental resource for quantum information processing and serves as a critical benchmark for quantum hardware performance. Cluster states are a special class of entangled states that serve as universal resources for measurement-based quantum computation and possess an intrinsic symmetry-protected topological order, which confers robustness against symmetry-respecting noise. Here we report the scalable preparation and verification of genuine multipartite cluster states on the 105-qubit Zuchongzhi 3.1 superconducting processor. We achieve one-dimensional cluster states of up to 95 qubits and two-dimensional cluster states of up to 72 qubits. The symmetry-protected topological cluster states exhibit input-state-dependent robustness under symmetry-breaking perturbations due to an operational parity structure that enhances the performance of measurement-based quantum computation. Furthermore, we use our two-dimensional cluster states to implement the Deutsch-Jozsa algorithm within the measurement-based quantum computation framework, achieving higher output-state fidelity compared with traditional circuit-based models and a query efficiency advantage over classical approaches. Our work establishes a scalable platform that combines large-scale entanglement generation, symmetry-protected topological order and practical quantum algorithms to enable robust, fault-tolerant measurement-based quantum computation.

quant-ph

Quantum Advantage with Adaptive Shallow Circuits

Quantum advantage is widely expected to require sufficiently deep circuits, where correlations and global computational structure can grow beyond the reach of efficient classical simulation. This expectation is especially stark for constant-depth circuits with local readout: the expectation value of any fixed local observable lies within a bounded backward lightcone and is therefore classically tractable. Here we show that measurement feedback changes this picture. We establish a strict hierarchy of computational power: at fixed coherent depth, increasing the number of feedback outcomes strictly enlarges the class of functions accessible through a local expectation value. The two ends of this hierarchy exhibit distinct computational regimes. With logarithmic feedback, local expectation values for product-state inputs are efficiently classically simulable. Polynomial feedback, by contrast, enables an explicit family of adaptive shallow circuits to encode prime-field discrete logarithm problem~(DLP) into a fixed single-qubit expectation. Assuming the standard worst-case classical hardness of DLP, estimating this expectation value is classically hard. These results reveal a feedback-driven complexity transition, with further implications for resource lower bounds on DLP and the complexity of local-observable estimation under area-law entanglement. Our results open a new route to quantum advantage with shallow quantum circuits.

quant-ph

Efficient Lindbladian Learning from Constant-Time Pauli Responses

Learning the generator of an open many-body system is more challenging than Hamiltonian learning: local responses, which can directly reveal coherent interaction terms in closed-system dynamics, may also contain dissipative contributions in open-system dynamics. In this paper, we address this challenge by developing an efficient Lindbladian learning framework for a known local candidate generator dictionary with bounded dissipative support and either bounded dual-interaction-graph degree or bounded unweighted local strength. The framework resolves the coherent-dissipative ambiguity by treating local Pauli responses as a linear system over both types of generator terms. Inverting this response system separates their contributions and makes the individual Lindbladian coefficients accessible from local response data in a fixed short-time window. Within this framework, we develop two efficient learning algorithms: Chebyshev--Lobatto response interpolation, which uses logarithmically many short evolution times and has a post-mean cost linear in $M$, with the stated dependence on $ε$, and Single-time projected response contraction, which uses a single fixed evolution time and globally inverts a truncated response function. Both procedures estimate $M$ candidate coefficients to entrywise accuracy $ε$ using $\widetilde{\mathcal{O}}(M/ε^2)$ sample and classical post-processing complexity. Our theoretical results establish local response inversion as a scalable paradigm for learning, calibrating, and diagnosing complex quantum systems from experimentally accessible short-time data.

quant-ph

CausalMix: Data Mixture as Causal Inference for Language Model Training

In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.

cs.LG

Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents

AI agents increasingly operate inside digital accounts by exercising privileges that users already hold, raising a new control question: whether an existing account entitlement must be exercised manually or may be exercised through a user-authorized automated proxy. We define \emph{delegation rights} as the revocable, identity-preserving, scope-limited, and mode-specific authority of an account holder to authorize such proxy execution. We develop a three-party incomplete-contracts model with a User, an AI Agent provider, and a Platform. The contested object is not platform ownership, account transferability, data portability, or unrestricted API access, but residual control over the mode of account execution. Under Platform Control, the platform can protect infrastructure, identity systems, privacy boundaries, and third parties, but its discretionary veto weakens the User--Agent coalition's disagreement payoff and depresses relationship-specific investment. Under User Control, hold-up is reduced, but security, privacy, congestion, and third-party risks may remain insufficiently internalized. We then analyze \emph{Certified Delegation}, under which access protection is conditional on verifiable authorization, revocability, auditability, rate-limit compliance, data minimization, and risk mitigation. Certification is therefore not merely a technical safety screen; it is a conditional allocation of residual control. Illustrative mechanism simulations show how this regime can reduce deadweight loss by restoring delegation incentives while bounding residual risk.

econ.EM

Simultaneous Estimation of Partial-Transpose Moments with Active Memory Independent of the Moment Order

We study the simultaneous estimation of partial-transpose moments $p_j(ρ_{AB})=\mathrm{Tr}[(ρ_{AB}^{T_B})^j]$, $j=2,\ldots,K$, of an unknown bipartite $n$-qubit state from independent copies under an explicit active-memory constraint. We give a sequential qubit-reuse realization of the partial-transpose permutation that uses at most $2n+1$ active qubits, independent of $K$, and estimates all moments $p_2,\ldots,p_K$ to uniform additive error $ε$ with total copy complexity $O(K\log K/ε^2)$. We also prove two converse bounds. First, any uniformly accurate simultaneous estimator requires $Ω(K/ε^2)$ copies in the worst case. Second, the same scaling holds on an explicit isospectral two-qubit negative-partial-transpose (NPT) family whose ordinary moments are constant while the partial-transpose moments vary. These results characterize the copy complexity of the partial-transpose moment hierarchy up to a logarithmic factor and extend simultaneous nonlinear-functional estimation from ordinary state powers to partial-transpose spectral data under active quantum memory independent of the target moment order.

quant-ph

PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting

Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables remains challenging due to scale differences. To address this, we introduce patch-mean decoupling (PMD), which separates the trend and residual shape information by subtracting the mean of each patch, preserving the original structure and ensuring that the attention mechanism captures true shape similarities. Futhermore, to more effectively model long-range dependencies and capture cross-variable relationships, we propose Trend Restoration Attention (TRA) and Proximal Variable Attention (PVA). The former module reintegrates the decoupled trend from PMD while calculating attention output. And the latter focuses cross-variable attention on the most relevant, recent time segments to avoid overfitting on outdated correlations. Combining these components, we propose PMDformer, a model designed to effectively capture shape similarity in long-term forecasting scenarios. Extensive experiments indicate that PMDformer outperforms existing state-of-the-art methods in stability and accuracy across multiple LTSF benchmarks. The code is available at https://github.com/aohu1105/PMDformer.

cs.AI

Efficient Noisy Quantum State and Process Tomography

Efficiently characterizing large quantum states and processes is a central yet notoriously challenging task in quantum information science, as conventional tomography methods typically require resources that grow exponentially with system size. Here, we introduce a structure-agnostic learning framework for noisy $n$-qubit quantum circuits under~i.i.d.~single-qubit noise. We first prove that quantum states with unital noise channels admit an efficient learnable representation in the logarithmic-depth regime. We then extend this framework to quantum process tomography under constant noise, deriving a unified protocol that applies to both unital and non-unital noisy channels and retains efficient guarantees for logarithmic-depth circuits. This process-learning formulation is input-agnostic and imposes no distributional assumptions on the input quantum states. We further study a more general regime with arbitrary noise strength. In this setting, low-weight Pauli propagation induces a terminal truncation whose threshold depends logarithmically on the inverse accuracy, leading to quasi-polynomial complexity and near-unit success probability in the average case. In contrast to the preceding two results, this arbitrary-noise guarantee does not impose any restriction on the circuit depth, and therefore covers arbitrary-depth circuits, including both the noiseless limit ($γ= 0$) and the strong-decoherence regime ($γ= Θ(1)$). Numerical simulations of two-dimensional Hamiltonian dynamics further demonstrate the accuracy and robustness of the approach, including for structured circuits beyond the random-circuit setting assumed in the theoretical analysis. These results provide a scalable and practically relevant route toward characterizing large-scale noisy quantum devices, addressing a key bottleneck in the development of quantum technologies.

quant-ph

Hardware-Efficient and Performance-Enhanced Joint Pulse Shaping and Dispersion Compensation for Coherent Data Center Interconnects

With the explosion of data traffic triggered by 5G/6G and Generative artificial intelligence, coherent optical communication is moving towards higher baud rates and more complex modulation formats. This leads to a significant increase in the computational complexity and power consumption of digital signal processing (DSP) at the transmitter and receiver ends, especially in the chromatic dispersion(CD) Compensation and low roll-off shaping filter modules. We propose a joint shaping filtering and CD compensation (JFS-CD) algorithm. This algorithm moves the CD compensation to the transmitter side and utilizes the characteristics of discrete fourier transform and the spectral features of shaping filtering for integrated processing. Aiming at the high peak-to-average power ratio (PAPR) problem caused by chromatic dispersion pre-compensation, we propose a low-complexity square boundary clipping algorithm(SBC). Simulation results show that, under the premise of maintaining unchanged performance, JFS-CD can reduce the real multiplication complexity by about 46%. Meanwhile, benefiting from the suppression of the effects of system nonlinearity and receiver IQ imbalance, the joint JFS-CD and SBC scheme improves the Q-factor by about 0.3 dB in experiments compared to the traditional post-chromatic dispersion compensation scheme. This research provides a highly potential transmitter DSP solution for next-generation low-power and high-performance data center interconnects (DCI).

cs.IT

Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation

Understanding how latent representations evolve during generation is a central open problem in large language model interpretability. We introduce \textbf{Dynamical Manifold Evolution Theory} (DMET), a phenomenological framework that models LLM generation as a controlled dynamical system evolving along a trajectory on a low-dimensional semantic manifold. DMET formalizes the structural correspondence between Transformer components and a first-order ODE governed by a semantic potential $V$, and characterizes trajectory geometry through three falsifiable proxy metrics: state continuity $C$, attractor clustering quality $Q$, and topological persistence $P$, targeting local smoothness, meso-scale basin structure, and global topological organization, respectively. Across six model architectures, four task types, and 1,080 experimental runs, all three metrics consistently predict text quality outcomes -- log-perplexity, grammaticality, and cross-sentence coherence -- after controlling for decoding parameters, with associations surviving Benjamini--Hochberg correction. Ablation and sanity-check experiments confirm that the effects arise from genuine trajectory structure rather than static distributional artefacts. Furthermore, online monitoring of $C$ drives an adaptive decoding controller that reduces perplexity from 48.5 to 14.6 relative to a fixed-parameter baseline, demonstrating that latent dynamics characterization translates directly into actionable generation control.

cs.CL

Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models

Visual autoregressive (VAR) models have recently emerged as a promising family of generative models, enabling a wide range of downstream vision tasks such as text-guided image editing. By shifting the editing paradigm from noise manipulation in diffusion-based methods to token-level operations, VAR-based approaches achieve better background preservation and significantly faster inference. However, existing VAR-based editing methods still face two key challenges: accurately localizing editable tokens and maintaining structural consistency in the edited results. In this work, we propose a novel text-guided image editing framework rooted in an analysis of intermediate feature distributions within VAR models. First, we introduce a coarse-to-fine token localization strategy that can refine editable regions, balancing editing fidelity and background preservation. Second, we analyze the intermediate representations of VAR models and identify structure-related features, by which we design a simple yet effective feature injection mechanism to enhance structural consistency between the edited and source images. Third, we develop a reinforcement learning-based adaptive feature injection scheme that automatically learns scale- and layer-specific injection ratios to jointly optimize editing fidelity and structure preservation. Extensive experiments demonstrate that our method achieves superior structural consistency and editing quality compared with state-of-the-art approaches, across both local and global editing scenarios.

cs.CV

How Vision Becomes Language: A Layer-wise Information-Theoretic Analysis of Multimodal Reasoning

When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how does this structure evolve across layers? We address this question with a layer-wise framework based on Partial Information Decomposition (PID) that decomposes the predictive information at each Transformer layer into redundant, vision-unique, language-unique, and synergistic components. To make PID tractable for high-dimensional neural representations, we introduce \emph{PID Flow}, a pipeline combining dimensionality reduction, normalizing-flow Gaussianization, and closed-form Gaussian PID estimation. Applying this framework to LLaVA-1.5-7B and LLaVA-1.6-7B across six GQA reasoning tasks, we uncover a consistent \emph{modal transduction} pattern: visual-unique information peaks early and decays with depth, language-unique information surges in late layers to account for roughly 82\% of the final prediction, and cross-modal synergy remains below 2\%. This trajectory is highly stable across model variants (layer-wise correlations $>$0.96) yet strongly task-dependent, with semantic redundancy governing the detailed information fingerprint. To establish causality, we perform targeted Image$\rightarrow$Question attention knockouts and show that disrupting the primary transduction pathway induces predictable increases in trapped visual-unique information, compensatory synergy, and total information cost -- effects that are strongest in vision-dependent tasks and weakest in high-redundancy tasks. Together, these results provide an information-theoretic, causal account of how vision becomes language in multimodal Transformers, and offer quantitative guidance for identifying architectural bottlenecks where modality-specific information is lost.

cs.AI

Where to Add PDE Diffusion in Transformers

Transformers enable powerful content-based global routing via self-attention, but they lack an explicit local geometric prior along the sequence axis. As a result, the placement of locality-inducing modules in hybrid architectures has largely been empirical. We study a simple deterministic PDE diffusion layer implemented as one explicit Euler step of one-dimensional heat smoothing using a discrete Neumann Laplacian under a spectral stability constraint, and ask a structural question: where should diffusion be inserted relative to attention? Our central claim is that diffusion and attention generally do not commute, so inserting the same local operator before versus after attention leads to qualitatively different behaviors. We develop a three-layer operator-theoretic framework that (1) establishes unconditional guarantees for the diffusion subsystem, including spectral non-expansiveness and monotone Dirichlet-energy dissipation when the diffusion step size is smaller than one half, (2) derives compositional perturbation bounds linking insertion effects to representation roughness and downstream amplification, and (3) uses diffusion-attention non-commutativity as a diagnostic for structural double-mixing conflicts. Guided by theory, we evaluate seven insertion positions on the Long Range Arena benchmark. Early diffusion acts as effective pre-regularization, improving average accuracy by 4.1 percentage points when applied after embedding, while post-attention diffusion degrades performance by 2.5 percentage points, consistent with the predicted conflict. A multi-scale diffusion variant yields consistent gains under the same global stability constraint. Our analysis provides a general template for reasoning about local-global compositions in sequence models by separating provable guarantees, compositional bounds, and mechanistic diagnostics.

cs.LG

SMES: Towards Scalable Multi-Task Recommendation via Expert Sparsity

Industrial recommender systems typically rely on multi-task learning to estimate diverse user feedback signals and aggregate them for ranking. Recent advances in model scaling have shown promising gains in recommendation. However, naively increasing model capacity imposes prohibitive online inference costs and often yields diminishing returns for sparse tasks with skewed label distributions. This mismatch between uniform parameter scaling and heterogeneous task capacity demands poses a fundamental challenge for scalable multi-task recommendation. In this work, we investigate parameter sparsification as a principled scaling paradigm and identify two critical obstacles when applying sparse Mixture-of-Experts (MoE) to multi-task recommendation: exploded expert activation that undermines instance-level sparsity and expert load skew caused by independent task-wise routing. To address these challenges, we propose SMES, a scalable sparse MoE framework with progressive expert routing. SMES decomposes expert activation into a task-shared expert subset jointly selected across tasks and task-adaptive private experts, explicitly bounding per-instance expert execution while preserving task-specific capacity. In addition, SMES introduces a global multi-gate load-balancing regularizer that stabilizes training by regulating aggregated expert utilization across all tasks. SMES has been deployed in Kuaishou large-scale short-video services, supporting over 400 million daily active users. Extensive online experiments demonstrate stable improvements, with GAUC gain of 0.29% and a 0.31% uplift in user watch time.

cs.IR

The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox

This paper develops a micro-founded economic theory of the AI industry by modeling large language models as a distinct asset class-Digital Intelligence Capital-characterized by data-compute complementarities, increasing returns to scale, and relative (rather than absolute) valuation. We show that these features fundamentally reshape industry dynamics along three dimensions. First, because downstream demand depends on relative capability, innovation by one firm endogenously depreciates the economic value of rivals' existing capital, generating a persistent innovation pressure we term the Red Queen Effect. Second, falling inference prices induce downstream firms to adopt more compute-intensive agent architectures, rendering aggregate demand for compute super-elastic and producing a structural Jevons paradox. Third, learning from user feedback creates a data flywheel that can destabilize symmetric competition: when data accumulation outpaces data decay, the market bifurcates endogenously toward a winner-takes-all equilibrium. We further characterize conditions under which expanding upstream capabilities erode downstream application value (the Wrapper Trap). A calibrated agent-based model confirms these mechanisms and their quantitative implications. Together, the results provide a unified framework linking intelligence production upstream with agentic demand downstream, offering new insights into competition, scalability, and regulation in the AI economy.

econ.GN