arXiv ScienceSearch

arXiv subjects

Hongliang Zhang

Publications and source records attributed to Hongliang Zhang.

At least 19 recordsLinked to original sources

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising. To address these issues, we propose Proxy Avatar Meets Low-Rank Caching, a cascaded framework for real-time one-shot emotion-controllable portrait animation. Instead of directly generating the target portrait from audio, our method uses a Gaussian-based emotion proxy avatar as a reusable motion generator, which is trained once on a single identity to produce expressive driving videos from audio and emotion labels. Since the proxy avatar only provides motion rather than target appearance or geometry, a large-scale one-shot retargeting model further extracts identity-independent motion from the proxy performance and adapts it to arbitrary target portraits. To improve inference efficiency, we introduce zero-shot appearance reuse with low-rank caching, which caches reference appearance features at the initial denoising step and models subsequent feature variations using lightweight low-rank adapters. Extensive experiments demonstrate that our method achieves stronger emotional expressiveness, better identity-preserving animation, and substantially reduced inference cost, enabling real-time one-shot portrait animation.

cs.CV

FL-OA: A Byzantine-Robust Federated Learning Framework with Outsourced Auditing for Intelligent Devices

Federated learning (FL) enables multiple intelligent devices to collaboratively train a high-accuracy model without sharing raw data. However, due to its distributed nature, FL is vulnerable to Byzantine attacks. Existing defense methods rely on strong assumptions, such as the proportion of malicious devices not exceeding 50\%, or the server having an additional root dataset that matches the training task. Moreover, they show limited efficacy as they overlook $(i)$ the divergence among benign updates and $(ii)$ the curse of dimensionality involved in comparing two high-dimensional updates. To solve these concerns, we propose FL-OA, a Byzantine-robust federated learning framework utilizing outsourced auditing. In FL-OA, the server collaborates with third-party organization that holds an additional root dataset to perform outsourced auditing, thereby enabling the server to achieve robust aggregation without strong assumptions. Additionally, FL-OA introduces a gradient ascent step and a correction term during local training to mitigate the divergence among benign updates, and designs a parameter importance indicator to extract critical parameters for auditing, alleviating the curse of dimensionality. We further provide a detailed theoretical analysis of FL-OA. Extensive experiments demonstrate that FL-OA outperforms existing defense methods against Byzantine attacks.

cs.LG

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by statistical heterogeneity and the stealthiness of backdoor attacks. To tackle these issues, we propose FedDAB, a two-phase method that combines local contrastive regularization with alignment checking, to defend against backdoor attacks. In the first phase, FedDAB introduces a novel model-contrastive term into the local objective to enhance direction and magnitude consistency among benign updates. In the second phase, FedDAB employs an alignment checking strategy to evaluate each local update in terms of overall-direction alignment and parameter-level alignment with historical information, excluding updates that exhibit abnormal alignment patterns from global aggregation. We theoretically prove FedDAB's robustness with a convergence rate of $\mathcal{O}(1/T)$. Extensive experiments show that FedDAB outperforms existing defense methods against backdoor attacks.

cs.CR

Theoretical Analysis of Diffusion Models for Radio Map Estimation with Ultra-low Sampling Rates

Radio maps, which characterize the spatial distribution of radio frequency metrics such as received signal strength, are essential for a wide range of wireless applications. The problem of radio map estimation involves constructing a radio map from sparse sensor measurements at multiple locations. This problem is particularly challenging due to ultra-low sampling rates, where available sensor measurements are far fewer than the high resolution requirement of radio maps to be estimated. Recently, diffusion models have been increasingly adopted for this problem, yet its theoretical performance remains unexamined. This paper bridges this gap by formulating radio map estimation as a non-linear matrix completion problem. Based on this formulation, we first derive a theoretical lower bound on the minimum estimation error achievable by diffusion models, which is fundamentally governed by the discrepancy between the deployment distribution and the true underlying radio propagation law. We then extend this bound to incorporate the effect of sampling sparsity, capturing the additional error introduced by ultra-low sampling rates. Furthermore, we establish a critical sampling rate threshold necessary for diffusion models to achieve performance convergence. Finally, considering that the derived error bounds depend on certain information that is difficult to obtain in practice, we propose empirical approximations that are readily computable from observable data. Extensive simulations based on real-world traces demonstrate that these empirical formulas tightly approximate the theoretical error bounds, validating their effectiveness for practical deployment.

eess.SP

Holographic Beamforming for Semantic Communication

Holographic beamforming enabled by metamaterial antennas has been proposed to facilitate spatial multiplexing at low hardware cost and low power consumption. However, existing holographic beamforming schemes are mainly developed for conventional bit-communication systems, which have not considered semantic-level importance and thus cannot be directly applied to support semantic communication. Specifically, in conventional bit communication, all bits are treated as equally important. In contrast, in semantic communication, different semantic information contribute unequally to task completion and therefore has different degrees of importance, with more important information requiring higher transmission quality. Ignoring semantic importance in holographic beamforming causes mismatches between importance of semantic information and its received SNR, thus degrading performances. In this paper, we propose a semantic-importance-aware holographic beamforming scheme enabled by metamaterial antennas with tunable radiated amplitudes to support semantic communication. It is challenging to design semantic-aware holographic beamforming schemes due to non-trivial modeling of the impact of semantic importance and unique amplitude-controlled structures of holographic beamforming. To address this, we characterize the dependence of semantic communication performance on semantic importance and received SNR via data fitting, and design a semantic-aware holographic beamforming algorithm to ensure reliable delivery of highly important semantic information. Simulation results validate effectiveness of the proposed method.

cs.IT

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, whereas Ring-2.6 is tailored for deeper reasoning and more advanced agentic workflows. Instead of training from scratch, we upgrade the Ling-2.0 base model through architectural migration pre-training and large-scale post-training. This upgrade is guided by a unified co-design of model architecture, optimization objectives, serving systems, and agent training environments, enabling improvements in both model capability and deployment efficiency. At the architectural level, we introduce a hybrid linear attention design that integrates Lightning Attention with MLA, improving the efficiency of long-context training and decoding. To further enhance token efficiency, we optimize capability per output token through Evolutionary Chain-of-Thought, Linguistic Unit Policy Optimization, bidirectional preference alignment, and shortest-correct-response distillation. For agentic capabilities, we propose KPop, a reinforcement learning framework designed to support stable training of Ring-2.6-1T on large-scale environment-grounded data. KPop improves training efficiency through asynchronous scheduling across coding, search, tool use, and workflow execution, enabling scalable learning from complex agent-environment interactions. Together, Ling-2.6 and Ring-2.6 provide a practical pathway toward efficient, scalable, and open agentic systems. We open-source all checkpoints in the 2.6 family to support further research and development in practical agentic intelligence.

cs.CL

Holographic Surface Enabled Integrated Sensing and Communications

Integrated sensing and communications (ISAC) is an essential 6G capability for joint data transmission and environmental sensing. To support 6G scenarios with stringent ISAC performance requirements, existing massive-MIMO-based systems are expected to scale toward ultra-massive MIMO. However, this scaling incurs prohibitive cost and power consumption when realized using widely adopted phased arrays with complex phase shifters and feeding networks. Recently, holographic integrated sensing and communications (HISAC) has emerged as a promising paradigm to address this issue. It employs reconfigurable holographic surfaces (RHSs), a type of leaky-wave antenna, as a cost- and energy-efficient implementation of ultra-massive MIMO-based ISAC, and offers enhanced flexibility for ISAC beam synthesis through holographic beamforming. In this paper, we provide a comprehensive tutorial on HISAC, focusing on how RHS-enabled holographic beamforming can be exploited to jointly support communication and sensing under practical hardware constraints. We first introduce the fundamentals of RHSs and discuss the unique leakage power constraint of holographic beamforming. We then present a general optimization framework for HISAC and show how HISAC enhances joint communication and sensing, sensing-assisted communication, and communication-assisted sensing. We further present HISAC system implementations and experimental results. Finally, we outline promising research directions for HISAC, highlighting the potential of HISAC in advancing efficient, flexible, and high-performance ISAC networks.

eess.SP

Token Economics for LLM Agents: A Dual-View Study from Computing and Economics

As LLM agents evolve, tokens have emerged as the core economic primitives of Agentic AI. However, their exponential consumption introduces severe computational, collaborative, and security bottlenecks. Current surveys remain fragmented across system optimization, architecture design, and trust, lacking a unified framework to evaluate the fundamental trade-off between output quality and economic cost. To bridge this gap, this survey presents the first comprehensive survey of Token Economics. By unifying computer science and economics, we conceptualize tokens as production factors, exchange mediums, and units of account. We synthesize existing literature across a four-dimensional taxonomy: (1) Micro-level (Single Agent): Optimizing budget-constrained factor substitution via neoclassical firm theory. (2) Meso-level (Multi-Agent Systems): Minimizing collaboration friction using transaction cost and principal-agent theories. (3) Macro-level (Agent Ecosystems): Addressing congestion externalities and pricing via mechanism design. (4) Security: Internalizing adversarial threats as endogenous economic constraints. Finally, we outline frontier directions, including differentiable token budgets and dynamic markets, to lay the theoretical foundation for scalable next-generation agent systems.

cs.AI

CellScientist: Dual-Space Hierarchical Orchestration for Closed-Loop Refinement of Virtual Cell Models

Virtual Cell Modeling (VCM) requires models that not only predict perturbation responses, but also support targeted revision when predictions fail. Current LLM-assisted modeling workflows face a refinement-routing problem: prediction discrepancies are observed through executable implementations, but the relevant revision may involve the modeling assumption, representation design, implementation, or task constraint. Without structured feedback propagation across these levels, iterative refinement may repair code while failing to revise the assumption responsible for the discrepancy. We propose CellScientist, a dual-space hierarchical framework that couples a high-level hypothesis space with a low-level executable implementation space. CellScientist represents modeling decisions as structured states, realizes them as admissible programs under task and interface constraints, and routes execution discrepancies back to targeted hypothesis or implementation updates. This enables a closed Hypothesis -> Implementation -> Hypothesis loop where failures become structured signals for model refinement rather than debugging events. Across morphology and transcriptomic benchmarks, with additional single-cell perturbation evaluations, the final executable models selected by CellScientist improve over reference baselines under fixed split and evaluation protocols, while the workflow produces auditable refinement traces.

cs.LG

Bayesian Cram\'er-Rao Bound for Sensing Performance in Meta-Backscatter Systems

Meta-backscatter system that utilizes meta-material sensors is a promising enabler for future environmental sensing, offering distinct advantages such as low cost, zero-power consumption, and robustness. Specifically, the electromagnetic response of the sensor, typically characterized by a frequency-selective absorption profile, is affected by the environmental conditions, allowing the estimation of these conditions from the reflected signal. However, it remains unclear what estimation accuracy can be achieved fundamentally. Motivated by this gap, we quantify this accuracy limit using the Bayesian Cram\'er-Rao bound (BCRB), which provides a lower bound on the mean-squared error for the environmental condition. Establishing this limit is challenging because the electromagnetic response of the sensor is distorted by the channel fading, while the channel estimation is infeasible since the sensors cannot be configured to predefined states to generate training data. To address this challenge, we consider the joint BCRB of the channel coefficient and the environmental condition in a multicarrier framework. The BCRB of the environmental condition is then obtained by selecting the corresponding element from the joint BCRB. An analysis of the derived BCRB reveals the impact of the absorption peak shape and the number of subcarriers. The derivation and analysis of the BCRB are verified through simulations.

eess.SP

Aerial IRS Deployment-Aided Secure Computation Offloading Against DISCO Jamming Attacks

With the rapid growth of Multi-access Edge Computing (MEC), secure and efficient computation offloading from user equipment (UEs) to edge access points (APs) is critical. However, DISCO intelligent reflective surface-based fully-passive jammers (DIRS-based FPJs) use random time-varying phase shifts to launch DISCO jamming attacks, disrupting offloading performance. This paper leverages an aerial intelligent reflective surface (AIRS) to enable secure computation offloading against DISCO jamming by jointly optimizing offloading ratios, AIRS phase shifts, and deployment. A two-timescale (2Ts) framework is proposed to address the optimization challenge caused by the distinct update frequencies of different strategies. Specifically, AIRS deployment is adjusted on a long timescale to boost antijamming capability due to the impracticality of frequent physical adjustment, while offloading ratios and phase shifts are optimized on a short timescale to adapt to DIRS-jammed dynamic channel conditions. We propose a dual-agent deep reinforcement learning (DRL)-based AIRS deployment-aided secure computation offloading (DDADSO) scheme to maximize the secure offloading utility under DISCO jamming. Simulation results verify that the proposed DDADSO scheme outperforms benchmark schemes, demonstrating the effectiveness of AIRS deployment in improving offloading performance against DISCO jamming attacks.

eess.SP

Bistatic Integrated Sensing and Communication in the Presence of a Disco Reconfigurable Intelligent Surface: Disruption, Enhancement, or Both?

Integrated sensing and communication (ISAC) is widely regarded as one of the key enabling technologies for future sixth-generation (6G) wireless communication systems. In this work, we investigate a bistatic ISAC system in the presence of a disco reconfigurable intelligent surface (DRIS), whose random and time-varying reflection coefficients emulate a "disco ball." The introduction of the DRIS breaks the underlying assumption in existing ISAC systems that the sensing and communication channels remain static or quasi-static within the channel coherence time. We first develop a bistatic system model incorporating the DRIS and characterize all involved wireless channels. Then, an ISAC waveform design that balances sensing and communication performance is proposed by formulating a Pareto optimization problem, where the trade-off is controlled through a tunable factor. Communication and sensing performance in the bistatic ISAC system are quantified by the signal-to-interference-plus-noise ratio (SINR) and the Cramer-Rao lower bound (CRLB), respectively. To quantify the impact of the DRIS on the bistatic ISAC system, we derive the statistical characteristics of DRIS-induced active channel aging (ACA) channels for communications and the cascaded DRIS-based sensing channel. Then, we establish a theoretical lower bound on the SINR and closed-form CRLB expressions in the presence of a DRIS. The analysis reveals several distinctive properties of the DRIS in bistatic ISAC systems. In particular, the DRIS degrades communication performance significantly due to the introduction of ACA interference. In contrast, with respect to sensing performance, the DRIS decreases the estimation accuracy of the angle of departure (AoD) while concurrently enhancing that of the angle of arrival (AoA). Numerical results validate the derived theoretical analysis and confirm these DRIS-induced behaviors.

eess.SP

Digital Twin-Enabled Mobility-Aware Cooperative Caching in Vehicular Edge Computing

With the advancement of vehicle-to-vehicle (V2V) ad hoc networks and wireless communication technologies, mobile edge caching has become a key enabler for enhancing network performance and user experience. However, traditional federated learning-based collaborative caching approaches in vehicular scenarios suffer from inadequate client selection mechanisms and limited prediction accuracy, which result in suboptimal cache hit ratios and increased content transmission latency. To address these challenges, we propose a Digital Twin-based Asynchronous Federated Learning-driven Predictive Edge Caching with Deep Reinforcement Learning (DAPR) framework. DAPR employs an intelligent client selection strategy based on asynchronous federated learning, which leverages mobility prediction and data quality assessment to avoid selecting highly mobile clients or clients with low-quality data, thereby significantly improving model convergence efficiency. In addition, we design a GRU-VAE prediction model that uses a Variational Autoencoder (VAE) to capture latent data distribution features and Gated Recurrent Units (GRUs) to model temporal dependencies, thereby substantially enhancing the accuracy of content request prediction. The predicted content popularities are then fed into a deep reinforcement learning-driven caching decision engine to dynamically optimize edge caching resource allocation. Extensive experiments demonstrate that DAPR achieves superior performance in terms of average reward, cache hit ratio, and transmission latency, thereby effectively improving the overall efficiency of vehicular edge caching systems.

cs.NI

Towards Performance-Enhanced Model-Contrastive Federated Learning using Historical Information in Heterogeneous Scenarios

Federated Learning (FL) enables multiple nodes to collaboratively train a model without sharing raw data. However, FL systems are usually deployed in heterogeneous scenarios, where nodes differ in both data distributions and participation frequencies, which undermines the FL performance. To tackle the above issue, this paper proposes PMFL, a performance-enhanced model-contrastive federated learning framework using historical training information. Specifically, on the node side, we design a novel model-contrastive term into the node optimization objective by incorporating historical local models to capture stable contrastive points, thereby improving the consistency of model updates in heterogeneous data distributions. On the server side, we utilize the cumulative participation count of each node to adaptively adjust its aggregation weight, thereby correcting the bias in the global objective caused by different node participation frequencies. Furthermore, the updated global model incorporates historical global models to reduce its fluctuations in performance between adjacent rounds. Extensive experiments demonstrate that PMFL achieves superior performance compared with existing FL methods in heterogeneous scenarios.

cs.LG

Unsupervised Semi-Parametric Plug-in Likelihood-Ratio Detection for Covert Communications in the Presence of Disco Reconfigurable Intelligent Surfaces

Covert communications, also referred to as low probability of detection (LPD) communications, provide a higher level of privacy protection than cryptography and physical-layer security (PLS) by hiding transmissions in the ambient environment. In this work, we investigate covert communications in the presence of a disco reconfigurable intelligent surface (DRIS) deployed by the warden Willie, which reduces Willie's detection error probability (DEP), i.e., the sum of the false alarm rate (FAR) and the miss detection rate (MDR), and degrades the communication performance between Alice and Bob, without relying on either channel state information (CSI) or additional jamming power. However, the introduction of the DRIS makes it analytically intractable for Willie to construct the Neyman-Pearson (NP) detector, which is the optimal detector for monitoring potential covert transmissions between Alice and Bob. To this end, we develop an unsupervised semi-parametric plug-in likelihood-ratio detector for Willie. The proposed detector retains the parametric Gamma reference model under the silent hypothesis without requiring prior knowledge of noise, and learns from unlabeled data a one-dimensional monotone normalizing flow model for the analytically intractable distribution under the transmission hypothesis. In particular, it exploits the structural prior inherent in covert communications that Willie's observations reduce to noise only when Alice and Bob are silent. The monitoring performance at Willie is evaluated in terms of DEP, while the communication impact on Alice and Bob is quantified by the signal-to-jamming-plus-noise ratio (SJNR). Simulation results verify the analysis and show that the proposed unsupervised plug-in likelihood-ratio detector achieves monitoring performance close to that of its supervised counterpart.

eess.SP

RIS-Aided Wireless Amodal Sensing for Single-View 3D Reconstruction

Amodal sensing is critical for various real-world sensing applications because it can recover the complete shapes of partially occluded objects in complex environments. Among various amodal sensing paradigms, wireless amodal sensing is a potential solution due to its advantages of environmental robustness, privacy preservation, and low cost. However, the sensing data obtained by wireless system is sparse for shape reconstruction because of the low spatial resolution, and this issue is further intensified in complex environments with occlusion. To address this issue, we propose a Reconfigurable Intelligent Surface (RIS)-aided wireless amodal sensing scheme that leverages a large-scale RIS to enhance the spatial resolution and create reflection paths that can bypass the obstacles. A generative learning model is also employed to reconstruct the complete shape based on the sensing data captured from the viewpoint of the RIS. In such a system, it is challenging to optimize the RIS phase shifts because the relationship between RIS phase shifts and amodal sensing accuracy is complex and the closed-form expression is unknown. To tackle this challenge, we develop an error prediction model that learns the mapping from RIS phase shifts to amodal sensing accuracy, and optimizes RIS phase shifts based on this mapping. Experimental results on the benchmark dataset show that our method achieves at least a 56.73% reduction in reconstruction error compared to conventional schemes under the same number of RIS configurations.

eess.SP

Mapping optical, chemical, structural features in ZrO2 via cross-sectional SEM-Cathodoluminescence correlation microscopy

Understanding how nanoscale heterogeneities influence charge transport and mass transfer in oxides is critical for developing advanced materials for energy and electronic uses. In high-temperature applications, the formation of thermal oxides with complex chemical and structural features plays a central role in material lifetime. While thermally grown zirconia (ZrO2) on zirconium alloys exhibits strong chemical and microstructural gradients across the oxide thickness, linking these heterogeneities to electronic-defect landscapes remains challenging. We demonstrate cross-sectional scanning electron microscope-cathodoluminescence (SEM-CL) as a mesoscale probe of spatial variations in luminescence in zirconia and establish correlations with co-registered electron backscatter diffraction (EBSD) and electron probe micro-analysis (EPMA) on the same region. The SEM-CL signal is dominated by the ~2.7 eV defect band, but its intensity varies strongly across the oxide cross section. Correlative EBSD-CL analysis reveals that CL intensity increases with grain area and decreases at the grain boundaries, consistent with enhanced non-radiative recombination associated with microstructural disorder. EPMA mapping shows that a substantial fraction of CL-dark features co-localize with secondary phase precipitates enriched in iron. These results show that SEM-CL contrast in corrosion-grown ZrO2 is controlled by both chemical heterogeneity and microstructural disorder, underscoring the need for correlative registration to interpret CL images. This multi-modal approach provides an efficient route to connect electronic properties and luminescence signatures across complex oxide cross sections to underlying chemistry and microstructure, thereby providing a pathway to relate local defect landscapes to regions likely to bias electronic/ionic transport during oxidation.

cond-mat.mtrl-sci

Meta-Backscatter: Long-Distance Battery-Free Metamaterial-Backscatter Sensing and Communication

Battery-free Internet of Things (BF-IoT) enabled by backscatter communication is a rapidly evolving technology offering advantages of low cost, ultra-low power consumption, and robustness. However, the practical deployment of BF-IoT is significantly constrained by the limited communication range of common backscatter tags, which typically operate with a range of merely a few meters due to inherent round-trip path loss. Meta-backscatter systems that utilize metamaterial tags present a promising solution, retaining the inherent advantages of BF-IoT while breaking the critical communication range barrier. By leveraging densely paved sub-wavelength units to concentrate the reflected signal power, metamaterial tags enable a significant communication range extension over existing BF-IoT tags that employ omni-directional antennas. In this paper, we synthesize the principles and paradigms of metamaterial sensing to establish a unified design framework and a forward-looking research roadmap. Specifically, we first provide an overview of backscatter communication, encompassing its development history, working principles, and tag classification. We then introduce the design methodology for both metamaterial tags and their compatible transceivers. Moreover, we present the implementation of a meta-backscatter system prototype and report the experimental results based on it. Finally, we conclude by highlighting key challenges and outlining potential avenues for future research.

eess.SP