arXiv ScienceSearch

arXiv subjects

Hai Lin

Publications and source records attributed to Hai Lin.

At least 19 recordsLinked to original sources

Robust Monitoring of Arc Welding Processes: A Generalizable Framework with DVAE and Particle Filter

Arc welding processes are essential for continuous fabrication but prone to disturbances that impair weld quality, making real-time monitoring critical yet difficult due to complex visual patterns and nonlinear, time-varying dynamics. Deep learning shows promise but faces scalability limits because of its dependence on large labeled datasets and application-specific tuning. We explore whether a unified approach can characterize major arc welding processes across applications and improve scalability through consistent state monitoring. This paper introduces a robust and generalizable monitoring framework for arc welding. It combines unsupervised deep latent representation learning, which extracts compact features from weld pool images, with Bayesian filtering to handle persistent and fluctuating disturbances such as arc radiation and specular reflections. Specifically, a Dynamic Variational Autoencoder (DVAE), consisting of a CNN-based encoder-decoder and an LSTM-based transition model, jointly learns latent representations and their evolution under control inputs. For robust real-time inference, a specialized Particle Filter (PF) propagates the latent and LSTM hidden states, preserving process history while suppressing sensor noise. This design is well suited to welding's slow and inertial dynamics. Validation on GTAW and GMAW without process-specific tuning demonstrates the framework's generalizability and robustness.

eess.IV

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscuring whether visual input helped, hurt, or was ignored. We introduce \textbf{MM-IssueLoc}, a controlled benchmark and evaluation protocol for repository-level localization with visual evidence. MM-IssueLoc contains 652 issue-PR instances across 23 languages, with annotations for 7 image categories and 4 relevance levels. It provides file-level and function-level gold labels, paired text-only and with-image evaluation, and VCE-based diagnostics that convert images into structured textual evidence. We evaluate LLM-based and retrieval-based systems, including MM-IssueLoc-VL-Emb as a controlled multimodal retriever. Results show that existing systems remain far from reliable multimodal repository localization: the strongest agent reaches 38.96 file Acc@5 and 22.45 function Acc@10, while the strongest retriever reaches 33.86 function Acc@10. Cross-benchmark comparisons show that high localization scores on text-dominant SWE benchmarks do not transfer cleanly to multimodal issue localization. MM-IssueLoc turns visual evidence into an explicit evaluation variable, enabling future work to test whether systems improve by using visual evidence for localization, rather than by relying on text-only cues or downstream patch-generation effects.

cs.SE

Coexistence and manipulation of multiple singularities in a reconfigurable non-Hermitian metasurface

Non-Hermitian frameworks extend conventional Hermitian physics, offering a powerful paradigm for describing open systems. Central to this field are various singularities within the complex parameter space, such as exceptional points (EPs) and scattering zeros, which dictate exotic physical behaviors. As research shifts from isolated singularities toward multi-singularity interactions, conventional planar metasurfaces remain constrained by limited tuning dimensions. Here, we propose a mirror-coupled design that maps a metasurface into a quasi-high-dimensional parameter space. By employing a metallic plane to generate image resonators, this scheme multiplies the system degrees of freedom without increasing the number of physical resonators. Its implementation on a reconfigurable platform integrated with PIN diodes yields the coexistence and manipulation of an EP and multiple reflection zeros. Through simulations and microwave experiments, we characterize the dynamic evolution of these singularities and exploit their synergistic effects for two distinct applications. First, for tunable absorption, multiple reflection zeros are spectrally coordinated to achieve a near-perfect absorption band exceeding $99.9\%$ across the X-band, thereby dynamically suppressing target scattering. Second, for enhanced sensing, a reflection zero couples with the EP to form a hybrid singularity. This hybrid state inherits the power-law sensitivity of the EP while substantially boosting robustness against fluctuations, resolving the conventional trade-off between sensitivity and stability and simplifying detection to direct peak tracking rather than complex multimode eigenvalue fitting. Our work provides a general methodology to circumvent parameter competition among non-Hermitian singularities, opening new avenues for multifunctional metadevices across the electromagnetic spectrum.

physics.optics

Model-Based and Data-Driven Hierarchical Control and Topology Co-Design for Robust Networked Systems

In this paper, we consider a class of networked systems comprising an interconnected set of linear subsystems, disturbance inputs, and performance outputs. Using dissipativity theory, we first propose a model-based hierarchical control design strategy to ensure the closed-loop networked system is dissipative from its disturbance inputs to performance outputs. This involves designing local controllers for each subsystem to enforce local dissipativity guarantees, which are then exploited to co-design distributed global controllers and the interconnection topology to enforce global dissipativity guarantees while optimizing interconnection topology costs. The overall design process requires only solving a sequence of linear matrix inequality (LMI) problems, thereby retaining compositionality and decentralizability while avoiding non-convex, iterative design processes that are inefficient and centralized. This model-based hierarchical control design strategy assumes the knowledge of the subsystem dynamics, which may not hold in many real-world networked systems. Motivated by this, we also propose a data-driven hierarchical control design strategy that assumes only the availability of rich input-state-output trajectory data from the subsystems. The proposed data-driven design process assumes that the unknown disturbances affecting the subsystem dynamics are bounded by a quadratic matrix inequality (relaxing conventional bounds) and accounts for this by using the matrix S-lemma. Finally, the effectiveness of the proposed model-based and data-driven hierarchical control designs is illustrated for a networked system representing a DC microgrid, with the aim of enforcing robust (dissipative) voltage regulation and current sharing.

eess.SY

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture

Large language models are undergoing a transition from model technology to system technology. Engineering challenges like cache reuse, context capacity, agent scheduling, and permission control resemble classical computer systems problems. This raises a question: if we treat the LLM as a CPU, KV cache as processor cache, context window as main memory, and agent framework as an operating system, can decades of computer architecture wisdom guide next generation model native systems? This paper pursues this analogy as a visionary survey. We map computer architecture concepts onto the emerging model native stack, survey literature across LLM as OS, memory management, agent frameworks, tool protocols, multi agent coordination, cognitive architectures, and safety governance, finding that each addresses a different layer without a unifying model. We propose the Intelligent Computing Architecture (ICA): six functional layers with interface contracts and design axioms. We resolve the tension over whether the LLM resembles a CPU or OS via a dual plane architecture a probabilistic execution plane (what can be computed) and a deterministic control plane (what should be computed), with every layer passing through as a graded crossover. We propose three Amdahl style design heuristics Semantic Locality, Context Budget, and Agent Speedup as organizing back of envelope models, illustrate their parameter ranges with published data, and identify predictive validation as the principal open task. We articulate analogy boundaries, note differences between silicon and model era architectures, and propose a research roadmap. This is a conceptual and survey contribution with no new experimental results.

cs.AI

Data-Driven Linear Quadratic Control Using Output-Feedback via Non-Minimal Realization

In this paper, we investigate a continuous-time linear quadratic control problem for systems with unknown matrices, where only input-output data are available. We propose an output-feedback learning framework based on a canonical nonminimal realization constructed through Kreisselmeier's adaptive filter. The filter admits an observer interpretation, which leads to an augmented system that preserves the input-output response of the realization and provides accessible state trajectories. We show that the optimal gain of this augmented system explicitly recovers the optimal gain associated with the canonical non-minimal realization, and hence achieves the optimal state-feedback solution of the original plant. Exploiting this relation and the known structure of the augmented input matrix, we develop a data-driven value iteration algorithm within the adaptive dynamic programming framework. The resulting controller is implementable from input-output data, and its performance is validated via simulations.

math.OC

Informative Graph Structure Learning

The quality of graph-structured data is fundamental to the success of modern graph analysis techniques such as Graph Neural Networks (GNNs). However, real-world graph data is often suboptimal, suffering from issues such as noise and incomplete connections. Graph Structure Learning (GSL) has emerged as a promising technique that adaptively optimizes node connections. However, we observe that the effectiveness of GSL often comes at the cost of a dramatic expansion in edge count, resulting in significant storage and computational overhead. In this work, we reveal that this limitation stems from the prevalent use of similarity-based edge construction, which predominantly connects highly similar neighbors based on their embeddings, introducing substantial structure redundancy. To address this, we propose a novel Informative Graph Structure Learning method (InGSL), which jointly considers both similarity and diversity in edge construction by incorporating a mutual-information-guided learning strategy. Notably, InGSL serves as a plug-in module that can be seamlessly integrated into existing GSL frameworks. Through extensive experiments on six representative GSL methods, we demonstrate that InGSL achieves significant performance improvements at a reduced number of edges.

cs.LG

Salted Fisher Information for Hybrid Systems

Discrete events alter how parameter influence propagates in hybrid systems. Prevailing Fisher information formulations assume that sensitivities evolve smoothly according to continuous-time variational equations and therefore neglect the sensitivity updates induced by discrete events. This paper derives a Fisher information matrix formulation compatible with hybrid systems. To do so, we use the saltation matrix, which encodes the first order transformation of sensitivities induced by discrete events. The resulting formulation is referred to as the salted Fisher information matrix (SFIM). The proposed framework unifies continuous information accumulation during flows with discrete updates at event times. We further establish that hybrid persistence of excitation provides a sufficient condition for positive definiteness of the SFIM. Examples are provided to demonstrate the merit of the proposed approach, including a three bus generator wind turbine differential algebraic power system

eess.SY

Anisotropic non-Hermitian skin effect in a two-dimensional Lieb photonic crystal

In this contribution paper, we construct a two-dimensional non-Hermitian (NH) photonic crystal (PhC) to prototype its anisotropic non-Hermitian skin effect (NHSE) for experimental proposal. Based on the tight-binding model for Lieb lattice with NH coupling, a nontrivial spectral winding number is pinpointed for certain eigenstates, which translates to geometry-dependent skin modes with tilt boundaries. For ease of implementation, complex refractive indices are employed for the Lieb unit cell of PhC to emulate the NH coupling. Validated by full wave simulation, our work underscores the boundary dependence of skin effect, and provides a concrete prototype design of NHSE implementable by state-of-the-art of topological metamaterial platforms.

physics.optics

SemantiCache: Efficient KV Cache Compression via Semantic Chunking and Clustered Merging

Existing KV cache compression methods generally operate on discrete tokens or non-semantic chunks. However, such approaches often lead to semantic fragmentation, where linguistically coherent units are disrupted, causing irreversible information loss and degradation in model performance. To address this, we introduce SemantiCache, a novel compression framework that preserves semantic integrity by aligning the compression process with the semantic hierarchical nature of language. Specifically, we first partition the cache into semantically coherent chunks by delimiters, which are natural semantic boundaries. Within each chunk, we introduce a computationally efficient Greedy Seed-Based Clustering (GSC) algorithm to group tokens into semantic clusters. These clusters are further merged into semantic cores, enhanced by a Proportional Attention mechanism that rebalances the reduced attention contributions of the merged tokens. Extensive experiments across diverse benchmarks and models demonstrate that SemantiCache accelerates the decoding stage of inference by up to 2.61 times and substantially reduces memory footprint, while maintaining performance comparable to the original model.

cs.CL

A Unified Pulse-Shaped OFDM Framework for Chirp-Domain Waveforms: Continuous-Time Modeling and Practical I/O Analysis

A unified framework for chirp-domain waveforms, including orthogonal chirp division multiplexing (OCDM) and affine frequency division multiplexing (AFDM), is developed. Their continuous-time representations are shown to fall within the conventional Weyl-Heisenberg (WH) framework for multicarrier waveforms, with the root chirp as the prototype pulse. Since the root chirp has constant envelope and is transparent to subcarrier orthogonality, these waveforms can be further interpreted as pulse-shaped (PS) orthogonal frequency division multiplexing (OFDM) signals, whose power spectral density is derived analytically. The derived spectrum reveals that implementations based on the discrete affine Fourier transform rely on sub-Nyquist samples and exhibit frequency aliasing. We prove that the corresponding aliased chirps are only conditionally orthogonal, whereas sample-wise root-Nyquist pulse shaping of the discrete-time AFDM (DT-AFDM) sequence produces mutually orthogonal pulse-shaped chirps, resulting in the pulse-shaped AFDM (PS-AFDM) waveform. We then derive an exact waveform-level input-output (I/O) relation for PS-AFDM over delay-Doppler (DD) channels, showing that the effective channel at a practical receiver is generally not a superposition of pure path-wise DD components. Waveform simulations verify the derived relation to machine precision, while the conventional sequence-level I/O relation for DT-AFDM exhibits a substantial mismatch with waveform behavior for practical channels with continuous-valued delays.

cs.IT

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical ``spatial intelligence gap,'' where models fail to construct coherent 3D mental representations from 2D observations. We uncover this gap via diagnostic analyses showing the bottleneck is a missing view-consistent spatial interface rather than insufficient visual features or weak reasoning. To bridge this, we introduce \textbf{3ViewSense}, a framework that grounds spatial reasoning in Orthographic Views. Drawing on engineering cognition, we propose a ``Simulate-and-Reason'' mechanism that decomposes complex scenes into canonical orthographic projections to resolve geometric ambiguities. By aligning egocentric perceptions with these allocentric references, our method facilitates explicit mental rotation and reconstruction. Empirical results on spatial reasoning benchmarks demonstrate that our method significantly outperforms existing baselines, with consistent gains on occlusion-heavy counting and view-consistent spatial reasoning. The framework also improves the stability and consistency of spatial descriptions, offering a scalable path toward stronger spatial intelligence in multimodal systems.~\footnote{https://github.com/Jasaxion/3ViewSense}

cs.CV

Doppler Shift Keying Modulation for Uplink Multiple Access over Doubly-Dispersive Channels

The delay-Doppler (DD) domain modulation has been regarded as one of the most competitive candidates to support wireless communications for emerging high-mobility applications in the sixth-generation mobile networks. Unfortunately, most of the existing designs for DD domain modulation suffer from high peak-to-average power ratio (PAPR) and unbearable detection complexity under uplink transmission since large time duration and bandwidth are required to guarantee high DD resolutions. To address these issues, the Doppler shift keying (DSK) modulation based on the orthogonal delay Doppler division multiplexing modulator is proposed in this paper, where the input-output characterization in the DD domain is fully exploited. The principle of the DSK transceiver is first established with the one-hot mapper and low-complexity iterative successive interference cancellation-maximum ratio combining detector for point-to-point scenarios. The proposed scheme is then generalized to the zero auto-correlation sequence-based implementation, which benefits the extension of multi-user (MU) uplink DSK frameworks. For uplink DSK transmission, Zadoff-Chu (ZC) sequences are adopted as the basis sequences. We optimize the assignment of ZC roots to different user equipments (UEs) by minimizing the maximum inter-user interference. This optimization process, which analyzes the root allocation, directly assigns a specific ZC sequence to each UE. The PAPR and bit error rate performance of the proposed DSK modulation with the low-complexity detector is finally verified by extensive simulation results under doubly-dispersive channels, which demonstrates the superiority of DSK modulation especially for uplink multiple access over doubly dispersive channels.

eess.SP

Channel Estimation with Hierarchical Sparse Bayesian Learning for ODDM Systems

Orthogonal delay-Doppler division multiplexing (ODDM) is a promising modulation technique for reliable communications in high-mobility scenarios. However, the existing channel estimation frameworks for ODDM systems cannot achieve both high accuracy and low complexity simultaneously, due to the inherent coupling of delay and Doppler parameters. To address this problem, a two-dimensional (2D) hierarchical sparse Bayesian learning (HSBL) based channel estimation framework is proposed in this paper. Specifically, we address the inherent coupling between delay and Doppler dimensions in ODDM by developing a partially-decoupled 2D sparse signal recovery (SSR) formulation on a virtual sampling grid defined in the delay-Doppler (DD) domain. With the help of the partially-decoupled formulation, the proposed 2D HSBL framework first performs low-complexity coarse on-grid 2D sparse Bayesian learning (SBL) estimation to identify potential channel paths. Then, high-resolution fine grids are constructed around these regions, where an off-grid 2D SBL estimation is applied to achieve accurate channel estimation. Simulation results demonstrate that the proposed framework achieves performance superior to conventional off-grid 2D SBL with significantly reduced computational complexity.

econ.EM

Primal-dual algorithm for distributed optimization: A dissipativity-based perspective

We study a continuous-time primal-dual algorithm for distributed optimization with nonconvex local cost functions over weight-unbalanced digraphs, and analyze its performance from a dissipativity-based perspective. We first reformulate the algorithm as a Lure type system, consisting of a linear subsystem that relies on the communication topology and the algorithm gains, and a static nonlinear gradient feedback. We then show that the linear subsystem is dissipative with respect to a suitable supply rate, while the nonlinear feedback is not passive. Finally, we establish that, by properly selecting the gains or appropriately designing the communication network, this algorithm converges to an equilibrium at an exponential rate, and thus, achieves an optimal solution to the distributed problem. This work provides new insights into the roles of the network topology, algorithm gains, and cost functions in the performance of a distributed algorithm, and complements existing results from a different viewpoint.

math.OC

Electrical detection of high-order optical orbital angular momentum

The orbital angular momentum (OAM) of light provides an unbounded set of orthogonal modes for ultrahigh-capacity optical information processing. However, current OAM detection schemes typically rely on light interference or diffraction, which require bulky optical components and pose a major obstacle to on-chip integration. Here, we demonstrate a fully integrated silicon-based photodetector that enables direct electrical detection of light OAM. This photodetector can resolve vortex beams with topological charges from m = -9 to 9, achieving a record-high mode number resolution among on-chip devices. By integrating plasmonic gratings onto the device electrodes, incident vortex beams can be converted into surface plasmon polaritons with OAM-dependent splitting angles, which in turn produce photocurrents that vary monotonically with the OAM order. Further incorporation of a surface dielectric lens can enhance mode resolution, and a split-electrode architecture enables OAM chirality discrimination. Owing to its CMOS-compatibility and spectral scalability, this platform provides a compact and robust solution for integrated OAM detection, opening new opportunities for on-chip optical communication and computing systems based on structured light.

physics.optics

Auxiliary-Hyperparameter-Free Sampling: Entropy Equilibrium for Text Generation

Token sampling strategies critically influence text generation quality in large language models (LLMs). However, existing methods introduce additional hyperparameters, requiring extensive tuning and complicating deployment. We present Entropy Equilibrium Sampling (EES), an auxiliary hyperparameter-free approach inspired by information theory that can dynamically adjust candidate sets by balancing normalized entropy with probability mass. We evaluate EES on both reasoning and generation tasks across a range of model architectures. Our results show that EES consistently performs well across temperature settings, delivering competitive accuracy and coherence while maintaining diversity. By eliminating the need for hyperparameter tuning, EES greatly simplifies deployment while improving performance. Code is available at https://github.com/shuanncai/EES

cs.CL

MA-enhanced Mixed Near-field and Far-field Covert Communications

In this paper, we propose to employ a modular-based movable extremely large-scale array (XL-array) at Alice for enhancing covert communication performance. Compared with existing work that mostly considered either far-field or near-field covert communications, we consider in this paper a more general and practical mixed-field scenario, where multiple Bobs are located in either the near-field or far-field of Alice, in the presence of multiple near-field Willies. Specifically, we first consider a two-Bob-one-Willie system and show that conventional fixed-position XL-arrays suffer degraded sum-rate performance due to the energy-spread effect in mixed-field systems, which, however, can be greatly improved by subarray movement. On the other hand, for transmission covertness, it is revealed that sufficient angle difference between far-field Bob and Willie as well as adequate range difference between near-field Bob and Willie are necessary for ensuring covertness in fixed-position XL-array systems, while this requirement can be relaxed in movable XL-array systems thanks to flexible channel correlation control between Bobs and Willie. Next, for general system setups, we formulate an optimization problem to maximize the achievable sum-rate under covertness constraint. To solve this non-convex optimization problem, we first decompose it into two subproblems, corresponding to an inner problem for beamforming optimization given positions of subarrays and an outer problem for subarray movement optimization. Although these two subproblems are still non-convex, we obtain their high-quality solutions by using the successive convex approximation technique and devising a customized differential evolution algorithm, respectively. Last, numerical results demonstrate the effectiveness of proposed movable XL-array in balancing sum-rate and covert communication requirements.

eess.SP