arXiv ScienceSearch

arXiv subjects

Weihao Wang

Publications and source records attributed to Weihao Wang.

15 recordsLinked to original sources

GPU-Accelerated Gate-Level Time-Based Power Analysis via Event-Density-Aware Partitioning and Kernel Fusion

Power analysis is crucial in modern chip design flow. Particularly, time-based power analysis can provide fine-grained power consumption information to facilitate the diagnosis of power issues and guide power optimization accordingly. However, it may take tens of hours to conduct time-based power analysis on modern large-scale circuits, which greatly slows down the power optimization flow. In this paper, we present the first GPU-accelerated gate-level time-based power analysis framework. We propose a novel data structure to enable efficient state-dependent power retrieval. To accommodate the imbalanced event distribution across gates, we propose an event-density-aware partitioning strategy that allocates GPU threads based on gate event density. Finally, we fuse the power computation into a single kernel invocation to reduce redundant work in separate kernels. Experimental results show that our proposed framework achieves high accuracy while delivering up to 37.63x end-to-end speedup compared to multi-threaded Synopsys PrimeTime PX.

cs.AR

RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations

Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the generative latent space discretized and reconstruction-oriented. We propose RaPD (Resolution-agnostic Pixel Diffusion), which performs diffusion in a continuous Neural Image Field (NIF) latent space. RaPD bridges this reconstruction-generation gap with Semantic Representation Guidance for generation-aware latent learning and a Coordinate-Queried Attention Renderer for coordinate-conditioned, scale-aware rendering. A single denoised latent can be rendered at arbitrary resolutions by changing only the query coordinates, keeping diffusion cost fixed. Experiments demonstrate superior generation quality and resolution scalability.

cs.CV

VisiFold: Long-Term Traffic Forecasting via Temporal Folding Graph and Node Visibility

Traffic forecasting is a cornerstone of intelligent transportation systems. While existing research has made significant progress in short-term prediction, long-term forecasting remains a largely uncharted and challenging frontier. Extending the prediction horizon intensifies two critical issues: escalating computational resource consumption and increasingly complex spatial-temporal dependencies. Current approaches, which rely on spatial-temporal graphs and process temporal and spatial dimensions separately, suffer from snapshot-stacking inflation and cross-step fragmentation. To overcome these limitations, we propose \textit{VisiFold}. Our framework introduces a novel temporal folding graph that consolidates a sequence of temporal snapshots into a single graph. Furthermore, we present a node visibility mechanism that incorporates node-level masking and subgraph sampling to overcome the computational bottleneck imposed by large node counts. Extensive experiments show that VisiFold not only drastically reduces resource consumption but also outperforms existing baselines in long-term forecasting tasks. Remarkably, even with a high mask ratio of 80\%, VisiFold maintains its performance advantage. By effectively breaking the resource constraints in both temporal and spatial dimensions, our work paves the way for more realistic long-term traffic forecasting. The code is available at~ https://github.com/PlanckChang/VisiFold.

cs.AI

EMAformer: Enhancing Transformer through Embedding Armor for Time Series Forecasting

Multivariate time series forecasting is crucial across a wide range of domains. While presenting notable progress for the Transformer architecture, iTransformer still lags behind the latest MLP-based models. We attribute this performance gap to unstable inter-channel relationships. To bridge this gap, we propose EMAformer, a simple yet effective model that enhances the Transformer with an auxiliary embedding suite, akin to armor that reinforces its ability. By introducing three key inductive biases, i.e., \textit{global stability}, \textit{phase sensitivity}, and \textit{cross-axis specificity}, EMAformer unlocks the further potential of the Transformer architecture, achieving state-of-the-art performance on 12 real-world benchmarks and reducing forecasting errors by an average of 2.73\% in MSE and 5.15\% in MAE. This significantly advances the practical applicability of Transformer-based approaches for multivariate time series forecasting. The code is available on https://github.com/PlanckChang/EMAformer.

cs.LG

Pre-equalization Design for ISAC-OTFS Air-Ground Transmission: A Deep Learning Approach

Despite the strong Doppler resilience capability, orthogonal time-frequency space (OTFS) modulation suffers from high channel estimation and equalization complexity at the receiver, hindering its applicability in air-ground transmission. In this paper, we propose a pre-equalization-based integrated sensing and communications-OTFS downlink transmission framework in which the terrestrial access point executes pre-equalization using the predicted channel state information (CSI), so that the unmanned aerial vehicle can perform direct symbol detection without channel equalization. In particular, the mean square error of OTFS symbol demodulation and Cramer-Rao lower bound of sensing parameter estimation are considered, with their weighted sum utilized as the metric for optimizing the pre-equalization matrix. To address the time-varying CSI, we develop a deep learning based framework composed of channel prediction and pre-equalization. In particular, a parameter-level channel prediction module is utilized to decouple OTFS channel parameters, and a low-dimensional prediction network is leveraged to correct outdated CSI, which is then used to initialize the input of the pre-equalization module. Finally, a dual-branch residual-structured deep neural network is cascaded to execute pre-equalization. Simulation results show that the proposed channel prediction-based pre-equalization framework significantly reduces receiver complexity and pilot overhead while achieving symbol detection performance close to minimum mean square error equalization with perfect CSI under high mobility, as well as substantially improving sensing accuracy.

eess.SP

CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models

Chain-of-thought (CoT) reasoning boosts large language models' (LLMs) performance on complex tasks but faces two key limitations: a lack of reliability when solely relying on LLM-generated reasoning chains and lower reasoning performance from natural language prompts compared with code prompts. To address these issues, we propose CoT-RAG, a novel reasoning framework with three key designs: (i) Knowledge Graph-driven CoT Generation, featuring knowledge graphs to modulate reasoning chain generation of LLMs, thereby enhancing reasoning credibility; (ii) Learnable Knowledge Case-aware RAG, which incorporates retrieval-augmented generation (RAG) into knowledge graphs to retrieve relevant sub-cases and sub-descriptions, providing LLMs with learnable information; (iii) Pseudo Program Prompting Execution, which promotes greater logical rigor by guiding LLMs to execute reasoning tasks as pseudo-programs. Evaluations on nine public datasets spanning three reasoning tasks reveal significant accuracy gains-ranging from 4.0% to 44.3%-over state-of-the-art methods. Furthermore, tests on four domain-specific datasets demonstrate exceptional accuracy and efficient execution, underscoring its practical applicability and scalability. Our code and data are available at https: //github.com/hustlfy123/CoT-RAG.

cs.CL

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively handle inputs and outputs of various and mixed modalities. The unified model flexibly supports a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation. Across various benchmarks, it demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation. This significantly highlights its potential as a next-generation foundation model. Code and models are released at https://github.com/showlab/Show-o.

cs.CV

Field-free perpendicular magnetization switching of low critical current density at room temperature in TaIrTe4/ferromagnet heterostructures

Spin-orbit torque-induced perpendicular magnetization switching has attracted much attention due to the advantages of nonvolatility, high density, infinite read/write counts, and low power consumption in spintronic applications. To achieve field-free deterministic switching of perpendicular magnetization, additional magnetic field, magnetic layer assistance, or artificially designed structural symmetry breaking are usually required, which are not conducive to the high-density integration and application of low-power devices. However, 2D Weyl semimetals with low-symmetry structures have recently been found to generate z-spin-polarized currents, which may induce out-of-plane damping-like torques to their neighboring ferromagnetic layers, and realize deterministic perpendicular magnetization switching at zero magnetic field. In this Letter, we report that current-induced field-free magnetization switching at room temperature can be achieved in a perpendicularly magnetized TaIrTe4/Pt/Co/Pt device, and the critical switching current density can be lowered to be about 2.64*105 Acm-2. This study suggests that TaIrTe4 has great potential for the design of room-temperature efficient spintronic devices.

cond-mat.mtrl-sci

CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models

Large language models (LLMs) have achieved remarkable performance on various NLP tasks, yet their potential in more challenging and domain-specific task, such as finance, has not been fully explored. In this paper, we present CFinBench: a meticulously crafted, the most comprehensive evaluation benchmark to date, for assessing the financial knowledge of LLMs under Chinese context. In practice, to better align with the career trajectory of Chinese financial practitioners, we build a systematic evaluation from 4 first-level categories: (1) Financial Subject: whether LLMs can memorize the necessary basic knowledge of financial subjects, such as economics, statistics and auditing. (2) Financial Qualification: whether LLMs can obtain the needed financial qualified certifications, such as certified public accountant, securities qualification and banking qualification. (3) Financial Practice: whether LLMs can fulfill the practical financial jobs, such as tax consultant, junior accountant and securities analyst. (4) Financial Law: whether LLMs can meet the requirement of financial laws and regulations, such as tax law, insurance law and economic law. CFinBench comprises 99,100 questions spanning 43 second-level categories with 3 question types: single-choice, multiple-choice and judgment. We conduct extensive experiments of 50 representative LLMs with various model size on CFinBench. The results show that GPT4 and some Chinese-oriented models lead the benchmark, with the highest average accuracy being 60.16%, highlighting the challenge presented by CFinBench. The dataset and evaluation code are available at https://cfinbench.github.io/.

cs.CL

Touch100k: A Large-Scale Touch-Language-Vision Dataset for Touch-Centric Multimodal Representation

Touch holds a pivotal position in enhancing the perceptual and interactive capabilities of both humans and robots. Despite its significance, current tactile research mainly focuses on visual and tactile modalities, overlooking the language domain. Inspired by this, we construct Touch100k, a paired touch-language-vision dataset at the scale of 100k, featuring tactile sensation descriptions in multiple granularities (i.e., sentence-level natural expressions with rich semantics, including contextual and dynamic relationships, and phrase-level descriptions capturing the key features of tactile sensations). Based on the dataset, we propose a pre-training method, Touch-Language-Vision Representation Learning through Curriculum Linking (TLV-Link, for short), inspired by the concept of curriculum learning. TLV-Link aims to learn a tactile representation for the GelSight sensor and capture the relationship between tactile, language, and visual modalities. We evaluate our representation's performance across two task categories (namely, material property identification and robot grasping prediction), focusing on tactile representation and zero-shot touch understanding. The experimental evaluation showcases the effectiveness of our representation. By enabling TLV-Link to achieve substantial improvements and establish a new state-of-the-art in touch-centric multimodal representation learning, Touch100k demonstrates its value as a valuable resource for research. Project page: https://cocacola-lab.github.io/Touch100k/.

cs.RO

π-π Interaction-facilitated formation of interwoven trimeric cage-catenanes with topological chirality

Catenanes as interlocked molecules with a nonplanar graph have gained increasing attention for their unique features such as topological chirality. To date, the majority of research in this field has been focusing on catenanes comprising monocyclic rings. Due to the lack of rational synthetic strategy, catenanes of cage-like monomers are hardly accessible. Here we report on the construction of an interwoven trimeric catenane that is composed of achiral organic cages, which exhibits topological chirality. Our rational design begins with a pure mathematical analysis, revealing that the formation probability of the interwoven trimeric catenane surpasses that of its chain-like analogue by 20%; while driven by efficient template effect provided by strong π-π stacking of aromatic panels, the interwoven structure emerges as the dominant species, almost ruling out the formation of the chain-like isomer. Its topological chirality is unambiguously unravelled by chiral-HPLC, CD spectroscopy and X-ray diffraction. Our probability analysis-aided rational design strategy would pave a new venue for the efficient synthesis of topologically sophisticated structures in one pot.

physics.chem-ph

Outage Performance of Multi-tier UAV Communication with Random Beam Misalignment

By exploiting the degree of freedom on the altitude, unmanned aerial vehicle (UAV) communication can provide ubiquitous communication for future wireless networks. In the case of concurrent transmission of multiple UAVs, the directional beamforming formed by multiple antennas is an effective way to reduce co-channel interference. However, factors such as airflow disturbance or estimation error for UAV communications can cause the occurrence of beam misalignment. In this paper, we investigate the system performance of a multi-tier UAV communication network with the consideration of unstable beam alignment. In particular, we propose a tractable random model to capture the impacts of beam misalignment in the 3D space. Based on this, by utilizing stochastic geometry, an analytical framework for obtaining the outage probability in the downlink of a multi-tier UAV communication network for the closest distance association scheme and the maximum average power association scheme is established. The accuracy of the analysis is verified by Monte-Carlo simulations. The results indicate that in the presence of random beam misalignment, the optimal number of UAV antennas needs to be adjusted to be relatively larger when the density of UAVs increases or the altitude of UAVs becomes higher.

eess.SY

Spheroid Model for Molecular Packing in Crystalline Phase

Dense packing of particles has provided important models to study the structure of matter in various systems such as liquid, glassy and crystalline phase, etc. The simplest sphere packing models are able to represent and capture salient properties of the building blocks for covalent, metallic and ionic crystals; it however becomes insufficient to reflect the broken symmetry of the commonly anisotropic molecules in complex molecular crystals. Here we develop spheroid models with the minimal degree of anisotropy, which serve as a simple geometrical representation for a rich spectrum of molecules--including both isotropic and anisotropic, convex and concave ones--in crystalline phases. Our models are determined via an inverse packing approach: given a molecular crystal, an optimal spheroid model is constructed using a contact diagram, which depicts packing relationship between neighboring molecules within the crystal. The spheroid models are capable of accurately capturing the broken symmetry and characterizing the equivalent volume of molecules in the crystalline phases. Our model also allows to retrieve such molecular information from poor-quality crystal X-ray diffraction data that otherwise would be simply discarded.

cond-mat.soft

3D Part Assembly Generation with Instance Encoded Transformer

It is desirable to enable robots capable of automatic assembly. Structural understanding of object parts plays a crucial role in this task yet remains relatively unexplored. In this paper, we focus on the setting of furniture assembly from a complete set of part geometries, which is essentially a 6-DoF part pose estimation problem. We propose a multi-layer transformer-based framework that involves geometric and relational reasoning between parts to update the part poses iteratively. We carefully design a unique instance encoding to solve the ambiguity between geometrically-similar parts so that all parts can be distinguished. In addition to assembling from scratch, we extend our framework to a new task called in-process part assembly. Analogous to furniture maintenance, it requires robots to continue with unfinished products and assemble the remaining parts into appropriate positions. Our method achieves far more than 10% improvements over the current state-of-the-art in multiple metrics on the public PartNet dataset. Extensive experiments and quantitative comparisons demonstrate the effectiveness of the proposed framework.

cs.RO

Robust Optimization on Unrelated Parallel Machine Scheduling with Setup Times

The parallel machine scheduling problem has been a popular topic for many years due to its theoretical and practical importance. This paper addresses the robust makespan optimization problem on unrelated parallel machine scheduling with sequence-dependent setup times, where the processing times are uncertain, and the only knowledge is the intervals they take values from. We propose a robust optimization model with min-max regret criterion to formulate this problem. To solve this problem, we prove that the worst-case scenario with the maximum regret for a given solution belongs to a finite set of extreme scenarios. Based on this theoretical analysis, the procedure to obtain the maximum regret is proposed and an enhanced regret evaluation method (ERE) is designed to accelerate this process. A multi-start decomposition-based heuristic algorithm (MDH) is proposed to solve this problem. High-quality initial solutions and an upper bound are examined to help better solve the problem. Computational experiments are conducted to justify the performance of these methods.

math.OC