arXiv ScienceSearch

arXiv subjects

Zehao Fan

Publications and source records attributed to Zehao Fan.

6 recordsLinked to original sources

Exploring causal relationships in elastomeric fracture

The conventional interpretation of elastomeric fracture treats the tearing energy G as an output determined by tear speed, taken to equal crack velocity vc, concluding that higher vc produced greater viscoelastic dissipation to increase G. Based on spatially and temporally resolved polarized optical microscopic (str-POM) measurements of tip stress tip stress, we use prenotched pure shear experiments to clarify the causality: Local crack-tip stress determines the network lifetime, making vc the passive kinetic output. In the elastic limit that can be readily achieved for a well crosslinked elastomer in a wide range of temperatures, Rivlin-Thomas scaling holds, and vc depends only on the applied strain, independent of stretch rate, indicating one-to-one correspondence between vc and tip stress. In a highly stretchable elastomer made with reduced crosslink density, vc no longer correlates well with the far-field load as rate-dependent viscoelastic processes emerge to affect network lifetime. Because of the rheological effects on how chain tension builds in the elastomeric network, vc is higher at a common nominal strain when it is imposed with higher stretch rate. confirm that different stretch rates produce different levels of tip stress. The same tip stress still produces the same vc. With str-POM measurements we are also able to elucidate the nature of temperature dependence: at a given tip stress,, crack growth is slower at lower temperatures; conversely it requires higher tip stress to produce the same vc at lower temperatures.

cond-mat.soft

Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation

Mixture-of-Experts (MoE) models scale capacity via sparse activation but stress memory and bandwidth. Offloading alleviates GPU memory by fetching experts on demand, yet token-level routing causes irregular transfers that make inference I/O-bound. Static uniform quantization reduces traffic but degrades accuracy under aggressive compression by ignoring expert heterogeneity. We present Bandwidth-Efficient Adaptive Mixture-of-Experts via Low-Rank Compensation, which performs router-guided precision restoration using precomputed low-rank compensators. At inference time, our method transfers compact low-rank factors with Top-n (n<k) experts per token and applies compensation to them, keeping others low-bit. Integrated with offloading on GPU and GPU-NDP systems, our method delivers a superior bandwidth-accuracy trade-off and improved throughput.

cs.LG

Context-Aware Mixture-of-Experts Inference on CXL-Enabled GPU-NDP Systems

Mixture-of-Experts (MoE) models scale large language models through conditional computation, but inference becomes memory-bound once expert weights exceed the capacity of GPU memory. In this case, weights must be offloaded to external memory, and fetching them incurs costly and repeated transfers. We address this by adopting CXL-attached near-data processing (CXL-NDP) as the offloading tier to execute cold experts in place, converting expensive parameter movement into cheaper activation movement. Unlike prior GPU-NDP systems that are largely context-agnostic and reactive, we develop a context-aware MoE system that uses prefill-stage activation statistics to guide decoding-stage expert placement, dynamically pins hot experts in GPU-side HBM, and maps the remainder to CXL-NDP. To meet NDP's limited compute throughput, we introduce context-aware mixed-precision quantization that allocates per-expert bitwidths (1-4 bit) based on prefill stage. The resulting MoE inference system overlaps GPU and NDP execution while minimizing cross-device movement. The evaluation on the GPU-NDP system shows that our approach achieves up to an 8.7-fold decoding throughput improvement over the state-of-the-art method, while incurring only a 0.13% average accuracy drop.

cs.LG

RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence

Recent advances in video generation have enabled the synthesis of videos with strong temporal consistency and impressive visual quality, marking a crucial step toward vision foundation models. To evaluate these video generation models, existing benchmarks primarily focus on factors related to visual perception and understanding, like visual aesthetics, instruction adherence, and temporal coherence. However, the rule-based reasoning capabilities of video generation models remain largely unexplored. Although recent studies have carried out preliminary explorations into whether video models can serve as zero-shot learners, they still lack a fine-grained decomposition of reasoning capabilities and a comprehensive evaluation protocol. To address this gap, we introduce RULER-Bench, a benchmark designed to evaluate the reasoning ability of video generation models from the perspective of cognitive rules. Built upon two fundamental paradigms: text-to-video and image-to-video, RULER-Bench covers 40 representative tasks spanning six rule categories with 622 high-quality annotated instances. For the evaluation of each generated video, we construct a checklist covering four metrics and leverage GPT-o3 to assign scores to each question, achieving 85% alignment with human judgements. Extensive experiments show that the state-of-the-art model achieves only 48.87% on the rule coherence metric, highlighting significant room for improvement in the reasoning capability of next-level video models. We expect that the insight obtained from RULER-Bench will facilitate further development of reasoning-aware video generation, advancing video generation models toward vision foundation intelligence.

cs.CV

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM

Transformer-based models are the foundation of modern machine learning, but their execution, particularly during autoregressive decoding in large language models (LLMs), places significant pressure on memory systems due to frequent memory accesses and growing key-value (KV) caches. This creates a bottleneck in memory bandwidth, especially as context lengths increase. Processing-in-memory (PIM) architectures are a promising solution, offering high internal bandwidth and compute parallelism near memory. However, current PIM designs are primarily optimized for dense attention and struggle with the dynamic, irregular access patterns introduced by modern KV cache sparsity techniques. Consequently, they suffer from workload imbalance, reducing throughput and resource utilization. In this work, we propose STARC, a novel sparsity-optimized data mapping scheme tailored specifically for efficient LLM decoding on PIM architectures. STARC clusters KV pairs by semantic similarity and maps them to contiguous memory regions aligned with PIM bank structures. During decoding, queries retrieve relevant tokens at cluster granularity by matching against precomputed centroids, enabling selective attention and parallel processing without frequent reclustering or data movement overhead. Experiments on the HBM-PIM system show that, compared to common token-wise sparsity methods, STARC reduces attention-layer latency by 19%--31% and energy consumption by 19%--27%. Under a KV cache budget of 1024, it achieves up to 54%--74% latency reduction and 45%--67% energy reduction compared to full KV cache retrieval. Meanwhile, STARC maintains model accuracy comparable to state-of-the-art sparse attention methods, demonstrating its effectiveness in enabling efficient and hardware-friendly long-context LLM inference on PIM architectures.

cs.CL

Resolving stress state at crack tip to elucidate nature of elastomeric fracture

Based on spatial-temporal resolved measurements of the stress field at crack tip based on polarized optical microscopy (str-POM), the stress analysis approach to elastomeric fracture uncovers new insights. We show new phenomenology in contrast to the standard description of linear elastic fracture mechanics (LEFM). First, str-POM measurements show emergence of a stress saturation zone whose dimension r_ss is independent of the stress intensity factor K. This elastic zone is plastic zone whose size would scale quadratically with K. The absence of stress divergence allows us to measure tip stress s_tip at the onset of fracture, identified as inherent material strength, i.e., s_tip(F) = s_F(inh). We are able to explain why LEFM applies well to elastomers, i.e., why toughness (either given as critical energy release rate Gc or critical stress intensity factor Kc) is a material constant, and we have identified parameters that determine the magnitude of toughness. Second, the popular Rivlin-Thomas energy balance description of elastomeric fracture in pure shear has acquired a fresh and different interpretation based on str-POM observations, which show that the stress buildup at cut tip explicitly scales with specimen height h0, leading to Gc = wch0 being constant. Third, the str-POM observations reveal how elastomeric fracture occurs at a common Kc independent of specimen thickness. At a given load there is weaker stress buildup for a thicker specimen due to greater stress saturation at cut tip, and fracture is observed to occur at lower tip stress for a thicker specimen.

cond-mat.soft