arXiv ScienceSearch

arXiv subjects

Lu Lu

Publications and source records attributed to Lu Lu.

At least 19 recordsLinked to original sources

Equality Cases for the Face-Degree Majorization Theorem on Simplicial Complexes

The Grone--Merris--Bai theorem states that the Laplacian spectrum of a simple graph is majorized by its conjugate degree sequence. Recently, Zhang, Song, and Fan extended this result to simplicial complexes by establishing a majorization relation between the spectrum of the $(r-1)$-dimensional up-Laplacian and the conjugate $(r-1)$-degree sequence. In this paper, we characterize all equality cases in the partial-sum inequalities of this higher-dimensional majorization theorem. For every $r$-dimensional simplicial complex $X$ with $r\ge2$, we prove that \[ \sum_{i=1}^{q}\lambda_{r-1,i}(X) = \sum_{i=1}^{q}d_{r-1,i}^{\top}(X) \] if and only if \[ q\ge \max\{\operatorname{rank}B_r(X),\Delta_{r-1}(X)\}. \] Thus, unlike the graph case, equality can occur only after both sequences have exhausted all their nonzero terms. As consequences, equality in the first partial sum and equality between the entire sequences are both equivalent to $X$ containing a unique $r$-simplex. The proof is based on the local down-Laplacian decomposition and the equality case of the Ky Fan inequality.

math.CO

A Spectral Hilton--Milner--Frankl Theorem for $t$-Intersecting Families

Keevash, Lenz, and Mubayi proved a spectral Erd\H{o}s--Ko--Rado theorem, showing that, for sufficiently large $n$, the complete $t$-star uniquely maximizes the adjacency-tensor spectral radius among all $t$-intersecting $k$-uniform families. In this paper, we establish a spectral Hilton--Milner--Frankl theorem for nontrivial $t$-intersecting families in the explicit range $1\le t\le k-2$ and $n\ge 100\cdot 2^k k^7$. More precisely, we prove that, for every nontrivial $t$-intersecting $k$-uniform family $\mathcal F$, the spectral radius satisfies \[ \rho(\mathcal F)\le \max\{\rho(\mathcal H_{n,k,t}),\rho(\mathcal A_{n,k,t})\}, \] where $\mathcal H_{n,k,t}$ and $\mathcal A_{n,k,t}$ are the two extremal families appearing in the classical Hilton--Milner--Frankl theorem. Moreover, equality holds only for the extremal candidates attaining the maximum, up to isomorphism. We further compare the two candidates asymptotically. For each fixed $t$, the unique real solution $x=x_t$ of \[ (t+2)^{x-t-1}(t+1)^{t+1}=(x-t+1)^{x-1} \] determines, as $k$ varies, which of $\mathcal H_{n,k,t}$ and $\mathcal A_{n,k,t}$ has the larger asymptotic spectral radius.

math.CO

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are poorly represented by examination-oriented medical benchmarks. Objective: To develop RESPClinBench, a real-world scenario-based benchmark for respiratory clinical decision-making, and evaluate seven contemporary large language models across AECOPD-PIM and PNBIM. Methods: RESPClinBench cases were adapted from de-identified respiratory clinical data. Three attending-level respiratory physicians revised cases, reference answers, and atomic clinical-action points, while one senior respiratory specialist performed cross-review and final adjudication. AECOPD-PIM comprised 427 open-ended COPD cases, and PNBIM comprised 196 multimodal pulmonary nodule cases combining chest CT with structured clinical information. Seven models generated 4,361 responses through standardized API inference with temperature 0 and a maximum output length of 8192 tokens. An automated framework calculated the final score as the arithmetic mean of atomic-action recall and rubric-based LLM-as-a-Judge assessment. Results: Across 623 cases, the mean final score was 68.58. Qwen3.6-27B ranked first overall at 71.22, Qwen3.5-397B-A17B led PNBIM at 72.48, and Qwen3.6-27B led AECOPD-PIM at 71.11. Imaging hallucination and serious medical risk occurred in 31.85% and 8.16% of PNBIM responses; medication-safety risk and serious medical risk occurred in 26.93% and 1.44% of AECOPD-PIM responses. Conclusions: RESPClinBench identifies task-specific limitations in multimodal pulmonary nodule assessment and longitudinal COPD management. Combining explicit clinical-action coverage, holistic evaluation, and independent safety flags provides a clinically grounded basis for model selection and prospective validation.

cs.CL

An improved range for the maximum critically $t$-intersecting hypergraphs

Let $k>t\ge 1$ be integers and set $d=k-t$. A $k$-uniform hypergraph $\mathcal F$ is called $t$-intersecting if any two edges intersect in at least $t$ vertices, and is called $t$-critical if its minimum $t$-transversal has size $k$. Frankl proved that, for $k\ge d^4$,$|\mathcal F|\le \binom{k+d}{d},$ with equality only for the complete $k$-graph on $k+d$ vertices, and conjectured that the same conclusion should hold when $k>c d^2$ for some constant $c$. In this paper we confirm this conjecture for $c=30$. The proof relies on Frankl's fixed-edge decomposition and F\"{u}redi's pseudo-sunflower method.

math.CO

CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop CardioBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: CardioBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Open-ended tasks were evaluated using key-point coverage and holistic clinical quality, while CardioEthics was scored by accuracy. Results: GPT-5.4 achieved the highest macro-average (62.55) and item-weighted mean (62.19), followed by Gemini 3.1 Pro (59.95) and Qwen 3.6 27B (59.72). GPT-5.4 ranked first in all three dimensions. CardioAuxReport performed best (86.38), whereas CardioECGRead (17.25) and CardioEthics (17.34) were lowest. The largest gaps between holistic clinical quality and key-point coverage occurred in CardioComm (52.71), CardioEmergRescue (52.05), and CardioTreatPlan (48.80). Conclusions: To our knowledge, CardioBench is the largest real-world, multi-task benchmark for LLM evaluation across the cardiovascular care continuum and offers the broadest coverage of clinically authentic cardiology scenarios reported to date. It provides a rigorous framework for identifying model strengths, clinically important omissions, and priorities for future development.

cs.CL

SkillComm: Skill-Driven Semantic Communication for Sequential Workflows via Incremental Token Transmission

As wireless visual intelligence evolves from isolated task inference to ordered skill workflows, the communication bottleneck shifts from transmitting a single semantic representation to coordinating reusable skill states under channel constraints. Existing DeepJSCC and prompt-guided visual transmitters usually treat each task as an independent full-token transmission, with limited reuse of execution memory across semantic workflows. This is inefficient for workflows such as Detect, Segment, and Keypoint, where later stages often require only state-relevant semantic updates. To this end, we propose SkillComm, a skill-driven semantic communication framework that uses reusable skill states as shared context for workflow-aware token prioritization and memory-assisted token-grid reconstruction. A shared Skill-Book maps a high-level visual intent into a synchronized executable skill sequence at the transmitter and receiver. Conditioned on this workflow, adaptive token selection exploits cross-step memory to transmit only state-active tokens through joint source-channel coding, while the receiver reconstructs a task-ready token grid by combining decoded tokens with local historical memory. Experiments on the MS COCO 2017 validation set for the Detect-Segment-Keypoint workflow show that SkillComm reduces token transmission cost by 51.2% while retaining 99.4% upper-bound-normalized average precision at high SNR. These results demonstrate that reusable skill states enable selective semantic update delivery for future agentic and embodied visual intelligence.

cs.NI

Neural Operator-enabled Topology-informed Evolutionary Strategy for PDE-Constrained Optimization

The inverse design of physical systems governed by partial differential equations is computationally demanding due to the high dimensionality and non-convexity of design spaces. Generative models for inverse design often lack robustness and transferability, whereas evolutionary strategies are robust but struggle in high-dimensional spaces. This paper introduces a Neural Operator-enabled Topology-informed Evolutionary Strategy (NOTES) that integrates dimensionality reduction, representation learning, and evolutionary optimization for efficient and transferable inverse design. NOTES couples a DeepONet-based neural operator with the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) to perform global optimization in a compact latent space that encodes topology-aware priors while discovering high-performance designs for unseen operating conditions. Applied to nanophotonic beam-deflector inverse design governed by Maxwell's equations, NOTES reduces the design dimensionality from 256 to 25 and consistently achieves over 95 percent efficiency, outperforming CMA-ES, topology optimization, and other baselines. Applied to structural optimization, NOTES discovers designs that achieve compliance down to 246. By decoupling topology learning of a DeepONet from the governing physics in a PDE solver, NOTES provides a flexible and transferable framework for the inverse design of physical systems.

cs.LG

On a conjecture regarding the product version of the Hilton-Milner theorem

Recently, Frankl and Wang considered a product version of the classical Hilton-Milner theorem. They conjectured that, if $\mathcal{F} \subset \binom{[n]}{k}$ and $\mathcal{G} \subset \binom{[n]}{\ell}$ are non-trivial cross-intersecting families with $n \geq 2k > 2\ell \geq 4$, the maximum of $|\mathcal{F}||\mathcal{G}|$ is attained by the natural Hilton-Milner-type configurations. In this paper, we present two main results concerning this conjecture. Firstly, we show that the conjecture does not hold in general. By introducing a two-center construction, we prove that for every fixed integer $\ell \geq 3$ and all sufficiently large $k$, the conjecture is false in a linear range $2k+1 \leq n \leq (c_\ell - \epsilon)k$ for any $0 < \epsilon < c_\ell - 2$, where $c_\ell > 2$ is an explicit constant. Secondly, we prove that the conjecture holds when $n > 100\ell k^2$ and $3 \leq \ell < k$, and we completely characterize the extremal families. Our proofs rely on the size of minimal covers and analyzing the structural properties of $2$-cover graphs.

math.CO

The Second Largest Eigenvalue of Stiffness Matrices of Normalized Complete Frameworks

Let $R(G,p)$ be the normalized rigidity matrix of a framework $(G,p)$ in $\mathbb R^d$, and let \[ L(G,p)=R(G,p)R(G,p)^{T} \] be the associated stiffness matrix. We study the extremal eigenvalues of $L(K_n,p)$ for complete frameworks whose vertices lie on the unit sphere and have centroid at the origin. Our main result shows that, whenever $d\ge2$ and the image of $p$ contains at least three distinct points, the second largest eigenvalue of $L(K_n,p)$ is exactly $n/2$. This settles the eigenvalue part of a conjecture of Lew et al. [Israel J. Math. 256, 2023]. We further construct an infinite family of examples, given by regular polygons embedded in a two-dimensional subspace, for which the eigenvalue $n/2$ has multiplicity $2n-4$. Consequently, the multiplicity predicted in the conjecture is not correct in general. Our results reveal a dichotomy: the value of the second largest eigenvalue is universal, while its multiplicity is sensitive to the geometry of the underlying point configuration.

math.CO

Thresholds for the Frankl-Wang $3/7$ conjecture on maximum-degree ratios

Let $\mathcal{F}\subset\binom{[n]}{k}$ be an intersecting family, $\Delta(\mathcal{F})=\max_{x\in[n]}|\{F\in\mathcal{F}:x\in F\}|$, and $\varrho(\mathcal{F})=\Delta(\mathcal{F})/|\mathcal{F}|$. Frankl and Wang conjectured that if $n>100k$ and $|\mathcal{F}|>\binom{n-3}{k-3}$, then $\varrho(\mathcal{F})\ge 3/7$; the constant $3/7$ is sharp because of the Fano-plane construction. In this note we obtain three results. First, we show that no linear threshold $n>Ck$ can be sufficient: using a truncated Fano-plane construction we exhibit, for every constant $C$ and all large $k$, an intersecting family with $n>Ck$, $|\mathcal{F}|>\binom{n-3}{k-3}$, yet $\varrho(\mathcal{F})<3/7$. In particular, the original condition $n>100k$ does not guarantee the conclusion. Second, for $k=3$ we prove that $\varrho(\mathcal{F})\ge 3/7$ holds for every nonempty intersecting $3$-uniform family; the proof is nontrivial and does not rely on any assumption on $n$ or $|\mathcal{F}|$. Third, using the classical pseudo-sunflower bound $|\mathcal{F}|\le t^k$ (for families containing no pseudo-sunflower of size $t+1$), we obtain a completely explicit polynomial threshold for all $k\ge4$: if $n>(k-3)(7k^4+k)+3$ and $|\mathcal{F}|>\binom{n-3}{k-3}$, then $\varrho(\mathcal{F})\ge 3/7$. In particular, the simplified bound $n>7k^5$ is sufficient for every $k\ge4$.

math.CO

FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering limited flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic-frame-rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame-rate controllability. However, this technique has not yet been applied to SLMs. We introduce FlexiSLM, the first SLM with dynamic and controllable frame rates. FlexiSLM uses the pretrained FlexiCodec to obtain dynamic speech output tokens. The main contributions of this work are threefold: (1) integrating and validating this dynamic-rate representation within a multi-task, speech-to-speech SLM architecture; (2) extending it to frame compression on the input side; and (3) introducing direct frame-rate conditioning to enable accurate and controllable SLM inference. FlexiSLM outperforms fixed-frame-rate 7B models including Qwen2.5-Omni and Kimi-Audio at its 12.5 Hz and 6.25 Hz operating points. We further verify that FlexiSLM can be accurately steered down to 4.0 Hz; at 6.25 Hz, it roughly halves inference time relative to 12.5 Hz while retaining strong speech-to-speech quality. Audio samples are available at: https://flexislm.github.io. Code and data are available at: https://github.com/AmphionTeam/FlexiSLM.

cs.SD

Building a Scalable, Reproducible, Evaluatable, and Closed-Loop Simulation Environment Foundation for Embodied Intelligence

This paper presents a cloud-native simulation infrastructure framework for embodied intelligence that supports large-scale training, standardized evaluation, and simulation-based data collection. The framework unifies simulation environment generation, task execution, trajectory collection, model evaluation, data management, and cloud services into a scalable and reproducible platform. To address the high cost, limited scalability, and poor reproducibility of real-world robotic data collection, the framework adopts cloud-native technologies including elastic resource scheduling, containerized simulation, unified data management, and service-oriented system design, enabling efficient large-scale simulation for multi-model and multi-task workloads. Built on a four-layer architecture, the framework provides standardized environment assets, automated task generation, trajectory collection, benchmark evaluation, and closed-loop data optimization. It further integrates representative systems including D-VLA, RL-VLA3, Sword, and Pre-VLA to support scalable simulation, dynamic scheduling, visual augmentation, and real-time data filtering. We argue that cloud-native simulation infrastructure provides a unified foundation for data generation, model training, standardized evaluation, and real-world deployment, and will play a key role in the future development of embodied intelligence.

cs.RO

On the sum of the two largest eigenvalues of the curl-curl operator on graphs

The Grone--Merris conjecture, proved by Bai in~2011, states that the spectrum of the graph Laplacian $\Delta_0 = -\operatorname{div}\operatorname{grad}$ is majorized by the conjugate of the vertex degree sequence. Duval and Reiner proposed a simplicial complex analogue of this statement. On a graph, where triangles serve as $2$-simplices, their conjecture reduces to the assertion that the spectrum of $\operatorname{curl}^*\operatorname{curl}$ is majorized by the conjugate of the second-order degree sequence, which records the number of triangles containing each vertex. We prove that the sum of the two largest eigenvalues of $\operatorname{curl}^*\operatorname{curl}$ does not exceed the sum of the first two entries of that conjugate sequence. This confirms the first two majorization inequalities predicted by Duval and Reiner for $\operatorname{curl}^*\operatorname{curl}$. As a corollary, we obtain upper bounds for the two largest eigenvalues of the full graph Helmholtzian $\Delta_1 = -\operatorname{grad}\operatorname{div} + \operatorname{curl}^*\operatorname{curl}$. The same result extends to the up-Laplacian of any $3$-family, yielding a concrete step towards the Duval--Reiner conjecture in dimension~$1$.

math.CO

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinical multimodal models (language, vision-language, and agent systems) that moves from static QA to dynamic, process-oriented evaluation. MedBench v5 features: (1) a dual-dimensional framework combining Clinical Cognitive Responsiveness (13 sub-dimensions) and Medical Atomic Skills (4 agent environments), covering 63 tasks; (2) three switchable information-flow stressors (omission, contradiction, evidence delay) for factorized degradation analysis; (3) a dynamic process audit protocol with five reasoning nodes that produces model-specific failure fingerprints; (4) hallucination propagation monitoring across initiation, propagation, anchoring, and contradiction interaction-capturing silent hallucination. Experiments on frontier models show that strong overall task performance does not guarantee process stability: stressors mainly disrupt contradiction detection, diagnosis updating, hallucination propagation, and contradiction-based self-correction, while final evidence grounding can remain superficially stable. MedBench v5 provides a unified infrastructure for capability profiling, controllable stress testing, process auditing, and hallucination trajectory analysis in clinical AI evaluation.

cs.CL

Data-driven discovery of governing differential equations across physical systems

Differential equations play a critical role in scientific discovery because they provide a mathematical framework to describe the behaviour of physical phenomena. As a promising alternative to traditional first principles, data-driven differential equation discovery has attracted increasing attention for its ability to infer governing laws directly from experimental or simulated data, especially when the underlying physics is unclear. However, the field has expanded rapidly along diverse methodological directions, particularly with the emergence of AI-based approaches, and still lacks a clear organizing perspective. In this Review, we propose a problem-oriented perspective on data-driven differential equation discovery. We first introduce a two-dimensional phase diagram of equation discoverability, where discovery problems are organized according to structural complexity and coefficient complexity. This phase diagram shows how the field has moved from the discovery of sparse equations with simple coefficients toward more complex governing laws with richer structures and more flexible parameterizations. It also clarifies why different methodological families succeed or fail in different problem settings. We then present the representation-evaluation-optimization (REO) framework as a fundamental abstraction of the discovery process. By identifying the core problems of equation discovery that persist across algorithmic variations, REO shifts the discussion from individual algorithms to the fundamental principles that determine discoverability. We connect these perspectives to applications across physics and adjacent sciences, and argue that the next challenge is not merely recovering equations, but using them to revise existing theories, distil mechanisms and form new scientific concepts.

cs.LG

Foundation Models for Wireless Communications: From PHY Intelligence to Network Autonomy

6G networks will introduce unprecedented complexity, which calls for a paradigm shift in network optimization and management. Artificial intelligence (AI)-based solutions, especially those enabled by the recently developed foundation models, have been recognized as promising candidates. Foundation models are large-scale AI models with general-purpose feature extraction capabilities, and once trained on massive amounts of data, they can be adapted to solve a wide range of downstream tasks, either in a zero-shot manner or with few-shot fine-tuning. This article provides a comprehensive overview of how foundation models are reshaping physical-layer processing and wireless resource management across three progressive paradigms. First, we examine the adaptation of off-the-shelf pre-trained foundation models to various wireless tasks. Second, we explore wireless-native foundation models, built from scratch on wireless data to bridge cross-domain modality gaps and capture universal wireless-domain physical characteristics. Third, we highlight agentic foundation models, which elevate static data processing into autonomous, reasoning-driven network orchestration. Furthermore, we discuss the impact of applying foundation models to emerging 6G frontiers, including integrated sensing and communications (ISAC), new multiple-input multiple-output (MIMO) architectures, semantic communications, and system-level network autonomy. Finally, we identify critical open challenges and opportunities, charting a promising path toward fully intelligent and adaptive wireless networks.

eess.SP

EarlyTom: Early Token Compression Completes Fast Video Understanding

Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is still hindered by the inefficiency introduced by processing massive amounts of visual tokens. Although recent approaches achieve extremely low token retention ratios while maintaining accuracy comparable to full-token baselines, most of them perform compression only at the late stage of prefilling, leaving the efficiency of the vision encoder unoptimized. In this paper, we first show that vision encoding contributes a large portion to the time-to-first-token (TTFT). Therefore, instead of compressing visual tokens only after the vision encoder, performing compression inside the encoder still leaves substantial room for exploration. Based on this insight, we propose EarlyTom, a training-free token compression framework that performs early-stage visual token compression inside the vision encoder, enabling significantly better TTFT reduction and higher throughput. In addition, we introduce a decoupled spatial token selection strategy that improves the overall compression effectiveness. EarlyTom reduces TTFT by up to 2.65x and FLOPs by up to 61% on a single NVIDIA A100 GPU for the LLaVA-OneVision-7B model, while maintaining accuracy comparable to the full-token baseline. These improvements substantially enhance the practicality of deploying Video-LLMs in real-world production scenarios.

cs.CV

Helmholzian Spectra of Graphs: Novel Properties

Let $\grad$, $\curl$, and $\dv$ be the graph-theoretic analogues of the gradient, curl, and divergence operators from multivariate calculus. The graph Laplacian $-\dv \grad$ gives rise to the celebrated Laplacian matrix, while the matrix representation of the graph Helmholtzian $\grad \grad^* + \curl^* \curl$ is called the Helmholtzian matrix. In this paper, we present a new graph-theoretic proof that the Helmholtzian matrix indeed represents the graph Helmholtzian. We then investigate the spectral properties of this matrix. Our main results are as follows: (i) a classification of graphs having exactly two distinct Helmholtzian eigenvalues; (ii) the nullity of the Helmholtzian matrix; and (iii) a combinatorial interpretation of the coefficients of the Helmholtzian polynomial. Furthermore, we determine the Helmholtzian spectrum for certain graph products and characterize Helmholtzian integral graphs, as well as derive bounds for the smallest Helmholtzian eigenvalue. Meanwhile, we pose some open problems for future research.

math.CO