arXiv ScienceSearch

arXiv subjects

Zhenhao Zhang

Publications and source records attributed to Zhenhao Zhang.

At least 19 recordsLinked to original sources

Baryogenesis and CMB spectral distortion from Axions

We discuss a mechanism for generating the baryon asymmetry in the early universe. We show that an axion-like particle can modify the related gauge field configurations in the Standard Model, thereby altering their dispersion relations. This change in the Chern-Simons number can source a violation of baryon number. We derive the relationship between the resulting baryon number and the evolution of the axion background. We estimate the baryon asymmetry produced via this mechanism and show that the observed value can be naturally achieved. We also show that axion photon coupling produces Cosmic Microwave Background spectral distortion. Our results show that the resulting distortion approaches a constant at low frequencies, unlike the conventional y-type and $\mu$-type distortions.

hep-ph

CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions

Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the level of task decisions. A connected route can expand cross-modal reach while changing an established native retrieval capability. We introduce CertBind, a multiscale theory of certifiable composition for frozen multimodal connector graphs. At the node scale, native anchors establish the exact task identification boundary under the stated chart model. At the edge scale, contract-aware conformal ranks provide graph-wide family-wise error control. At the path scale, an overlap-aware budget and clean calibration yield a finite-sample recovery radius under declared conditions. At the query scale, this radius yields a covered top-k candidate set that becomes a point certificate when its size equals k. CertBind therefore retains supported routes as Direct, sends only flagged routes to recovery, returns Certified for decisive recovery, and returns Abstain for unresolved queries. The evaluated C-MCR shared route reduced native CLIP R@1 from 0.524 to 0.290. The production fallback recovered 0.963 +- 0.002 of clean retrieval, while the passing branch recorded a no-harm value of 1.000. CertBind extends multimodal composability from connected representations to certifiable task decisions.

cs.LG

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.

cs.CV

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.

cs.CV

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Existing indoor layout generators produce globally plausible layouts yet may retain local violations such as collisions, out-of-bounds placements, obstructed openings, and blocked circulation. Most prior work focuses on full-scene synthesis or scene-level optimization, with limited support for identifying responsible objects and locally repairing affected regions. We present Roomer, a reflective repair framework that casts these violations as sparse, object-grounded repair problems. Roomer encodes layouts as ``RoState'' and uses ``RoReview'' to bind measured violations to implicated objects. A geometry-conditioned vision-language model planner proposes a structured local edit, while a deterministic solver validates it and generates a finite set of candidate edits when needed. Each candidate is committed only if full-scene verification confirms that it resolves the target violation without new hard violations or broken protected constraints. We train the planner on Roomer-CC, a controlled-corruption dataset that pairs faulty layouts with object-grounded violation evidence and known-feasible inverse StatePatches. Since existing benchmarks rarely assess whether physically valid layouts are usable, we introduce Roomer-Eval to assess distributional quality, physical validity, and practical usability. Experiments show that Roomer repairs residual violations while preserving valid regions, improves physical validity and usability, and transfers across external generators.

cs.RO

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-interface tasks and how we train them with reinforcement learning. This judgment has long relied on rule-based evaluation, which struggles to align with human intention and goes stale when an app updates or its online content drifts. Existing model-based judges attempt to address these problems but still leave a performance gap to the rule-based evaluation. We propose the \textbf{SeekJudge} framework, in which four role-specialized agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek--Analyze loop over the trajectory. A seed-calibrated distillation pipeline trains one specialized $9$B model to serve as the shared backbone for all four agents. Measured by downstream success rate on held-out RL test goals, SeekJudge is the first practical model-based reward to match or surpass native rule-based supervision in online RL. Beyond accuracy, SeekJudge provides step-level judgments, runs far cheaper than a closed-source large model, and keeps a small per-call context that scales to much longer trajectories. We further contribute a general architectural improvement to the reward server that speeds up judging in RL. Together these make model-based reward a practical drop-in for rule-based supervision in CUA reinforcement learning.

cs.AI

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-world embodied manipulation may involve degraded RGB observations caused by illumination shifts, posing critical challenges for robust robotic manipulation. To address this gap, we propose \textbf{Event-VLA}, an event-enhanced VLA framework for generalizable manipulation across varying illumination conditions. We formulate VLA-based manipulation under degraded visibility as a practical robustness problem for RGB-centric policies, and introduce event streams as an illumination-robust, motion-sensitive complementary observation to improve robustness across visibility levels. Specifically, unlike conventional multimodal fusion that directly merges event features into the global semantic token space, Event-VLA injects event information through an action-query routing pathway. It uses learnable action queries to extract task-relevant semantics from the VLA reasoning process, and selectively aggregates event tokens via gated cross-attention to construct event-aware action representations. This design preserves the pretrained RGB-language semantic priors while effectively leveraging event information for robust action prediction. Experiments in simulation and real-world deployment show that Event-VLA maintains strong manipulation performance under normal lighting and improves success rates under low-light degradation and near-dark real-world settings.

cs.CV

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on annotation conventions, span boundaries, entity granularity, and type schemas. Multi-LLM agreement is a salience signal, not corpus-convention correctness. We introduce a candidate-level panel-output benchmark for panel-surfaced candidate verification, where the unit is an aligned candidate surfaced by an explicitly defined multi-model panel rather than a standalone extractor output. The benchmark aligns eight LLMs' predictions over five public biomedical NER datasets into a candidate master table. BioConCal is an in-domain supervised scorer that instantiates this layer with inference-time gold-free agreement, mention, surface-availability, and document features for a fixed candidate stream. In domain, BioConCal improves AUROC from 0.753 for raw agreement to 0.910. At a validation-selected 0.95 precision target it selects 1,340 candidates at empirical test precision 0.939, compared with 293 for raw agreement. This corresponds to candidate-level recall 0.592 and corpus-level recall 0.523 against a within-panel row-label ceiling of 0.883. The main benefit is not recovering entities missed by every panel member, but reshaping a noisy panel stream into a higher-yield review queue. Under entity-type shift, thresholds require target-domain validation, and exact character localization remains a separate deterministic post-processing step.

cs.CL

Size and spectral conditions for a graph with given minimum degree to be $k$-$d$-critical

A $k$-matching in a graph $G$ is defined as a function $f:E(G) \rightarrow \{0,1,\ldots,k\}$ satisfying $\sum_{e\in E_G(v)} f(e)$ $\leq k$ for each vertex $v\in V(G)$, where $E_G(v)$ denotes the set of edges incident to $v$ in $G$. For $1\leq d\leq k$ and $d \equiv |V(G)|~(\mathrm{mod}~2)$, if for any $ v \in V(G)$, there exists a $k$-matching $f$ such that $\sum_{e\in E_G(v)}f(e)=k-d$ and $\sum_{e\in E_G(u)}f(e)=k \text{ for any } u\in V(G)-\{v\}$, then $G$ is $k$-$d$-critical. A graph $G$ of odd order (resp. even order) is generalized factor-critical (resp. generalized bicritical) if the empty set is the unique set attaining the maximum value in $k$-Berge-Tutte-formula of $G$. In this paper, we provide sharp sufficient conditions in terms of size or spectral radius respectively for a graph $G$ to be $k$-$d$-critical, generalized factor-critical and generalized bicritical with minimum degree.

math.CO

AuraDesk: Data Physicalization through Olfaction Metaphors for Representing and Mitigating Workplace Stress

Workplace stress is often addressed through visual or auditory interventions, yet these modalities can compete with attention and contribute to sensory overload. We explore olfaction as an alternative ambient medium for representing stress-related physiological signals in office settings. We present AuraDesk, an olfactory data physicalization system that translates wearable-derived physiological cues into situated scent expressions at the workstation. The system combines local physiological state inference with a constrained actuation strategy to produce temporally regulated and spatially localized scent output suitable for everyday work environments. To examine the feasibility and experiential qualities of this approach, we conducted a one-day in-situ field deployment with 25 knowledge workers at their actual workstations. Our findings show that participants often interpreted the scent output not as an explicit alert, but as a subtle atmospheric cue that supported momentary awareness, micro-break taking, and perceived environmental attunement. At the same time, participants raised important concerns regarding scent preference, habituation, and contextual appropriateness in shared offices. This work contributes (1) an olfactory interface for physiologically driven ambient feedback in the workplace, (2) a hybrid mapping approach for coupling continuous biosignal interpretation with constrained scent actuation, and (3) empirical insights into how workers perceive, negotiate, and appropriate ambient olfactory feedback in real office contexts. Rather than claiming therapeutic efficacy, we position AuraDesk as a probe into the design space of olfactory data physicalization for workplace wellbeing and attention-sensitive interaction.

cs.HC

NEP-CG and NEP-AACG: Efficient coarse-grained and multiscale all-atom-coarse-grained neuroevolution potentials

Machine-learned coarse-grained (CG) models often suffer from noisy training data, limiting their accuracy and transferability. We propose a method to generate low-noise training data based on the potential of mean force by constraining CG beads during atomistic simulations and accumulating time-averaged forces. Implemented within the neuroevolution potential (NEP) framework, our approach achieves training accuracy comparable to atomistic models trained on density functional theory data. For liquid water, the NEP-CG model accurately reproduces densities from 1 bar to 1 GPa, successfully extrapolating beyond the 0.5 GPa training limit, with a virial correction essential for the correct equation of state. For an anisotropic C$_{60}$ monolayer, distinguishing crystallographically distinct bead types reduces stress errors by an order of magnitude and captures directional thermal conductivity. We further introduce a multiscale NEP-AACG model integrating all-atom (AA) and CG degrees of freedom, demonstrated for gold nanowire fracture at an experimentally relevant strain rate. Computational speeds for NEP-CG models reach hundreds to thousands of ns/day using a single consumer-grade GPU. This work provides a robust framework for constructing accurate, transferable, and efficient CG models across diverse systems.

physics.comp-ph

UniHM: Unified Dexterous Hand Manipulation with Vision Language Model

Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or precise hand-object interaction sequences, foregoing the rich, compositional guidance of open-vocabulary instruction. We introduce UniHM, the first framework for unified dexterous hand manipulation guided by free-form language commands. We propose a Unified Hand-Dexterous Tokenizer that maps heterogeneous dexterous-hand morphologies into a single shared codebook, improving cross-dexterous hand generalization and scalability to new morphologies. Our vision language action model is trained solely on human-object interaction data, eliminating the need for massive real-world teleoperation datasets, and demonstrates strong generalizability in producing human-like manipulation sequences from open-ended language instructions. To ensure physical realism, we introduce a physics-guided dynamic refinement module that performs segment-wise joint optimization under generative and temporal priors, yielding smooth and physically feasible manipulation sequences. Across multiple datasets and real-world evaluations, UniHM attains state-of-the-art results on both seen and unseen objects and trajectories, demonstrating strong generalization and high physical feasibility. Our project page at \href{https://unihm.github.io/}{https://unihm.github.io/}.

cs.RO

Distance spectral radius conditions for perfect $k$-matching, generalized factor-criticality (bicriticality) and $k$-$d$-criticality of graphs

Let $G$ be a simple connected graph with vertex set $V(G)$ and edge set $E(G)$. A $k$-matching of a graph $G$ is a function $f:E(G)\rightarrow \{0,1,\ldots, k\}$ satisfying $\sum_{e \in E_G(v)} f(e) \leq k$ for every vertex $v \in V(G)$, where $E_G(v)$ is the set of edges incident with $v$ in $G$. A $k$-matching of a graph $G$ is perfect if $ \sum_{e \in E_G(v) } f(e) = k $ for any vertex $v \in V(G)$. The $k$-Berge-Tutte-formula of a graph $G$ is defined as: \[ \defk(G) = \max_{S \subseteq V(G)} \begin{cases} k \cdot i(G - S) - k|S|, & k \text{ is even;} \\[6pt] \odd(G - S) + k \cdot i(G - S) - k|S|, & k \text{ is odd.} \end{cases} \] A $k$-barrier of the graph $G$ is the subset $S \subseteq V(G)$ that reaches the maximum value in $k$-Berge-Tutte-formula. A connected graph \( G \) of odd (even) order is a {generalized factor-critical (generalized bicritical) graph about integer \( k \)-matching}, abbreviated as a \( \mathrm{GFC}_k (\mathrm{GBC}_k)\) graph, if $\emptyset$ is a unique $k$-barrier. When $k$ is odd, let \( 1 \leq d \leq k \) and \( |V(G)| \equiv d \pmod{2} \). If for any \( v \in V(G) \), there exists a \( k \)-matching \( h \) such that $\sum_{e \in E_G(v)} h(e) = k - d$ {and} $\sum_{e \in E_G(u)} h(e) = k$ for any \( u \in V(G) - \{v\} \), then \( G \) is said to be \( k \)-\( d \)-critical. In this paper, we provide sufficient conditions in terms of distance spectral radius to ensure that a graph has a perfect $k$-matching and a graph is \( k \)-\( d \)-critical, $\mathrm{GFC}_k$ or $\mathrm{GBC}_k$, respectively.

math.CO

Mitigating Conversational Inertia in Multi-Turn Agents

Large language models excel as few-shot learners when provided with appropriate demonstrations, yet this strength becomes problematic in multiturn agent scenarios, where LLMs erroneously mimic their own previous responses as few-shot examples. Through attention analysis, we identify conversational inertia, a phenomenon where models exhibit strong diagonal attention to previous responses, which is associated with imitation bias that constrains exploration. This reveals a tension when transforming few-shot LLMs into agents: longer context enriches environmental feedback for exploitation, yet also amplifies conversational inertia that undermines exploration. Our key insight is that for identical states, actions generated with longer contexts exhibit stronger inertia than those with shorter contexts, enabling construction of preference pairs without environment rewards. Based on this, we propose Context Preference Learning to calibrate model preferences to favor low-inertia responses over highinertia ones. We further provide context management strategies at inference time to balance exploration and exploitation. Experimental results across eight agentic environments and one deep research scenario validate that our framework reduces conversational inertia and achieves performance improvements.

cs.AI

Size conditions and spectral conditions for generalized factor-critical (bicritical) graphs and $k$-$d$-critical graphs

Let $\mbox{odd}(G)$ and $i(G)$ denote the number of nontrivial odd components and the number of isolated vertices of a graph $G$, respectively. The $k$-Berge-Tutte-formula of a graph $G$ is defined as: $\mbox{def}_k(G)=\mathop{\text{max}}\limits_{S\subseteq V(G)}\{k\cdot i(G-S)-k|S|\} $ for even $k$; $\mbox{def}_k(G)=\mathop{\mbox{max}}\limits_{S\subseteq V(G)}\{\mbox{odd}(G-S)+k\cdot i(G-S)-k|S|\} $ for odd $k$. A $k$-barrier of a graph $G$ is the subset $S\subseteq V(G)$ that reaches the maximum value in the $k$-Berge-Tutte-formula of $G$. A graph $G$ of odd order (resp. even order) is generalized factor-critical (resp. generalized bicritical) if $\emptyset$ is its only $k$-barrier. Denote by $E_G(v)$ the set of all edges incident to a vertex $v$ in $G$. A $k$-matching of a graph $G$ is a function $f:E(G) \rightarrow \{0,1,...,k\}$ such that $\sum_{e\in E_G(v)} f(e)$ $\leq k$ for every vertex $v\in V(G)$. For $1\leq d\leq k$ and $d \equiv |V(G)|$(mod 2), if for any $ v \in V(G)$, there exists a $k$-matching $f$ such that $\sum_{e\in E_G(v)}f(e)=k-d$ and $\sum_{e\in E_G(u)}f(e)=k \text{ for any } u\in V(G)-\{v\}$. Then $G$ is $k$-$d$-critical. In this paper, we establish tight sufficient conditions in terms of size or spectral radius respectively for a graph $G$ to be generalized factor-critical, generalized bicritical, and $k$-$d$-critical. Furthermore, we prove the equivalence of the existence of four factors (namely, $\{K_2,\{C_t: t\geq 3\}\}$-factor, $\{K_2,\{C_{2t+1}:t\geq 1 \}\}$-factor, fractional perfect matching, perfect $k$-matching with even $k$) in a graph. Thus we also give size conditions and spectral radius conditions for a graph $G-v$ to have one of the four factors for any $v\in V(G)$.

math.CO

Unlocking the Power of Orbital-Free Density Functional Theory to Explore the Electronic Structure Under Extreme Conditions

Recent advances in X-ray free-electron laser diagnostics have enabled direct probing of the electronic structure under extreme pressures and temperatures, such as those encountered in stellar interiors and inertial confinement fusion experiments, challenging theoretical models for interpreting experimental data. Kohn-Sham density functional theory (KSDFT) has been successfully applied to analyze experimental X-ray scattering measurements, but its high computational cost renders routine application impractical. Orbital-free DFT (OFDFT) is a substantially more efficient alternative, with computational cost scaling linearly with system size and a weak temperature dependence, yet it often lacks the accuracy required for electronic structure description. Overcoming this limitation, we present a non-empirical Kohn-Sham-assisted orbital-free density functional framework for calculations at extreme conditions, which enables efficient OFDFT simulations with KSDFT-level accuracy for electron densities, electron-ion structure factors, and equations of state across a broad range of conditions. Benchmark comparisons with quantum Monte Carlo data for dense hydrogen and validation against Rayleigh weight measurements of hot dense beryllium demonstrate the reliability of the framework and speedups of up to several hundred times compared with KSDFT. We further show that even at temperatures on the order of 100 eV, quantum non-locality remains essential for correctly describing the electronic structure of dense hydrogen.

cond-mat.mtrl-sci

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in extended interaction settings that traditional single-turn safety evaluations fail to capture. We systematically study these interactional dynamics using a controlled LLM-to-LLM simulation framework for automated red-teaming across bilingual social engineering scenarios. Evaluating eight state-of-the-art models in English and Chinese, we analyze dialogue-level outcomes, annotate attacker and defender strategy families, and model interaction dynamics between them. Results show that multi-turn adversarial dialogues follow recurrent escalation patterns, while defensive responses frequently rely on verification, delay, and channel control. We further find statistically significant cross-model and cross-lingual differences in outcome distributions, and transition analysis reveals systematic structural variation in how defender strategies respond to attacker tactics across languages. These findings highlight the importance of studying interactional structure in multi-turn adversarial dialogue settings and demonstrate how controlled LLM-to-LLM simulations can support mechanistic analysis of adversarial conversational dynamics.

cs.CL

Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis

Vision-language models achieve expert-level performance on medical imaging tasks but exhibit significant diagnostic accuracy disparities across demographic groups. We introduce fairness-aware Low-Rank Adaptation for medical VLMs, combining parameter efficiency with explicit fairness optimization. Our key algorithmic contribution is a differentiable MaxAccGap loss that enables end-to-end optimization of accuracy parity across demographic groups. We propose three methods: FR-LoRA integrates MaxAccGap regularization into the training objective, GR-LoRA applies inverse frequency weighting to balance gradient contributions, and Hybrid-LoRA combines both mechanisms. Evaluated on 10,000 glaucoma fundus images, GR-LoRA reduces diagnostic accuracy disparities by 69% while maintaining 53.15% overall accuracy. Ablation studies reveal that strong regularization strength achieves optimal fairness with minimal accuracy trade-off, and race-specific optimization yields 60% disparity reduction. Our approach requires only 0.24% trainable parameters, enabling practical deployment of fair medical AI in resource-constrained healthcare settings.

cs.CV