arXiv ScienceSearch

arXiv subjects

Tianhao Wu

Publications and source records attributed to Tianhao Wu.

At least 19 recordsLinked to original sources

Scenario MPC with STL Specifications and Pareto-Based Feasibility Repair

Temporal logic is a formal language for reasoning about system behaviors over time. Signal temporal logic (STL), in particular, has been used to encode spatio-temporal requirements for control synthesis in multi-agent systems, often under the assumption that agents are cooperative and their dynamics are known. However, real-world multi-agent applications, such as autonomous driving, typically involve stochastic and uncontrollable agents. Recent work explored robust control with worst-case or probabilistic formulations, but remains limited in that it either (1) certifies strict satisfaction of STL constraints without addressing feasibility recovery, or (2) relaxes infeasible constraints with ego-centric objectives. In this paper, we propose a model predictive control (MPC) framework that treats feasibility repair as a Pareto optimization problem to explicitly characterize tradeoffs among agent objectives. We further provide a probabilistic certificate on STL violation rate to formally quantify uncertainty under stochastic and uncontrollable agents. The proposed framework is evaluated on two autonomous driving scenarios. Results show that the framework recovers feasible control with demonstrated safe behaviors.

cs.RO

STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleoperation system to collect a 200-hour bimanual dexterous manipulation dataset with synchronized visual, tactile, and language annotations, comprising 10,576 trajectories across 65 tasks, 69.5% of which involve dexterous multi-finger manipulation. We further propose STAR, an integrated training recipe for vision-tactile-language-action (VTLA) models that addresses the spatial, temporal, and informational sparsity of tactile signals through visual-tactile joint pre-training, sparse-global tactile token representation, and sparse future tactile prediction. Trained on this dataset, STAR achieves a 61% average success rate across four real-world tasks with 100 post-training trajectories per task, demonstrating dexterous performance under task-specific post-training.

cs.RO

Exact Black Branes in General Dimensions in Higher-Curvature Scalar-Tensor Gravity

We obtain two families of neutral planar black branes in a potential-free dilaton-coupled Lovelock--Horndeski theory. Heterotic compactification relates higher-dimensional curvature to scalar dynamics and motivates the interaction structure. At fixed scalar zero mode, the horizon radius and temperature vary while the complete stationary Wald entropy remains constant and positive. The exponential scalar factor at the horizon exactly compensates the area growth, giving this behavior a direct geometric explanation. An exact coefficient analysis determines the two four-dimensional branches and their continuation to every higher dimension through closed coupling relations. The linear family is conformally flat and leaves the AdS scale free at fixed couplings; the fractional family is Weyl curved and fixes that scale. The common kinetic relation also excludes their realization through a real one-scalar flat-torus reduction of a metric--dilaton parent.

hep-th

Exact Hairy Black Holes in Higher-Curvature Scalar--Tensor Gravity

We construct exact hairy black holes in Lovelock--Horndeski gravity and establish a relation between horizon topology, scalar hair and covariant charges. Rooted in the dimensional reduction of higher-curvature gravity, this framework connects scalar--tensor dynamics to the low-energy description of heterotic strings. In five dimensions, a single fixed theory admits a continuous planar family and two isolated hyperbolic black holes. The hyperbolic solutions share a locally AdS metric with distinct scalar dressings, while the planar family is Weyl curved and has a freely varying horizon scale. Both branches carry scalar hair regular on the future horizon. The hyperbolic construction extends above four dimensions; compatibility of the displayed planar profile selects five and seven dimensions. Time translations of the scalar act through a global shift, which enters the complete covariant charge balance. Along the planar family, a changing bulk shift-charge density supplies this balance despite zero radial flux and vanishing horizon charge variation. The hyperbolic roots admit no horizon-scale variation at fixed couplings.

gr-qc

Holographic Renormalization of String-Derived Lovelock--Horndeski Theory

String-derived higher-curvature scalar--tensor gravities encode microscopic coupling data in boundary response, raising the question of whether holographic observables can reconstruct the underlying higher-dimensional parameters. We answer this question for the five-dimensional string-derived Lovelock--Horndeski (SDLH) theory on its exact linear-dilaton asymptotically locally AdS branch, constructing the renormalized generating functional for an arbitrary boundary metric and spacetime-dependent scalar source. A boundary-covariant radial hierarchy unifies the variational problem, local backreaction, logarithmic obstruction, finite one-point functions, and Ward identities. Two response determinants organize the recursion, resonant obstructions, and metric--scalar mixing. The Weyl anomaly condenses into an Euler density, a Weyl-squared density, and a single curvature--scalar square whose paired variations generate the metric and scalar obstructions. The resulting renormalized functional carries string-selected coupling data into boundary geometry, operator response, anomaly coefficients, and a calculable interface with gravitational observables. On the regular branch, four scalar-normalization-invariant holographic combinations admit a global rational inverse to the continuous reduced couplings. At fixed compactification dimension the map has maximal rank, while the curvature-anomaly sum reconstructs the higher-dimensional Gauss--Bonnet coefficient without sign ambiguity. Holographic response thus provides an explicit, overdetermined boundary fingerprint of the underlying string reduction.

hep-th

SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation

American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.

cs.CV

Sharp Plucker Geometry for Three-Copy Werner Distillation

Whether negative-partial-transpose entanglement can remain undistillable is a longstanding problem in quantum information theory. We analyze the first unresolved three-copy Werner endpoint using a sharp dimension-free inequality for complementary partial traces and an optimal exterior-square inequality for orthonormal tripartite vectors. The latter identifies local SWAP- parity statistics with metric data of a decomposable Plucker bivector. Together these inequalities prove endpoint nonnegativity for every positive semidefinite rank-two coefficient operator and for the complete normal rank-two sector in arbitrary finite local dimensions. For genuinely nonnormal operators, an exact crossed-Gram criterion proves nonnegativity when one local outpu-input support overlap is at most two, including every system with a qubit-sized factor, and when either support plane contains a product ray. Explicit anti-state and rank-boundary families establish optimality of the constants and the rank restriction.

quant-ph

Autodata: An agentic data scientist to create high quality synthetic data

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.

cs.AI

An Algebraic Framework for Quantitative Semantics of Spatio-Temporal Logic with Graph Operators

Spatio-Temporal Logic with Graph Operators (STL-GO) extends Signal Temporal Logic (STL) to multi-agent systems via graph operators that count neighboring agents satisfying a property, together with multi-agent quantifiers. While Boolean semantics for STL-GO are well-defined, quantitative semantics have not yet been developed and existing quantitative semantics for spatio-temporal logics such as STREL cannot capture the counting constraints in STL-GO's graph operators. We develop quantitative semantics for STL-GO as a layered algebraic construction that separates temporal aggregation from graph-operator aggregation (governed by an abstract accumulator with a monotone fold and readout). We prove that soundness and completeness reduce to monotonicity conditions on these components. We implement the framework and evaluate it on two multi-agent environments: a 2D bounded region with stochastic Dubins-car dynamics and a 3D Earth-satellite system, under four semantic instantiations (Boolean, min-max, signed-deficit, and a hybrid), demonstrating the tradeoffs between accumulator choices and reporting scalability in the number of agents and time horizon.

cs.LO

Lyapunov-Based Sample Complexity Analysis for Weakly-Coupled MDPs

We study the sample complexity of learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits (RBs) under a generative model. Naive reduction to a tabular MDP leads to high complexity bounds as the state-action space is exponentially large in the number of arms $N$. By exploiting the weakly coupled structure, we show that near-optimal policies can be learned with sample and computational complexities that are polynomial in $N$. Specifically, we analyze the plug-in approach, which applies an efficient planning algorithm to an empirical model estimated from data. For fully heterogeneous WCMDPs, we establish the first finite-sample PAC guarantee with polynomial complexity and an $O(1/\sqrt{N})$ optimality gap. For homogeneous RBs, we further prove that a smaller optimality gap is achievable under mild structural assumptions. A primary technical contribution of our work is a novel Lyapunov-based analysis framework. Unlike classical approaches that rely on the difficult-to-control bias function, our framework uses an explicitly constructed Lyapunov function along with a drift transfer technique between the true and empirical models. A key step of independent interest in our framework is a fine-grained perturbation analysis for the underlying linear programming (LP) relaxation, which provides a general tool for analyzing LP-based policies and weakly-coupled systems.

cs.LG

CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation

Urban trajectory generation is a fundamental task for transportation simulation, urban planning, and mobility analytics. However, systematic comparison across trajectory generation methods remains difficult because existing studies often rely on different datasets, preprocessing pipelines, trajectory representations, and evaluation metrics. This fragmentation makes it unclear whether reported performance differences arise from the generation mechanism itself or from inconsistent experimental protocols. To address this issue, we present CityTrajBench, a unified benchmark framework and protocol for city-scale vehicle trajectory generation. CityTrajBench standardizes data ingestion, trajectory normalization, feature construction, model adaptation, map-aware post-processing, model selection, and multi-level evaluation under a common setting. It supports heterogeneous generators, including statistical baselines, VAE-based, GAN-based, diffusion-based, and flow-matching-based models, and evaluates them on three real-world urban trajectory datasets. The benchmark measures global spatial realism, trip-level distribution fidelity, trajectory-level geometric similarity, conditional mobility consistency, and efficiency. Experiments reveal clear trade-offs across model families: DiffTraj is strongest on trajectory-level geometric fidelity, DiffRNTraj is competitive on structure-sensitive global realism, and TrajFlow provides a strong balance across realism, quality, conditional consistency, and efficiency. Meanwhile, a simple Markov baseline remains competitive on coarse-grained trip and local-movement statistics. These findings show that urban trajectory generation quality is inherently multi-objective, that no single model dominates all criteria equally, and that CityTrajBench provides a reproducible benchmark protocol and testbed for future research on urban mobility generation.

cs.LG

WOW-Seg: A Word-free Open World Segmentation Model

Open world image segmentation aims to achieve precise segmentation and semantic understanding of targets within images by addressing the infinitely open set of object categories encountered in the real world. However, traditional closed-set segmentation approaches struggle to adapt to complex open world scenarios, while foundation segmentation models such as SAM exhibit notable discrepancies between their strong segmentation capabilities and relatively weaker semantic understanding. To bridge these discrepancies, we propose WOW-Seg, a Word-free Open World Segmentation model for segmenting and recognizing objects from open-set categories. Specifically, WOW-Seg introduces a novel visual prompt module, Mask2Token, which transforms image masks into visual tokens and ensures their alignment with the VLLM feature space. Moreover, we introduce the Cascade Attention Mask to decouple information across different instances. This approach mitigates inter-instance interference, leading to a significant improvement in model performance. We further construct an open world region recognition test benchmark: the Region Recognition Dataset (RR-7K). With 7,662 classes, it represents the most extensive category-rich region recognition dataset to date. WOW-Seg attains strong results on the LVIS dataset, achieving a semantic similarity of 89.7 and a semantic IoU of 82.4. This performance surpasses the previous SOTA while using only one-eighth the parameter count. These results underscore the strong open world generalization capabilities of WOW-Seg. The code and related resources are available at https://github.com/AAwcAA/WOW-Seg-Meta.

cs.CV

The Impact of Spin Priors on Parameterized Tests of General Relativity

Spin priors play a fundamental role in gravitational-wave parameter estimation, yet their impact on parameterized tests of General Relativity (GR) remains insufficiently understood. In this work, we systematically investigate how spin prior choices affect the 1.5PN deviation parameter $δ\hatϕ_3$ using real gravitational-wave events. We quantify prior-induced effects through the Jensen--Shannon divergence (JSD) and median shifts of posterior distributions. We find that the effective precession spin parameter $χ_p$ exhibits significantly stronger prior sensitivity than the effective inspiral spin $χ_{\rm eff}$. While $δ\hatϕ_3$ is generally robust across most events, GW231123\_135430 exhibits a noticeable discrepancy, with a JSD at the $\mathcal{O}(0.4)$ level. Examining the median shift, we note that events with very short inspiral durations, such as GW231028\_153006, GW231123\_135430, and GW191109\_010717, show more pronounced shifts, indicating increased sensitivity in low-information regimes. We further explore the relationship between the prior sensitivity of spin parameters and that of $δ\hatϕ_3$. No significant correlation is observed when spin parameters are inferred within the standard GR framework. However, when $δ\hatϕ_3$ is included in the analysis, a strong correlation emerges between $χ_{\rm eff}$ and $δ\hatϕ_3$, which we attribute to partial parameter degeneracy at the 1.5PN order. A leave-one-out test shows that the observed correlation is sensitive to the inclusion of specific events, indicating that it is partially driven by a subset of high-sensitivity events. Our results demonstrate that spin-prior choices can propagate into parameterized tests of GR in a non-trivial and model-dependent manner, and may mimic or reshape apparent deviations from GR.

gr-qc

Feasibility Restoration under Conflicting STL Specifications with Pareto-Optimal Refinement

Signal Temporal Logic (STL) is expressive formal language that specifies spatio-temporal requirements in robotics. Its quantitative robustness semantics can be easily integrated with optimization-based control frameworks. However, STL specifications may become conflicting in real-world applications, where safety rules, traffic regulations, and task objectives can be cannot be satisfied together. In these situations, traditional STL-constrained Model Predictive Control (MPC) becomes infeasible and default to conservative behaviors such as freezing, which can largely increase risks in safety-critical scenarios. In this paper, we proposes a unified two-stage framework that first restores feasibility via minimal relaxation, then refine the feasible solution by formulating it as a value-aware multi-objective optimization problem. Using $\varepsilon$-constraint method, we approximate the Pareto front of the multi-objective optimization, which allows analysis of tradeoffs among competing objectives and counterfactual analysis of alternative actions. We demonstrate that the proposed approach avoids deadlock under conflicting STL specifications and enables interpretable decision-making in safety-critical applications by conducting a case study in autonomous driving.

cs.RO

A unified foundational framework for knowledge injection and evaluation of Large Language Models in Combustion Science

To advance foundation Large Language Models (LLMs) for combustion science, this study presents the first end-to-end framework for developing domain-specialized models for the combustion community. The framework comprises an AI-ready multimodal knowledge base at the 3.5 billion-token scale, extracted from over 200,000 peer-reviewed articles, 8,000 theses and dissertations, and approximately 400,000 lines of combustion CFD code; a rigorous and largely automated evaluation benchmark (CombustionQA, 436 questions across eight subfields); and a three-stage knowledge-injection pathway that progresses from lightweight retrieval-augmented generation (RAG) to knowledge-graph-enhanced retrieval and continued pretraining. We first quantitatively validate Stage 1 (naive RAG) and find a hard ceiling: standard RAG accuracy peaks at 60%, far surpassing zero-shot performance (23%) yet well below the theoretical upper bound (87%). We further demonstrate that this stage's performance is severely constrained by context contamination. Consequently, building a domain foundation model requires structured knowledge graphs and continued pretraining (Stages 2 and 3).

cs.CL

Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization

Large language models (LLMs) in the direction of task adaptation and capability enhancement for professional fields demonstrate significant application potential. Nevertheless, for complex physical systems such as combustion science, general-purpose LLMs often generate severe hallucinations due to insufficient domain knowledge and the inability to adhere to physical conservation laws. To address this issue, we propose the first full-stack domain-enhanced LLM workflow tailored for the field of combustion science, which integrates automated domain corpus construction, incremental pre-training, instruction fine-tuning, and verifiable reward-based reinforcement learning. This workflow ensures that the model truly internalizes physical laws rather than merely learning textual statistical patterns. We also release FlameBench, a standardized evaluation benchmark specifically designed for complex reasoning tasks in combustion science. Experimental results demonstrate that the model developed in this work significantly outperforms state-of-the-art general-purpose closed-source models and traditional retrieval-augmented generation methods on combustion science reasoning tasks. This work lays a solid technical and resource foundation for the subsequent development of domain-specific scientific research agents with reliable scientific reasoning capabilities.

cs.CL

PoseCraft: Tokenized 3D Body Landmark and Camera Conditioning for Photorealistic Human Image Synthesis

Digitizing humans and synthesizing photorealistic avatars with explicit 3D pose and camera controls are central to VR, telepresence, and entertainment. Existing skinning-based workflows require laborious manual rigging or template-based fittings, while neural volumetric methods rely on canonical templates and re-optimization for each unseen pose. We present PoseCraft, a diffusion framework built around tokenized 3D interface: instead of relying only on rasterized geometry as 2D control images, we encode sparse 3D landmarks and camera extrinsics as discrete conditioning tokens and inject them into diffusion via cross-attention. Our approach preserves 3D semantics by avoiding 2D re-projection ambiguity under large pose and viewpoint changes, and produces photorealistic imagery that faithfully captures identity and appearance. To train and evaluate at scale, we also implement GenHumanRF, a data generation workflow that renders diverse supervision from volumetric reconstructions. Our experiments show that PoseCraft achieves significant perceptual quality improvement over diffusion-centric methods, and attains better or comparable metrics to latest volumetric rendering SOTA while better preserving fabric and hair details.

cs.CV

Boosting high-current alkaline water electrolysis and carbon dioxide reduction with novel CuNiFe-based anodes

The transition to a green hydrogen economy demands robust, scalable, and sustainable anodes for alkaline water electrolysis operating at industrial current densities (>1 A/cm2). However, achieving high activity and long-term stability under such conditions remains a formidable challenge with conventional catalysts. Here, we report a novel trimetallic CuNiFe anode fabricated through a rapid, single-step electrodeposition process at room temperature without organic additives. The catalyst exhibits an exceptionally low overpotential of <270 mV at 100 mA cm(-2) and operates stably for over 500 hours at 1 A cm(-2) in 30 wt% KOH. In a practical anion exchange membrane water electrolyzer (AEM-WE), the CuNiFe anode enables a current density of 2.5 A cm(-2) at only 2.5 V, with a voltage efficiency of 66.8%. Beyond water splitting, this anode also significantly enhances CO2 electrolysis, tripling the CO2 reduction current density and steering selectivity toward valuable multi-carbon products when paired with commercial copper cathodes. A cradle-to-gate life cycle assessment confirms that the CuNiFe anode reduces the carbon footprint by an order of magnitude and decreases environmental impacts by 40-60% across multiple categories compared to benchmark IrRuO2. Our work establishes a scalable, high-performance, and environmentally benign anode technology, paving the way for cost-effective electrochemical production of green hydrogen and carbon-neutral chemicals.

cond-mat.mtrl-sci