arXiv ScienceSearch

arXiv subjects

Yiming Hao

Publications and source records attributed to Yiming Hao.

15 recordsLinked to original sources

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle complex spatial movements (e.g., intricate trajectories involving rotations or flips) that are essential for robot manipulation. Motivated by this gap, we adopt an implicit, data-driven approach guided by fine-grained hand-pose motions. To this end, we introduce a scalable acquisition pipeline to collect hand-gripper paired demonstrations, governed by a rigorous protocol that prioritizes motion complexity and leverages a handheld gripper for seamless action mimicry. This yields a large-scale paired dataset comprising 6,189 episodes across 1,254 unique objects, exhibiting significantly higher spatial complexity than existing benchmarks. However, learning such complex mappings remains challenging. We observe that naive end-to-end generation of full gripper pose sequences is insufficient, as minor trajectory deviations compound rapidly under intricate dynamics. To address this, we propose a two-stage framework: Stage I predicts sparse gripper keyframes (initial and terminal) to simplify the mapping objective, while Stage II generates the full continuous action sequence conditioned on these keyframes. Furthermore, to mitigate cumulative drift, we keep the gripper's orientation being learned while post-optimizing its translation based on the grasping heuristic and kinematic consistency. In both simulation and real-robot experiments, our framework enables stable and precise hand-to-gripper transfer of complex spatial manipulations, significantly outperforming traditional baselines. Project page: https://cosmoh2g.github.io.

cs.RO

A linear bound for nested cycles without geometric crossings

Cycles $C_1,\ldots,C_k$ in a graph are called nested without geometric crossings if they are pairwise edge-disjoint, $V(C_k)\subseteq\cdots\subseteq V(C_1)$, and each pair of consecutive cycles induces the same cyclic order on the vertices of the inner cycle, up to reversal. Let $f_k(n)$ be the least number of edges that forces such a family in every $n$-vertex graph. Gil Fern\'andez, Kim, Kim and Liu proved that $f_2(n)=O(n)$, answering a question of Erd\H{o}s, and asked whether $f_k(n)=O_k(n)$ for every fixed $k$. We prove this for all $k$. The proof selects the inner cycles together with a disjoint subgraph that supplies their external neighbours. A reselection argument gives disjoint paths from every inner-cycle vertex to any sufficiently large target set. Sublinear expansion and a rooted clique minor then allow the vertices to be joined in the required cyclic order.

math.CO

Bounds on Odd and Odd-Even Induced Subgraphs

Let $G$ be an $n$-vertex graph and let $\ell:V(G)\to\mathbb{F}_2$ prescribe degree parities. A set $S\subseteq V(G)$ is $\ell$-admissible if every $v\in S$ has degree congruent to $\ell(v)$ modulo $2$ in $G[S]$. Let $h_\ell(G)$ be the maximum order of an $\ell$-admissible set, set $f_{\mathrm{oe}}(G):=\min_\ell h_\ell(G)$, and write $f_o(G):=h_{\mathbf{1}}(G)$, where $\mathbf{1}(v)=1$ for every $v\in V(G).$ We prove three main results for graphs without isolated vertices. First, by extending Zeng's odd-cut method to arbitrary parity prescriptions an introducing a one-sided completion lemma, we show that $h_\ell(G)\ge n/6$ for every $\ell$. Consequently, $f_{\mathrm{oe}}(G)\ge n/6$, improving the previous bound $2n/21$. Second, for bipartite graphs we derive lower bounds on $f_o(G)$ in terms of the $\mathbb{F}_2$-rank of the bipartite adjacency matrix and combine them to obtain \[ f_o(G)\ge \left(\frac14+\frac1{256}\right)n=\frac{65}{256}n. \] Thus, in the bipartite case, the factor $2$ in Scott's bound $f_o(G)\ge n/(2\chi(G))$ can be replaced by $128/65<2$. Finally, writing $\alpha=\alpha(G)$, a fourth-moment argument gives, for $\alpha\ge2$, \[ f_o(G)\ge \frac{\alpha}{2}+\frac{\log_3\alpha}{8} -\frac14\log_3\log_3\sqrt{\alpha}. \] We also construct bipartite graphs satisfying \[ f_o(G)\le \frac{\alpha(G)}2+\log_2\!\bigl(\alpha(G)+1\bigr)+\frac12, \] showing that the logarithmic additive improvement over Scott's bound $f_o(G)\ge\alpha(G)/2$ has the optimal order of magnitude.

math.CO

Weighted Counting Formula and $2n/21$ Lower Bound for Induced Subgraphs with Prescribed Degree Parities

Let $G=(V,E)$ be a finite simple graph of order $n\geq 1$, and let $\ell:V\to\{0,1\}$ be a prescribed parity labeling. A set $S\subseteq V$ is called $\ell$-admissible if $d_S(v)\equiv \ell(v)\pmod 2$ for every $v\in S$, where $d_S(v)=|N_G(v)\cap S|$. Let $h_\ell(G)$ be the maximum order of an $\ell$-admissible set and let $f_{\rm oe}(G)=\min_\ell h_\ell(G)$. For $x\in\mathbb R$, define the weighted counting polynomial $$ M_{\ell,x}(G)=\sum_{S\in {\cal A}_\ell(G)}x^{|S|}, $$ where ${\cal A}_\ell(G)$ is the collection of all $\ell$-admissible sets in $G$. For $R\subseteq V$, let $z_\ell(R)$ be the number of vertices $v\in V\setminus R$ for which $d_R(v)\equiv\ell(v)\pmod 2$. We prove the exact identity $$ M_{\ell,x}(G) =2^{-n}\sum_{R\subseteq V} x^{|R|}(2+x)^{z_\ell(R)}(2-x)^{n-z_\ell(R)-|R|}. $$ If $G$ has no isolated vertices, then, for every $\ell$ and every $x\in(0,2)$, $ M_{\ell,x}(G)>x^{n/2}(4-x^2)^{n/4}. $ Combining this estimate with a binary-entropy upper bound and optimizing $x$ gives $$ f_{\rm oe}(G)>c_*n>\frac{2n}{21}, $$ where $c_*\approx0.095862615$. Ferber and Krivelevich (Adv. Math. 2022) proved that $h_{\mathbf{1}}(G)\ge 10^{-4}n$, where $\mathbf{1}$ is the all-one labeling. Since $h_{\mathbf{1}}(G)\ge f_{\rm oe}(G)$, our result improves coefficient in their bound by almost three orders of magnitude, and does so simultaneously for every labeling.

math.CO

Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation

Recent multimodal large language models have achieved strong performance in unified text and image understanding and generation, yet extending such native capability to 3D remains challenging due to limited data. Compared to abundant 2D imagery, high-quality 3D assets are scarce, making 3D synthesis under-constrained. Existing methods often rely on indirect pipelines that edit in 2D and lift results into 3D via optimization, sacrificing geometric consistency. We present Omni123, a 3D-native foundation model that unifies text-to-2D and text-to-3D generation within a single autoregressive framework. Our key insight is that cross-modal consistency between images and 3D can serve as an implicit structural constraint. By representing text, images, and 3D as discrete tokens in a shared sequence space, the model leverages abundant 2D data as a geometric prior to improve 3D representations. We introduce an interleaved X-to-X training paradigm that coordinates diverse cross-modal tasks over heterogeneous paired datasets without requiring fully aligned text-image-3D triplets. By traversing semantic-visual-geometric cycles (e.g., text to image to 3D to image) within autoregressive sequences, the model jointly enforces semantic alignment, appearance fidelity, and multi-view geometric consistency. Experiments show that Omni123 significantly improves text-guided 3D generation and editing, demonstrating a scalable path toward multimodal 3D world models.

cs.CV

LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models

Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to their demand for task-specific data and lengthy optimization. While a few hypernetwork-based approaches attempt to predict adaptation weights directly, they struggle to map fine-grained user prompts to complex LoRA distributions, limiting their practical applicability. To bridge this gap, we propose LoFA, a general framework that efficiently predicts personalized priors for fast model adaptation. We first identify a key property of LoRA: structured distribution patterns emerge in the relative changes between LoRA and base model parameters. Building on this, we design a two-stage hypernetwork: first predicting relative distribution patterns that capture key adaptation regions, then using these to guide final LoRA weight prediction. Extensive experiments demonstrate that our method consistently predicts high-quality personalized priors within seconds, across multiple tasks and user prompts, even outperforming conventional LoRA that requires hours of processing. Project page: https://jaeger416.github.io/lofa/.

cs.CV

Odd Induced Subgraphs in Graphs of Maximum Degree Four

A graph is called odd if all of its vertex degrees are odd. A long-standing conjecture asked whether there exists a positive constant $c$ such that every $n$-vertex graph without isolated vertices contains an odd induced subgraph on at least $cn$ vertices. In 2022, Ferber and Krivelevich resolved this conjecture affirmatively with $c=10^{-4}$. A natural question is to determine the largest possible constant $c$. In 1994, Caro remarked that if $2/7$ is a valid value for $c$, then it is the largest possible one. To the best of our knowledge, the bound $c\ge 2/7$ has not been improved. Previous research has established tight bounds for specific graph classes -- for instance, $c = 2/5$ for graphs with maximum degree at most $3$ and without isolated vertices. In this paper, we prove that $c=2/7$ is the tight bound for graphs with maximum degree at most $4$ and without isolated vertices. Our result provides some support for $2/7$ being the largest value of $c$.

math.CO

VC-Agent: An Interactive Agent for Customized Video Dataset Collection

Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the way to expedite this collection process and propose VC-Agent, the first interactive agent that is able to understand users' queries and feedback, and accordingly retrieve/scale up relevant video clips with minimal user input. Specifically, considering the user interface, our agent defines various user-friendly ways for the user to specify requirements based on textual descriptions and confirmations. As for agent functions, we leverage existing multi-modal large language models to connect the user's requirements with the video content. More importantly, we propose two novel filtering policies that can be updated when user interaction is continually performed. Finally, we provide a new benchmark for personalized video dataset collection, and carefully conduct the user study to verify our agent's usage in various real scenarios. Extensive experiments demonstrate the effectiveness and efficiency of our agent for customized video dataset collection. Project page: https://allenyidan.github.io/vcagent_page/.

cs.AI

Large induced subgraphs with prescribed degree parity

A long-standing conjecture of Caro (Discrete Math, 1994), confirmed by Ferber and Krivelevich (Adv Math, 2022), states that every $n$-vertex graph $G$ without isolated vertices contains an induced subgraph of order linear in $n$ in which every vertex has odd degree. We generalize this result to graphs $G$ whose vertices are labeled by $\ell: V(G)\to \{0,1\}$. We require, in an induced subgraph, all $0$-labeled vertices to have even degree and all $1$-labeled vertices to have odd degree. Let $h_{\ell}(G)$ denote the maximum order of such a subgraph. Let $f_{oe}(G)=\min_{\ell} h_{\ell}(G)$ be the worst-labeling parameter. We establish a pointwise lower bound for $h_{\ell}(G)$ that immediately yields a linear lower bound in $|V(G)|$ for $f_{oe}(G)$, where $G$ has no isolated vertices. For an $n$-vertex connected graph, we obtain a sharp lower bound for $f_{oe}(G)$: $f_{oe}(G)\ge \lceil (n-1)/{\chi}_{mm}{(G)} \rceil ,$ where ${\chi}_{mm}{(G)}$ is the maximum chromatic number of a minor of $G.$ Using proved cases of Hadwiger's Conjecture, we show that for $t\in \{3,4,5,6\}$, if an $n$-vertex connected graph $G$ is $K_t$-minor-free, then $f_{oe}(G)\ge \lceil (n-1)/(t-1)\rceil$ and this bound is sharp for each $t\in \{3,4,5,6\}$. Finally, we conjecture that $f_{oe}(G)\ge f_o(G)/2$ for all graphs $G$ and confirm the conjecture for all trees and complete multipartite graphs.

math.CO

IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects

Estimating albedo (a.k.a., intrinsic image decomposition) from single RGB images captured in real-world environments (e.g., the MVImgNet dataset) presents a significant challenge due to the absence of paired images and their ground truth albedos. Therefore, while recent methods (e.g., IntrinsicAnything) have achieved breakthroughs by harnessing powerful diffusion priors, they remain predominantly trained on large-scale synthetic datasets (e.g., Objaverse) and applied directly to real-world RGB images, which ignores the large domain gap between synthetic and real-world data and leads to suboptimal generalization performance. In this work, we address this gap by proposing IntrinsicReal, a novel domain adaptation framework that bridges the above-mentioned domain gap for real-world intrinsic image decomposition. Specifically, our IntrinsicReal adapts IntrinsicAnything to the real domain by fine-tuning it using its high-quality output albedos selected by a novel dual pseudo-labeling strategy: i) pseudo-labeling with an absolute confidence threshold on classifier predictions, and ii) pseudo-labeling using the relative preference ranking of classifier predictions for individual input objects. This strategy is inspired by human evaluation, where identifying the highest-quality outputs is straightforward, but absolute scores become less reliable for sub-optimal cases. In these situations, relative comparisons of outputs become more accurate. To implement this, we propose a novel two-phase pipeline that sequentially applies these pseudo-labeling techniques to effectively adapt IntrinsicAnything to the real domain. Experimental results show that our IntrinsicReal significantly outperforms existing methods, achieving state-of-the-art results for albedo estimation on both synthetic and real-world datasets.

cs.GR

TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation

We address key limitations in existing datasets and models for task-oriented hand-object interaction video generation, a critical approach of generating video demonstrations for robotic imitation learning. Current datasets, such as Ego4D, often suffer from inconsistent view perspectives and misaligned interactions, leading to reduced video quality and limiting their applicability for precise imitation learning tasks. Towards this end, we introduce TASTE-Rob -- a pioneering large-scale dataset of 100,856 ego-centric hand-object interaction videos. Each video is meticulously aligned with language instructions and recorded from a consistent camera viewpoint to ensure interaction clarity. By fine-tuning a Video Diffusion Model (VDM) on TASTE-Rob, we achieve realistic object interactions, though we observed occasional inconsistencies in hand grasping postures. To enhance realism, we introduce a three-stage pose-refinement pipeline that improves hand posture accuracy in generated videos. Our curated dataset, coupled with the specialized pose-refinement framework, provides notable performance gains in generating high-quality, task-oriented hand-object interaction videos, resulting in achieving superior generalizable robotic manipulation. The TASTE-Rob dataset is publicly available to foster further advancements in the field, TASTE-Rob dataset and source code will be made publicly available on our website https://taste-rob.github.io.

cs.CV

Arc-disjoint in- and out-branchings in semicomplete split digraphs

An \emph{out-tree (in-tree)} is an oriented tree where every vertex except one, called the \emph{root}, has in-degree (out-degree) one. An \emph{out-branching $B^+_u$ (in-branching $B^-_u$)} of a digraph $D$ is a spanning out-tree (in-tree) rooted at $u$. A \emph{good $(u,v)$-pair} in $D$ is a pair of branchings $B^+_u, B^-_v$ which are arc-disjoint. Thomassen proved that deciding whether a digraph has any good pair is NP-complete. A \emph{semicomplete split digraph} is a digraph where the vertex set is the disjoint union of two non-empty sets, $V_1$ and $V_2$, such that $V_1$ is an independent set, the subdigraph induced by $V_2$ is semicomplete, and every vertex in $V_1$ is adjacent to every vertex in $V_2$. In this paper, we prove that every $2$-arc-strong semicomplete split digraph $D$ contains a good $(u, v)$-pair for any choice of vertices $u, v$ of $D$, thereby confirming a conjecture by Bang-Jensen and Wang [Bang-Jensen and Wang, J. Graph Theory, 2024].

math.CO

Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models

Optimizing a text-to-image diffusion model with a given reward function is an important but underexplored research area. In this study, we propose Deep Reward Tuning (DRTune), an algorithm that directly supervises the final output image of a text-to-image diffusion model and back-propagates through the iterative sampling process to the input noise. We find that training earlier steps in the sampling process is crucial for low-level rewards, and deep supervision can be achieved efficiently and effectively by stopping the gradient of the denoising network input. DRTune is extensively evaluated on various reward models. It consistently outperforms other algorithms, particularly for low-level control signals, where all shallow supervision methods fail. Additionally, we fine-tune Stable Diffusion XL 1.0 (SDXL 1.0) model via DRTune to optimize Human Preference Score v2.1, resulting in the Favorable Diffusion XL 1.0 (FDXL 1.0) model. FDXL 1.0 significantly enhances image quality compared to SDXL 1.0 and reaches comparable quality compared with Midjourney v5.2.

cs.CV

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Recent text-to-image generative models can generate high-fidelity images from text inputs, but the quality of these generated images cannot be accurately evaluated by existing evaluation metrics. To address this issue, we introduce Human Preference Dataset v2 (HPD v2), a large-scale dataset that captures human preferences on images from a wide range of sources. HPD v2 comprises 798,090 human preference choices on 433,760 pairs of images, making it the largest dataset of its kind. The text prompts and images are deliberately collected to eliminate potential bias, which is a common issue in previous datasets. By fine-tuning CLIP on HPD v2, we obtain Human Preference Score v2 (HPS v2), a scoring model that can more accurately predict human preferences on generated images. Our experiments demonstrate that HPS v2 generalizes better than previous metrics across various image distributions and is responsive to algorithmic improvements of text-to-image generative models, making it a preferable evaluation metric for these models. We also investigate the design of the evaluation prompts for text-to-image generative models, to make the evaluation stable, fair and easy-to-use. Finally, we establish a benchmark for text-to-image generative models using HPS v2, which includes a set of recent text-to-image models from the academic, community and industry. The code and dataset is available at https://github.com/tgxs002/HPSv2 .

cs.CV

Perception Imitation: Towards Synthesis-free Simulator for Autonomous Vehicles

We propose a perception imitation method to simulate results of a certain perception model, and discuss a new heuristic route of autonomous driving simulator without data synthesis. The motivation is that original sensor data is not always necessary for tasks such as planning and control when semantic perception results are ready, so that simulating perception directly is more economic and efficient. In this work, a series of evaluation methods such as matching metric and performance of downstream task are exploited to examine the simulation quality. Experiments show that our method is effective to model the behavior of learning-based perception model, and can be further applied in the proposed simulation route smoothly.

cs.RO