arXiv ScienceSearch

arXiv subjects

Zhong Ren

Publications and source records attributed to Zhong Ren.

10 recordsLinked to original sources

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

We propose ILP-CoT, a method that bridges Inductive Logic Programming (ILP) and Multimodal Large Language Models (MLLMs) for abductive logical rule induction. The task involves both discovering logical facts and inducing logical rules from a small number of unstructured textual or visual inputs, which still remain challenging when solely relying on ILP, due to the requirement of specified background knowledge and high computational cost, or MLLMs, due to the appearance of perceptual hallucinations. Based on the key observation that MLLMs could propose structure-correct rules even under hallucinations, our approach automatically builds ILP tasks with pruned search spaces based on the rule structure proposals from MLLMs, and utilizes ILP system to output rules built upon rectified logical facts and formal inductive reasoning. Its effectiveness is verified through challenging logical induction benchmarks, as well as a potential application of our approach, namely text-to-image customized generation with rule induction. Our code and data are released at https://github.com/future-item/ILP-CoT.

cs.LG

ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation

We present ROVI, a high-quality synthetic dataset for instance-grounded text-to-image generation, created by labeling 1M curated web images. Our key innovation is a strategy called re-captioning, focusing on the pre-detection stage, where a VLM (Vision-Language Model) generates comprehensive visual descriptions that are then processed by an LLM (Large Language Model) to extract a flat list of potential categories for OVDs (Open-Vocabulary Detectors) to detect. This approach yields a global prompt inherently linked to instance annotations while capturing secondary visual elements humans typically overlook. Evaluations show that ROVI exceeds existing detection datasets in image quality and resolution while containing two orders of magnitude more categories with an open-vocabulary nature. For demonstrative purposes, a text-to-image model GLIGEN trained on ROVI significantly outperforms state-of-the-art alternatives in instance grounding accuracy, prompt fidelity, and aesthetic quality. Our dataset and reproducible pipeline are available at https://github.com/CihangPeng/ROVI.

cs.CV

Shape Adaptation for 3D Hairstyle Retargeting

It is demanding to author an existing hairstyle for novel characters in games and VR applications. However, it is a non-trivial task for artists due to the complicated hair geometries and spatial interactions to preserve. In this paper, we present an automatic shape adaptation method to retarget 3D hairstyles. We formulate the adaptation process as a constrained optimization problem, where all the shape properties and spatial relationships are converted into individual objectives and constraints. To make such an optimization on high-resolution hairstyles tractable, we adopt a multi-scale strategy to compute the target positions of the hair strands in a coarse-to-fine manner. The global solving for the inter-strands coupling is restricted to the coarse level, and the solving for fine details is made local and parallel. In addition, we present a novel hairline edit tool to allow for user customization during retargeting. We achieve it by solving physics-based deformations of an embedded membrane to redistribute the hair roots with minimal distortion. We demonstrate the efficacy of our method through quantitative and qualitative experiments on various hairstyles and characters.

cs.GR

Conformally-coated Nickel-Carbon Nitride on Structured Pyrolytic Carbon Electrodes for Enhanced Hydrogen Evolution

This study presents a three-dimensional electrode for hydrogen evolution reaction (HER), tackling challenges in corrosion resistance, catalytic activity, and durability. The electrode features a rod-connected diamond structure made of pyrolytic carbon (PyC), with sequential Ni and CN conformal coatings, resulting in an increase of its electrochemical active surface area more than 80 times their geometric surface area, marking the first application of CN/Ni/RCD-PyC electrodes for HER. The CN layer enhances corrosion resistance under acidic and alkaline conditions, while synergizing with Ni boosts catalytic activity. We demonstrate that the CN/Ni/RCD-PyC electrode exhibited superior HER activity, achieving lower overpotentials of 0.6 V and 0.4 V for a current density of $\mathrm{10~mA\cdot cm^{-2}}$ in acidic and alkaline media, respectively. The electrode demonstrates excellent durability with performance over 2000 cycles of cyclic voltammetry at high scan rates with no significant decay. Therefore, this electrode offers a scalable, cost-effective solution for sustainable hydrogen production.

physics.chem-ph

QPoser: Quantized Explicit Pose Prior Modeling for Controllable Pose Generation

Explicit pose prior models compress human poses into latent representations for using in pose-related downstream tasks. A desirable explicit pose prior model should satisfy three desirable abilities: 1) correctness, i.e. ensuring to generate physically possible poses; 2) expressiveness, i.e. ensuring to preserve details in generation; 3) controllability, meaning that generation from reference poses and explicit instructions should be convenient. Existing explicit pose prior models fail to achieve all of three properties, in special controllability. To break this situation, we propose QPoser, a highly controllable explicit pose prior model which guarantees correctness and expressiveness. In QPoser, a multi-head vector quantized autoencoder (MS-VQVAE) is proposed for obtaining expressive and distributed pose representations. Furthermore, a global-local feature integration mechanism (GLIF-AE) is utilized to disentangle the latent representation and integrate full-body information into local-joint features. Experimental results show that QPoser significantly outperforms state-of-the-art approaches in representing expressive and correct poses, meanwhile is easily to be used for detailed conditional generation from reference poses and prompting instructions.

cs.CV

Generating by Understanding: Neural Visual Generation with Logical Symbol Groundings

Making neural visual generative models controllable by logical reasoning systems is promising for improving faithfulness, transparency, and generalizability. We propose the Abductive visual Generation (AbdGen) approach to build such logic-integrated models. A vector-quantized symbol grounding mechanism and the corresponding disentanglement training method are introduced to enhance the controllability of logical symbols over generation. Furthermore, we propose two logical abduction methods to make our approach require few labeled training data and support the induction of latent logical generative rules from data. We experimentally show that our approach can be utilized to integrate various neural generative models with logical reasoning systems, by both learning from scratch or utilizing pre-trained models directly. The code is released at https://github.com/future-item/AbdGen.

cs.AI

BEDRF: Bidirectional Edge Diffraction Response Function for Interactive Sound Propagation

We introduce bidirectional edge diffraction response function (BEDRF), a new approach to model wave diffraction around edges with path tracing. The diffraction part of the wave is expressed as an integration on path space, and the wave-edge interaction is expressed using only the localized information around points on the edge similar to a bidirectional scattering distribution function (BSDF) for visual rendering. For an infinite single wedge, our model generates the same result as the analytic solution. Our approach can be easily integrated into interactive geometric sound propagation algorithms that use path tracing to compute specular and diffuse reflections. Our resulting propagation algorithm can approximate complex wave propagation phenomena involving high-order diffraction, and is able to handle dynamic, deformable objects and moving sources and listeners. We highlight the performance of our approach in different scenarios to generate smooth auralization.

cs.SD

A Psychoacoustic Quality Criterion for Path-Traced Sound Propagation

In developing virtual acoustic environments, it is important to understand the relationship between the computation cost and the perceptual significance of the resultant numerical error. In this paper, we propose a quality criterion that evaluates the error significance of path-tracing-based sound propagation simulators. We present an analytical formula that estimates the error signal power spectrum. With this spectrum estimation, we can use a modified Zwicker's loudness model to calculate the relative loudness of the error signal masked by the ideal output. Our experimental results show that the proposed criterion can explain the human perception of simulation error in a variety of cases.

cs.SD

Observation of Complete Photonic Bandgap in Low Refractive Index Contrast Inverse Rod-Connected Diamond Structured Chalcogenides

Three-dimensional complete photonic bandgap materials or photonic crystals block light propagation in all directions. The rod-connected diamond structure exhibits the largest photonic bandgap known to date and supports a complete bandgap for the lowest refractive index contrast ratio down to $n_{high}/n_{low} \sim 1.9$. We confirm this threshold by measuring a complete photonic bandgap in the infrared region in Sn--S--O $(n\sim1.9)$ and Ge--Sb--S--O $(n\sim2)$ inverse rod-connected diamond structures. The structures were fabricated using a low-temperature chemical vapor deposition process via a single-inversion technique. This provides a reliable fabrication technique of complete photonic bandgap materials and expands the library of backfilling materials, leading to a wide range of future photonic applications.

physics.optics

MPO: An Efficient and Low-cost Peer-to-Peer Overlay for Autonomic Communications

The term Autonomic Communication (AC) refers to self-managing systems which are capable of supporting self-configuration, self-healing and self-optimization. However, information reflection and collection, lack of centralized control, non-cooperation and so on are just some of the challenges within AC systems. We have considered these problems in theory and practice and reached the following conclusion; in order to build an ideal system for autonomic communication, there are three key problems to be solved. Motivated by the need for AC, we have designed an efficient and low-cost Peer-to-Peer (P2P) overlay called Maya-Pyramid overlay (MPO) and combined merits of unstructured P2P with those of structured P2P overlays. Differing from the traditional hierarchical P2P (i.e. tree-like structure) overlay, (1) MPO is composed of levels and layers, which uses small world characteristic to improve efficiency, and the maintenance cost is decreased because update and backup only take place in two neighboring levels or layers instead of recursively perform in higher levels. (2) Unlike normal redundant mechanisms for solving the single fault problem: Tri-Information Center (Tri-IC) mechanism is presented in order to improve robustness by alleviating the load of cluster heads in a hierarchical P2P overlay. (3) A source ranking mechanism is proposed in order to discourage free riding and whitewashing and to encourage frequent information exchanges between peers. (4) Inspired by Pastry's ID structure for a structured DHT algorithm, a 3D unique ID structure is presented in the unstructured P2P overlay. This will guarantee anonymity in routing, and will be, not only more efficient because it applies the DHT-like routing algorithm in the unstructured P2P overlay, but also more adaptive to suit AC. Evaluation proved that MPO is robust, highly efficient and of a low-cost.

cs.DC