arXiv ScienceSearch

arXiv subjects

Yanbing Zhang

Publications and source records attributed to Yanbing Zhang.

At least 19 recordsLinked to original sources

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

cs.CV

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static observation. We propose Thinking with Novel Views (TwNV), a paradigm that integrates generative novel-view synthesis into the reasoning loop: a Reasoner LMM identifies spatial ambiguity, instructs a Painter to synthesize an alternative viewpoint, and re-examines the scene with the additional evidence. Through systematic experiments we address three research questions. (1) Instruction format: numerical camera-pose specifications yield more reliable view control than free-form language. (2) Generation fidelity: synthesized view quality is tightly coupled with downstream spatial accuracy. (3) Inference-time visual scaling: iterative multi-turn view refinement further improves performance, echoing recent scaling trends in language reasoning. Across four spatial subtask categories and four LMM architectures (both closed- and open-source), TwNV consistently improves accuracy by +1.3 to +3.9 pp, with the largest gains on viewpoint-sensitive subtasks. These results establish novel-view generation as a practical lever for advancing spatial intelligence of LMMs.

cs.CV

TextLDM: Language Modeling with Continuous Latent Diffusion

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding (text generation) is to apply this framework to language modeling. We propose TextLDM, which transfers the visual latent diffusion recipe to text generation with minimal architectural modification. A Transformer-based VAE maps discrete tokens to continuous latents, enhanced by Representation Alignment (REPA) with a frozen pretrained language model to produce representations effective for conditional denoising. A standard DiT then performs flow matching in this latent space, identical in architecture to its visual counterpart. The central challenge we address is obtaining high-quality continuous text representations: we find that reconstruction fidelity alone is insufficient, and that aligning latent features with a pretrained language model via REPA is critical for downstream generation quality. Trained from scratch on OpenWebText2, TextLDM substantially outperforms prior diffusion language models and matches GPT-2 under the same settings. Our results establish that the visual DiT recipe transfers effectively to language, taking a concrete step toward unified diffusion architectures for multimodal generation and understanding.

cs.CL

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing. JoyAI-Image couples a spatially enhanced Multimodal Large Language Model (MLLM) with a Multimodal Diffusion Transformer (MMDiT), allowing perception and generation to interact through a shared multimodal interface. Around this architecture, we build a scalable training recipe that combines unified instruction tuning, long-text rendering supervision, spatially grounded data, and both general and spatial editing signals. This design gives the model broad multimodal capability while strengthening geometry-aware reasoning and controllable visual synthesis. Experiments across understanding, generation, long-text rendering, and editing benchmarks show that JoyAI-Image achieves state-of-the-art or highly competitive performance. More importantly, the bidirectional loop between enhanced understanding, controllable spatial editing, and novel-view-assisted reasoning enables the model to move beyond general visual competence toward stronger spatial intelligence. These results suggest a promising path for unified visual models in downstream applications such as vision-language-action systems and world models.

cs.GR

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence

Spatial understanding is a fundamental cornerstone of human-level intelligence. Nonetheless, current research predominantly focuses on domain-specific data production, leaving a critical void: the absence of a principled, open-source engine capable of fully unleashing the potential of high-quality spatial data. To bridge this gap, we elucidate the design principles of a robust data generation system and introduce OpenSpatial -- an open-source data engine engineered for high quality, extensive scalability, broad task diversity, and optimized efficiency. OpenSpatial adopts 3D bounding boxes as the fundamental primitive to construct a comprehensive data hierarchy across five foundational tasks: Spatial Measurement (SM), Spatial Relationship (SR), Camera Perception (CP), Multi-view Consistency (MC), and Scene-Aware Reasoning (SAR). Leveraging this scalable infrastructure, we curate OpenSpatial-3M, a large-scale dataset comprising 3 million high-fidelity samples. Extensive evaluations demonstrate that versatile models trained on our dataset achieve state-of-the-art performance across a wide spectrum of spatial reasoning benchmarks. Notably, the best-performing model exhibits a substantial average improvement of 19 percent, relatively. Furthermore, we provide a systematic analysis of how data attributes influence spatial perception. By open-sourcing both the engine and the 3M-scale dataset, we provide a robust foundation to accelerate future research in spatial intelligence.

cs.CL

An efficient treatment of heat-flux boundary conditions in GSIS for rarefied gas flows

Heat-flux boundary conditions are challenging to implement efficiently in rarefied gas flow simulations because the wall-reflected gas temperature and density must be determined dynamically during the computation. This paper aims to tackle this problem within the general synthetic iterative scheme (GSIS), where the Boltzmann kinetic equation is solved deterministically in an outer loop and macroscopic synthetic equations are solved in an inner loop. To avoid kinetic-macroscopic boundary-flux mismatch and the resulting convergence bottlenecks, for the macroscopic boundary flux at every inner iteration, the incident increment is estimated using a Maxwellian distribution, and then the reflected contribution is obtained by boundary conditions consistent with those in the kinetic solver. In addition to retaining the fast-converging and asymptotic-preserving properties of GSIS, the proposed method significantly reduces the iterations required to determine the wall-reflected gas parameters. Numerical simulations of rarefied gas flows in and around a 3D nozzle, a 2D adiabatic cylinder, and a 2D annular heat-transfer configuration show good agreement with the direct simulation Monte Carlo method, while achieving substantial efficiency gains over conventional iterative schemes.

physics.comp-ph

Accelerated simulation of multiscale gas-radiation coupling flows via a general synthetic iterative scheme

Gas-radiation coupling critically influences hypersonic reentry flows, where extreme temperatures induce pronounced non-equilibrium gas and radiative heat transport. Accurate and efficient simulation of radiative gas dynamics is therefore indispensable for reliable design of thermal protection systems for atmospheric entry vehicles. In this study, a Boltzmann-type kinetic model for radiative gas flows is solved across a broad spectrum of flow and radiation transport regimes using the general synthetic iterative scheme (GSIS). The approach integrates an unstructured finite-volume discrete velocity method with a set of macroscopic synthetic equations. Within this framework, the kinetic model provides high-order closures for the constitutive relations in the synthetic equations. Simultaneously, the macroscopic synthetic equations drive the evolution of the mesoscopic kinetic system, significantly accelerating steady-state convergence in near-continuum regimes, as substantiated by linear Fourier stability analysis. Crucially, the algorithm is proven to be asymptotic-preserving, correctly recovering the continuum and optically thick limits, represented by the radiative Navier-Stokes-Fourier equations governing distinct translational, rotational, vibrational, and radiative temperatures, on coarse meshes independent of the mean free path. Numerical simulations of challenging benchmarks, including three-dimensional hypersonic flow over an Apollo reentry capsule, demonstrate that GSIS achieves orders-of-magnitude speedup over conventional iterative schemes in multiscale simulations of radiative gas flows while accurately capturing non-equilibrium effects and radiative heat transfer in hypersonic environments.

physics.comp-ph

Surrogate-assisted airfoil optimization in rarefied gas flows

With growing interest in space exploration, optimized airfoil design has become increasingly important. However, airfoil design in rarefied gas flows remains underexplored because solving the Boltzmann equation formulated in a six dimensional phase space is time consuming. To address this problem, a solver-in-the-loop Bayesian optimization framework for symmetric, thickness-only airfoils is developed. First, airfoils are parameterized using a class shape transformation that enforce geometric admissibility. Second, a Gaussian process expected improvement surrogate is coupled in batches to a fast converging, asymptotic preserving Boltzmann solver for sample efficient exploration. Drag minimizing airfoils are identified in a wide range of gas rarefaction. It is found that, at Mach numbers Ma=2 and 4, the streamwise force increases with the gas rarefaction and shifts from pressure dominated to shear dominated drag, while optimization reduces drag at all conditions. The benefit of optimization peaks in the weakly rarefied regime, about 30% at Ma=2 and 40 to 50% at Ma=4, and falls to a few percent in transition and free-molecular flow regimes. Drag decomposition shows that these gains come mainly from reduced pressure drag, with viscous drag almost unchanged. The optimal airfoils form a coherent rarefaction-aware family: they retain a smooth, single-peaked thickness profile, are aft-loaded at low gas rarefaction, and exhibit a forward shift of maximum thickness and thickness area toward mid-chord as gas rarefaction increases. These trends provide a physically interpretable map that narrows the design space.

physics.flu-dyn

A fast-converging and asymptotic-preserving method for adjoint shape optimization of rarefied gas flows

Adjoint based shape optimization is a powerful technique in fluid-dynamics optimization, capable of identifying an optimal shape within only dozens of design iterations. However, when extended to rarefied gas flows, the computational cost becomes enormous because both the six dimensional primal and adjoint Boltzmann equations must be solved for each candidate shape. Building on the general synthetic iterative scheme (GSIS) for solving the primal Boltzmann model equation, this paper presents a fast converging and asymptotic preserving method for solving the adjoint kinetic equation. The GSIS accelerates the convergence of the adjoint kinetic equation by incorporating solutions of macroscopic synthetic equations, whose constitutive relations include the Newtonian stress law along with higher order terms capturing rarefaction effects. As a result, the method achieves asymptotic preservation (allowing the use of large spatial cell sizes in the continuum limit) while maintaining accuracy in highly rarefied regimes. Numerical tests demonstrate exceptional performance on drag minimization problems for 3D bodies, achieving drag reductions of 34.5% in the transition regime and 61.1% in the slip-flow regime within roughly ten optimization iterations. For each candidate shape, converged solutions of the primal and adjoint Boltzmann equation are obtained with only a few dozen updates of the velocity distribution function, dramatically reducing computational cost compared with conventional methods.

physics.comp-ph

FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

In light of recent breakthroughs in text-to-image (T2I) generation, particularly with diffusion transformers (DiT), subject-driven technologies are increasingly being employed for high-fidelity customized production that preserves subject identity from reference inputs, enabling thrilling design workflows and engaging entertainment. Existing alternatives typically require either per-subject optimization via trainable text embeddings or training specialized encoders for subject feature extraction on large-scale datasets. Such dependencies on training procedures fundamentally constrain their practical applications. More importantly, current methodologies fail to fully leverage the inherent zero-shot potential of modern diffusion transformers (e.g., the Flux series) for authentic subject-driven synthesis. To bridge this gap, we propose FreeCus, a genuinely training-free framework that activates DiT's capabilities through three key innovations: 1) We introduce a pivotal attention sharing mechanism that captures the subject's layout integrity while preserving crucial editing flexibility. 2) Through a straightforward analysis of DiT's dynamic shifting, we propose an upgraded variant that significantly improves fine-grained feature extraction. 3) We further integrate advanced Multimodal Large Language Models (MLLMs) to enrich cross-modal semantic representations. Extensive experiments reflect that our method successfully unlocks DiT's zero-shot ability for consistent subject synthesis across diverse contexts, achieving state-of-the-art or comparable results compared to approaches that require additional training. Notably, our framework demonstrates seamless compatibility with existing inpainting pipelines and control modules, facilitating more compelling experiences. Our code is available at: https://github.com/Monalissaa/FreeCus.

cs.CV

General synthetic iterative scheme for multiscale radiative transfer in the finite-volume framework

Achieving efficient and accurate simulation of the radiative transfer has long been a research challenge. Here we introduce the general synthetic iterative scheme as an easy-to-implement approach to address this issue. First, a macroscopic synthetic equation, which combines the asymptotic equation at the diffusion limit and the "high-order terms" extracted from the transport equation to account for transport effects, is introduced to accelerate the simulation of the radiative transfer equation. Second, the asymptotic preserving property is directly provided by the macroscopic process, eliminating the need for fine spatial discretization in optically thick media, as well as the need for consistency enforcement. Third, to address the issue of opacity discontinuity in the finite volume method, an adaptive least square method for gradient approximation is proposed. Numerical results on several canonical tests demonstrate that, in optically thick problems, our method achieves significant speed-up over the conventional iterative schemes. Finally, with our newly developed method, we reveal the importance of resolving the Knudsen layer in the initial stage of Tophat problem, while in steady-state the Knudsen layer can be under-resolved.

physics.comp-ph

InstantCharacter: Personalize Any Characters with a Scalable Diffusion Transformer Framework

Current learning-based subject customization approaches, predominantly relying on U-Net architectures, suffer from limited generalization ability and compromised image quality. Meanwhile, optimization-based methods require subject-specific fine-tuning, which inevitably degrades textual controllability. To address these challenges, we propose InstantCharacter, a scalable framework for character customization built upon a foundation diffusion transformer. InstantCharacter demonstrates three fundamental advantages: first, it achieves open-domain personalization across diverse character appearances, poses, and styles while maintaining high-fidelity results. Second, the framework introduces a scalable adapter with stacked transformer encoders, which effectively processes open-domain character features and seamlessly interacts with the latent space of modern diffusion transformers. Third, to effectively train the framework, we construct a large-scale character dataset containing 10-million-level samples. The dataset is systematically organized into paired (multi-view character) and unpaired (text-image combinations) subsets. This dual-data structure enables simultaneous optimization of identity consistency and textual editability through distinct learning pathways. Qualitative experiments demonstrate the advanced capabilities of InstantCharacter in generating high-fidelity, text-controllable, and character-consistent images, setting a new benchmark for character-driven image generation. Our source code is available at https://github.com/Tencent/InstantCharacter.

cs.CV

Adaptive GSIS for rarefied gas flow simulations

The parallel solver of the general synthetic iterative scheme (GSIS), as recently developed by Zhang \textit{et. al.} in Comput. Fluids 281 (2024) 106374, is an efficient method to find the solution of the Boltzmann equation deterministically. However, it consumes a significant computational memory due to the discretization of molecular velocity space in hypersonic flows. In this paper, we address this issue by introducing the adaptive GSIS, where the Boltzmann equation is applied only in rarefied regions when the local Knudsen number exceeds a reference value, $\text{Kn}{ref}$. In contrast, the Navier-Stokes equations, with and without the high-order corrections to the constitutive relations, are applied in the continuum and rarefied regimes, respectively. Numerical results indicate that setting $\text{Kn}{ref}=0.01$ yields acceptable outcomes. With the adaptive GSIS, the computational memory and time can be significantly reduced in near-continuum flows, e.g. 24 and 7 times, respectively. in the simulation of rarefied gas flow passing the International Space Station.

physics.comp-ph

Multiscale simulation of neutral particle flows in the plasma edge

The plasma edge flow, situated at the intricate boundary between plasma and neutral particles, plays a pivotal role in the design of nuclear fusion devices such as divertors and pumps. Traditional numerical simulation methods, such as the direct simulation Monte Carlo approach and the discrete velocity method, are hindered by extensive computation times when dealing with near-continuum flow conditions. This paper presents a general synthetic iterative scheme to deterministically simulate the plasma edge flows. By alternately solving the kinetic equations and macroscopic synthetic equations, our method substantially decreases the number of iterations, while maintains asymptotic-preserving properties even when the spatial cell size is much larger than the mean free path. Consequently, our approach achieves rapid convergence and high accuracy in plasma edge flow simulations, particularly in near-continuum flow regimes. This advancement provides a robust and efficient computational tool, essential for the advancement of next-generation nuclear fusion reactors.

physics.comp-ph

Multiscale simulation of rarefied gas flows in Divertor Tokamak Test facility

Simulating gas flow within the divertor, which is a crucial component in nuclear fusion reactors, is essential for assessing and enhancing its design and performance. Traditional methods, such as the direct simulation Monte Carlo and the discrete velocity method, often fall short in efficiency for these simulations. In this study, we utilize the general synthetic iterative scheme to simulate a simplified Tokamak divertor model, demonstrating its fast convergence and asymptotic-preserving properties in complex three-dimensional scenarios. A conservative estimate of speedup by three orders of magnitude is achieved by the general synthetic iterative scheme when compared to the direct simulation Monte Carlo method. We further investigate the relationship between pumping efficiency and factors like temperature, absorptivity, and the Knudsen number, providing valuable insights to guide the design and optimization of divertor structures.

physics.plasm-ph

GSIS-ALE for moving boundary problems in rarefied gas flows

Multiscale rarefied gas flows with moving boundaries pose significant challenges to the numerical simulation, where the primary difficulties involve robustly managing the mesh movement and ensuring computational efficiency across all flow regimes. Build upon recent advancements of the general synthetic iterative scheme (GSIS), this paper presents an efficient solver to simulate the large displacement of rigid-body in rarefied gas flows. The newly developed solver utilizes a dual time step method to solve the mesoscopic kinetic and macroscopic synthetic equations alternately, in an arbitrary Lagrangian-Eulerian framework. Additionally, the overset mesh is used and the six degree-of-freedom rigid body dynamics equation is integrated to track the motion of solids. Four moving boundary problems encompassing a wide range of flow velocities and gas rarefaction are simulated, including the periodic pitching of airfoil, particle motion in lid-driven cavity flow, two-body separation in supersonic flow, and three-dimensional lunar landing, demonstrating the accuracy and efficiency of the GSIS in handling multi-scale moving boundary problems within an overset framework.

physics.comp-ph

Attention Calibration for Disentangled Text-to-Image Personalization

Recent thrilling progress in large-scale text-to-image (T2I) models has unlocked unprecedented synthesis quality of AI-generated content (AIGC) including image generation, 3D and video composition. Further, personalized techniques enable appealing customized production of a novel concept given only several images as reference. However, an intriguing problem persists: Is it possible to capture multiple, novel concepts from one single reference image? In this paper, we identify that existing approaches fail to preserve visual consistency with the reference image and eliminate cross-influence from concepts. To alleviate this, we propose an attention calibration mechanism to improve the concept-level understanding of the T2I model. Specifically, we first introduce new learnable modifiers bound with classes to capture attributes of multiple concepts. Then, the classes are separated and strengthened following the activation of the cross-attention operation, ensuring comprehensive and self-contained concepts. Additionally, we suppress the attention activation of different classes to mitigate mutual influence among concepts. Together, our proposed method, dubbed DisenDiff, can learn disentangled multiple concepts from one single image and produce novel customized images with learned concepts. We demonstrate that our method outperforms the current state of the art in both qualitative and quantitative evaluations. More importantly, our proposed techniques are compatible with LoRA and inpainting pipelines, enabling more interactive experiences.

cs.CV

Further acceleration of multiscale simulation of rarefied gas flow via a generalized boundary treatment

The recently-developed general synthetic iterative scheme (GSIS) is efficient in simulating multiscale rarefied gas flows due to the coupling of mesoscopic kinetic equation and macroscopic synthetic equation: for linearized Poiseuille flow where the boundary flux is fixed at each iterative step, the steady-state solutions are found within dozens of iterations in solving the gas kinetic equations, while for general nonlinear flows the iteration number is increased by about one order of magnitude, caused by the incompatible treatment of the boundary flux for the macroscopic synthetic equation. In this paper, we propose a generalized boundary treatment (GBT) to further accelerate the convergence of GSIS. The main idea is, the truncated velocity distribution function at the boundary, similar to that used in the Grad 13-moment equation, is reconstructed by the macroscopic conserved quantities from the synthetic equation, and the high-order correction of non-equilibrium stress and heat flux from the kinetic equation; therefore, in each inner iteration solving the synthetic equation, the explicit constitutive relations facilitate real-time updates of the macroscopic boundary flux, driving faster information exchange in the flow field, and consequently achieving quicker convergence. Moreover, the high-order correction derived from the kinetic equation can compensate the approximation by the truncation and ensure the boundary accuracy. The accuracy of GSIS-GBT is validated by the direct simulation Monte Carlo method, the previous versions of GSIS, and the unified gas-kinetic wave-particle method. For the efficiency, in the near-continuum flow regime and slip regime, GSIS-GBT can be faster than the conventional iteration scheme in the discrete velocity method and the previous versions of GSIS by two- and one-order of magnitude, respectively.

physics.flu-dyn