arXiv ScienceSearch

arXiv subjects

Size Wu

Publications and source records attributed to Size Wu.

At least 19 recordsLinked to original sources

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

cs.CV

Layer-Number-Controlled Symmetry Breaking and Surface-State Transport in Rhombohedral Graphene Multilayers

Rhombohedral multilayer graphene hosts layer-polarized flat bands, providing an intriguing platform for correlated and topological electronic states; however, the role of layer number in governing symmetry breaking and surface screening remains elusive. Here we prepare rhombohedral graphene multilayers and systematically conduct electrical transport measurements. We uncover an unconventional layer dependence of phase transitions: the critical displacement field (D$_{c}$) for the layer-antiferromagnetic (LAF)-to-semimetal transitions remains constant across tetralayer to hexalayer graphene, whereas the D$_{c}$ for semimetal-to-layer-polarized-insulator (LPI) transition increases with layer number, defying unscreened Coulomb interaction models. In hexalayer graphene, surface-state-dominated transport emerges, with Landau levels (LLs) and resistive peaks selectively controlled by adjacent gates, a signature of strong interlayer screening absent in thinner stacks. High magnetic fields reveal valley-layer-locked LLs and dissipative states possibly from interlayer backscattering, highlighting the presence of decoupled surface states. Our findings establish layer number as a key tuning knob for engineering correlated and topological phases in rhombohedral graphene multilayers.

cond-mat.mes-hall

Prisma-World: Camera-Controllable Multi-Agent Video World Model

Video world models have made rapid progress in generating controllable visual experiences, but most of them still simulate the world from a single observer. Extending such models to multiple agents raises a central challenge: if each agent's future state is generated independently, overlapping views may instantiate different versions of the same scene, leading to inconsistent objects, layouts, and appearances across agents. Conventional camera conditioning controls individual trajectories, but it does not explicitly couple the generation of views that should agree under shared scene geometry. We introduce Prisma-World, a camera-controllable multi-agent world model that formulates multi-agent generation as a joint geometry-aware denoising process for cross-view consistency. Prisma-World processes all agent videos within one full-attention sequence, uses a multi-agent RoPE design to distinguish agent identities while preserving synchronized temporal coordinates, and injects relative camera geometry into attention to bias overlapping viewpoints toward shared scene evidence. To further strengthen multi-view consistency and enhance global spatial perception, we augment our framework with an overlap-decaying curriculum training paradigm alongside minimap-conditioned structural guidance. To facilitate the training and evaluation of multi-agent models, we introduce PrismaDataset, a large-scale UE5 dataset with panoramic acquisition across diverse scenes, composable multi-agent view groups with flexible agent counts and complex camera trajectories, and precise camera/action annotations for consistency training and evaluation. Experiments show that a single Prisma-World model can generate high-fidelity multi-agent videos with flexible agent numbers, camera controllability, improved cross-view consistency, and spatial grounding under minimap guidance.

cs.CV

Field-induced asymmetric band flattening and ideal quantum geometry in rhombohedral graphene

Rhombohedral graphene exhibits an exceptionally diverse array of correlated phases that depend sensitively on the displacement field. Compiling reported phases into a unified phase diagram reveals a pronounced field-dependent electron-hole asymmetry: correlated states on the hole-doped side emerge at small displacement fields, whereas the fractional quantum anomalous Hall effect (FQAHE) is observed exclusively on the electron-doped side under large displacement fields. This stark asymmetry highlights the need to understand how flat bands evolve with displacement fields. Here, we directly visualize the field-induced electron-hole asymmetric band flattening in rhombohedral pentalayer graphene (R5G) using nanospot angle-resolved photoemission spectroscopy with electrostatic gating. Beyond gap opening and spectral weight redistribution indicative of layer polarization, the gating field drives a strongly asymmetric modification of the flat bands: the flat valence band (FVB) evolves into an M-shaped dispersion at high field, whereas the flat conduction band (FCB) progressively flattens with increasing field. Comparison with calculations identifies critical parameters governing the band curvature of R5G, from which the resulting finite Berry curvature and near-ideal quantum geometry support the emergence of topological phases under electron doping at large fields. These results establish a direct link between the asymmetric phase diagram, band structure evolution, and quantum geometry, providing a microscopic framework for understanding correlated and topological phases in rhombohedral graphene.

cond-mat.mes-hall

Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation

Recent text-to-image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliability in knowledge-intensive settings remains largely unexplored. Unlike natural image generation, knowledge visualization requires not only semantic alignment but also strict adherence to domain knowledge, structural constraints, and symbolic conventions, exposing a critical gap between visual plausibility and scientific correctness. To systematically study this problem, we introduce KVBench, a curriculum-grounded benchmark for evaluating knowledge-intensive T2I generation. KVBench covers six senior high-school subjects: Biology, Chemistry, Geography, History, Mathematics, and Physics. The benchmark consists of 1,800 expert-curated prompts derived from over 30 authoritative textbooks. Using this benchmark, we evaluate 14 state-of-the-art open- and closed-source models, revealing substantial deficiencies in logical reasoning, symbolic precision, and multilingual robustness, with open-source models consistently underperforming proprietary systems. To address these limitations, we further propose KE-Check, a two-stage framework that improves scientific fidelity via (1) Knowledge Elaboration for structured prompt enrichment, and (2) Checklist-Guided Refinement for explicit constraint enforcement through violation identification and constraint-guided editing. KE-Check effectively mitigates scientific hallucinations, narrowing the performance gap between open-source and leading closed-source models. Data and codes are publicly available at https://github.com/zhaoran66/KVBench.

cs.CV

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address this, we propose UniReason, a unified framework that harmonizes these two tasks through two complementary reasoning paradigms. We incorporate world knowledge-enhanced textual reasoning into generation to infer implicit knowledge, and leverage editing capabilities for fine-grained editing-like visual refinement to further correct visual errors via self-reflection. This approach unifies generation and editing within a shared architecture, mirroring the human cognitive process of planning followed by refinement. We support this framework by systematically constructing a large-scale reasoning-centric dataset (~300k samples) covering five major knowledge domains (e.g., cultural commonsense, physics, etc.) for textual reasoning, alongside an agent-generated corpus for visual refinement. Extensive experiments demonstrate that UniReason achieves advanced performance on reasoning-intensive benchmarks such as WISE, KrisBench and UniREditBench, while maintaining superior general synthesis capabilities.

cs.CV

Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling

The recent surge in popularity of Nano-Banana and Seedream 4.0 underscores the community's strong interest in multi-image composition tasks. Compared to single-image editing, multi-image composition presents significantly greater challenges in terms of consistency and quality, yet existing models have not disclosed specific methodological details for achieving high-quality fusion. Through statistical analysis, we identify Human-Object Interaction (HOI) as the most sought-after category by the community. We therefore systematically analyze and implement a state-of-the-art solution for multi-image composition with a primary focus on HOI-centric tasks. We present Skywork UniPic 3.0, a unified multimodal framework that integrates single-image editing and multi-image composition. Our model supports an arbitrary (1~6) number and resolution of input images, as well as arbitrary output resolutions (within a total pixel budget of 1024x1024). To address the challenges of multi-image composition, we design a comprehensive data collection, filtering, and synthesis pipeline, achieving strong performance with only 700K high-quality training samples. Furthermore, we introduce a novel training paradigm that formulates multi-image composition as a sequence-modeling problem, transforming conditional generation into unified sequence synthesis. To accelerate inference, we integrate trajectory mapping and distribution matching into the post-training stage, enabling the model to produce high-fidelity samples in just 8 steps and achieve a 12.5x speedup over standard synthesis sampling. Skywork UniPic 3.0 achieves state-of-the-art performance on single-image editing benchmark and surpasses both Nano-Banana and Seedream 4.0 on multi-image composition benchmark, thereby validating the effectiveness of our data pipeline and training paradigm. Code, models and dataset are publicly available.

cs.CV

RecTok: Reconstruction Distillation along Rectified Flow

Visual tokenizers play a crucial role in diffusion models. The dimensionality of latent space governs both reconstruction fidelity and the semantic expressiveness of the latent feature. However, a fundamental trade-off is inherent between dimensionality and generation quality, constraining existing methods to low-dimensional latent spaces. Although recent works have leveraged vision foundation models to enrich the semantics of visual tokenizers and accelerate convergence, high-dimensional tokenizers still underperform their low-dimensional counterparts. In this work, we propose RecTok, which overcomes the limitations of high-dimensional visual tokenizers through two key innovations: flow semantic distillation and reconstruction--alignment distillation. Our key insight is to make the forward flow in flow matching semantically rich, which serves as the training space of diffusion transformers, rather than focusing on the latent space as in previous works. Specifically, our method distills the semantic information in VFMs into the forward flow trajectories in flow matching. And we further enhance the semantics by introducing a masked feature reconstruction loss. Our RecTok achieves superior image reconstruction, generation quality, and discriminative performance. It achieves state-of-the-art results on the gFID-50K under both with and without classifier-free guidance settings, while maintaining a semantically rich latent space structure. Furthermore, as the latent dimensionality increases, we observe consistent improvements. Code and model are available at https://shi-qingyu.github.io/rectok.github.io.

cs.CV

Generative Photographic Control for Scene-Consistent Video Cinematic Editing

Cinematic storytelling is profoundly shaped by the artful manipulation of photographic elements such as depth of field and exposure. These effects are crucial in conveying mood and creating aesthetic appeal. However, controlling these effects in generative video models remains highly challenging, as most existing methods are restricted to camera motion control. In this paper, we propose CineCtrl, the first video cinematic editing framework that provides fine control over professional camera parameters (e.g., bokeh, shutter speed). We introduce a decoupled cross-attention mechanism to disentangle camera motion from photographic inputs, allowing fine-grained, independent control without compromising scene consistency. To overcome the shortage of training data, we develop a comprehensive data generation strategy that leverages simulated photographic effects with a dedicated real-world collection pipeline, enabling the construction of a large-scale dataset for robust model training. Extensive experiments demonstrate that our model generates high-fidelity videos with precisely controlled, user-specified photographic camera effects.

cs.CV

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and diffusion-based generation to interpret and create scenes from arbitrary viewpoints. To bridge the modality gap between cameras and vision-language, we introduce a novel paradigm that treats camera as language, enabling thinking with camera. This guides the model to align spatially grounded visual cues with photographic terminology while reasoning across geometric context. Puffin is trained on Puffin-4M, a large-scale dataset of 4 million vision-language-camera triplets. We incorporate both global camera parameters and pixel-wise camera maps, yielding flexible and reliable spatial generation. Experiments demonstrate Puffin superior performance over specialized models for camera-centric generation and understanding. With instruction tuning, Puffin generalizes to diverse cross-view tasks such as spatial imagination, world exploration, and photography guidance. We will release the code, models, dataset pipeline, and benchmark to advance multimodal spatial intelligence research.

cs.CV

Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model

Recent advances in multimodal models have demonstrated impressive capabilities in unified image generation and editing. However, many prominent open-source models prioritize scaling model parameters over optimizing training strategies, limiting their efficiency and performance. In this work, we present UniPic2-SD3.5M-Kontext, a 2B-parameter DiT model based on SD3.5-Medium, which achieves state-of-the-art image generation and editing while extending seamlessly into a unified multimodal framework. Our approach begins with architectural modifications to SD3.5-Medium and large-scale pre-training on high-quality data, enabling joint text-to-image generation and editing capabilities. To enhance instruction following and editing consistency, we propose a novel Progressive Dual-Task Reinforcement strategy (PDTR), which effectively strengthens both tasks in a staged manner. We empirically validate that the reinforcement phases for different tasks are mutually beneficial and do not induce negative interference. After pre-training and reinforcement strategies, UniPic2-SD3.5M-Kontext demonstrates stronger image generation and editing capabilities than models with significantly larger generation parameters-including BAGEL (7B) and Flux-Kontext (12B). Furthermore, following the MetaQuery, we connect the UniPic2-SD3.5M-Kontext and Qwen2.5-VL-7B via a connector and perform joint training to launch a unified multimodal model UniPic2-Metaquery. UniPic2-Metaquery integrates understanding, generation, and editing, achieving top-tier performance across diverse tasks with a simple and scalable training paradigm. This consistently validates the effectiveness and generalizability of our proposed training paradigm, which we formalize as Skywork UniPic 2.0.

cs.CV

Matrix-game 2.0: An open-source real-time and streaming interactive world model

Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynamics and interactive behaviors. However, existing interactive world models depend on bidirectional attention and lengthy inference steps, severely limiting real-time performance. Consequently, they are hard to simulate real-world dynamics, where outcomes must update instantaneously based on historical context and current actions. To address this, we present Matrix-Game 2.0, an interactive world model generates long videos on-the-fly via few-step auto-regressive diffusion. Our framework consists of three key components: (1) A scalable data production pipeline for Unreal Engine and GTA5 environments to effectively produce massive amounts (about 1200 hours) of video data with diverse interaction annotations; (2) An action injection module that enables frame-level mouse and keyboard inputs as interactive conditions; (3) A few-step distillation based on the casual architecture for real-time and streaming video generation. Matrix Game 2.0 can generate high-quality minute-level videos across diverse scenes at an ultra-fast speed of 25 FPS. We open-source our model weights and codebase to advance research in interactive world modeling.

cs.CV

Controllable Human-centric Keyframe Interpolation with Generative Prior

Existing interpolation methods use pre-trained video diffusion priors to generate intermediate frames between sparsely sampled keyframes. In the absence of 3D geometric guidance, these methods struggle to produce plausible results for complex, articulated human motions and offer limited control over the synthesized dynamics. In this paper, we introduce PoseFuse3D Keyframe Interpolator (PoseFuse3D-KI), a novel framework that integrates 3D human guidance signals into the diffusion process for Controllable Human-centric Keyframe Interpolation (CHKI). To provide rich spatial and structural cues for interpolation, our PoseFuse3D, a 3D-informed control model, features a novel SMPL-X encoder that transforms 3D geometry and shape into the 2D latent conditioning space, alongside a fusion network that integrates these 3D cues with 2D pose embeddings. For evaluation, we build CHKI-Video, a new dataset annotated with both 2D poses and 3D SMPL-X parameters. We show that PoseFuse3D-KI consistently outperforms state-of-the-art baselines on CHKI-Video, achieving a 9% improvement in PSNR and a 38% reduction in LPIPS. Comprehensive ablations demonstrate that our PoseFuse3D model improves interpolation fidelity.

cs.CV

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector. With a minimalist choice of architecture, we demonstrate that OpenUni can: 1) generate high-quality and instruction-aligned images, and 2) achieve exceptional performance on standard benchmarks such as GenEval, DPG- Bench, and WISE, with only 1.1B and 3.1B activated parameters. To support open research and community advancement, we release all model weights, training code, and our curated training datasets (including 23M image-text pairs) at https://github.com/wusize/OpenUni.

cs.CV

Intrinsic layer polarization and multi-flatband transport in non-centrosymmetric mixed-stacked multilayer graphene

Graphene multilayers exhibit electronic spectra that depend sensitively on both the number of layers and their stacking order. Beyond trilayer graphene, mixed stacking sequences (alternating Bernal and rhombohedral layers) give rise to multiple coexisting low-energy bands. Here we investigate ABCBC-stacked pentalayer graphene, a less-studied non-centrosymmetric mixed sequence. This stacking can be regarded as an ABC (rhombohedral) trilayer on top of an AB (Bernal) bilayer, so its low-energy band structure contains both a cubic band and a parabolic band that hybridize. In transport measurements, we observe an intrinsic band gap at charge neutrality whose magnitude changes asymmetrically under an applied perpendicular displacement field. This behavior reflects the spontaneous layer polarization inherent to the broken inversion symmetry and mirror symmetry. By tuning the displacement field and carrier density, we drive multiple Lifshitz transitions in the Fermi surface topology and realize Landau levels with different degeneracies arising from the multi-flatband system. Remarkably, a v = -6 quantum Hall state emerges at an exceptionally low magnetic field (~20 mT), indicating the interplay between spontaneous symmetry breaking and Berry curvatures. Our results establish mixed-stacked multilayer graphene as a tunable platform with various broken symmetries and multiple flatbands, suitable for exploring emergent correlated electronic states.

cond-mat.mes-hall

Moir\'e enhanced flat band in rhombohedral graphene

The fractional quantum anomalous Hall effect (FQAHE) is a fascinating emergent quantum state characterized by fractionally charged excitations in the absence of magnetic field,which could arise from the intricate interplay between electron correlation, nontrivial topology and spontaneous time-reversal symmetry breaking. Recently, FQAHE has been realized in aligned rhombohedral pentalayer graphene on BN superlattice (aligned R5G/BN), where the topological flat band is modulated by the moir\'e potential. However, intriguingly, the FQAHE is observed only when electrons are pushed away from the moir\'e interface. The apparently opposite implications from these experimental observations, along with different theoretical models, have sparked intense debates regarding the role of the moir\'e potential. Unambiguous experimental observation of the topological flat band as well as moir\'e bands with energy and momentum resolved information is therefore critical to elucidate the underlying mechanism. Here by performing nanospot angle-resolved photoemission spectroscopy (NanoARPES) measurements, we directly reveal the topological flat band electronic structures of R5G, from which key hopping parameters essential for determining the fundamental electronic structure of rhombohedral graphene are extracted. Moreover, a comparison of electronic structures between aligned and non-aligned samples reveals that the moir\'e potential plays a pivotal role in enhancing the topological flat band in the aligned sample. Our study provides experimental guiding lines to narrow down the phase space of rhombohedral graphene, laying an important foundation for understanding exotic quantum phenomena in this emerging platform.

cond-mat.mes-hall

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations at different levels of granularity. Current approaches that utilize vector quantization (VQ) or variational autoencoders (VAE) for unified visual representation prioritize intrinsic imagery features over semantics, compromising understanding performance. In this work, we take inspiration from masked image modelling (MIM) that learns rich semantics via a mask-and-reconstruct pre-training and its successful extension to masked autoregressive (MAR) image generation. A preliminary study on the MAR encoder's representation reveals exceptional linear probing accuracy and precise feature response to visual concepts, which indicates MAR's potential for visual understanding tasks beyond its original generation role. Based on these insights, we present \emph{Harmon}, a unified autoregressive framework that harmonizes understanding and generation tasks with a shared MAR encoder. Through a three-stage training procedure that progressively optimizes understanding and generation capabilities, Harmon achieves state-of-the-art image generation results on the GenEval, MJHQ30K and WISE benchmarks while matching the performance of methods with dedicated semantic encoders (e.g., Janus) on image understanding benchmarks. Our code and models will be available at https://github.com/wusize/Harmon.

cs.CV

Switchable Chern insulator, isospin competitions and charge density waves in rhombohedral graphene moire superlattices

Graphene-based moire superlattices provide a versatile platform for exploring novel correlated and topological electronic states, driven by enhanced Coulomb interactions within flat bands. The intrinsic tunability of graphene s multiple degrees of freedom enables precise control over these complex quantum phases. In this study, we observe a range of competing phases and their transitions in rhombohedrally stacked hexalayer graphene on hexagonal boron nitride (r-6G/hBN) moire superlattices. When electrons are polarized away from the moire superlattice, we firstly identify a Chern insulator with reversible Chern numbers at v = 1 (one electron per moire cell), attributed to the competition between bulk and edge magnetizations.Then, we detect transitions between three distinct insulating states at v = 2, driven by vertical displacement field D and vertical magnetic field B. These insulating phases are distinguished as spin-antiferromagnetic, spin-polarized, and valley-polarized insulators, based on their responses to parallel and perpendicular magnetic fields. When electrons are polarized toward the moire superlattice, in a device with large twist angle, insulating states appear at v = 1/3 and 2/3 at zero magnetic field, and v = 1/2 in a magnetic field. Our findings reveal a rich interplay of charge, isospin, topology and magnetic field in rhombohedral graphene moire superlattices.

cond-mat.str-el