arXiv ScienceSearch

arXiv subjects

Jian Mao

Publications and source records attributed to Jian Mao.

7 recordsLinked to original sources

MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling

Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token length and memory. We introduce MegaParts, a scalable autoregressive 3D generation framework to address this challenge by combining structured sequence modeling with a token-efficient vector-quantized shape tokenizer. Our tokenizer learns discrete latent representations for part-level geometry by minimizing token usage subject to high-fidelity reconstruction, enabling adaptive-length tokenization based on geometric complexity. On top of this compact representation, we train a large language model to generate object bounding boxes, part bounding boxes, and part shape tokens within a unified structured sequence. Combined with efficient long-context training strategy, our token-efficient formulation scales to objects with up to 300 parts and sequence lengths up to 256k tokens. This substantially extends the scale of part-aware 3D generation while preserving compositional structure and enabling fine-grained part-level control. Our method achieves higher mesh quality than baseline autoregressive and diffusion models, showing that compressed discrete part tokens improve not only scalability but also the achievable fidelity of generated geometry. These results suggest that LLM native token-efficient autoregressive modeling is a compelling alternative to diffusion for large-scale part-aware 3D generation. The project page is available at https://expmaster.github.io/megaparts_webpage.

cs.CV

MergeDJD: A Fast Constructive Algorithm with Piece Merging for the Two-Dimensional Irregular Bin Packing Problem

The two-dimensional irregular bin packing problem (2DIBPP) aims to pack a given set of irregular polygons, referred to as pieces, into fixed-size rectangular bins without overlap, while maximizing bin utilization. Although numerous metaheuristic algorithms have been proposed for the 2DIBPP, many industrial applications favor simpler constructive heuristics due to their deterministic behavior and low computational overhead. Among such methods, the DJD algorithm proposed by L'opez-Camacho et al. is one of the most competitive constructive heuristics for the 2DIBPP. However, DJD is less effective for cutting instances, in which many pieces can be seamlessly combined into larger polygons. To address the issue, we propose MergeDJD, a novel constructive algorithm that integrates and extends the DJD framework. MergeDJD first preprocesses the instance by iteratively identifying groups of pieces that can be combined into larger and more regular piece. It then employs an improved version of DJD, in which the placement strategy is enhanced to better handle non-convex and combined shapes, to pack all resulting pieces into bins. Computational experiments on 1,089 well-known benchmark instances show that MergeDJD consistently outperforms DJD on 1,083 instances while maintaining short runtimes. Notably, MergeDJD attains new best known values on 515 instances. Ablation studies further confirm the effectiveness of the proposed components. To facilitate reproducibility and future research, we have open-sourced the complete implementation and provided interfaces for visualizing packing results.

cs.CG

Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising

Diffusion-based text-to-video models are increasingly capable, but mask-based editing over hundreds of frames remains challenging: na\"ive long-video generation suffers from memory blow-up, window seams, and temporal drift, while existing editors often require specialized modules or heavy fine-tuning. We present Overlapping High-Order Co-Denoising, a lightweight framework that turns a single pre-trained text-to-video model into a unified inpainting-outpainting editor. We train only LoRA adapters using mixed interior and border masks together with a dual-region loss that improves synthesis inside the mask while explicitly preserving known content. At inference, we denoise long latent sequences using overlapping windows, apply second-order Heun sampling within each window, and fuse overlaps with Hamming-weighted blending to reduce boundary artifacts and improve temporal coherence. On InpaintBench (30 real-world videos, 81--300 frames), our method outperforms Wan 2.1 variants and VACE in background faithfulness (SSIM/LPIPS), temporal consistency (tLPIPS), and text alignment (CLIP), and scales to long horizons, demonstrated up to 800 frames, with memory bounded by the chosen window size.

cs.CV

MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration

Real-world videos often suffer from complex degradations, such as noise, compression artifacts, and low-light distortions, due to diverse acquisition and transmission conditions. Existing restoration methods typically require professional manual selection of specialized models or rely on monolithic architectures that fail to generalize across varying degradations. Inspired by expert experience, we propose MoA-VR, the first \underline{M}ixture-\underline{o}f-\underline{A}gents \underline{V}ideo \underline{R}estoration system that mimics the reasoning and processing procedures of human professionals through three coordinated agents: Degradation Identification, Routing and Restoration, and Restoration Quality Assessment. Specifically, we construct a large-scale and high-resolution video degradation recognition benchmark and build a vision-language model (VLM) driven degradation identifier. We further introduce a self-adaptive router powered by large language models (LLMs), which autonomously learns effective restoration strategies by observing tool usage patterns. To assess intermediate and final processed video quality, we construct the \underline{Res}tored \underline{V}ideo \underline{Q}uality (Res-VQ) dataset and design a dedicated VLM-based video quality assessment (VQA) model tailored for restoration tasks. Extensive experiments demonstrate that MoA-VR effectively handles diverse and compound degradations, consistently outperforming existing baselines in terms of both objective metrics and perceptual quality. These results highlight the potential of integrating multimodal intelligence and modular reasoning in general-purpose video restoration systems.

cs.CV

Physical-Layer Signal Injection Attacks on EV Charging Ports: Bypassing Authentication via Electrical-Level Exploits

The proliferation of electric vehicles in recent years has significantly expanded the charging infrastructure while introducing new security risks to both vehicles and chargers. In this paper, we investigate the security of major charging protocols such as SAE J1772, CCS, IEC 61851, GB/T 20234, and NACS, uncovering new physical signal spoofing attacks in their authentication mechanisms. By inserting a compact malicious device into the charger connector, attackers can inject fraudulent signals to sabotage the charging process, leading to denial of service, vehicle-induced charger lockout, and damage to the chargers or the vehicle's charge management system. To demonstrate the feasibility of our attacks, we propose PORTulator, a proof-of-concept (PoC) attack hardware, including a charger gun plugin device for injecting physical signals and a wireless controller for remote manipulation. By evaluating PORTulator on multiple real-world chargers, we identify 7 charging standards used by 20 charger piles that are vulnerable to our attacks. The root cause is that chargers use simple physical signals for authentication and control, making them easily spoofed by attackers. To address this issue, we propose enhancing authentication circuits by integrating non-resistive memory components and utilizing dynamic high-frequency Pulse Width Modulation (PWM) signals to counter such physical signal spoofing attacks.

cs.CR

Visualizing Nanodomain Superlattices in Halide Perovskites Giving Picosecond Quantum Transients

The high optoelectronic quality of halide perovskites lends them to be utilized in optoelectronic devices and recently in emerging quantum emission applications. Advancements in perovskite nanomaterials have led to the discovery of processes in which luminescence decay times are sub-100 picoseconds, stimulating the exploration of even faster radiative rates for advanced quantum applications, which have only been prominently realised in III-V materials grown through costly epitaxial growth methods. Here, we discovered ultrafast quantum transients of time scales ~2 picoseconds at low temperature in bulk formamidinium lead iodide films grown through scalable solution or vapour approaches. Using a multimodal strategy, combining ultrafast spectroscopy, optical and electron microscopy, we show that these transients originate from quantum tunnelling in nanodomain superlattices. The outcome of the transient decays, photoluminescence, mirrors the photoabsorption of the states, with an ultra-narrow linewidth at low temperature as low as <2 nm (~4 meV). Localized correlation of the emission and structure reveals that the nanodomain superlattices are formed by alternating ordered layers of corner sharing and face sharing octahedra. This discovery opens new applications leveraging intrinsic quantum properties and demonstrates powerful multimodal approaches for quantum investigations.

cond-mat.mtrl-sci

Extending the Defect Tolerance of Halide Perovskite Nanocrystals to Hot Carrier Cooling Dynamics

Defect tolerance is a critical enabling factor for efficient lead-halide perovskite materials, but the current understanding is primarily on band-edge (cold) carriers, with significant debate over whether hot carriers (HCs) can also exhibit defect tolerance. Here, this important gap in the field is addressed by investigating how internationally-introduced traps affect HC relaxation in CsPbX3 nanocrystals (X = Br, I, or mixture). Using femtosecond interband and intraband spectroscopy, along with energy-dependent photoluminescence measurements and kinetic modelling, it is found that HCs are not universally defect tolerant in CsPbX3, but are strongly correlated to the defect tolerance of cold carriers, requiring shallow traps to be present (as in CsPbI3). It is found that HCs are directly captured by traps, instead of going through an intermediate cold carrier, and deeper traps cause faster HC cooling, reducing the effects of the hot phonon bottleneck and Auger reheating. This work provides important insights into how defects influence HCs, which will be important for designing materials for hot carrier solar cells, multiexciton generation, and optical gain media.

physics.app-ph