arXiv ScienceSearch

arXiv subjects

Han Zhou

Publications and source records attributed to Han Zhou.

At least 19 recordsLinked to original sources

A Conditional Denoising Diffusion Probabilistic Model for RFI Mitigation in Synthetic Aperture Interferometric Radiometer

In Earth remote sensing, spatial-frequency domain visibility samples are inversely transformed into spatial-domain brightness temperature (BT) images through the signal processing pipeline of synthetic aperture interferometric radiometers (SAIR). However, L-band radio-frequency interference (RFI) contaminates the measured visibilities and severely degrades BT image quality, thereby impairing geophysical parameter retrieval. To address this issue, we propose VFDM, a Visibility-Function Diffusion Model based on Denoising Diffusion Probabilistic Models (DDPM), to mitigate RFI in the spatial-frequency domain while preserving fine-scale structures consistent with natural scene statistics. Furthermore, we construct a comprehensive dataset comprising more than ten thousand pairs of RFI-free natural scene visibility sample sets and their corresponding simulated contaminated counterparts, categorized by varying RFI intensities, numbers, and distributions. Finally, comprehensive experiments on both simulated and real-world data demonstrate the effectiveness and robustness of the proposed VFDM-based approach.

eess.SP

Open Inextensible Filaments in Planar Stokes Flow: Well-Posedness, Endpoint Asymptotics, and Straightening

We study an inextensible open filament with free ends in a planar Stokes fluid. The system reduces to a third-order nonlocal curvature equation coupled to an elliptic equation for the tension. We prove local well-posedness for nearly critical initial data in supported Sobolev spaces $\widetilde H^s$, $-1/2<s\le0$, satisfying the arc-chord condition. For positive times, we prove improved Sobolev regularity and derive a $d^{3/2}$-type expansion at each free end using Wiener--Hopf factorization, where $d$ denotes the distance to that endpoint. We further prove global existence and exponential convergence to a straight filament for sufficiently small initial data and for finite-energy initial data satisfying $E(0)<π^2/4$. The finite-energy result follows from an energy identity and a geometric estimate relating the bending energy to the arc-chord constant. More generally, any finite-time breakdown must be accompanied by loss of the arc-chord condition, while every global solution either converges exponentially to a straight filament or has arc-chord constants tending to zero along a sequence of times tending to infinity.

math.AP

EvoGS: Modeling Deformation Evolution for Dynamic Gaussian Splatting

Recent extensions of 3D Gaussian Splatting (3DGS) enable real-time novel view synthesis in dynamic scenes by learning time-conditioned Gaussian deformations. However, existing MLP-based methods typically estimate deformations independently at each timestamp, making them less robust to large or abrupt motions. To address this issue, we propose \textbf{EvoGS}, a 3DGS-based dynamic reconstruction framework that models Gaussian deformation as a temporal evolution process. EvoGS maintains persistent deformation states for each Gaussian, extrapolates future states from historical deformation states, and corrects the predictions with MLP-derived observations. The correction is adaptively weighted using a temporal residual memory and evolution statistics such as deformation velocity and trajectory deviation. To further improve reconstruction quality, EvoGS introduces deformation-aware densification. Clone and split operations are performed along corrected deformation directions, while an uncertainty-aware strategy suppresses densification for Gaussians with unstable deformation histories. Experiments show that EvoGS improves dynamic novel view synthesis quality and achieves competitive performance across benchmarks.

cs.CV

An Ultra-Compact Differential V-Band Power Amplifier Using EDMOS Transistors With 18.1 dBm P1dB and 21% PAE in 22nm FD-SOI CMOS

This paper presents a compact, fully differential, two-stage millimeter-wave (mm-wave) cascode power amplifier (PA) designed and implemented in a 22nm FD-SOI CMOS process (22FDX+). The PA employs the newly introduced extended-drain MOS (EDMOS) device in 22FDX+, together with a carefully engineered device core and transformer baluns. At 50 GHz, the prototype achieves 18.8 dBm saturated output power (PSAT), 18.1 dBm 1-dB compression output power P1dB, and 21% power-added efficiency (PAE) at P1dB. To the best of our knowledge, this work achieves the highest reported power density of 2.6 W/mm2 among single-way, two-stage CMOS cascode PAs.

eess.SP

An Efficient W-/D-Band Power Amplifier in a 130 nm SiGe BiCMOS Process

This paper presents a wideband power amplifier (PA) designed and implemented in Infineon Technologies' 130-nm SiGe BiCMOS process for upper W-band and lower D-band applications. A complete load-pull simulation methodology is carried out, and a band pass filter (BPF)-based matching strategy is employed for the design of the output and inter-stage matching networks. The fabricated PA prototype achieves a small-signal gain 3-dB bandwidth of 71-133 GHz. Moreover, it maintains a relatively flat gain of approximately 15.7 dB over 75-128 GHz, with less than 1-dB fluctuation. The measured saturated output power is 8.8-11.7 dBm, while the measured peak power-added efficiency (PAE) is 7.2-11.1%. These results demonstrate the potential of SiGe BiCMOS technology for wideband and integrated transmitter front ends operating across the W-/D-band frequency range.

eess.SP

Deep Learning-Driven Inverse Design of Doherty Power Amplifiers Using Pixelated Combiners and Dual-State Impedance Synthesis

The output combiner of a Doherty power amplifier (PA) integrates load modulation, impedance matching, and phase compensation within a single network, making its design and synthesis highly challenging. In this paper, we propose a three-port Doherty combiner design methodology that combines deep convolutional neural networks (CNNs), pixelated layout representations, and genetic algorithms (GA) with dual-state impedance synthesis to address both peak and back-off power conditions. As a proof of concept, two GaN HEMT Doherty PA prototypes incorporating three-port pixelated combiners are designed and fabricated. Both prototypes achieve a measured saturated output power exceeding 44.2 dBm with peak drain efficiency above 71.2% within 2.6-2.8 GHz. Furthermore, a drain efficiency as high as 64% is measured at the 6-dB back-off level. After applying digital predistortion, each prototype achieves an adjacent channel leakage ratio (ACLR) better than -51.3 dBc.

eess.SP

Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements

Traditional microwave filter design typically relies on iterative parameter tuning and predefined topologies, which limits design space and increases development time. This study uses a deep learning approach combining convolutional neural networks with genetic algorithms to automate pixelated microwave filter synthesis. To validate the approach experimentally, both S-parameter and spatial electric-field measurements were analyzed. The synthesized low-pass filter demonstrated excellent agreement between simulated and measured performance, achieving a 7 GHz passband with over 20 dB suppression beyond 9.5 GHz. Electro-optical measurements, for the first time, revealed electric field patterns that resemble coupled transmission-lines or stub structures, providing insight into the emergent characteristics of AI-generated designs.

eess.SP

Bulk-surface coupled PDE with an open boundary

We study a bulk-surface coupled Laplace system involving an embedded open boundary. The problem is reformulated as an integro-differential equation using boundary integral representations, for which we establish existence and uniqueness of the solution. A Wiener-Hopf technique is employed to study the solution regularity and derive asymptotic expressions for the edge singularity. Building on these results, we develop a finite element method that incorporates the singularity structure and provide a rigorous error analysis. Numerical experiments confirm the theoretical convergence rates.

math.NA

Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

cs.CL

Leveraging AI for fine-grained food safety risk forecasting in sparse data conditions

Ensuring food safety represents a critical public health challenge, particularly when inspection resources are limited and regional sampling data are sparse. This study proposes a Transformer-based framework capable of forecasting fine-grained, city-level food safety risks by unifying over 11 million inspection records with supplemental demographic, economic, and environmental indicators extracted from the Statistical Yearbook. A three-stage pretraining design leverages partial supervision from the Wilson interval (capturing both safety and risk rankings), together with semi-supervised label refinement, to effectively utilize historical records even when local sample sizes are insufficient. Experimental evaluations on data from 2022 show that the proposed approach outperforms baselines significantly. A subsequent field experiment in collaboration with the Zhejiang Provincial Administration for Market Regulation further demonstrates improved detection rates and more efficient allocation of inspection resources compared to a manually developed plan. Observations of regulatory decision-making reveal a threshold-based heuristic employed by inspectors, hinting that additional training or decision-support interfaces could further enhance the impact of AI-generated risk scores. Overall, these findings underscore that a rigorous integration of large-scale public inspection data, Wilson interval-based confidence modeling, and advanced deep learning can facilitate earlier and more granular identification of food safety threats. By reducing reliance on reactive measures alone, the proposed framework has the potential to advance proactive, data-driven oversight of the global food supply.

cs.AI

A correction function-based kernel-free boundary integral method for elliptic PDEs with implicitly defined interfaces

This work addresses a novel version of the kernel-free boundary integral (KFBI) method for solving elliptic PDEs with implicitly defined irregular boundaries and interfaces. We focus on boundary value problems and interface problems, which are reformulated into boundary integral equations and solved with the matrix-free GMRES method. In the KFBI method, evaluating boundary and volume integrals only requires solving equivalent but much simpler interface problems in a bounding box, for which fast solvers such as FFTs and geometric multigrid methods are applicable. For the simple interface problem, a correction function is introduced for both the evaluation of right-hand side correction terms and the interpolation of a non-smooth potential function. A mesh-free collocation method is proposed to compute the correction function near the interface. The new method avoids complicated derivation for derivative jumps of the solution and is easy to implement, especially for the fourth-order method in three space dimensions. Various numerical examples are presented, including challenging cases such as high-contrast coefficients, arbitrarily close interfaces and heterogeneous interface problems. The reported numerical results verify that the proposed method is both accurate and efficient.

math.NA

Robostral Navigate

Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view. Operating purely in image space, rather than robot-specific coordinates, makes the policy naturally robust to changes in camera intrinsics and scene scale, enabling deployment across wheeled, legged, and aerial robots without recalibration. We generate 2.4 million trajectories across 350k simulated scenes to reduce the reliance on real-world data collection and scale easily. We further introduce a prefix-caching training recipe that packs entire episodes into single training sequences, reducing training tokens by 22x and cutting training time from months to days. A tree-based attention mask prevents conditioning on previous ground-truth actions, encouraging visually grounded action prediction, and reinforcement learning is used to further improve exploration and recovery capabilities. On the Room-to-Room and Room-Across-Room in Continuous Environments (R2R-CE and RxR-CE) benchmarks, Robostral Navigate sets a new state of the art. On R2R-CE, it achieves a 77.4% success rate, surpassing the best monocular method by 10.5 points and the strongest depth- or multi-camera system by 5.3 points despite using only a single RGB camera. On RxR-CE, it reaches 75.1% success rate, outperforming all monocular baselines.

cs.RO

A Cartesian Grid Method for Advection-Diffusion Equations with Robin Boundary Conditions on Moving Domains

We develop a Cartesian grid method for advection--diffusion equations with Robin boundary conditions on moving domains. The moving-domain problem is reformulated as an interface problem on a box, with an unknown density introduced on the moving interface to enforce the Robin condition. The bulk equation is discretized by a cell-centered finite-difference scheme on the Cartesian grid, while interface corrections are obtained from local problems in a narrow band around the interface. The resulting method requires only modest computational geometry, avoids remeshing and cut cells, and is compatible with geometric multigrid and matrix-free GMRES. The GMRES iteration count is essentially independent of the mesh size, and the computational cost scales linearly with the number of bulk degrees of freedom. For the one-dimensional scheme, first-order convergence in time and second-order convergence in space are proved. Numerical examples in one and two dimensions, including manufactured solutions and an active transport problem without an exact solution, demonstrate the accuracy and efficiency of the method.

math.NA

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

Multimodal large language models (MLLMs) require huge memory and computational costs, which limits their practical deployment. Post-training quantization (PTQ) techniques offer an efficient solution for model compression and inference acceleration. Yet, the quantized model faces performance degradation due to outlier channels, which are highly sensitive to quantization and substantially impair activation fidelity and task accuracy. To protect these salient channels during quantization, existing PTQ methods leverage modality- or token-level metrics to guide channel-wise scaling (CWS) of LLM decoders. However, these orthogonal measurements fail to capture channel-wise impacts on task-specific loss, and the misalignment between importance and scaling factors ultimately leads to suboptimal performance. To address this issue, we propose C-PTQ, a unified channel-wise PTQ method that harmonizes task-specific loss perturbation and quantization error. Motivated by second-order derivatives, we design a Fisher-weighted objective as a tractable Hessian approximation, seamlessly injecting task sensitivity into the scaling process. Notably, we achieve state-of-the-art performance without auxiliary modules like LoRA, thereby maintaining high efficiency. Experiments on Qwen2.5VL, InternVL2 and LLaVA-OV across 8 benchmarks demonstrate our effectiveness in both weight-only and weight-activation settings.

cs.CV

ClawBench: Can AI Agents Complete Everyday Online Tasks?

AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework comprising 153 everyday online tasks that people need to accomplish regularly in their lives and work, spanning 144 platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and interaction challenges of real-world web environments. An interception layer captures and blocks the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 8 frontier models show that both proprietary and open-source models complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%, which exposes gaps in current AI agents. Progress on ClawBench brings us closer to AI agents that can function as general-purpose assistants.

cs.CL

Interactive Training 2: Auditable Control Plane for Live Model Training

Experiment trackers show how training is progressing, but changing a live run still usually requires trainer-specific code. We present Interactive Training 2, an open-source control plane for steering training through a shared protocol. Training applications declare which settings and actions they expose, humans and automated controllers submit requests through the same interface, and the training loop validates and applies them at safe control points. A customized Aim workspace combines live metrics and controls with a chronological record of requests and outcomes. We demonstrate the system across five NLP and reinforcement-learning workflows. The released code and traces provide a reusable foundation for auditable human- and agent-guided training.

cs.LG

GlobalForge: Towards Robust AI-Generated Image Detection

AI-generated image (AIGI) detectors achieve strong accuracy on clean benchmarks, but their performance drops sharply after images are propagated through real-world channels. We trace this fragility to what these detectors actually learn: they overfit to local artifacts left by generators in small spatial neighborhoods, which are easily destroyed by common propagation degradations such as JPEG compression and blur. Instead, we shift the discriminative cue from fragile local artifacts to more robust global structure. Building on this, we propose GlobalForge, a framework with two complementary modules. The Local Information Bottleneck (LIB) suppresses local components to block shortcut learning, while the Global Structural Reasoning (GSR) module forces every token to gather evidence from distant regions. Both modules are trained jointly under a contrastive structural loss based on degradation that keeps the resulting features stable under degradation. To support fine-grained robustness evaluation, we further introduce RealDeg-Bench, covering 7 common degradation operators and multi-step compound chains. GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by $\mathbf{5.89\%}$ over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations. Code is available at https://anonymous.4open.science/r/GlobalForge-BE0F/.

cs.CV

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection

AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In practice, however, video backbones with standard global readouts often fail to outperform strong image pretrained probes on AIGV benchmarks. We attribute this gap to excessive spatiotemporal aggregation in the readout. Video pretrained backbones tend to compress each frame into a single global descriptor. This compression suppresses local patch level temporal dynamics and discards inter patch relations, which are precisely the cues that AIGV detection most reliably depends on. Based on this, we propose Velocity Gated Patch Velocity Profiling (V-PVP), a lightweight readout that replaces only the aggregation layer with two parallel streams over the patch velocity field, adding only about $0.5$M trainable parameters. V-PVP serves as a general plug-and-play module that consistently improves performance across diverse video backbones under both end-to-end fine-tuning and linear probing settings. Our method reaches \textbf{95.28} AUC on AIGVDBench while keeping the backbone fully frozen. The results show that simply replacing the aggregation layer reactivates the temporal potential of frozen video backbones, restoring their advantage on AIGV detection. Code is available at https://anonymous.4open.science/r/PVP-81B3/.

cs.CV