arXiv ScienceSearch

arXiv subjects

Yiting Zhao

Publications and source records attributed to Yiting Zhao.

10 recordsLinked to original sources

MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting

High-resolution volumetric medical imaging is critical for clinical diagnosis, yet acquisition is often limited by scanner hardware, scan time, and for CT, radiation dose. Medical 3D Super-Resolution (Med3DSR) offers a computational alternative, but existing methods commonly rely on per-subject optimization, pretrained priors, or coordinate-based implicit representations, which compromise anatomical fidelity and limit efficiency. To address these limitations, we present MedGSSR, a fully end-to-end feed-forward framework that represents volumes as an explicit 3D Gaussian field for Med3DSR. Unlike coordinate-based implicit functions, our explicit 3D Gaussian representation naturally enhances signal continuity and local high-frequency fidelity. Specifically, MedGSSR explicitly decouples the reconstruction process into coarse-grained structural preservation and fine-grained textural refinement through the proposed Pyramid Anatomical Encoder and a Hierarchical Gaussian Projector. To support arbitrary-scale super-resolution, we introduce sub-voxel Gaussian decomposition and a Differentiable Gaussian Voxelizer that directly queries the continuous 3D intensity field, reducing discretization artifacts. Extensive experiments on MRI and CT benchmarks demonstrate that MedGSSR significantly outperforms state-of-the-art methods. Notably, our framework exhibits robust generalizability across unseen datasets without requiring per-subject optimization, enabling fast inference and high-fidelity volumetric super-resolution in practical clinical settings. Our project webpage, including code, is at https://william2ai.github.io/medgssr

eess.IV

Dual-polarized, mid-infrared nonreciprocal absorption

The emission and absorption of thermal radiation are usually coupled via Kirchhoff's law or reciprocity, stated as the equality of spectral directional emissivity and absorptivity. Magneto-optical materials have recently been identified as a promising route to lifting the constraint of reciprocity, with multiple experimental demonstrations using doped InAs. However, these demonstrations have been limited to p-polarized light in the Voigt configuration, whereas thermal radiation from a blackbody is unpolarized. Therefore, to break reciprocity in both polarization channels, we design a nanophotonic, dual-polarized nonreciprocal absorber operating in the mid-infrared spectral range (11-20 $\unicode{x03BC}$m), consisting of an a-Si photonic crystal slab on top of a doped InAs substrate described by an antisymmetric, nonreciprocal dielectric tensor under an applied magnetic field. The photonic crystal slab supports eigenmodes that couple to both s- and p-polarized light, resulting in absorption peaks that frequency shift in opposite directions for forward- and backward-propagating light$\unicode{x2014}$a signature of nonreciprocity in planar, subwavelength systems. We fabricate our design, then measure its room-temperature absorptance using magnetic-field-integrated absorptance spectroscopy, experimentally demonstrating nonreciprocal absorption for both polarizations. Our design is a step toward the complete control of light as heat, which could improve photonic energy conversion, thermal management, and mid-infrared optical isolation and circulation.

physics.optics

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).

cs.AI

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the multifaceted nature of complex memory systems, such as dynamic state tracking and hierarchical reasoning in continuous interactions. To overcome these limitations, we propose MemGround, a rigorous long-term memory benchmark natively grounded in rich, gamified interactive scenarios. To systematically assess these capabilities, MemGround introduces a three-tier hierarchical framework that evaluates Surface State Memory, Temporal Associative Memory, and Reasoning-Based Memory through specialized interactive tasks. Furthermore, to comprehensively quantify both memory utilization and behavioral trajectories, we propose a multi-dimensional metric suite comprising Question-Answer Score (QA Overall), Memory Fragments Unlocked (MFU), Memory Fragments with Correct Order (MFCO), and Exploration Trajectory Diagrams (ETD). Extensive experiments reveal that state-of-the-art LLMs and memory agents still struggle with sustained dynamic tracking, temporal event association, and complex reasoning derived from long-term accumulated evidence in interactive environments.

cs.CL

SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian Splatting

3D super-resolution (3DSR) aims to reconstruct high-resolution (HR) 3D scenes from low-resolution (LR) multi-view images. Existing methods rely on dense LR inputs and per-scene optimization, which restricts the high-frequency priors for constructing HR 3D Gaussian Splatting (3DGS) to those inherited from pretrained 2D super-resolution (2DSR) models. This severely limits reconstruction fidelity, cross-scene generalization, and real-time usability. We propose to reformulate 3DSR as a direct feed-forward mapping from sparse LR views to HR 3DGS representations, enabling the model to autonomously learn 3D-specific high-frequency geometry and appearance from large-scale, multi-scene data. This fundamentally changes how 3DSR acquires high-frequency knowledge and enables robust generalization to unseen scenes. Specifically, we introduce SR3R, a feed-forward framework that directly predicts HR 3DGS representations from sparse LR views via the learned mapping network. To further enhance reconstruction fidelity, we introduce Gaussian offset learning and feature refinement, which stabilize reconstruction and sharpen high-frequency details. SR3R is plug-and-play and can be paired with any feed-forward 3DGS reconstruction backbone: the backbone provides an LR 3DGS scaffold, and SR3R upscales it to an HR 3DGS. Extensive experiments across three 3D benchmarks demonstrate that SR3R surpasses state-of-the-art (SOTA) 3DSR methods and achieves strong zero-shot generalization, even outperforming SOTA per-scene optimization methods on unseen scenes.

cs.CV

An Overall Real-Time Mechanism for Classification and Quality Evaluation of Rice

Rice is one of the most widely cultivated crops globally and has been developed into numerous varieties. The quality of rice during cultivation is primarily determined by its cultivar and characteristics. Traditionally, rice classification and quality assessment rely on manual visual inspection, a process that is both time-consuming and prone to errors. However, with advancements in machine vision technology, automating rice classification and quality evaluation based on its cultivar and characteristics has become increasingly feasible, enhancing both accuracy and efficiency. This study proposes a real-time evaluation mechanism for comprehensive rice grain assessment, integrating a one-stage object detection approach, a deep convolutional neural network, and traditional machine learning techniques. The proposed framework enables rice variety identification, grain completeness grading, and grain chalkiness evaluation. The rice grain dataset used in this study comprises approximately 20,000 images from six widely cultivated rice varieties in China. Experimental results demonstrate that the proposed mechanism achieves a mean average precision (mAP) of 99.14% in the object detection task and an accuracy of 97.89% in the classification task. Furthermore, the framework attains an average accuracy of 97.56% in grain completeness grading within the same rice variety, contributing to an effective quality evaluation system.

cs.CV

AirFogSim: A Light-Weight and Modular Simulator for UAV-Integrated Vehicular Fog Computing

Vehicular Fog Computing (VFC) is significantly enhancing the efficiency, safety, and computational capabilities of Intelligent Transportation Systems (ITS), and the integration of Unmanned Aerial Vehicles (UAVs) further elevates these advantages by incorporating flexible and auxiliary services. This evolving UAV-integrated VFC paradigm opens new doors while presenting unique complexities within the cooperative computation framework. Foremost among the challenges, modeling the intricate dynamics of aerial-ground interactive computing networks is a significant endeavor, and the absence of a comprehensive and flexible simulation platform may impede the exploration of this field. Inspired by the pressing need for a versatile tool, this paper provides a lightweight and modular aerial-ground collaborative simulation platform, termed AirFogSim. We present the design and implementation of AirFogSim, and demonstrate its versatility with five key missions in the domain of UAV-integrated VFC. A multifaceted use case is carried out to validate AirFogSim's effectiveness, encompassing several integral aspects of the proposed AirFogSim, including UAV trajectory, task offloading, resource allocation, and blockchain. In general, AirFogSim is envisioned to set a new precedent in the UAV-integrated VFC simulation, bridge the gap between theoretical design and practical validation, and pave the way for future intelligent transportation domains. Our code will be available at https://github.com/ZhiweiWei-NAMI/AirFogSim.

cs.NI

Cooper minimum of high-order harmonic spectra from MgO crystal in an ultrashort laser pulse

Cooper minimum structure of high-order harmonic spectra from atoms or molecules has been extensively studied. In this paper, we demonstrate that the crystal harmonic spectra from an ultrashort mid-infrared laser pulse also exhibit the Cooper minimum characteristic. Based on the accurate band dispersion and k-dependent transition dipole moment (TDM) from the first-principle calculations, it can be found that the harmonic spectra from MgO crystal along {\Gamma}-X direction present a dip structure in the plateau, which is originated from the valley of TDM by examining the distribution of the harmonic intensity at the k-space. The Cooper minimum feature in crystal HHG will pave a new way to retrieve the band information of solid materials by using HHG from the ultrashort mid-infrared laser pulse.

physics.atm-clus

Graph Computing based Fast Screening in Contingency Analysis

During last decades, contingency analysis has been facing challenges from significant load demand increase and high penetrations of intermittent renewable energy, fluctuant responsive loads and non-linear power electronic interfaces. It requires an advanced approach for high-performance contingency analysis as a safeguard of the power system operation. In this paper, a graph-based method is employed for N-1 contingency analysis (CA) fast screening. At first, bi-directional breadth-first search (BFS) is proposed and adopted on graph model to detect the potential shedding component in contingency analysis. It implements hierarchical parallelism of the graph traverse and speedup its process. Then, the idea of evolving graph is introduced in this paper to improve computation performance. For each contingency scenario, N-1 contingency graph quickly derives from system graph in basic status, and parallelly analyzes each contingency scenario using graph computing. The efficiency and effectiveness of the proposed approach have been tested and verified by IEEE 118-bus system and a practical case SC 2645-bus system.

cs.DC

Graph-based Preconditioning Conjugate Gradient Algorithm for N-1 Contingency Analysis

Contingency analysis (CA) plays a critical role to guarantee operation security in the modern power systems. With the high penetration of renewable energy, a real-time and comprehensive N-1 CA is needed as a power system analysis tool to ensure system security. In this paper, a graph-based preconditioning conjugate gradient (GPCG) approach is proposed for the nodal parallel computing in N-1 CA. To pursue a higher performance in the practical application, the coefficient matrix of the base case is used as the incomplete LU (ILU) preconditioner for each N-1 scenario. Additionally, the re-dispatch strategy is employed to handle the islanding issues in CA. Finally, computation performance of the proposed GPCG approach is tested on a real provincial system in China.

cs.DC