arXiv Science⌕ Search

arXiv subjects

Ming Sun

Publications and source records attributed to Ming Sun.

At least 91 records · Page 5Linked to original sources

Long working distance portable smartphone microscopy for metallic mesh defect detection

Metallic mesh is a transparent electromagnetic shielding film with a fine metal line structure. However, it can develop defects that affect the optoelectronic performance whether in the production preparation or in actual use. The development of in-situ non-destructive testing (NDT) devices for metallic mesh requires long working distances, reflective optical path design, and miniaturization. To address the limitations of existing smartphone microscopes, which feature short working distances and inadequate transmission imaging for industrial in-situ inspection, we propose a novel long-working distance reflective smartphone microscopy system (LD-RSM). LD-RSM builds a 4f optical imaging system with external optical components and a smartphone, utilizing a beam splitter to achieve reflective imaging with the illumination system and imaging system on the same side of the sample. It achieves an optical resolution of 4.92$μ$m and a working distance of up to 22.23 mm. Additionally, we introduce a dual prior weighted Robust Principal Component Analysis (DW-RPCA) for defect detection. This approach leverages spectral filter fusion and Hough transform to model different defect types, enhancing the accuracy and efficiency of defect identification. Coupled with an optimized threshold segmentation algorithm, DW-RPCA method achieves a pixel-level accuracy of 84.8%. Our work showcases strong potential for growth in the field of in-situ on-line inspection of industrial products.

cs.CV↗

A New Dataset and Framework for Real-World Blurred Images Super-Resolution

Recent Blind Image Super-Resolution (BSR) methods have shown proficiency in general images. However, we find that the efficacy of recent methods obviously diminishes when employed on image data with blur, while image data with intentional blur constitute a substantial proportion of general data. To further investigate and address this issue, we developed a new super-resolution dataset specifically tailored for blur images, named the Real-world Blur-kept Super-Resolution (ReBlurSR) dataset, which consists of nearly 3000 defocus and motion blur image samples with diverse blur sizes and varying blur intensities. Furthermore, we propose a new BSR framework for blur images called Perceptual-Blur-adaptive Super-Resolution (PBaSR), which comprises two main modules: the Cross Disentanglement Module (CDM) and the Cross Fusion Module (CFM). The CDM utilizes a dual-branch parallelism to isolate conflicting blur and general data during optimization. The CFM fuses the well-optimized prior from these distinct domains cost-effectively and efficiently based on model interpolation. By integrating these two modules, PBaSR achieves commendable performance on both general and blur data without any additional inference and deployment cost and is generalizable across multiple model architectures. Rich experiments show that PBaSR achieves state-of-the-art performance across various metrics without incurring extra inference costs. Within the widely adopted LPIPS metrics, PBaSR achieves an improvement range of approximately 0.02-0.10 with diverse anchor methods and blur types, across both the ReBlurSR and multiple common general BSR benchmarks. Code here: https://github.com/Imalne/PBaSR.

cs.CV↗

XPSR: Cross-modal Priors for Diffusion-based Image Super-Resolution

Diffusion-based methods, endowed with a formidable generative prior, have received increasing attention in Image Super-Resolution (ISR) recently. However, as low-resolution (LR) images often undergo severe degradation, it is challenging for ISR models to perceive the semantic and degradation information, resulting in restoration images with incorrect content or unrealistic artifacts. To address these issues, we propose a \textit{Cross-modal Priors for Super-Resolution (XPSR)} framework. Within XPSR, to acquire precise and comprehensive semantic conditions for the diffusion model, cutting-edge Multimodal Large Language Models (MLLMs) are utilized. To facilitate better fusion of cross-modal priors, a \textit{Semantic-Fusion Attention} is raised. To distill semantic-preserved information instead of undesired degradations, a \textit{Degradation-Free Constraint} is attached between LR and its high-resolution (HR) counterpart. Quantitative and qualitative results show that XPSR is capable of generating high-fidelity and high-realism images across synthetic and real-world datasets. Codes are released at \url{https://github.com/qyp2000/XPSR}.

cs.CV↗

Consistency Model is an Effective Posterior Sample Approximation for Diffusion Inverse Solvers

Diffusion Inverse Solvers (DIS) are designed to sample from the conditional distribution $p_θ(X_0|y)$, with a predefined diffusion model $p_θ(X_0)$, an operator $f(\cdot)$, and a measurement $y=f(x'_0)$ derived from an unknown image $x'_0$. Existing DIS estimate the conditional score function by evaluating $f(\cdot)$ with an approximated posterior sample drawn from $p_θ(X_0|X_t)$. However, most prior approximations rely on the posterior means, which may not lie in the support of the image distribution, thereby potentially diverge from the appearance of genuine images. Such out-of-support samples may significantly degrade the performance of the operator $f(\cdot)$, particularly when it is a neural network. In this paper, we introduces a novel approach for posterior approximation that guarantees to generate valid samples within the support of the image distribution, and also enhances the compatibility with neural network-based operators $f(\cdot)$. We first demonstrate that the solution of the Probability Flow Ordinary Differential Equation (PF-ODE) with an initial value $x_t$ yields an effective posterior sample $p_θ(X_0|X_t=x_t)$. Based on this observation, we adopt the Consistency Model (CM), which is distilled from PF-ODE, for posterior sampling. Furthermore, we design a novel family of DIS using only CM. Through extensive experiments, we show that our proposed method for posterior sample approximation substantially enhance the effectiveness of DIS for neural network operators $f(\cdot)$ (e.g., in semantic segmentation). Additionally, our experiments demonstrate the effectiveness of the new CM-based inversion techniques. The source code is provided in the supplementary material.

cs.CV↗

PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild

Video quality assessment (VQA) is a challenging problem due to the numerous factors that can affect the perceptual quality of a video, \eg, content attractiveness, distortion type, motion pattern, and level. However, annotating the Mean opinion score (MOS) for videos is expensive and time-consuming, which limits the scale of VQA datasets, and poses a significant obstacle for deep learning-based methods. In this paper, we propose a VQA method named PTM-VQA, which leverages PreTrained Models to transfer knowledge from models pretrained on various pre-tasks, enabling benefits for VQA from different aspects. Specifically, we extract features of videos from different pretrained models with frozen weights and integrate them to generate representation. Since these models possess various fields of knowledge and are often trained with labels irrelevant to quality, we propose an Intra-Consistency and Inter-Divisibility (ICID) loss to impose constraints on features extracted by multiple pretrained models. The intra-consistency constraint ensures that features extracted by different pretrained models are in the same unified quality-aware latent space, while the inter-divisibility introduces pseudo clusters based on the annotation of samples and tries to separate features of samples from different clusters. Furthermore, with a constantly growing number of pretrained models, it is crucial to determine which models to use and how to use them. To address this problem, we propose an efficient scheme to select suitable candidates. Models with better clustering performance on VQA datasets are chosen to be our candidates. Extensive experiments demonstrate the effectiveness of the proposed method.

cs.CV↗

X-ray Cool Core Remnants Heated by Strong Radio AGN Feedback

Strong AGN heating provides an alternative means for the disruption of cluster cool cores (CCs) to cluster mergers. In this work we present a systematic Chandra study of a sample of 108 nearby ($z<0.1$) galaxy clusters, to investigate the effect of AGN heating on CCs. About 40% of clusters with small offsets between the BCG and the X-ray centre ($\le50$ kpc) have small CCs. For comparison, 14 of 17 clusters with large offsets have small CCs, which suggests that mergers or sloshing can be efficient in reducing the CC size. Relaxed, small CC clusters generally have weak radio AGNs ($P_{1.4\rm GHz}<10^{23}$ W Hz$^{-1}$), and they show a lack of systems hosting a radio AGN with intermediate radio power ($2\times10^{23}<P_{1.4\rm GHz}<2\times10^{24}$ W Hz$^{-1}$). We found that the strongest circumnuclear ($<1$ kpc) X-ray emission only exists in clusters with strong radio AGN. The duty cycle of relaxed, small CC clusters is less than half of that for large CC clusters. It suggests that the radio activity of BCGs is affected by the properties of the surrounding gas beyond the central $\sim10$ kpc, and strong radio AGNs in small X-ray CCs fade more rapidly than those embedded in large X-ray CCs. A scenario is also presented for the transition of large CCs and coronae due to radio AGN feedback. We also present a detailed analysis of galaxy cluster 3C 129.1 as an example of a CC remnant possibly disrupted by radio AGN.

astro-ph.GA↗

NTIRE 2024 Challenge on Short-form UGC Video Quality Assessment: Methods and Results

This paper reviews the NTIRE 2024 Challenge on Shortform UGC Video Quality Assessment (S-UGC VQA), where various excellent solutions are submitted and evaluated on the collected dataset KVQ from popular short-form video platform, i.e., Kuaishou/Kwai Platform. The KVQ database is divided into three parts, including 2926 videos for training, 420 videos for validation, and 854 videos for testing. The purpose is to build new benchmarks and advance the development of S-UGC VQA. The competition had 200 participants and 13 teams submitted valid solutions for the final testing phase. The proposed solutions achieved state-of-the-art performances for S-UGC VQA. The project can be found at https://github.com/lixinustc/KVQChallenge-CVPR-NTIRE2024.

eess.IV↗

CasSR: Activating Image Power for Real-World Image Super-Resolution

The objective of image super-resolution is to generate clean and high-resolution images from degraded versions. Recent advancements in diffusion modeling have led to the emergence of various image super-resolution techniques that leverage pretrained text-to-image (T2I) models. Nevertheless, due to the prevalent severe degradation in low-resolution images and the inherent characteristics of diffusion models, achieving high-fidelity image restoration remains challenging. Existing methods often exhibit issues including semantic loss, artifacts, and the introduction of spurious content not present in the original image. To tackle this challenge, we propose Cascaded diffusion for Super-Resolution, CasSR , a novel method designed to produce highly detailed and realistic images. In particular, we develop a cascaded controllable diffusion model that aims to optimize the extraction of information from low-resolution images. This model generates a preliminary reference image to facilitate initial information extraction and degradation mitigation. Furthermore, we propose a multi-attention mechanism to enhance the T2I model's capability in maximizing the restoration of the original image content. Through a comprehensive blend of qualitative and quantitative analyses, we substantiate the efficacy and superiority of our approach.

cs.CV↗

KVQ: Kwai Video Quality Assessment for Short-form Videos

Short-form UGC video platforms, like Kwai and TikTok, have been an emerging and irreplaceable mainstream media form, thriving on user-friendly engagement, and kaleidoscope creation, etc. However, the advancing content-generation modes, e.g., special effects, and sophisticated processing workflows, e.g., de-artifacts, have introduced significant challenges to recent UGC video quality assessment: (i) the ambiguous contents hinder the identification of quality-determined regions. (ii) the diverse and complicated hybrid distortions are hard to distinguish. To tackle the above challenges and assist in the development of short-form videos, we establish the first large-scale Kaleidoscope short Video database for Quality assessment, termed KVQ, which comprises 600 user-uploaded short videos and 3600 processed videos through the diverse practical processing workflows, including pre-processing, transcoding, and enhancement. Among them, the absolute quality score of each video and partial ranking score among indistinguishable samples are provided by a team of professional researchers specializing in image processing. Based on this database, we propose the first short-form video quality evaluator, i.e., KSVQE, which enables the quality evaluator to identify the quality-determined semantics with the content understanding of large vision language models (i.e., CLIP) and distinguish the distortions with the distortion understanding module. Experimental results have shown the effectiveness of KSVQE on our KVQ database and popular VQA databases.

eess.IV↗

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.

eess.AS↗

FADI-AEC: Fast Score Based Diffusion Model Guided by Far-end Signal for Acoustic Echo Cancellation

Despite the potential of diffusion models in speech enhancement, their deployment in Acoustic Echo Cancellation (AEC) has been restricted. In this paper, we propose DI-AEC, pioneering a diffusion-based stochastic regeneration approach dedicated to AEC. Further, we propose FADI-AEC, fast score-based diffusion AEC framework to save computational demands, making it favorable for edge devices. It stands out by running the score model once per frame, achieving a significant surge in processing efficiency. Apart from that, we introduce a novel noise generation technique where far-end signals are utilized, incorporating both far-end and near-end signals to refine the score model's accuracy. We test our proposed method on the ICASSP2023 Microsoft deep echo cancellation challenge evaluation dataset, where our method outperforms some of the end-to-end methods and other diffusion based echo cancellation methods.

eess.AS↗

Dynamics of a $2$-dimensional slow-fast Belousov-Zabotinsky model

For the reduced two-dimensional Belousov-Zhabotinsky slow-fast differential system, the known results are the existence of one limit cycle and its stability for particular values of the parameters. Here, we characterize all dynamics of this system except one degenerate case. The results include global stability of the positive equilibrium, supercritical and subcritical Hopf bifurcations, the existence of a canard explosion and relaxation oscillation, and the coexistence of one nest of two limit cycles with the outer one originating from the supercritical Hopf bifurcation at one canard point and the inner one from the subcritical Hopf bifurcation at another canard point. This last one is a new dynamical phenomenon.

math.DS↗

Blind Image Super-resolution with Rich Texture-Aware Codebooks

Blind super-resolution (BSR) methods based on high-resolution (HR) reconstruction codebooks have achieved promising results in recent years. However, we find that a codebook based on HR reconstruction may not effectively capture the complex correlations between low-resolution (LR) and HR images. In detail, multiple HR images may produce similar LR versions due to complex blind degradations, causing the HR-dependent only codebooks having limited texture diversity when faced with confusing LR inputs. To alleviate this problem, we propose the Rich Texture-aware Codebook-based Network (RTCNet), which consists of the Degradation-robust Texture Prior Module (DTPM) and the Patch-aware Texture Prior Module (PTPM). DTPM effectively mines the cross-resolution correlation of textures between LR and HR images by exploiting the cross-resolution correspondence of textures. PTPM uses patch-wise semantic pre-training to correct the misperception of texture similarity in the high-level semantic regularization. By taking advantage of this, RTCNet effectively gets rid of the misalignment of confusing textures between HR and LR in the BSR scenarios. Experiments show that RTCNet outperforms state-of-the-art methods on various benchmarks by up to 0.16 ~ 0.46dB.

cs.CV↗

Accelerating Monte Carlo Tree Search with Probability Tree State Abstraction

Monte Carlo Tree Search (MCTS) algorithms such as AlphaGo and MuZero have achieved superhuman performance in many challenging tasks. However, the computational complexity of MCTS-based algorithms is influenced by the size of the search space. To address this issue, we propose a novel probability tree state abstraction (PTSA) algorithm to improve the search efficiency of MCTS. A general tree state abstraction with path transitivity is defined. In addition, the probability tree state abstraction is proposed for fewer mistakes during the aggregation step. Furthermore, the theoretical guarantees of the transitivity and aggregation error bound are justified. To evaluate the effectiveness of the PTSA algorithm, we integrate it with state-of-the-art MCTS-based algorithms, such as Sampled MuZero and Gumbel MuZero. Experimental results on different tasks demonstrate that our method can accelerate the training process of state-of-the-art algorithms with 10%-45% search space reduction.

cs.AI↗

Exploring chemical enrichment of the intracluster medium with the Line Emission Mapper

Synthesized in the cores of stars and supernovae, most metals disperse over cosmic scales and are ultimately deposited well outside the gravitational potential of their host galaxies. Since their presence is well visible through their X-ray emission lines in the hot gas pervading galaxy clusters, measuring metal abundances in the intracluster medium (ICM) offers us a unique view of chemical enrichment of the Universe as a whole. Despite extraordinary progress in the field thanks to four decades of X-ray spectroscopy using CCD (and gratings) instruments, understanding the precise stellar origins of the bulk of metals, and when the latter were mixed on Mpc scales, requires an X-ray mission capable of spatial, non-dispersive high resolution spectroscopy covering at least the soft X-ray band over a large field of view. In this White Paper, we demonstrate how the Line Emission Mapper (LEM) probe mission concept will revolutionize our current picture of the ICM enrichment. Specifically, we show that LEM will be able to (i) spatially map the distribution of ten key chemical elements out to the virial radius of a nearby relaxed cluster and (ii) measure metal abundances in serendipitously discovered high-redshift protoclusters. Altogether, these key observables will allow us to constrain the chemical history of the largest gravitationally bound structures of the Universe. They will also solve key questions such as the universality of the initial mass function (IMF) and the initial metallicity of the stellar populations producing these metals, as well as the relative contribution of asymptotic giant branch (AGB) stars, core-collapse, and Type Ia supernovae to enrich the cosmic web over Mpc scales. Concrete observing strategies are also briefly discussed.

astro-ph.GA↗

The strongest cool core in REXCESS: Missing X-ray cavities in RXC J2014.8-2430

We present a multiwavelength study of RXC J2014.8-2430, the most extreme cool-core cluster in the Representative $XMM-Newton$ Cluster Structure Survey (REXCESS), using $Chandra$ X-ray, Southern Astrophysical Research (SOAR) Telescope, Atacama Large Millimeter/submillimeter Array (ALMA), Very Large Array (VLA), and Giant Metrewave Radio Telescope (GMRT) observations. While feedback from an active galactic nucleus (AGN) is thought to be the dominant mechanism by which a cooling flow is suppressed, the $Chandra$ imaging observations surprisingly do not reveal the bi-lateral X-ray cavities expected in the intracluster medium (ICM) of an extreme cool core hosting a powerful radio source. We discuss the limits on the presence of any radio bubbles associated with any undetected X-ray cavities. We place upper limits on any significant X-ray AGN in the brightest cluster galaxy, and show that the X-ray peak is offset from the central radio source, which exhibits a steep low frequency radio spectrum indicative of electron ageing. The SOAR data reveal an extended, luminous emission line source. From our narrowband H$α$ imaging of the BCG, the central H$α$ peak is coincident with the radio observations, yet offset from the X-ray peak, consistent with sloshing found previously in this cluster. ALMA observations reveal a large reservoir of molecular gas that traces the extended H$α$ emission. We conclude either that the radio source and its cavities in the X-ray gas are nearly aligned along the line of sight, or that ram pressure induced by sloshing has significantly displaced the cool molecular gas feeding it, perhaps preempting the AGN feedback cycle. We argue that the sloshing near the core is likely subsonic, as expected, given the co-location of the H$α$, CO(1-0), radio continuum, and stellar emission peaks and their proximity to the intact cool core seen in X-ray.

astro-ph.HE↗

NuSTAR Observations of Abell 665 and 2146: Constraints on Non-Thermal Emission

Observations from past missions such as RXTE and Beppo-SAX suggested the presence of inverse Compton (IC) scattering at hard X-ray energies within the intracluster medium of some massive galaxy clusters. In subsequent years, observations by, e.g., Suzaku, and now NuSTAR, have not been able to confirm these detections. We report on NuSTAR hard X-ray searches for IC emission in two massive galaxy clusters, Abell 665 and Abell 2146. To constrain the global IC flux in these two clusters, we fit global NuSTAR spectra with three models: single (1T) and two-temperature (2T) models, and a 1T plus power law component (T$+$IC). The temperature components are meant to characterize the thermal ICM emission, while the power law represents the IC emission. We find that the 3-30 keV Abell 665 and 3-20 keV Abell 2146 spectra are best described by thermal emission alone, with average global temperatures of $kT = (9.15\pm 0.1)$ keV for Abell 665 and $kT = (8.29\pm 0.1)$ keV for Abell 2146. We constrain the IC flux to $F_{\rm NT} < 0.60 \times 10^{-12}$ erg s$^{-1}$ cm$^{-2}$ and $F_{\rm NT} < 0.85 \times 10^{-12}$ erg s$^{-1}$ cm$^{-2}$ (20-80 keV) for Abell 665 and Abell 2146, respectively both at the 90% confidence level. When we couple the IC flux limits with 1.4 GHz diffuse radio data from the VLA, we set lower limits on the average magnetic field strengths of $>$0.14 $μ$G and $>$0.011 $μ$G for Abell 665 and Abell 2146, respectively.

astro-ph.HE↗

Ada-DQA: Adaptive Diverse Quality-aware Feature Acquisition for Video Quality Assessment

Video quality assessment (VQA) has attracted growing attention in recent years. While the great expense of annotating large-scale VQA datasets has become the main obstacle for current deep-learning methods. To surmount the constraint of insufficient training data, in this paper, we first consider the complete range of video distribution diversity (\ie content, distortion, motion) and employ diverse pretrained models (\eg architecture, pretext task, pre-training dataset) to benefit quality representation. An Adaptive Diverse Quality-aware feature Acquisition (Ada-DQA) framework is proposed to capture desired quality-related features generated by these frozen pretrained models. By leveraging the Quality-aware Acquisition Module (QAM), the framework is able to extract more essential and relevant features to represent quality. Finally, the learned quality representation is utilized as supplementary supervisory information, along with the supervision of the labeled quality score, to guide the training of a relatively lightweight VQA model in a knowledge distillation manner, which largely reduces the computational cost during inference. Experimental results on three mainstream no-reference VQA benchmarks clearly show the superior performance of Ada-DQA in comparison with current state-of-the-art approaches without using extra training data of VQA.

cs.CV↗