arXiv ScienceSearch

arXiv subjects

Ling Fu

Publications and source records attributed to Ling Fu.

13 recordsLinked to original sources

CycleGuardian: A Framework for Automatic RespiratorySound classification Based on Improved Deep clustering and Contrastive Learning

Auscultation plays a pivotal role in early respiratory and pulmonary disease diagnosis. Despite the emergence of deep learning-based methods for automatic respiratory sound classification post-Covid-19, limited datasets impede performance enhancement. Distinguishing between normal and abnormal respiratory sounds poses challenges due to the coexistence of normal respiratory components and noise components in both types. Moreover, different abnormal respiratory sounds exhibit similar anomalous features, hindering their differentiation. Besides, existing state-of-the-art models suffer from excessive parameter size, impeding deployment on resource-constrained mobile platforms. To address these issues, we design a lightweight network CycleGuardian and propose a framework based on an improved deep clustering and contrastive learning. We first generate a hybrid spectrogram for feature diversity and grouping spectrograms to facilitating intermittent abnormal sound capture.Then, CycleGuardian integrates a deep clustering module with a similarity-constrained clustering component to improve the ability to capture abnormal features and a contrastive learning module with group mixing for enhanced abnormal feature discernment. Multi-objective optimization enhances overall performance during training. In experiments we use the ICBHI2017 dataset, following the official split method and without any pre-trained weights, our method achieves Sp: 82.06 $\%$, Se: 44.47$\%$, and Score: 63.26$\%$ with a network model size of 38M, comparing to the current model, our method leads by nearly 7$\%$, achieving the current best performances. Additionally, we deploy the network on Android devices, showcasing a comprehensive intelligent respiratory sound auscultation system.

cs.SD

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks (4x more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios (31 diverse scenarios), and thorough evaluation metrics, with 10,000 human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with 1,500 manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below 50 (100 in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The project website is at: https://99franklin.github.io/ocrbench_v2/

cs.CV

SceneGenAgent: Precise Industrial Scene Generation with Coding Agent

The modeling of industrial scenes is essential for simulations in industrial manufacturing. While large language models (LLMs) have shown significant progress in generating general 3D scenes from textual descriptions, generating industrial scenes with LLMs poses a unique challenge due to their demand for precise measurements and positioning, requiring complex planning over spatial arrangement. To address this challenge, we introduce SceneGenAgent, an LLM-based agent for generating industrial scenes through C# code. SceneGenAgent ensures precise layout planning through a structured and calculable format, layout verification, and iterative refinement to meet the quantitative requirements of industrial scenarios. Experiment results demonstrate that LLMs powered by SceneGenAgent exceed their original performance, reaching up to 81.0% success rate in real-world industrial scene generation tasks and effectively meeting most scene generation requirements. To further enhance accessibility, we construct SceneInstruct, a dataset designed for fine-tuning open-source LLMs to integrate into SceneGenAgent. Experiments show that fine-tuning open-source LLMs on SceneInstruct yields significant performance improvements, with Llama3.1-70B approaching the capabilities of GPT-4o. Our code and data are available at https://github.com/THUDM/SceneGenAgent .

cs.CL

Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, facilitating better understanding and interaction with Urdu-language visual data. This initiative seeks to bridge the gap between textual and visual comprehension. We propose a new multi-task Urdu scene text dataset comprising over 1000 natural scene images, which can be used for text detection, recognition, and VQA tasks. We provide fine-grained annotations for text instances, addressing the limitations of previous datasets for facing arbitrary-shaped texts. By incorporating additional annotation points, this dataset facilitates the development and assessment of methods that can handle diverse text layouts, intricate shapes, and non-standard orientations commonly encountered in real-world scenarios. Besides, the VQA annotations make it the first benchmark for the Urdu Text VQA method, which can prompt the development of Urdu scene text understanding. The proposed dataset is available at: https://github.com/Hiba-MeiRuan/Urdu-VQA-Dataset-/tree/main

cs.CV

The First Swahili Language Scene Text Detection and Recognition Dataset

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has been done to improve the accuracy and performance of scene text detection and recognition models. However, most of this research has been conducted in the most common languages, English and Chinese. There is a significant gap in low-resource languages, especially the Swahili Language. Swahili is widely spoken in East African countries but is still an under-explored language in scene text recognition. No studies have been focused explicitly on Swahili natural scene text detection and recognition, and no dataset for Swahili language scene text detection and recognition is publicly available. We propose a comprehensive dataset of Swahili scene text images and evaluate the dataset on different scene text detection and recognition models. The dataset contains 976 images collected in different places and under various circumstances. Each image has its annotation at the word level. The proposed dataset can also serve as a benchmark dataset specific to the Swahili language for evaluating and comparing different approaches and fostering future research endeavors. The dataset is available on GitHub via this link: https://github.com/FadilaW/Swahili-STR-Dataset

cs.CV

Enhancing Scene Text Detectors with Realistic Text Image Synthesis Using Diffusion Models

Scene text detection techniques have garnered significant attention due to their wide-ranging applications. However, existing methods have a high demand for training data, and obtaining accurate human annotations is labor-intensive and time-consuming. As a solution, researchers have widely adopted synthetic text images as a complementary resource to real text images during pre-training. Yet there is still room for synthetic datasets to enhance the performance of scene text detectors. We contend that one main limitation of existing generation methods is the insufficient integration of foreground text with the background. To alleviate this problem, we present the Diffusion Model based Text Generator (DiffText), a pipeline that utilizes the diffusion model to seamlessly blend foreground text regions with the background's intrinsic features. Additionally, we propose two strategies to generate visually coherent text with fewer spelling errors. With fewer text instances, our produced text images consistently surpass other synthetic data in aiding text detectors. Extensive experiments on detecting horizontal, rotated, curved, and line-level texts demonstrate the effectiveness of DiffText in producing realistic text images.

cs.CV

Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition

Artistic text recognition is an extremely challenging task with a wide range of applications. However, current scene text recognition methods mainly focus on irregular text while have not explored artistic text specifically. The challenges of artistic text recognition include the various appearance with special-designed fonts and effects, the complex connections and overlaps between characters, and the severe interference from background patterns. To alleviate these problems, we propose to recognize the artistic text at three levels. Firstly, corner points are applied to guide the extraction of local features inside characters, considering the robustness of corner structures to appearance and shape. In this way, the discreteness of the corner points cuts off the connection between characters, and the sparsity of them improves the robustness for background interference. Secondly, we design a character contrastive loss to model the character-level feature, improving the feature representation for character classification. Thirdly, we utilize Transformer to learn the global feature on image-level and model the global relationship of the corner points, with the assistance of a corner-query cross-attention mechanism. Besides, we provide an artistic text dataset to benchmark the performance. Experimental results verify the significant superiority of our proposed method on artistic text recognition and also achieve state-of-the-art performance on several blurred and perspective datasets.

cs.CV

Lattice-Driven Chiral Charge Density Wave State in 1T-TaS$_{2}$

We use scanning tunneling microscopy to study the domain structure of the nearly-commensurate charge density wave (NC-CDW) state of 1T-TaS$_2$. In our sub-angstrom characterization of the state, we find a continual evolution of the CDW lattice from domain wall to domain center, instead of a fixed CDW arrangement within a domain. Further, we uncover an intradomain chirality characterizing the NC-CDW state. Unlike the orbital-driven chirality previously observed in related transition metal dichalcogenides, the chiral nature of the NC-CDW state in 1T-TaS$_2$ appears driven by a strong coupling of the NC-CDW state to the lattice.

cond-mat.str-el

The tipping times in an Arctic sea ice system under influence of extreme events

In light of the rapid recent retreat of Arctic sea ice, the extreme weather events triggering the variability in Arctic ice cover has drawn increasing attention. A non-Gaussian $\alpha$-stable L\'evy process is thought to be an appropriate model to describe such extreme event. The maximal likely trajectory, based on the nonlocal Fokker-Planck equation, is applied to a nonautonomous Arctic sea ice system under $\alpha$-stable L\'evy noise. Two types of tipping times, the early-warning tipping time and the disaster-happening tipping time, are used to predict the critical time for the maximal likely transition from a perennially ice-covered state to a seasonally ice-free one, and from a seasonally ice-free state to a perennially ice-free one, respectively. We find that the increased intensity of extreme events results in shorter warning time for sea ice melting, and that an enhanced greenhouse effect will intensify this influence, making the arrival of warning time significantly earlier. Meanwhile, for the enhanced greenhouse effect, we discover that increased intensity and frequency of extreme events will advance the disaster-happening tipping time, in which an ice-free state is maintained throughout the year in the Arctic Ocean. Finally, we identify values of L\'evy index $\alpha$ and noise intensity $\epsilon$ in $\alpha \epsilon$-space that can trigger a transition between the Arctic sea ice state. These results provide an effective theoretical framework for studying Arctic sea ice variations under the influence of extreme events.

physics.ao-ph

The maximum likelihood climate change for global warming under the influence of greenhouse effect and L\'evy noise

An abrupt climatic transition could be triggered by a single extreme event, an $\alpha$-stable non-Gaussian L\'evy noise is regarded as a type of noise to generate such extreme events. In contrast with the classic Gaussian noise, a comprehensive approach of the most probable transition path for systems under $\alpha$-stable L\'evy noise is still lacking. We develop here a probabilistic framework, based on the nonlocal Fokker-Planck equation, to investigate the maximum likelihood climate change for an energy balance system under the influence of greenhouse effect and L\'evy fluctuations. We find that a period of the cold climate state can be interrupted by a sharp shift to the warmer one due to larger noise jumps, and the climate change for warming $1.5\rm ^oC$ under an enhanced greenhouse effect generates a step-like growth process. These results provide important insights into the underlying mechanisms of abrupt climate transitions triggered by a L\'evy process.

cond-mat.stat-mech

Multiple Charge Density Wave States at the Surface of TbTe$_3$

We studied TbTe$_{3}$ using scanning tunneling microscopy (STM) in the temperature range of 298 - 355 K. As seen in previous STM measurements on RTe$_{3}$ compounds, our measurements detect a unidirectional charge density wave state (CDW) in the surface Te-layer with a wavevector consistent with that of the bulk, q$_{cdw}$ = 0.30 $\pm$ 0.01c$^{*}$. However, unlike previous STM measurements, and differing from measurements probing the bulk, we detect two perpendicular orientations for the unidirectional CDWs with no directional preference for the in-plane crystal axes (a- or c-axis) and no noticeable difference in wavevector magnitude. In addition, we find regions in which the bidirectional CDW states coexist. We propose that observation of two CDW states indicates a decoupling of the surface Te-layer from the rare-earth block layer below, and that strain variations in the Te surface layer drive the local CDW direction to the specific unidirectional or, in rare occurrences, bidirectional CDW orders observed. This indicates that similar driving mechanisms for CDW formation in the bulk, where anisotropic lattice strain energy is important, are at play at the surface. In our bias-dependent measurements, we find no contrast inversion for the CDW state between occupied and empty states. This finding differs from other quasi 2-dimensional materials containing a hidden 1-dimensional character which leads to a favorable Fermi surface nesting scenario. Our temperature-dependent measurements provide evidence for localized CDW formation above the bulk transition temperature, T$_{cdw}$.

cond-mat.mtrl-sci

Suppression of Superfluid Density and the Pseudogap State in the Cuprates by Impurities

We use scanning tunneling microscopy (STM) to study magnetic Fe impurities intentionally doped into the high-temperature superconductor Bi$_{2}$Sr$_{2}$Ca$_{2}$CuO$_{8+\delta}$. Our spectroscopic measurements reveal that Fe impurities introduce low-lying resonances in the density of states at \Omega$_{1}$ $\approx$ 4meV and \Omega$_{2}$ $\approx$ 15 meV allowing us to determine that, despite having a large magnetic moment, potential scattering of quasiparticles by Fe impurities dominates magnetic scattering. In addition, using high-resolution spatial characterizations of the local density of states near and away from Fe impurities, we detail the spatial extent of impurity affected regions as well as provide a local view of impurity-induced effects on the superconducting and pseudogap states. Our studies of Fe impurities, when combined with a reinterpretation of earlier STM work in the context of a two-gap scenario, allow us to present a unified view of the atomic-scale effects of elemental impurities on the pseudogap and superconducting states in hole-doped cuprates; this may help resolve a previously assumed dichotomy between the effects of magnetic and non-magnetic impurities in these materials.

cond-mat.supr-con

Origin of a Superlattice Observed in Li$_{0.9}$Mo$_{6}$O$_{17}$ by Scanning Tunneling Microscopy

We use scanning tunneling microscopy to study the lithium purple bronze (Li$_{0.9}$Mo$_{6}$O$_{17}$) at room temperature. Our measurements allow us to identify the single-crystal cleave plane and show that it is possible to obtain clean cleaved surfaces reflecting the crystal structure without the complications of nanoscale surface disorder. In addition to the crystal lattice, we observe a coexisting discommensurate superlattice with wavevectors q = 0.5a* $\pm$ 0.25b*. We propose that the origin of the superstructure is a surface reconstruction which is driven by cleaving along a crystal plane which contains in-plane MoO$_{4}$ tetrahedra connected to out-of-plane MoO$_{6}$ octahedra through corner-sharing oxygens. When combined with spectroscopic measurements, our studies show a promising avenue through which to study the complex physics within Li$_{0.9}$Mo$_{6}$O$_{17}$.

cond-mat.mtrl-sci