arXiv Science⌕ Search

arXiv subjects

Ping Chen

Publications and source records attributed to Ping Chen.

At least 91 records · Page 5Linked to original sources

EP250108a/SN 2025kg: A Jet-Driven Stellar Explosion Interacting With Circumstellar Material

We present optical, radio, and X-ray observations of EP250108a/SN 2025kg, a broad-line Type Ic supernova (SN Ic-BL) accompanying an Einstein Probe (EP) fast X-ray transient (FXT) at $z=0.176$. EP250108a/SN 2025kg possesses a double-peaked optical light curve and its spectrum transitions from a blue underlying continuum to a typical SN Ic-BL spectrum over time. We fit a radioactive decay model to the second peak of the optical light curve and find SN parameters that are consistent with the SNe Ic-BL population, while its X-ray and radio properties are consistent with those of low-luminosity GRB (LLGRB) 060218/SN 2006aj. We explore three scenarios to understand the system's multi-wavelength emission -- (a) SN ejecta interacting with an extended circumstellar medium (CSM), (b) the shocked cocoon of a collapsar-driven jet choked in its stellar envelope, and (c) the shocked cocoon of a collapsar-driven jet choked in an extended CSM. Models (b) and (c) can explain the optical light curve and are also consistent with the radio and X-ray observations. We favor model (c) because it can self-consistently explain both the X-ray prompt emission and first optical peak, but we do not rule out model (b). From the properties of the first peak in model (c), we find evidence that EP250108a/SN 2025kg interacts with an extended CSM, and infer an envelope mass $M_{\rm e} \sim 0.1\,\rm M_\odot$ and radius $R_{\rm e} \sim 4 \times 10^{13}$ cm. EP250108a/SN 2025kg's multi-wavelength properties make it a close analog to LLGRB 060218/SN 2006aj, and highlight the power of early follow-up observations in mapping the environments of massive stars prior to core collapse.

astro-ph.HE↗

GeSn 320 \times 256 Focal Plane Array for Silicon-Based Short-wave Infrared Imaging

Short-wave infrared (SWIR) imaging arrays have demonstrated great potential in applications spanning from military to civilian consumer electronics. However, the current focal plane arrays (FPAs), which are based on compound semiconductors, have limited applications in civilian circumstances due to elevated manufacturing costs and prolonged fabrication cycle time. To address this, a high-performance 320 $\times$ 256 focal plane array based on group-IV semiconductors has been designed and manufactured on a Si substrate using a complementary metal-oxide semiconductor (CMOS) compatible fabrication process. The optical absorption layer is composed of GeSn alloy, whose bandgap could be tailored by choosing the appropriate Sn concentration. In this work, a 10% Sn concentration was employed, yielding a response cutoff wavelength of 2308 nm for the Si-based photodetector, which was measured at 298 K. Moreover, a specific detectivity of 9.7 $\times$ 10$^{11}$ cm$\cdot$ Hz$^{1/2}$ $\cdot$ W$^{-1}$ has been achieved at 77 K, surpassing all previously reported GeSn devices, and rivals commercial extended InGaAs photodetectors. With the help of read-out circuits (ROIC), SWIR images have been successfully captured for the first time by using Si-based GeSn FPA. This work demonstrates the potential of group IV imaging arrays for various applications in the commercial SWIR imaging field.

physics.optics↗

RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models

Current image watermarking technologies are predominantly categorized into text watermarking techniques and image steganography; however, few methods can simultaneously handle text and image-based watermark data, which limits their applicability in complex digital environments. This paper introduces an innovative multi-modal watermarking approach, drawing on the concept of vector discretization in encoder-based vector quantization. By constructing adjacency matrices, the proposed method enables the transformation of text watermarks into robust image-based representations, providing a novel multi-modal watermarking paradigm for image generation applications. Additionally, this study presents a newly designed image restoration module to mitigate image degradation caused by transmission losses and various noise interferences, thereby ensuring the reliability and integrity of the watermark. Experimental results validate the robustness of the method under multiple noise attacks, providing a secure, scalable, and efficient solution for digital image copyright protection.

cs.MM↗

Strong well-posedness of the two-dimensional stochastic Navier-Stokes equation on moving domains

In this paper, we establish the strong($H^1$) well-posedness of the two dimensional stochastic Navier-Stokes equation with multiplicative noise on moving domains. Due to the nonlocality effect, this equation exhibits a ``piecewise" variational setting. Namely the global well-posedness of this equation is decomposed into the well-posedness of a family of stochastic partial differential equations(SPDEs) in the variational setting on each small time-interval. We first examine the well-posedness on each time interval, which does not have (nonhomogeneous) coercivity. Subsequently, we give an estimate of lower bound of length of the time-interval, which enables us to achieve the global well-posedness.

math.PR↗

A Dataset and Toolkit for Multiparameter Cardiovascular Physiology Sensing on Rings

Smart rings offer a convenient way to continuously and unobtrusively monitor cardiovascular physiological signals. However, a gap remains between the ring hardware and reliable methods for estimating cardiovascular parameters, partly due to the lack of publicly available datasets and standardized analysis tools. In this work, we present $τ$-Ring, the first open-source ring-based dataset designed for cardiovascular physiological sensing. The dataset comprises photoplethysmography signals (infrared and red channels) and 3-axis accelerometer data collected from two rings (reflective and transmissive optical paths), with 28.21 hours of raw data from 34 subjects across seven activities. $τ$-Ring encompasses both stationary and motion scenarios, as well as stimulus-evoked abnormal physiological states, annotated with four ground-truth labels: heart rate, respiratory rate, oxygen saturation, and blood pressure. Using our proposed RingTool toolkit, we evaluated three widely-used physics-based methods and four cutting-edge deep learning approaches. Our results show superior performance compared to commercial rings, achieving best MAE values of 5.18 BPM for heart rate, 2.98 BPM for respiratory rate, 3.22\% for oxygen saturation, and 13.33/7.56 mmHg for systolic/diastolic blood pressure estimation. The open-sourced dataset and toolkit aim to foster further research and community-driven advances in ring-based cardiovascular health sensing.

eess.IV↗

Optimizing Large Model Training through Overlapped Activation Recomputation

Large model training often uses recomputation to alleviate memory pressure and pipelines to exploit the parallelism of data, tensors, and devices. However, existing recomputation approaches may incur high overhead when training real-world models, as they are executed on demand in the critical training path. In this paper, we present Lynx, a new recomputation framework to reduce overhead by overlapping recomputation with communication in training pipelines. To reduce the large search space for recomputation strategies, we propose a heuristic-based recomputation scheduling algorithm, which is based on the observation that there are identical structures in large DNN models so that we can apply the same scheduling policy to all such structures. Additionally, we propose a recomputation-aware model partitioning method to balance each stage's execution time for improved training throughput. Our comprehensive evaluation using GPT models with 1.3B-23B parameters shows that Lynx outperforms existing recomputation approaches by up to 1.37x.

cs.DC↗

Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA.

cs.CV↗

Optimizing for the Shortest Path in Denoising Diffusion Model

In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Path Diffusion Model (ShortDF), treats the denoising process as a shortest-path problem aimed at minimizing reconstruction error. By optimizing the initial residuals, we improve the efficiency of the reverse diffusion process and the quality of the generated samples. Extensive experiments on multiple standard benchmarks demonstrate that ShortDF significantly reduces diffusion time (or steps) while enhancing the visual fidelity of generated samples compared to prior arts. This work, we suppose, paves the way for interactive diffusion-based applications and establishes a foundation for rapid data generation. Code is available at https://github.com/UnicomAI/ShortDF.

cs.CV↗

A Rule Based Solution to Co-reference Resolution in Clinical Text

Objective: The aim of this study was to build an effective co-reference resolution system tailored for the biomedical domain. Materials and Methods: Experiment materials used in this study is provided by the 2011 i2b2 Natural Language Processing Challenge. The 2011 i2b2 challenge involves coreference resolution in medical documents. Concept mentions have been annotated in clinical texts, and the mentions that co-refer in each document are to be linked by coreference chains. Normally, there are two ways of constructing a system to automatically discover co-referent links. One is to manually build rules for co-reference resolution, and the other category of approaches is to use machine learning systems to learn automatically from training datasets and then perform the resolution task on testing datasets. Results: Experiments show the existing co-reference resolution systems are able to find some of the co-referent links, and our rule based system performs well finding the majority of the co-referent links. Our system achieved 89.6% overall performance on multiple medical datasets. Conclusion: The experiment results show that manually crafted rules based on observation of training data is a valid way to accomplish high performance in this coreference resolution task for the critical biomedical domain.

cs.CL↗

"Stones from Other Hills can Polish Jade": Zero-shot Anomaly Image Synthesis via Cross-domain Anomaly Injection

Industrial image anomaly detection (IAD) is a pivotal topic with huge value. Due to anomaly's nature, real anomalies in a specific modern industrial domain (i.e. domain-specific anomalies) are usually too rare to collect, which severely hinders IAD. Thus, zero-shot anomaly synthesis (ZSAS), which synthesizes pseudo anomaly images without any domain-specific anomaly, emerges as a vital technique for IAD. However, existing solutions are either unable to synthesize authentic pseudo anomalies, or require cumbersome training. Thus, we focus on ZSAS and propose a brand-new paradigm that can realize both authentic and training-free ZSAS. It is based on a chronically-ignored fact: Although domain-specific anomalies are rare, real anomalies from other domains (i.e. cross-domain anomalies) are actually abundant and directly applicable to ZSAS. Specifically, our new ZSAS paradigm makes three-fold contributions: First, we propose a novel method named Cross-domain Anomaly Injection (CAI), which directly exploits cross-domain anomalies to enable highly authentic ZSAS in a training-free manner. Second, to supply CAI with sufficient cross-domain anomalies, we build the first Domain-agnostic Anomaly Dataset within our best knowledge, which provides ZSAS with abundant real anomaly patterns. Third, we propose a CAI-guided Diffusion Mechanism, which further breaks the quantity limit of real anomalies and enable unlimited anomaly synthesis. Our head-to-head comparison with existing ZSAS solutions justifies our paradigm's superior performance for IAD and demonstrates it as an effective and pragmatic ZSAS solution.

cs.CV↗

Toroidal graphs without $K_{5}^{-}$ and 6-cycles

Cai et al.\ proved that a toroidal graph $G$ without $6$-cycles is $5$-choosable, and proposed the conjecture that $\textsf{ch}(G) = 5$ if and only if $G$ contains a $K_{5}$ [J. Graph Theory 65 (2010) 1--15], where $\textsf{ch}(G)$ is the choice number of $G$. However, Choi later disproved this conjecture, and proved that toroidal graphs without $K_{5}^{-}$ (a $K_{5}$ missing one edge) and $6$-cycles are $4$-choosable [J. Graph Theory 85 (2017) 172--186]. In this paper, we provide a structural description, for toroidal graphs without $K_{5}^{-}$ and $6$-cycles. Using this structural description, we strengthen Choi's result in two ways: (I) we prove that such graphs have weak degeneracy at most three (nearly $3$-degenerate), and hence their DP-paint numbers and DP-chromatic numbers are at most four; (II) we prove that such graphs have Alon-Tarsi numbers at most $4$. Furthermore, all of our results are sharp in some sense.

math.CO↗

SteLLA: A Structured Grading System Using LLMs with RAG

Large Language Models (LLMs) have shown strong general capabilities in many applications. However, how to make them reliable tools for some specific tasks such as automated short answer grading (ASAG) remains a challenge. We present SteLLA (Structured Grading System Using LLMs with RAG) in which a) Retrieval Augmented Generation (RAG) approach is used to empower LLMs specifically on the ASAG task by extracting structured information from the highly relevant and reliable external knowledge based on the instructor-provided reference answer and rubric, b) an LLM performs a structured and question-answering-based evaluation of student answers to provide analytical grades and feedback. A real-world dataset that contains students' answers in an exam was collected from a college-level Biology course. Experiments show that our proposed system can achieve substantial agreement with the human grader while providing break-down grades and feedback on all the knowledge points examined in the problem. A qualitative and error analysis of the feedback generated by GPT4 shows that GPT4 is good at capturing facts while may be prone to inferring too much implication from the given text in the grading task which provides insights into the usage of LLMs in the ASAG system.

cs.CL↗

TreeMatch: A Fully Unsupervised WSD System Using Dependency Knowledge on a Specific Domain

Word sense disambiguation (WSD) is one of the main challenges in Computational Linguistics. TreeMatch is a WSD system originally developed using data from SemEval 2007 Task 7 (Coarse-grained English All-words Task) that has been adapted for use in SemEval 2010 Task 17 (All-words Word Sense Disambiguation on a Specific Domain). The system is based on a fully unsupervised method using dependency knowledge drawn from a domain specific knowledge base that was built for this task. When evaluated on the task, the system precision performs above the Most Frequent Selection baseline.

cs.CL↗

From Language To Vision: A Case Study of Text Animation

Information can be expressed in multiple formats including natural language, images, and motions. Human intelligence usually faces little difficulty to convert from one format to another format, which often shows a true understanding of encoded information. Moreover, such conversions have broad application in many real-world applications. In this paper, we present a text visualization system that can visualize free text with animations. Our system is illustrated by visualizing example sentences of elementary Physics laws.

cs.CL↗

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.

cs.CV↗

Unleashing the Potential of Model Bias for Generalized Category Discovery

Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known categories. The primary challenges stem from model bias induced by pre-training on only known categories and the lack of precise supervision for novel ones, leading to category bias towards known categories and category confusion among different novel categories, which hinders models' ability to identify novel categories effectively. To address these challenges, we propose a novel framework named Self-Debiasing Calibration (SDC). Unlike prior methods that regard model bias towards known categories as an obstacle to novel category identification, SDC provides a novel insight into unleashing the potential of the bias to facilitate novel category learning. Specifically, the output of the biased model serves two key purposes. First, it provides an accurate modeling of category bias, which can be utilized to measure the degree of bias and debias the output of the current training model. Second, it offers valuable insights for distinguishing different novel categories by transferring knowledge between similar categories. Based on these insights, SDC dynamically adjusts the output logits of the current training model using the output of the biased model. This approach produces less biased logits to effectively address the issue of category bias towards known categories, and generates more accurate pseudo labels for unlabeled data, thereby mitigating category confusion for novel categories. Experiments on three benchmark datasets show that SDC outperforms SOTA methods, especially in the identification of novel categories. Our code and data are available at \url{https://github.com/Lackel/SDC}.

cs.LG↗

An Eccentric Binary with a Misaligned Circumbinary Disk

We present spectroscopic and photometric observations of Bernhard-2, which was previously identified as a candidate system to host a misaligned circumbinary disk. Our spectroscopic measurements confirm that Bernhard-2 indeed contains an eccentric ($e=0.69 \pm 0.08$) binary and thus that the periodic variability in the photometric light curve is best explained by the occultation by the misaligned circumbinary disk. By modeling the spectral energy distributions at different phases, we infer the masses of the two binary components to be $\sim 1.1\,M_\odot$ and $\sim 0.9\,M_\odot$, respectively. The system age is determined to be $\lesssim$ 20 Myr by combining the stellar isochrone model with lithium abundance. Our new photometric observations show clear deviations from the model prediction based on the archival data, suggesting ongoing precession of the circumbinary disk. The H$α$ line of Bernhard-2 also shows an inverse P-Cygni profile at epochs close to the pericenter passage, which could be attributed to the pulsed accretion around the pericenter. Bernhard-2 therefore closely resembles the well studied KH 15D system. Further detailed observations and studies of such rare systems can provide useful information about disk physics and evolution.

astro-ph.SR↗

Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision Localization

As large-scale diffusion models continue to advance, they excel at producing high-quality images but often generate unwanted content, such as sexually explicit or violent content. Existing methods for concept removal generally guide the image generation process but can unintentionally modify unrelated regions, leading to inconsistencies with the original model. We propose a novel approach for targeted concept replacing in diffusion models, enabling specific concepts to be removed without affecting non-target areas. Our method introduces a dedicated concept localizer for precisely identifying the target concept during the denoising process, trained with few-shot learning to require minimal labeled data. Within the identified region, we introduce a training-free Dual Prompts Cross-Attention (DPCA) module to substitute the target concept, ensuring minimal disruption to surrounding content. We evaluate our method on concept localization precision and replacement efficiency. Experimental results demonstrate that our method achieves superior precision in localizing target concepts and performs coherent concept replacement with minimal impact on non-target areas, outperforming existing approaches.

cs.CV↗