arXiv ScienceSearch

arXiv subjects

Yingyu Yang

Publications and source records attributed to Yingyu Yang.

12 recordsLinked to original sources

Orientation-Robust Latent Motion Trajectory Learning for Annotation-free Cardiac Phase Detection in Fetal Echocardiography

Fetal echocardiography is essential for detecting congenital heart disease (CHD), facilitating pregnancy management, optimized delivery planning, and timely postnatal interventions. Among standard imaging planes, the four-chamber view (4CV) provides important information for CHD diagnosis, where clinicians carefully inspect the end-diastolic (ED) and end-systolic (ES) phases to evaluate cardiac structure and motion. Automated detection of these cardiac phases is thus a critical component towards fully automated CHD analysis. However, existing approaches typically rely on manual annotation of ED/ES frames, which is labour-intensive and time-consuming. We present ORBIT (Orientation-Robust Beat Inference from Trajectories), a self-supervised framework that identifies cardiac phases without manual annotations under various fetal heart orientation. ORBIT employs registration as self-supervision task and learns a latent motion trajectory of cardiac deformation, whose turning points capture transitions between cardiac relaxation and contraction, enabling accurate and orientation-robust localization of ED and ES frames across diverse fetal positions. Trained exclusively on normal fetal echocardiography videos, ORBIT achieves consistent performance on both normal (mean absolute error 1.9 frames for ED and 1.6 for ES) and CHD cases (mean absolute error 2.4 frames for ED and 2.1 for ES), outperforming existing annotation-free approaches constrained by fixed orientation assumptions. These results highlight the potential of ORBIT to facilitate robust cardiac phase detection directly from 4CV fetal echocardiography.

eess.IV

Asynchronous Federated Continual Segmentation with Evolving Clients and Label Spaces

Federated learning seeks to foster collaboration among distributed clients while preserving the privacy of their local data. Traditional federated learning methods typically assume a fixed setting, where participating clients, client data, and learning objectives remain unchanged. However, in real-world scenarios, a federation may evolve over time, with changes in both its client composition and target label space. In this evolving federated setting, conventional round-wise model aggregation becomes inflexible, as each federation update requires repeated communication, repeated local computation, and synchronized participation from all accumulated clients. To address this limitation, we propose CA-MMDS, a continual multiple-model distillation framework for federated continual segmentation with asynchronous clients and evolving label spaces. Instead of repeatedly aggregating model parameters from all clients, CA-MMDS maintains a server-side archive of client models and updates the global model through proxy-based distillation from multiple archived local models. When new clients join or existing clients evolve, only the newly added or updated local models need to be uploaded, while unchanged clients can remain offline and continue to contribute through their archived models. This design substantially reduces communication and computation costs while enabling flexible asynchronous cooperation among evolving clients. Using multi-class 3D abdominal CT segmentation as an application task, we demonstrate that CA-MMDS efficiently incorporates evolving client knowledge while achieving competitive segmentation performance.

eess.IV

Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning

Recent progress in deep learning has significantly advanced CT image analysis, particularly for segmentation tasks. However, these advances are largely confined to image-level pattern recognition, with most methods lacking explicit anatomical or contextual reasoning. Large vision-language models introduce linguistic context into image analysis, yet most approaches typically focus on a single task, which is insufficient for clinical workflow analysis that requires multiple fine-grained types of analysis, such as anatomy detection and segmentation. In this paper, we propose a unified autoregressive framework that integrates language-guided visual reasoning into CT interpretation. Our method introduces task-routing tokens that trigger detection and segmentation heads conditioned on the hidden states of a large vision-language model, enabling coherent generation of visual outputs (e.g., masks and bounding boxes) and textual reasonings. To progressively enhance localisation accuracy and semantic clarity, we further design a "closer-look" mechanism that allows the model to perform progressive coarse-to-fine visits to regions of interest under refined fields of view. To support model training and evaluation, we curated a new multimodal CT dataset containing pixel-wise masks, bounding boxes, spatial prompts, and structured descriptions for visual objects constructed through an AI-assisted annotation process with human verification. Experiments on public benchmarks demonstrate consistent improvements over the SoTA, achieving up to 1.0% Dice on BTCV and 1.7% Dice on MosMed+, while additionally providing appearance reasoning outputs. The code and dataset will be available.

cs.CV

Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew

Federated Learning (FL) enables collaborative model training across distributed clients while preserving data privacy, but remains challenging when client data are highly heterogeneous. These challenges are further amplified in multi-label scenarios, where inter-label dependencies and mismatches between local and global label relationships introduce additional optimization conflicts. While most FL studies focus on single-label classification, many real-world applications are inherently multi-label and often exhibit severe label skew across clients. To address this important yet underexplored problem, we propose FedNCA-ML, a novel FL framework that aligns client representations and learns discriminative, well-clustered features inspired by Neural Collapse (NC) theory. NC describes an ideal latent geometry where each class's features collapse to their mean, forming a maximally separated simplex. FedNCA-ML further introduces an attention-based module to extract class-specific representations, enabling more balanced learning under heavy label imbalance. These class-wise representations are then aligned via a shared NC-inspired structure, mitigating inter-client conflicts induced by heterogeneous local data and inconsistent label dependencies. In addition, we design regularisation losses to encourage compact and consistent feature clustering in the latent space. Experiments on five benchmark datasets under nine FL settings demonstrate the effectiveness of the proposed method, achieving improvements of up to 3.92% in class-wise AUC and 4.93% in class-wise F1 score.

cs.CV

Latent Motion Profiling for Annotation-free Cardiac Phase Detection in Adult and Fetal Echocardiography Videos

The identification of cardiac phase is an essential step for analysis and diagnosis of cardiac function. Automatic methods, especially data-driven methods for cardiac phase detection, typically require extensive annotations, which is time-consuming and labor-intensive. In this paper, we present an unsupervised framework for end-diastole (ED) and end-systole (ES) detection through self-supervised learning of latent cardiac motion trajectories from 4-chamber-view echocardiography videos. Our method eliminates the need for manual annotations, including ED and ES indices, segmentation, or volumetric measurements, by training a reconstruction model to encode interpretable spatiotemporal motion patterns. Evaluated on the EchoNet-Dynamic benchmark, the approach achieves mean absolute error (MAE) of 3 frames (58.3 ms) for ED and 2 frames (38.8 ms) for ES detection, matching state-of-the-art supervised methods. Extended to fetal echocardiography, the model demonstrates robust performance with MAE 1.46 frames (20.7 ms) for ED and 1.74 frames (25.3 ms) for ES, despite the fact that the fetal heart model is built using non-standardized heart views due to fetal heart positioning variability. Our results demonstrate the potential of the proposed latent motion trajectory strategy for cardiac phase detection in adult and fetal echocardiography. This work advances unsupervised cardiac motion analysis, offering a scalable solution for clinical populations lacking annotated data. Code will be released at https://github.com/YingyuYyy/CardiacPhase.

eess.IV

Half-wormholes in a complex SYK model

We compute the half-wormhole contribution in a complex SYK model with one time point. When the chemical potential is zero, the result is similar to two decoupled Majorana SYK models. There's a disk contribution in a single copy of the model, which is a bit subdominant to the unlinked half-wormhole. After removing out the disk we find out the linked half-wormhole which restores the factorization in the two copies of the complex SYK model. When the chemical potential is small, the disk gets more enhancement than the wormhole and the half-wormhole. When the chemical potential is finite comparing to the random coupling, thd disk dominates so that there's no wormhole and half-wormhole.

hep-th

Thermodynamics and Holographic RG Flow in 3D C-metric

In this paper, we investigate the microscopic derivation of the entropy and the holographic RG flow in 3D C-metric. We first discuss the case of a sector in BTZ (Banados-Teitelboim-Zanelli) black hole. By rescaling the Newton's constant we recover the area law of entropy of this sector by microstate counting. Then we apply this technique to all accelerating BTZ phases in 3D C-metric. Finally, for the boundary entropy in 3D C-metric, we study the monotonicity of the $g$-function of 3D C-metric in small acceleration limit and find that the $g$-theorem is satisfied only in $\rm I_{2}$.

hep-th

Diffusion based Zero-shot Medical Image-to-Image Translation for Cross Modality Segmentation

Cross-modality image segmentation aims to segment the target modalities using a method designed in the source modality. Deep generative models can translate the target modality images into the source modality, thus enabling cross-modality segmentation. However, a vast body of existing cross-modality image translation methods relies on supervised learning. In this work, we aim to address the challenge of zero-shot learning-based image translation tasks (extreme scenarios in the target modality is unseen in the training phase). To leverage generative learning for zero-shot cross-modality image segmentation, we propose a novel unsupervised image translation method. The framework learns to translate the unseen source image to the target modality for image segmentation by leveraging the inherent statistical consistency between different modalities for diffusion guidance. Our framework captures identical cross-modality features in the statistical domain, offering diffusion guidance without relying on direct mappings between the source and target domains. This advantage allows our method to adapt to changing source domains without the need for retraining, making it highly practical when sufficient labeled source domain data is not available. The proposed framework is validated in zero-shot cross-modality image segmentation tasks through empirical comparisons with influential generative models, including adversarial-based and diffusion-based models.

eess.IV

Zero-shot-Learning Cross-Modality Data Translation Through Mutual Information Guided Stochastic Diffusion

Cross-modality data translation has attracted great interest in image computing. Deep generative models (\textit{e.g.}, GANs) show performance improvement in tackling those problems. Nevertheless, as a fundamental challenge in image translation, the problem of Zero-shot-Learning Cross-Modality Data Translation with fidelity remains unanswered. This paper proposes a new unsupervised zero-shot-learning method named Mutual Information guided Diffusion cross-modality data translation Model (MIDiffusion), which learns to translate the unseen source data to the target domain. The MIDiffusion leverages a score-matching-based generative model, which learns the prior knowledge in the target domain. We propose a differentiable local-wise-MI-Layer ($LMI$) for conditioning the iterative denoising sampling. The $LMI$ captures the identical cross-modality features in the statistical domain for the diffusion guidance; thus, our method does not require retraining when the source domain is changed, as it does not rely on any direct mapping between the source and target domains. This advantage is critical for applying cross-modality data translation methods in practice, as a reasonable amount of source domain dataset is not always available for supervised training. We empirically show the advanced performance of MIDiffusion in comparison with an influential group of generative models, including adversarial-based and other score-matching-based models.

cs.CV

Unsupervised Echocardiography Registration through Patch-based MLPs and Transformers

Image registration is an essential but challenging task in medical image computing, especially for echocardiography, where the anatomical structures are relatively noisy compared to other imaging modalities. Traditional (non-learning) registration approaches rely on the iterative optimization of a similarity metric which is usually costly in time complexity. In recent years, convolutional neural network (CNN) based image registration methods have shown good effectiveness. In the meantime, recent studies show that the attention-based model (e.g., Transformer) can bring superior performance in pattern recognition tasks. In contrast, whether the superior performance of the Transformer comes from the long-winded architecture or is attributed to the use of patches for dividing the inputs is unclear yet. This work introduces three patch-based frameworks for image registration using MLPs and transformers. We provide experiments on 2D-echocardiography registration to answer the former question partially and provide a benchmark solution. Our results on a large public 2D echocardiography dataset show that the patch-based MLP/Transformer model can be effectively used for unsupervised echocardiography registration. They demonstrate comparable and even better registration performance than a popular CNN registration model. In particular, patch-based models better preserve volume changes in terms of Jacobian determinants, thus generating robust registration fields with less unrealistic deformation. Our results demonstrate that patch-based learning methods, whether with attention or not, can perform high-performance unsupervised registration tasks with adequate time and space complexity. Our codes are available https://gitlab.inria.fr/epione/mlp\_transformer\_registration

cs.CV

More on Half-Wormholes and Ensemble Average

We continue our study about the half-wormhole proposal. By generalizing the original proposal of half-wormhole we propose a new way to detect half-wormholes. The crucial idea is to decompose the observables into self-averaged sector and non-self-averaged sectors. We find the contributions from different sectors have interesting statistics in the semi-classical limit. In particular, dominant sectors tend to condense and the condensation explains the emergence of half-wormholes and we expect that the appearance of condensation is a signal of possible bulk description. We also initiate the study of multi-linked-half-wormholes using our approach.

hep-th

Half-Wormholes and Ensemble Averages

We study "half-wormhole-like" saddle point contributions to spectral correlators in a variety of ensemble average models, including various statistical models, generalized 0d SYK models, 1d Brownian SYK models and an extension of it. In statistical ensemble models, where more general distributions of the random variables could be studied in great details, we find the accuracy of the previously proposed approximation for the half-wormholes could be improved when the distribution of the random variables deviate significantly from Gaussian distributions. We propose a modified approximation scheme of the half-wormhole contributions that also work well in these more general theories. In various generalized 0d SYK models we identify new half-wormhole-like saddle point contributions. In the 0d SYK model and 1d Brownian SYK model, apart from the wormhole and half-wormhole saddles, we find new non-trivial saddles in the spectral correlators that would potentially give contributions of the same order as the trivial self-averaging saddles. However after a careful Lefschetz-thimble analysis we show that these non-trivial saddles should not be included. We also clarify the difference between "linked half-wormholes" and "unlinked half-wormholes" in some models.

hep-th