arXiv ScienceSearch

arXiv subjects

Jiawen Yang

Publications and source records attributed to Jiawen Yang.

11 recordsLinked to original sources

Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

Animatronic faces are a central component of socially interactive robots, enabling rich nonverbal communication through facial articulation. However, state-of-the-art animatronic faces are typically tailored systems: each new facial geometry requires extensive manual mechanical redesign, making large-scale personalization prohibitively slow and costly. In this work, we pursue automated and scalable mechanical face synthesis, aiming to rapidly generate a physically realizable facial mechanism for a wide range of facial geometries. We introduce a parametric, linkage-driven mechanical face template whose topology and actuator layout are explicitly parameterized to support systematic scaling and retargeting across diverse facial morphologies. Building on this template, we propose a hierarchical automatic design algorithm that takes a single 2D portrait as input, reconstructs a target 3D face, and synthesizes a collision-free, manufacturable internal mechanism. The algorithm combines anatomy-guided feasible motion volumes, Action Unit (AU)-derived trajectory-based expressiveness objectives, and a collision-driven outer-loop refinement strategy. Beyond hardware synthesis, we argue that future mechanical faces deployed at scale must engage in bidirectional, multi-turn conversation rather than functioning solely as speaking or listening heads. To this end, we develop a dual-identity conversational facial motion synthesis framework that jointly models speaking and listening behaviors from audio, producing temporally coherent 3D facial motion suitable for physical execution. We validate our system through extensive experiments, including (i) quantitative evaluation of automatic mechanism synthesis across diverse facial geometries, (ii) comparisons against manual mechanical design, (iii) benchmarks on conversational facial motion synthesis and real-time deployment, and (iv) perceptual user studies.

cs.RO

A GPU-Accelerated Transient Detection Pipeline for DECam Time-Domain Surveys

We present a GPU-accelerated transient detection pipeline developed for time-domain surveys with the Dark Energy Camera (DECam). It enables real-time-capable image processing, incorporating science-driven candidate filtering to support rapid transient identification in time-critical observing programs. The pipeline serves as the core transient discovery engine for multiple long-term DECam programs, including the GW-MMADS gravitational-wave follow-up campaign and the DESIRT survey for intermediate-redshift transients with DESI synergy. The pipeline ingests calibrated imaging products from the DECam Community Pipeline and performs image differencing using the SFFT algorithm, coupled with CNN-based real-bogus classification, to produce science-ready transient alerts and light curves that are delivered to community brokers. We validate the pipeline using archival DECam data from the DESIRT survey. The real-bogus classifier achieves a completeness of $\sim$ 99\% of real transients while rejecting $\sim$ 96\% of subtraction artifacts, and the workflow typically reduces the candidate load to a manageable level for survey operations. With GPU acceleration, the typical processing time per DECam exposure is $\sim$ 50 s from calibrated image processing to alert generation using a modest allocation of computing resources.

astro-ph.IM

NICE: Neural Implicit Craniofacial Model for Orthognathic Surgery Prediction

Orthognathic surgery is a crucial intervention for correcting dentofacial skeletal deformities to enhance occlusal functionality and facial aesthetics. Accurate postoperative facial appearance prediction remains challenging due to the complex nonlinear interactions between skeletal movements and facial soft tissue. Existing biomechanical, parametric models and deep-learning approaches either lack computational efficiency or fail to fully capture these intricate interactions. To address these limitations, we propose Neural Implicit Craniofacial Model (NICE) which employs implicit neural representations for accurate anatomical reconstruction and surgical outcome prediction. NICE comprises a shape module, which employs region-specific implicit Signed Distance Function (SDF) decoders to reconstruct the facial surface, maxilla, and mandible, and a surgery module, which employs region-specific deformation decoders. These deformation decoders are driven by a shared surgical latent code to effectively model the complex, nonlinear biomechanical response of the facial surface to skeletal movements, incorporating anatomical prior knowledge. The deformation decoders output point-wise displacement fields, enabling precise modeling of surgical outcomes. Extensive experiments demonstrate that NICE outperforms current state-of-the-art methods, notably improving prediction accuracy in critical facial regions such as lips and chin, while robustly preserving anatomical integrity. This work provides a clinically viable tool for enhanced surgical planning and patient consultation in orthognathic procedures.

cs.CV

Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control

Previous animatronic faces struggle to express emotions effectively due to hardware and software limitations. On the hardware side, earlier approaches either use rigid-driven mechanisms, which provide precise control but are difficult to design within constrained spaces, or tendon-driven mechanisms, which are more space-efficient but challenging to control. In contrast, we propose a hybrid actuation approach that combines the best of both worlds. The eyes and mouth-key areas for emotional expression-are controlled using rigid mechanisms for precise movement, while the nose and cheek, which convey subtle facial microexpressions, are driven by strings. This design allows us to build a compact yet versatile hardware platform capable of expressing a wide range of emotions. On the algorithmic side, our method introduces a self-modeling network that maps motor actions to facial landmarks, allowing us to automatically establish the relationship between blendshape coefficients for different facial expressions and the corresponding motor control signals through gradient backpropagation. We then train a neural network to map speech input to corresponding blendshape controls. With our method, we can generate distinct emotional expressions such as happiness, fear, disgust, and anger, from any given sentence, each with nuanced, emotion-specific control signals-a feature that has not been demonstrated in earlier systems. We release the hardware design and code at https://github.com/ZZongzheng0918/Morpheus-Hardware and https://github.com/ZZongzheng0918/Morpheus-Software.

cs.RO

SPIDER: Structure-Preferential Implicit Deep Network for Biplanar X-ray Reconstruction

Biplanar X-ray imaging is widely used in health screening, postoperative rehabilitation evaluation of orthopedic diseases, and injury surgery due to its rapid acquisition, low radiation dose, and straightforward setup. However, 3D volume reconstruction from only two orthogonal projections represents a profoundly ill-posed inverse problem, owing to the intrinsic lack of depth information and irreducible ambiguities in soft-tissue visualization. Some existing methods can reconstruct skeletal structures and Computed Tomography (CT) volumes, they often yield incomplete bone geometry, imprecise tissue boundaries, and a lack of anatomical realism, thereby limiting their clinical utility in scenarios such as surgical planning and postoperative assessment. In this study, we introduce SPIDER, a novel supervised framework designed to reconstruct CT volumes from biplanar X-ray images. SPIDER incorporates tissue structure as prior (e.g., anatomical segmentation) into an implicit neural representation decoder in the form of joint supervision through a unified encoder-decoder architecture. This design enables the model to jointly learn image intensities and anatomical structures in a pixel-aligned fashion. To address the challenges posed by sparse input and structural ambiguity, SPIDER directly embeds anatomical constraints into the reconstruction process, thereby enhancing structural continuity and reducing soft-tissue artifacts. We conduct comprehensive experiments on clinical head CT datasets and show that SPIDER generates anatomically accurate reconstructions from only two projections. Furthermore, our approach demonstrates strong potential in downstream segmentation tasks, underscoring its utility in personalized treatment planning and image-guided surgical navigation.

eess.IV

Heterogeneous-Modal Unsupervised Domain Adaptation via Latent Space Bridging

Unsupervised domain adaptation (UDA) methods effectively bridge domain gaps but become struggled when the source and target domains belong to entirely distinct modalities. To address this limitation, we propose a novel setting called Heterogeneous-Modal Unsupervised Domain Adaptation (HMUDA), which enables knowledge transfer between completely different modalities by leveraging a bridge domain containing unlabeled samples from both modalities. To learn under the HMUDA setting, we propose Latent Space Bridging (LSB), a specialized framework designed for the semantic segmentation task. Specifically, LSB utilizes a dual-branch architecture, incorporating a feature consistency loss to align representations across modalities and a domain alignment loss to reduce discrepancies between class centroids across domains. Extensive experiments conducted on six benchmark datasets demonstrate that LSB achieves state-of-the-art performance.

cs.CV

Newly Formed Dust within the Circumstellar Environment of SNIa-CSM 2018evt

Dust associated with various stellar sources in galaxies at all cosmic epochs remains a controversial topic, particularly whether supernovae (SNe) play an important role in dust production. We report evidence of dust formation in the cold, dense shell behind the ejecta-circumstellar medium (CSM) interaction in the Type Ia-CSM SN 2018evt three years after the explosion, characterized by a rise in the mid-infrared (MIR) emission accompanied by an accelerated decline in the optical radiation of the SN. Such a dust-formation picture is also corroborated by the concurrent evolution of the profiles of the Ha emission line. Our model suggests enhanced CSM dust concentration at increasing distances from the SN as compared to what can be expected from the density profile of the mass loss from a steady stellar wind. By the time of the last MIR observations at day +1041, a total amount of 1.2+-0.2x10^{-2} Msun of new dust has been formed by SN 2018evt, making SN 2018evt one of the most prolific dust factories among SNe with evidence of dust formation. The unprecedented witness of the intense production procedure of dust may shed light on the perceptions of dust formation in cosmic history.

astro-ph.HE

Deep Drilling in the Time Domain with DECam: Survey Characterization

This paper presents a new optical imaging survey of four deep drilling fields (DDFs), two Galactic and two extragalactic, with the Dark Energy Camera (DECam) on the 4 meter Blanco telescope at the Cerro Tololo Inter-American Observatory (CTIO). During the first year of observations in 2021, $>$4000 images covering 21 square degrees (7 DECam pointings), with $\sim$40 epochs (nights) per field and 5 to 6 images per night per filter in $g$, $r$, $i$, and/or $z$, have become publicly available (the proprietary period for this program is waived). We describe the real-time difference-image pipeline and how alerts are distributed to brokers via the same distribution system as the Zwicky Transient Facility (ZTF). In this paper, we focus on the two extragalactic deep fields (COSMOS and ELAIS-S1), characterizing the detected sources and demonstrating that the survey design is effective for probing the discovery space of faint and fast variable and transient sources. We describe and make publicly available 4413 calibrated light curves based on difference-image detection photometry of transients and variables in the extragalactic fields. We also present preliminary scientific analysis regarding Solar System small bodies, stellar flares and variables, Galactic anomaly detection, fast-rising transients and variables, supernovae, and active galactic nuclei.

astro-ph.IM

Using 1991T/1999aa-like Type Ia Supernovae as Standardizable Candles

We present the photometry of 16 91T/99aa-like Type Ia Supernovae (SNe Ia) observed by the Las Cumbres Observatory. We also use an additional set of 21 91T/99aa-like SNe Ia and 87 normal SNe Ia from the literature for an analysis of the standardizability of the luminosity of 91T/99aa-like SNe. We find that 91T/99aa-like SNe are 0.2 mag brighter than normal SNe Ia, even when fully corrected by the light curve shapes and colors. The weighted root-mean-square of 91T/99aa-like SNe (with $z_{CMB}>0.01$) Hubble residuals is $0.25\pm0.03$ mag, suggesting that 91T/99aa-like SNe are also excellent relative distance indicators to $\pm$12%. We compare the Hubble residuals with the pseudo-equivalent width (pEW) of Si II $\lambda\lambda$6355 around the date of maximum brightness. We find that there is a broken linear correlation in between those two measurements for our sample including both 91T/99aa-like and normal SNe Ia. As the $pEW_{max}$(Si II $\lambda\lambda$6355) increasing, the Hubble residual increases when $pEW_{max}$(Si II $\lambda\lambda$6355)$<55.6$ \r{A}. However, the Hubble residual stays constant beyond this. Given that 91T/99aa-like SNe possess shallower Si II lines than normal SNe Ia, the linear correlation at $pEW_{max}$(Si II $\lambda\lambda$6355)$<55.6$ \r{A} can account for the overall discrepancy of Hubble residuals derived from the two subgroups. Such a systematic effect needs to be taken into account when using SNe Ia to measure luminosity distances.

astro-ph.HE

Expansion-Squeeze-Excitation Fusion Network for Elderly Activity Recognition

This work focuses on the task of elderly activity recognition, which is a challenging task due to the existence of individual actions and human-object interactions in elderly activities. Thus, we attempt to effectively aggregate the discriminative information of actions and interactions from both RGB videos and skeleton sequences by attentively fusing multi-modal features. Recently, some nonlinear multi-modal fusion approaches are proposed by utilizing nonlinear attention mechanism that is extended from Squeeze-and-Excitation Networks (SENet). Inspired by this, we propose a novel Expansion-Squeeze-Excitation Fusion Network (ESE-FN) to effectively address the problem of elderly activity recognition, which learns modal and channel-wise Expansion-Squeeze-Excitation (ESE) attentions for attentively fusing the multi-modal features in the modal and channel-wise ways. Furthermore, we design a new Multi-modal Loss (ML) to keep the consistency between the single-modal features and the fused multi-modal features by adding the penalty of difference between the minimum prediction losses on single modalities and the prediction loss on the fused modality. Finally, we conduct experiments on a largest-scale elderly activity dataset, i.e., ETRI-Activity3D (including 110,000+ videos, and 50+ categories), to demonstrate that the proposed ESE-FN achieves the best accuracy compared with the state-of-the-art methods. In addition, more extensive experimental results show that the proposed ESE-FN is also comparable to the other methods in terms of normal action recognition task.

cs.CV

Image Subtraction in Fourier Space

Image subtraction is essential for transient detection in time-domain astronomy. The point spread function (PSF), photometric scaling, and sky background generally vary with time and across the field-of-view for imaging data taken with ground-based optical telescopes. Image subtraction algorithms need to match these variations for the detection of flux variability. An algorithm that can be fully parallelized is highly desirable for future time-domain surveys. Here we show the Saccadic Fast Fourier Transform (SFFT) algorithm for image differencing. SFFT uses $\delta$-function basis for kernel decomposition, and the image subtraction is performed in Fourier Space. This brings about a remarkable improvement of computational performance of about an order of magnitude compared to other published image subtraction codes. SFFT can accommodate the spatial variations in wide-field imaging data, including PSF, photometric scaling, and sky background. However, the flexibility of the $\delta$-function basis may also make it more prone to overfitting. The algorithm has been tested extensively in real astronomical data taken by a variety of telescopes. Moreover, the SFFT code allows for the spatial variations of the PSF and sky background to be fitted by spline functions.

astro-ph.IM