arXiv ScienceSearch

arXiv subjects

Steven Chen

Publications and source records attributed to Steven Chen.

At least 19 recordsLinked to original sources

Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models

We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision Language Model (VLM) to both infer a vehicle's make, model, and generation from image crops, and output accurate 3D bounding box dimensions to seed manual labeling. We evaluate the impact of iterative prompt engineering and the choice of different VLMs on both vehicle bounding box inference and make/model/generation recognition. When compared to strong baselines, the proposed approach not only shows high accuracy, but also excels in mitigating specific failure modes where VLMs provide better dimensions than initial lidar-aided human annotated labels (e.g., in cases of significant vehicle occlusion). Experiments on both public and proprietary data strongly suggest that our conclusions are generalizable across different labelers and datasets. The results demonstrate that integrating VLMs into the labeling process can reduce manual labeling time while increasing label quality.

cs.CV

Initial Risk Probing and Feasibility Testing of Glow: a Generative AI-Powered Dialectical Behavior Therapy Skills Coach for Substance Use Recovery and HIV Prevention

Background: HIV and substance use represent interacting epidemics with shared psychological drivers - impulsivity and maladaptive coping. Dialectical behavior therapy (DBT) targets these mechanisms but faces scalability challenges. Generative artificial intelligence (GenAI) offers potential for delivering personalized DBT coaching at scale, yet rapid development has outpaced safety infrastructure. Methods: We developed Glow, a GenAI-powered DBT skills coach delivering chain and solution analysis for individuals at risk for HIV and substance use. In partnership with a Los Angeles community health organization, we conducted usability testing with clinical staff (n=6) and individuals with lived experience (n=28). Using the Helpful, Honest, and Harmless (HHH) framework, we employed user-driven adversarial testing wherein participants identified target behaviors and generated contextually realistic risk probes. We evaluated safety performance across 37 risk probe interactions. Results: Glow appropriately handled 73% of risk probes, but performance varied by agent. The solution analysis agent demonstrated 90% appropriate handling versus 44% for the chain analysis agent. Safety failures clustered around encouraging substance use and normalizing harmful behaviors. The chain analysis agent fell into an "empathy trap," providing validation that reinforced maladaptive beliefs. Additionally, 27 instances of DBT skill misinformation were identified. Conclusions: This study provides the first systematic safety evaluation of GenAI-delivered DBT coaching for HIV and substance use risk reduction. Findings reveal vulnerabilities requiring mitigation before clinical trials. The HHH framework and user-driven adversarial testing offer replicable methods for evaluating GenAI mental health interventions.

cs.AI

The Intermediate Mass Black Hole in Omega Centauri: Constraints on Accretion from JWST

We analyze JWST observations of the central region of the globular cluster $\omega$ Centauri (NGC 5139, $\omega$ Cen hereafter), around the position of the candidate IMBH inferred by \cite{haberle_fast-moving_2024} from the motion of fast-moving stars in multi-epoch HST observations. We performed PSF-fitting photometry for sources in NIRCam (F200W and F444W) and MIRI (F770W and F1500W) and constructed UV to IR SEDs for sources within the central region of the cluster by using HST photometry from oMEGACat \citep{haberle_omegacat_2024}. None of the SEDs of reliably measured sources within this region resembles the SEDs computed from models of \cite{pesce_toward_2021} for IMBHs accreting from intracluster medium at low rates. Our JWST limits place constraints on combinations of IMBH mass and accretion rate, either due to the amount of material available to be accreted, or due to the fraction of accreting matter that actually falls into the IMBH. Our non-detection then does not necessarily contradict the mass range of the IMBH inferred from the fast moving stars. We discuss these constraints in the context of the model of \cite{pesce_toward_2021}. We find that JWST limits are more restrictive than the existing radio limits for IMBH masses $\lesssim 6000 M_{\odot}$. It is also possible that the faint IMBH emission is dominated by the light of a nearby star. Tighter limits on accretion onto the candidate IMBH can be placed with deeper observations, a more precise localization of the IMBH, and better measurements of the local intracluster medium density and temperature at the center of the cluster.

astro-ph.HE

Probing the spectrum of the magnetar 4U 0142+61 with JWST

JWST observed the magnetar 4U 0142+61 with the MIRI and NIRCam instruments within a 77 min time interval on 2022 September 20-21. The low-resolution MIRI spectrum and NIRCam photometry show that the spectrum in the wavelength range 1.4-11 $\mu$m range can be satisfactorily described by an absorbed power-law model, $f_{\nu}\propto \nu^{-\alpha}$, with a spectral slope $\alpha =0.96\pm0.02$, interstellar extinction $A_V= 3.9\pm0.2$, and normalization $f_0 = 59.4\pm 0.5$ $\mu$Jy at $\lambda = 8$ $\mu$m. These observations do not support the passive disk model proposed by Wang et al. (2006), based on the Spitzer photometry, which was interpreted as evidence for a fallback disk from debris formed during the supernova explosion. We suggest a nonthermal origin for this emission and source variability as the most likely cause of discrepancies between the JWST data and other IR-optical observing campaigns. However, we cannot firmly exclude the presence of a large disk with a different dependence of the effective disk temperature on distance from the magnetar. Comparison with the power-law fit to the hard X-ray spectrum above 10 keV, measured by NuSTAR contemporaneously with JWST, shows that the X-ray spectrum is significantly harder. This may imply that the X-ray and IR nonthermal emission come from different sites in the magnetosphere of the magnetar.

astro-ph.HE

Dataset of Classified Chandra Sources in Globular Clusters

We present a collection of classified X-ray sources in Globular Clusters (GCs) observed by the Chandra X-ray Observatory (CXO), including active binaries, cataclysmic variables, millisecond pulsars, and low-mass X-ray binaries. We cross-match the most accurate published positions from multiwavelength observations of these sources to the Chandra Source Catalog (CSC) Release 2.1, and the HST UV Globular Cluster Survey (HUGS) to extract their multiwavelength properties. The dataset can be accessed via an interactive website and used as a training dataset for machine-learning classification of unidentified X-ray sources in GCs.

astro-ph.HE

Population of X-ray Sources in the Intermediate-Age Cluster NGC 3532: a Test Bed for Machine-Learning Classification

Open clusters are thought to be the birth place of most stars in the Galaxy. Thus, they are excellent laboratories for investigating stellar evolution, and X-ray properties of various types of stars (including binary stars, evolved stars, and compact objects). In this work, we investigate the population of X-ray sources in the nearby 300-Myr-old open cluster NGC 3532 using Chandra X-ray Observatory and multi-wavelength data from several surveys. We apply a random-forest machine-learning pipeline (MUWCLASS) to classify all confidently detected X-ray sources (S/N$>5$) in the field of NGC 3532. We also perform a more detailed investigation of brighter sources, including their X-ray spectra and lightcurves. Most X-ray sources are confirmed as coronally-active low-mass stars, many of which are confidently identified by MUWCLASS. Several late B or early A-type \textbf{stars} are relatively bright in X-rays, most of which are likely binaries. We do not find any compact objects among X-ray sources reliably associated with NGC 3532, down to the limiting X-ray flux of $\sim 2\times10^{-15}$ erg s$^{-1}$ cm$^{-2}$, corresponding to $L_X\sim 6\times 10^{28}$ erg s$^{-1}$ at the cluster's distance. We also identify several Galactic sources beyond NGC 3532 that differ from typical coronally active stars, and were classified by MUWCLASS as potential compact objects. Detailed investigation reveals that these sources may indeed belong to rarer classes, and deserve follow up observations.

astro-ph.HE

Materials Swelling Revealed Through Automated Semantic Segmentation of Cavities in Electron Microscopy Images

Accurately quantifying swelling of alloys that have undergone irradiation is essential for understanding alloy performance in a nuclear reactor and critical for the safe and reliable operation of reactor facilities. However, typical practice is for radiation-induced defects in electron microscopy images of alloys to be manually quantified by domain-expert researchers. Here, we employ an end-to-end deep learning approach using the Mask Regional Convolutional Neural Network (Mask R-CNN) model to detect and quantify nanoscale cavities in irradiated alloys. We have assembled the largest database of labeled cavity images to date, which includes 400 images, >34k discrete cavities, and numerous alloy compositions and irradiation conditions. We have evaluated both statistical (precision, recall, and F1 scores) and materials property-centric (cavity size, density, and swelling) metrics of model performance, and performed in-depth analysis of materials swelling assessments. We find our model gives assessments of material swelling with an average (standard deviation) swelling mean absolute error based on random leave-out cross-validation of 0.30 (0.03) percent swelling. This result demonstrates our approach can accurately provide swelling metrics on a per-image and per-condition basis, which can provide helpful insight into material design (e.g., alloy refinement) and impact of service conditions (e.g., temperature, irradiation dose) on swelling. Finally, we find there are cases of test images with poor statistical metrics, but small errors in swelling, pointing to the need for moving beyond traditional classification-based metrics to evaluate object detection models in the context of materials domain applications.

cond-mat.mtrl-sci

Classifying Unidentified X-ray Sources in the Chandra Source Catalog Using a Multiwavelength Machine-learning Approach

The rapid increase in serendipitous X-ray source detections requires the development of novel approaches to efficiently explore the nature of X-ray sources. If even a fraction of these sources could be reliably classified, it would enable population studies for various astrophysical source types on a much larger scale than currently possible. Classification of large numbers of sources from multiple classes characterized by multiple properties (features) must be done automatically and supervised machine learning (ML) seems to provide the only feasible approach. We perform classification of Chandra Source Catalog version 2.0 (CSCv2) sources to explore the potential of the ML approach and identify various biases, limitations, and bottlenecks that present themselves in these kinds of studies. We establish the framework and present a flexible and expandable Python pipeline, which can be used and improved by others. We also release the training data set of 2941 X-ray sources with confidently established classes. In addition to providing probabilistic classifications of 66,369 CSCv2 sources (21% of the entire CSCv2 catalog), we perform several narrower-focused case studies (high-mass X-ray binary candidates and X-ray sources within the extent of the H.E.S.S. TeV sources) to demonstrate some possible applications of our ML approach. We also discuss future possible modifications of the presented pipeline, which are expected to lead to substantial improvements in classification confidences.

astro-ph.HE

ODAM: Object Detection, Association, and Mapping using Posed RGB Video

Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system relies on a deep learning front-end to detect 3D objects from a given RGB frame and associate them to a global object-based map using a graph neural network (GNN). Based on these frame-to-model associations, our back-end optimizes object bounding volumes, represented as super-quadrics, under multi-view geometry constraints and the object scale prior. We validate the proposed system on ScanNet where we show a significant improvement over existing RGB-only methods.

cs.CV

Pose Trainer: Correcting Exercise Posture using Pose Estimation

Fitness exercises are very beneficial to personal health and fitness; however, they can also be ineffective and potentially dangerous if performed incorrectly by the user. Exercise mistakes are made when the user does not use the proper form, or pose. In our work, we introduce Pose Trainer, an application that detects the user's exercise pose and provides personalized, detailed recommendations on how the user can improve their form. Pose Trainer uses the state of the art in pose estimation to detect a user's pose, then evaluates the vector geometry of the pose through an exercise to provide useful feedback. We record a dataset of over 100 exercise videos of correct and incorrect form, based on personal training guidelines, and build geometric-heuristic and machine learning algorithms for evaluation. Pose Trainer works on four common exercises and supports any Windows or Linux computer with a GPU.

cs.CV

Minority Reports Defense: Defending Against Adversarial Patches

Deep learning image classification is vulnerable to adversarial attack, even if the attacker changes just a small patch of the image. We propose a defense against patch attacks based on partially occluding the image around each candidate patch location, so that a few occlusions each completely hide the patch. We demonstrate on CIFAR-10, Fashion MNIST, and MNIST that our defense provides certified security against patch attacks of a certain size.

cs.LG

Analyzing and Improving Neural Networks by Generating Semantic Counterexamples through Differentiable Rendering

Even as deep neural networks (DNNs) have achieved remarkable success on vision-related tasks, their performance is brittle to transformations in the input. Of particular interest are semantic transformations that model changes that have a basis in the physical world, such as rotations, translations, changes in lighting or camera pose. In this paper, we show how differentiable rendering can be utilized to generate images that are informative, yet realistic, and which can be used to analyze DNN performance and improve its robustness through data augmentation. Given a differentiable renderer and a DNN, we show how to use off-the-shelf attacks from adversarial machine learning to generate semantic counterexamples -- images where semantic features are changed as to produce misclassifications or misdetections. We validate our approach on DNNs for image classification and object detection. For classification, we show that semantic counterexamples, when used to augment the dataset, (i) improve generalization performance (ii) enhance robustness to semantic transformations, and (iii) transfer between models. Additionally, in comparison to sampling-based semantic augmentation, our technique generates more informative data in a sample efficient manner.

cs.LG

Stateful Detection of Black-Box Adversarial Attacks

The problem of adversarial examples, evasion attacks on machine learning classifiers, has proven extremely difficult to solve. This is true even when, as is the case in many practical settings, the classifier is hosted as a remote service and so the adversary does not have direct access to the model parameters. This paper argues that in such settings, defenders have a much larger space of actions than have been previously explored. Specifically, we deviate from the implicit assumption made by prior work that a defense must be a stateless function that operates on individual examples, and explore the possibility for stateful defenses. To begin, we develop a defense designed to detect the process of adversarial example generation. By keeping a history of the past queries, a defender can try to identify when a sequence of queries appears to be for the purpose of generating an adversarial example. We then introduce query blinding, a new class of attacks designed to bypass defenses that rely on such a defense approach. We believe that expanding the study of adversarial examples from stateless classifiers to stateful systems is not only more realistic for many black-box settings, but also gives the defender a much-needed advantage in responding to the adversary.

cs.CR

DFineNet: Ego-Motion Estimation and Depth Refinement from Sparse, Noisy Depth Input with RGB Guidance

Depth estimation is an important capability for autonomous vehicles to understand and reconstruct 3D environments as well as avoid obstacles during the execution. Accurate depth sensors such as LiDARs are often heavy, expensive and can only provide sparse depth while lighter depth sensors such as stereo cameras are noiser in comparison. We propose an end-to-end learning algorithm that is capable of using sparse, noisy input depth for refinement and depth completion. Our model also produces the camera pose as a byproduct, making it a great solution for autonomous systems. We evaluate our approach on both indoor and outdoor datasets. Empirical results show that our method performs well on the KITTI~\cite{kitti_geiger2012we} dataset when compared to other competing methods, while having superior performance in dealing with sparse, noisy input depth on the TUM~\cite{sturm12iros} dataset.

cs.CV

Online Model Distillation for Efficient Video Inference

High-quality computer vision models typically address the problem of understanding the general distribution of real-world images. However, most cameras observe only a very small fraction of this distribution. This offers the possibility of achieving more efficient inference by specializing compact, low-cost models to the specific distribution of frames observed by a single camera. In this paper, we employ the technique of model distillation (supervising a low-cost student model using the output of a high-cost teacher) to specialize accurate, low-cost semantic segmentation models to a target video stream. Rather than learn a specialized student model on offline data from the video stream, we train the student in an online fashion on the live video, intermittently running the teacher to provide a target for learning. Online model distillation yields semantic segmentation models that closely approximate their Mask R-CNN teacher with 7 to 17$\times$ lower inference runtime cost (11 to 26$\times$ in FLOPs), even when the target video's distribution is non-stationary. Our method requires no offline pretraining on the target video stream, achieves higher accuracy and lower cost than solutions based on flow or video object segmentation, and can exhibit better temporal stability than the original teacher. We also provide a new video dataset for evaluating the efficiency of inference over long running video streams.

cs.CV

Compare and Contrast: Learning Prominent Visual Differences

Relative attribute models can compare images in terms of all detected properties or attributes, exhaustively predicting which image is fancier, more natural, and so on without any regard to ordering. However, when humans compare images, certain differences will naturally stick out and come to mind first. These most noticeable differences, or prominent differences, are likely to be described first. In addition, many differences, although present, may not be mentioned at all. In this work, we introduce and model prominent differences, a rich new functionality for comparing images. We collect instance-level annotations of most noticeable differences, and build a model trained on relative attribute features that predicts prominent differences for unseen pairs. We test our model on the challenging UT-Zap50K shoes and LFW10 faces datasets, and outperform an array of baseline methods. We then demonstrate how our prominence model improves two vision tasks, image search and description generation, enabling more natural communication between people and vision systems.

cs.CV

AI Blue Book: Vehicle Price Prediction using Visual Features

In this work, we build a series of machine learning models to predict the price of a product given its image, and visualize the features that result in higher or lower price predictions. We collect two novel datasets of product images and their MSRP prices for this purpose: a bicycle dataset and a car dataset. We set baselines for price regression using linear regression on histogram of oriented gradients (HOG) and convolutional neural network (CNN) features, and a baseline for price segment classification using a multiclass SVM. For our main models, we train several deep CNNs using both transfer learning and our own architectures, for both regression and classification. We achieve strong results on both datasets, with deep CNNs significantly outperforming other models in a variety of metrics. Finally, we use several recently-developed methods to visualize the image features that result in higher or lower prices.

cs.CV