arXiv ScienceSearch

arXiv subjects

Christopher Klein

Publications and source records attributed to Christopher Klein.

11 recordsLinked to original sources

Protected group bias and stereotypes in Large Language Models

As modern Large Language Models (LLMs) shatter many state-of-the-art benchmarks in a variety of domains, this paper investigates their behavior in the domains of ethics and fairness, focusing on protected group bias. We conduct a two-part study: first, we solicit sentence continuations describing the occupations of individuals from different protected groups, including gender, sexuality, religion, and race. Second, we have the model generate stories about individuals who hold different types of occupations. We collect >10k sentence completions made by a publicly available LLM, which we subject to human annotation. We find bias across minoritized groups, but in particular in the domains of gender and sexuality, as well as Western bias, in model generations. The model not only reflects societal biases, but appears to amplify them. The model is additionally overly cautious in replies to queries relating to minoritized groups, providing responses that strongly emphasize diversity and equity to an extent that other group characteristics are overshadowed. This suggests that artificially constraining potentially harmful outputs may itself lead to harm, and should be applied in a careful and controlled manner.

cs.CY

DELPHI: Data for Evaluating LLMs' Performance in Handling Controversial Issues

Controversy is a reflection of our zeitgeist, and an important aspect to any discourse. The rise of large language models (LLMs) as conversational systems has increased public reliance on these systems for answers to their various questions. Consequently, it is crucial to systematically examine how these models respond to questions that pertaining to ongoing debates. However, few such datasets exist in providing human-annotated labels reflecting the contemporary discussions. To foster research in this area, we propose a novel construction of a controversial questions dataset, expanding upon the publicly released Quora Question Pairs Dataset. This dataset presents challenges concerning knowledge recency, safety, fairness, and bias. We evaluate different LLMs using a subset of this dataset, illuminating how they handle controversial issues and the stances they adopt. This research ultimately contributes to our understanding of LLMs' interaction with controversial issues, paving the way for improvements in their comprehension and handling of complex societal debates.

cs.CL

Feedback Effect in User Interaction with Intelligent Assistants: Delayed Engagement, Adaption and Drop-out

With the growing popularity of intelligent assistants (IAs), evaluating IA quality becomes an increasingly active field of research. This paper identifies and quantifies the feedback effect, a novel component in IA-user interactions: how the capabilities and limitations of the IA influence user behavior over time. First, we demonstrate that unhelpful responses from the IA cause users to delay or reduce subsequent interactions in the short term via an observational study. Next, we expand the time horizon to examine behavior changes and show that as users discover the limitations of the IA's understanding and functional capabilities, they learn to adjust the scope and wording of their requests to increase the likelihood of receiving a helpful response from the IA. Our findings highlight the impact of the feedback effect at both the micro and meso levels. We further discuss its macro-level consequences: unsatisfactory interactions continuously reduce the likelihood and diversity of future user engagements in a feedback loop.

cs.HC

Improving Human-Labeled Data through Dynamic Automatic Conflict Resolution

This paper develops and implements a scalable methodology for (a) estimating the noisiness of labels produced by a typical crowdsourcing semantic annotation task, and (b) reducing the resulting error of the labeling process by as much as 20-30% in comparison to other common labeling strategies. Importantly, this new approach to the labeling process, which we name Dynamic Automatic Conflict Resolution (DACR), does not require a ground truth dataset and is instead based on inter-project annotation inconsistencies. This makes DACR not only more accurate but also available to a broad range of labeling tasks. In what follows we present results from a text classification task performed at scale for a commercial personal assistant, and evaluate the inherent ambiguity uncovered by this annotation strategy as compared to other common labeling strategies.

cs.CL

Generating Natural Questions from Images for Multimodal Assistants

Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in the images properly. The research in visual question answering (VQA) and visual question generation (VQG) is a great step. However, this research does not capture questions that a visually-abled person would ask multimodal assistants. Recently published datasets such as KB-VQA, FVQA, and OK-VQA try to collect questions that look for external knowledge which makes them appropriate for multimodal assistants. However, they still contain many obvious and common-sense questions that humans would not usually ask a digital assistant. In this paper, we provide a new benchmark dataset that contains questions generated by human annotators keeping in mind what they would ask multimodal digital assistants. Large scale annotations for several hundred thousand images are expensive and time-consuming, so we also present an effective way of automatically generating questions from unseen images. In this paper, we present an approach for generating diverse and meaningful questions that consider image content and metadata of image (e.g., location, associated keyword). We evaluate our approach using standard evaluation metrics such as BLEU, METEOR, ROUGE, and CIDEr to show the relevance of generated questions with human-provided questions. We also measure the diversity of generated questions using generative strength and inventiveness metrics. We report new state-of-the-art results on the public and our datasets.

cs.CV

Active Learning for Domain Classification in a Commercial Spoken Personal Assistant

We describe a method for selecting relevant new training data for the LSTM-based domain selection component of our personal assistant system. Adding more annotated training data for any ML system typically improves accuracy, but only if it provides examples not already adequately covered in the existing data. However, obtaining, selecting, and labeling relevant data is expensive. This work presents a simple technique that automatically identifies new helpful examples suitable for human annotation. Our experimental results show that the proposed method, compared with random-selection and entropy-based methods, leads to higher accuracy improvements given a fixed annotation budget. Although developed and tested in the setting of a commercial intelligent assistant, the technique is of wider applicability.

cs.LG

A Near-Infrared Period-Luminosity Relation for Miras in NGC 4258, an Anchor for a New Distance Ladder

We present year-long, near-infrared Hubble Space Telescope WFC3 observations of Mira variables in the water megamaser host galaxy NGC 4258. Miras are AGB variables that can be divided into oxygen- (O-) and carbon- (C-) rich subclasses. Oxygen-rich Miras follow a tight (scatter $\sim 0.14$ mag) Period-Luminosity Relation (PLR) in the near-infrared and can be used to measure extragalactic distances. The water megamaser in NGC 4258 gives a geometric distance to the galaxy accurate to 2.6% that can serve to calibrate the Mira PLR. We develop criteria for detecting and classifying O-rich Miras with optical and NIR data as well as NIR data alone. In total, we discover 438 Mira candidates that we classify with high confidence as O-rich. Our most stringent criteria produce a sample of 139 Mira candidates that we use to measure a PLR. We use the OGLE-III sample of O-rich Miras in the LMC to obtain a relative distance modulus, $\mu_{4258} - \mu_{LMC} = 10.95 \pm 0.01 $ (statistical) $\pm 0.06 $ (systematic) mag in good agreement with the relative distance determined using Cepheids. These results demonstrate the feasibility of discovering and characterizing Miras using the near-infrared with the Hubble Space Telescope and the upcoming James Webb Space Telescope and using them to measure extragalactic distances and determine the Hubble constant.

astro-ph.CO

Clustered Cepheid Variables 90 kiloparsec from the Galactic Center

Distant regions close to the plane of our Galaxy are largely unexplored by optical surveys as they are hidden by dust. We have used near-infrared data (that minimizes dust obscuration) from the ESO Public survey VISTA Variables of the Via Lactea (VVV) (Minniti et al. 2011; Saito et al. 2012; henceforth S12) to search for distant stars at low latitudes. We have discovered four Cepheid variables within an angular extent of one degree centered at Galactic longitude of $l = -27.4^\circ$ and Galactic latitude of $b = -1.08 ^\circ$. We use the tightly constrained period-luminosity relationship that these pulsating stars obey (Persson et al. 2004; Matsunaga et al. 2011) to derive distances. We infer an average distance to these Cepheid variables of 90 kpc. The Cepheid variables are highly clustered in angle (within one degree) and in distance (the standard deviation of the distances is 12 kpc). They are at an average distance of $\sim 2~\rm kpc$ from the plane and their maximum projected separation is $\sim 1~ \rm kpc$. These young ($\sim$ 100 Myr old), pulsating stars (Bono et al. 2005) are unexpected at such large distances from the Galactic disk, which terminates at $\sim$ 15 kpc (Minniti et al. 2011). The highly clustered nature in distance and angle of the Cepheid variables suggests that the stars may be associated with a dwarf galaxy, one that was earlier predicted by a dynamical analysis (Chakrabarti \& Blitz 2009).

astro-ph.GA

Distance determination from the Cepheids and RR Lyrae period-luminosity relations

Cepheids and RR Lyrae stars are important pulsating variable stars in distance scale work because they serve as standard candles. Cepheids follow well-defined period-luminosity (PL) relations defined for bands extending from optical to mid-infrared (MIR). On the other hand, RR Lyrae stars also exhibit PL relations in the near-infrared and MIR wavelengths. In this article, we review some of the recent developments and calibrations of PL relations for Cepheids and RR Lyrae stars. For Cepheids, we discuss the calibration of PL relations via the Galactic and the Large Magellanic Cloud routes. For RR Lyrae stars, we summarize some recent work in developing the MIR PL relations.

astro-ph.SR

The SEDs of Interacting Galaxies

The evolution of galaxies is greatly influenced by their interactions. As part of a program to study interacting galaxies, we have measured and modeled the spectral energy distri- butions (SEDs) from the ultraviolet (UV) to the far-infrared (FIR). We describe the constraints imposed on star formation histories by these SEDs, and the variations therein seen across the interaction sequence, and we compare the results of different star formation rate prescriptions applied to the data. The sample itself is based on the Spitzer Interacting Galaxy Survey (SIGS) of 111 galaxies in 50 systems, a project designed to probe a range of galaxy interaction parameters in the infrared. Our SEDs combine the Spitzer results with multiwavelength data from other missions, in particular GALEX and Herschel. The subset presented here is the sample for which FIR Herschel observations are currently publicly available.

astro-ph.GA

PTF10nvg: An Outbursting Class I Protostar in the Pelican/North American Nebula

During a synoptic survey of the North American Nebula region, the Palomar Transient Factory (PTF) detected an optical outburst (dubbed PTF10nvg) associated with the previously unstudied flat or rising spectrum infrared source IRAS 20496+4354. The PTF R-band light curve reveals that PTF10nvg brightened by more than 5 mag during the current outburst, rising to a peak magnitude of R~13.5 in 2010 Sep. Follow-up observations indicate PTF10nvg has undergone a similar ~5 mag brightening in the K band, and possesses a rich emission-line spectrum, including numerous lines commonly assumed to trace mass accretion and outflows. Many of these lines are blueshifted by ~175 km/s from the North American Nebula's rest velocity, suggesting that PTF10nvg is driving an outflow. Optical spectra of PTF10nvg show several TiO/VO bandheads fully in emission, indicating the presence of an unusual amount of dense (> 10^10 cm^-3), warm (1500-4000 K) circumstellar material. Near-infrared spectra of PTF10nvg appear quite similar to a spectrum of McNeil's Nebula/V1647 Ori, a young star which has undergone several brightenings in recent decades, and 06297+1021W, a Class I protostar with a similarly rich near--infrared emission line spectrum. While further monitoring is required to fully understand this event, we conclude that the brightening of PTF10nvg is indicative of enhanced accretion and outflow in this Class-I-type protostellar object, similar to the behavior of V1647 Ori in 2004-2005.

astro-ph.SR