arXiv ScienceSearch

arXiv subjects

Winston Wu

Publications and source records attributed to Winston Wu.

10 recordsLinked to original sources

MOKA: Moral Knowledge Augmentation for Moral Event Extraction

News media often strive to minimize explicit moral language in news articles, yet most articles are dense with moral values as expressed through the reported events themselves. However, values that are reflected in the intricate dynamics among participating entities and moral events are far more challenging for most NLP systems to detect, including LLMs. To study this phenomenon, we annotate a new dataset, MORAL EVENTS, consisting of 5,494 structured event annotations on 474 news articles by diverse US media across the political spectrum. We further propose MOKA, a moral event extraction framework with MOral Knowledge Augmentation, which leverages knowledge derived from moral words and moral scenarios to produce structural representations of morality-bearing events. Experiments show that MOKA outperforms competitive baselines across three moral event understanding tasks. Further analysis shows even ostensibly nonpartisan media engage in the selective reporting of moral events. Our data and codebase are available at https://github.com/launchnlp/MOKA.

cs.CL

Crossing the Aisle: Unveiling Partisan and Counter-Partisan Events in News Reporting

News media is expected to uphold unbiased reporting. Yet they may still affect public opinion by selectively including or omitting events that support or contradict their ideological positions. Prior work in NLP has only studied media bias via linguistic style and word usage. In this paper, we study to which degree media balances news reporting and affects consumers through event inclusion or omission. We first introduce the task of detecting both partisan and counter-partisan events: events that support or oppose the author's political ideology. To conduct our study, we annotate a high-quality dataset, PAC, containing 8,511 (counter-)partisan event annotations in 304 news articles from ideologically diverse media outlets. We benchmark PAC to highlight the challenges of this task. Our findings highlight both the ways in which the news subtly shapes opinion and the need for large language models that better understand events within a broader context. Our dataset can be found at https://github.com/launchnlp/Partisan-Event-Dataset.

cs.CL

You Are What You Annotate: Towards Better Models through Annotator Representations

Annotator disagreement is ubiquitous in natural language processing (NLP) tasks. There are multiple reasons for such disagreements, including the subjectivity of the task, difficult cases, unclear guidelines, and so on. Rather than simply aggregating labels to obtain data annotations, we instead try to directly model the diverse perspectives of the annotators, and explicitly account for annotators' idiosyncrasies in the modeling process by creating representations for each annotator (annotator embeddings) and also their annotations (annotation embeddings). In addition, we propose TID-8, The Inherent Disagreement - 8 dataset, a benchmark that consists of eight existing language understanding datasets that have inherent annotator disagreement. We test our approach on TID-8 and show that our approach helps models learn significantly better from disagreements on six different datasets in TID-8 while increasing model size by fewer than 1% parameters. By capturing the unique tendencies and subjectivity of individual annotators through embeddings, our representations prime AI models to be inclusive of diverse viewpoints.

cs.CL

EASE: An Easily-Customized Annotation System Powered by Efficiency Enhancement Mechanisms

The performance of current supervised AI systems is tightly connected to the availability of annotated datasets. Annotations are usually collected through annotation tools, which are often designed for specific tasks and are difficult to customize. Moreover, existing annotation tools with an active learning mechanism often only support limited use cases. To address these limitations, we present EASE, an Easily-Customized Annotation System Powered by Efficiency Enhancement Mechanisms. \sysname provides modular annotation units for building customized annotation interfaces and also provides multiple back-end options that suggest annotations using (1) multi-task active learning; (2) demographic feature based active learning; (3) a prompt system that can query the API of large language models. We conduct multiple experiments and user studies to evaluate our system's flexibility and effectiveness. Our results show that our system can meet the diverse needs of NLP researchers and significantly accelerate the annotation process.

cs.HC

Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models

Recent progress in large language models (LLMs) has enabled the deployment of many generative NLP applications. At the same time, it has also led to a misleading public discourse that ``it's all been solved.'' Not surprisingly, this has, in turn, made many NLP researchers -- especially those at the beginning of their careers -- worry about what NLP research area they should focus on. Has it all been solved, or what remaining questions can we work on regardless of LLMs? To address this question, this paper compiles NLP research directions rich for exploration. We identify fourteen different research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs. While we identify many research areas, many others exist; we do not cover areas currently addressed by LLMs, but where LLMs lag behind in performance or those focused on LLM development. We welcome suggestions for other research directions to include: https://bit.ly/nlp-era-llm

cs.CL

Institutional ownership and liquidity commonality: evidence from Australia

We study the liquidity commonality impact of local and foreign institutional investment in the Australian equity market in the cross-section and over time. We find that commonality in liquidity is higher for large stocks compared to small stocks in the cross-section of stocks, and the spread between the two has increased over the past two decades. We show that this divergence can be explained by foreign institutional ownership. This finding suggests that foreign institutional investment contributes to an increase in the exposure of large stocks to unexpected liquidity events in the local market. We find a positive association between foreign institutional ownership and commonality in liquidity across all stocks, particularly in large and mid-cap stocks. Correlated trading by foreign institutions explains this association. However, local institutional ownership is positively related to the commonality in liquidity for large-cap stocks only.

q-fin.PR

Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis

Prior work on ideology prediction has largely focused on single modalities, i.e., text or images. In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content. We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of US mainstream media and social media posts from Reddit and Twitter. We conduct in-depth analyses of news articles and reveal differences in image content and usage across the political spectrum. Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components. Our best-performing model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.

cs.CL

Trends in Silicates in the $\beta$ Pictoris Disk

While beta Pic is known to host silicates in ring-like structures, whether the properties of these silicate dust vary with stellocentric distance remains an open question. We re-analyze the beta Pictoris debris disk spectrum from the Spitzer Infrared Spectrograph (IRS) and a new IRTF/SpeX spectrum to investigate trends in Fe/Mg ratio, shape, and crystallinity in grains as a function of wavelength, a proxy for stellocentric distance. By analyzing a re-calibrated and re-extracted spectrum, we identify a new 18 micron forsterite emission feature and recover a 23 micron forsterite emission feature with a substantially larger line-to-continuum ratio than previously reported. We find that these prominent spectral features are primarily produced by small submicron-sized grains, which are continuously generated and replenished from planetesimal collisions in the disk and can elucidate their parent bodies' composition. We discover three trends about these small grains: as stellocentric distance increases, (1) small silicate grains become more crystalline (less amorphous), (2) they become more irregular in shape, and (3) for crystalline silicate grains, the Fe/Mg ratio decreases. Applying these trends to beta Pic's planetary architecture, we find that the dust population exterior to the orbits of beta Pic b and c differs substantially in crystallinity and shape. We also find a tentative 3-5 micron dust excess due to spatially unresolved hot dust emission close to the star. From our findings, we infer that the surfaces of large planetesimals are more Fe-rich and collisionally-processed closer to the star but more Fe-poor and primordial farther from the star.

astro-ph.EP

Evaluating Neural Machine Comprehension Model Robustness to Noisy Inputs and Adversarial Attacks

We evaluate machine comprehension models' robustness to noise and adversarial attacks by performing novel perturbations at the character, word, and sentence level. We experiment with different amounts of perturbations to examine model confidence and misclassification rate, and contrast model performance in adversarial training with different embedding types on two benchmark datasets. We demonstrate improving model performance with ensembling. Finally, we analyze factors that effect model behavior under adversarial training and develop a model to predict model errors during adversarial attacks.

cs.CL

Modeling Color Terminology Across Thousands of Languages

There is an extensive history of scholarship into what constitutes a "basic" color term, as well as a broadly attested acquisition sequence of basic color terms across many languages, as articulated in the seminal work of Berlin and Kay (1969). This paper employs a set of diverse measures on massively cross-linguistic data to operationalize and critique the Berlin and Kay color term hypotheses. Collectively, the 14 empirically-grounded computational linguistic metrics we design---as well as their aggregation---correlate strongly with both the Berlin and Kay basic/secondary color term partition (gamma=0.96) and their hypothesized universal acquisition sequence. The measures and result provide further empirical evidence from computational linguistics in support of their claims, as well as additional nuance: they suggest treating the partition as a spectrum instead of a dichotomy.

cs.CL