arXiv Science⌕ Search

arXiv subjects

Kevin Tang

Publications and source records attributed to Kevin Tang.

32 records · Page 2Linked to original sources

Connecting the Persian-speaking World through Transliteration

Despite speaking mutually intelligible varieties of the same language, speakers of Tajik Persian, written in a modified Cyrillic alphabet, cannot read Iranian and Afghan texts written in the Perso-Arabic script. As the vast majority of Persian text on the Internet is written in Perso-Arabic, monolingual Tajik speakers are unable to interface with the Internet in any meaningful way. Due to overwhelming similarity between the formal registers of these dialects and the scarcity of Tajik-Farsi parallel data, machine transliteration has been proposed as more a practical and appropriate solution than machine translation. This paper presents a transformer-based G2P approach to Tajik-Farsi transliteration, achieving chrF++ scores of 58.70 (Farsi to Tajik) and 74.20 (Tajik to Farsi) on novel digraphic datasets, setting a comparable baseline metric for future work. Our results also demonstrate the non-trivial difficulty of this task in both directions. We also provide an overview of the differences between the two scripts and the challenges they present, so as to aid future efforts in Tajik-Farsi transliteration.

cs.CL↗

Analysis of LLM as a grammatical feature tagger for African American English

African American English (AAE) presents unique challenges in natural language processing (NLP). This research systematically compares the performance of available NLP models--rule-based, transformer-based, and large language models (LLMs)--capable of identifying key grammatical features of AAE, namely Habitual Be and Multiple Negation. These features were selected for their distinct grammatical complexity and frequency of occurrence. The evaluation involved sentence-level binary classification tasks, using both zero-shot and few-shot strategies. The analysis reveals that while LLMs show promise compared to the baseline, they are influenced by biases such as recency and unrelated features in the text such as formality. This study highlights the necessity for improved model training and architectural adjustments to better accommodate AAE's unique linguistic characteristics. Data and code are available.

cs.CL↗

GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning

This paper investigates the ethical implications of aligning Large Language Models (LLMs) with financial optimization, through the case study of GreedLlama, a model fine-tuned to prioritize economically beneficial outcomes. By comparing GreedLlama's performance in moral reasoning tasks to a base Llama2 model, our results highlight a concerning trend: GreedLlama demonstrates a marked preference for profit over ethical considerations, making morally appropriate decisions at significantly lower rates than the base model in scenarios of both low and high moral ambiguity. In low ambiguity situations, GreedLlama's ethical decisions decreased to 54.4%, compared to the base model's 86.9%, while in high ambiguity contexts, the rate was 47.4% against the base model's 65.1%. These findings emphasize the risks of single-dimensional value alignment in LLMs, underscoring the need for integrating broader ethical values into AI development to ensure decisions are not solely driven by financial incentives. The study calls for a balanced approach to LLM deployment, advocating for the incorporation of ethical considerations in models intended for business applications, particularly in light of the absence of regulatory oversight.

cs.CL↗

VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence

Current diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignment. However, these approaches are often ineffective when the target edit involves a shape change. To embark on video editing with shape change, we explore customized video subject swapping in this work, where we aim to replace the main subject in a source video with a target subject having a distinct identity and potentially different shape. In contrast to previous methods that rely on dense correspondences, we introduce the VideoSwap framework that exploits semantic point correspondences, inspired by our observation that only a small number of semantic points are necessary to align the subject's motion trajectory and modify its shape. We also introduce various user-point interactions (\eg, removing points and dragging points) to address various semantic point correspondence. Extensive experiments demonstrate state-of-the-art video subject swapping results across a variety of real-world videos.

cs.CV↗

Disambiguation of morpho-syntactic features of African American English -- the case of habitual be

Recent research has highlighted that natural language processing (NLP) systems exhibit a bias against African American speakers. The bias errors are often caused by poor representation of linguistic features unique to African American English (AAE), due to the relatively low probability of occurrence of many such features in training data. We present a workflow to overcome such bias in the case of habitual "be". Habitual "be" is isomorphic, and therefore ambiguous, with other forms of "be" found in both AAE and other varieties of English. This creates a clear challenge for bias in NLP technologies. To overcome the scarcity, we employ a combination of rule-based filters and data augmentation that generate a corpus balanced between habitual and non-habitual instances. With this balanced corpus, we train unbiased machine learning classifiers, as demonstrated on a corpus of AAE transcribed texts, achieving .65 F$_1$ score disambiguating habitual "be".

cs.CL↗

The Lick Observatory Supernova Search follow-up program: photometry data release of 70 stripped-envelope supernovae

We present BVRI and unfiltered Clear light curves of 70 stripped-envelope supernovae (SESNe), observed between 2003 and 2020, from the Lick Observatory Supernova Search (LOSS) follow-up program. Our SESN sample consists of 19 spectroscopically normal SNe~Ib, two peculiar SNe Ib, six SN Ibn, 14 normal SNe Ic, one peculiar SN Ic, ten SNe Ic-BL, 15 SNe IIb, one ambiguous SN IIb/Ib/c, and two superluminous SNe. Our follow-up photometry has (on a per-SN basis) a mean coverage of 81 photometric points (median of 58 points) and a mean cadence of 3.6d (median of 1.2d). From our full sample, a subset of 38 SNe have pre-maximum coverage in at least one passband, allowing for the peak brightness of each SN in this subset to be quantitatively determined. We describe our data collection and processing techniques, with emphasis toward our automated photometry pipeline, from which we derive publicly available data products to enable and encourage further study by the community. Using these data products, we derive host-galaxy extinction values through the empirical colour evolution relationship and, for the first time, produce accurate rise-time measurements for a large sample of SESNe in both optical and infrared passbands. By modeling multiband light curves, we find that SNe Ic tend to have lower ejecta masses and lower ejecta velocities than SNe~Ib and IIb, but higher $^{56}$Ni masses.

astro-ph.HE↗

Investigating the Nature of the Luminous Ambiguous Nuclear Transient ASASSN-17jz

We present observations of the extremely luminous but ambiguous nuclear transient (ANT) ASASSN-17jz, spanning roughly 1200 days of the object's evolution. ASASSN-17jz was discovered by the All-Sky Automated Survey for Supernovae (ASAS-SN) in the galaxy SDSS J171955.84+414049.4 on UT 2017 July 27 at a redshift of $z=0.1641$. The transient peaked at an absolute $B$-band magnitude of $M_{B,{\rm peak}}=-22.81$, corresponding to a bolometric luminosity of $L_{\rm bol,peak}=8.3\times10^{44}$~erg~s$^{-1}$, and exhibited late-time ultraviolet emission that was still ongoing in our latest observations. Integrating the full light curve gives a total emitted energy of $E_{\rm tot}=(1.36\pm0.08)\times10^{52}$~erg, with $(0.80\pm0.02)\times10^{52}$~erg of this emitted within 200 days of peak light. This late-time ultraviolet emission is accompanied by increasing X-ray emission that becomes softer as it brightens. ASASSN-17jz exhibited a large number of spectral emission lines most commonly seen in active galactic nuclei (AGNs) with little evidence of evolution. It also showed transient Balmer features which became fainter and broader over time, and are still being detected $>1000$ days after peak brightness. We consider various physical scenarios for the origin of the transient, including supernovae (SNe), tidal disruption events (TDEs), AGN outbursts, and ANTs. We find that the most likely explanation is that ASASSN-17jz was an SN~IIn occurring in or near the disk of an existing AGN, and that the late-time emission is caused by the AGN transitioning to a more active state.

astro-ph.HE↗

Prosody leaks into the memories of words

The average predictability (aka informativity) of a word in context has been shown to condition word duration (Seyfarth, 2014). All else being equal, words that tend to occur in more predictable environments are shorter than words that tend to occur in less predictable environments. One account of the informativity effect on duration is that the acoustic details of probabilistic reduction are stored as part of a word's mental representation. Other research has argued that predictability effects are tied to prosodic structure in integral ways. With the aim of assessing a potential prosodic basis for informativity effects in speech production, this study extends past work in two directions; it investigated informativity effects in another large language, Mandarin Chinese, and broadened the study beyond word duration to additional acoustic dimensions, pitch and intensity, known to index prosodic prominence. The acoustic information of content words was extracted from a large telephone conversation speech corpus with over 400,000 tokens and 6,000 word types spoken by 1,655 individuals and analyzed for the effect of informativity using frequency statistics estimated from a 431 million word subtitle corpus. Results indicated that words with low informativity have shorter durations, replicating the effect found in English. In addition, informativity had significant effects on maximum pitch and intensity, two phonetic dimensions related to prosodic prominence. Extending this interpretation, these results suggest that predictability is closely linked to prosodic prominence, and that the lexical representation of a word includes phonetic details associated with its average prosodic prominence in discourse. In other words, the lexicon absorbs prosodic influences on speech production.

cs.CL↗

Distribution of Si II $λ$6355 Velocities of Type Ia Supernovae and Implications for Asymmetric Explosions

The ejecta velocity is a very important parameter in studying the structure and properties of Type Ia supernovae (SNe Ia). It is also a candidate key parameter in improving the utility of SNe Ia for cosmological distance determinations. Here we study the velocity distribution of a sample of 311 SNe Ia from the kaepora database. The velocities are derived from the Si II $λ$6355 absorption line in optical spectra measured at (or extrapolated to) the time of peak brightness. We statistically show that the observed velocity has a bimodal Gaussian distribution consisting of two groups of SNe Ia: Group I with a lower but narrower scatter ($μ_1 = 11000 \text{km s}^{-1}$, $σ_1 = 700 \text{km s}^{-1}$), and Group II with a higher but broader scatter ($μ_2 = 12300 \text{km s}^{-1}$, $σ_2 = 1800 \text{km s}^{-1}$). The population ratio of Group I to Group II is 201:110 (65%:35%). There is substantial degeneracy between the two groups, but for SNe Ia with velocity $v > 12000 \text{km s}^{-1}$, the distribution is dominated by Group II. The true origin of the two components is unknown, though there could be that naturally there exist two intrinsic velocity distributions as observed. However, we try to use asymmetric geometric models through statistical simulations to reproduce the observed distribution assuming all SNe Ia share the same intrinsic distribution. In the two cases we consider, 35\% of SNe Ia are considered to be asymmetric in Case 1, and all SNe Ia are asymmetric in Case 2. Simulations for both cases can reproduce the observed velocity distribution but require a significantly large portion ($>35\%$) of SNe Ia to be asymmetric. In addition, the Case 1 result is consistent with recent polarization observations that SNe Ia with higher Si II $λ$6355 velocity tend to be more polarized.

astro-ph.HE↗

SN 2017cfd: A Normal Type Ia Supernova Discovered Very Young

The Type~Ia supernova (SN~Ia) 2017cfd in IC~0511 (redshift z = 0.01209+- 0.00016$) was discovered by the Lick Observatory Supernova Search 1.6+-0.7 d after the fitted first-light time (FFLT; 15.2 d before B-band maximum brightness). Photometric and spectroscopic follow-up observations show that SN~2017cfd is a typical, normal SN~Ia with a peak luminosity MB ~ -19.2+-0.2 mag, Delta m15(B) = 1.16 mag, and reached a B-band maximum ~16.8 d after the FFLT. We estimate there to be moderately strong host-galaxy extinction (A_V = 0.39 +- 0.03 mag) based on MLCS2k2 fitting. The spectrum reveals a Si~II lambda 6355 velocity of ~11,200 kms at peak brightness. The analysis shows that SN~2017cfd is a very typical, normal SN Ia in nearly every aspect. SN~2017cfd was discovered very young, with multiband data taken starting 2 d after the FFLT, making it a valuable complement to the currently small sample (fewer than a dozen) of SNe~Ia with color data at such early times. We find that its intrinsic early-time (B - V)0 color evolution belongs to the "blue" population rather than to the distinct "red" population. Using the photometry, we constrain the companion star radius to be < 2.5 R_sun, thus ruling out a red-giant companion.

astro-ph.SR↗

Lick Observatory Supernova Search Follow-Up Program: Photometry Data Release of 93 Type Ia Supernovae

We present BVRI and unfiltered light curves of 93 Type Ia supernovae (SNe Ia) from the Lick Observatory Supernova Search (LOSS) follow-up program conducted between 2005 and 2018. Our sample consists of 78 spectroscopically normal SNe Ia, with the remainder divided between distinct subclasses (three SN 1991bg-like, three SN 1991T-like, four SNe Iax, two peculiar, and three super-Chandrasekhar events), and has a median redshift of 0.0192. The SNe in our sample have a median coverage of 16 photometric epochs at a cadence of 5.4 days, and the median first observed epoch is ~4.6 days before maximum B-band light. We describe how the SNe in our sample are discovered, observed, and processed, and we compare the results from our newly developed automated photometry pipeline to those from the previous processing pipeline used by LOSS. After investigating potential biases, we derive a final systematic uncertainty of 0.03 mag in BVRI for our dataset. We perform an analysis of our light curves with particular focus on using template fitting to measure the parameters that are useful in standardising SNe Ia as distance indicators. All of the data are available to the community, and we encourage future studies to incorporate our light curves in their analyses.

astro-ph.SR↗

Dense Captioning with Joint Inference and Visual Context

Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images, labeling each with a short descriptive phrase. We identify two key challenges of dense captioning that need to be properly addressed when tackling the problem. First, dense visual concept annotations in each image are associated with highly overlapping target regions, making accurate localization of each visual concept challenging. Second, the large amount of visual concepts makes it hard to recognize each of them by appearance alone. We propose a new model pipeline based on two novel ideas, joint inference and context fusion, to alleviate these two challenges. We design our model architecture in a methodical manner and thoroughly evaluate the variations in architecture. Our final model, compact and efficient, achieves state-of-the-art accuracy on Visual Genome for dense captioning with a relative gain of 73\% compared to the previous best algorithm. Qualitative experiments also reveal the semantic capabilities of our model in dense captioning.

cs.CV↗

Improving Image Classification with Location Context

With the widespread availability of cellphones and cameras that have GPS capabilities, it is common for images being uploaded to the Internet today to have GPS coordinates associated with them. In addition to research that tries to predict GPS coordinates from visual features, this also opens up the door to problems that are conditioned on the availability of GPS coordinates. In this work, we tackle the problem of performing image classification with location context, in which we are given the GPS coordinates for images in both the train and test phases. We explore different ways of encoding and extracting features from the GPS coordinates, and show how to naturally incorporate these features into a Convolutional Neural Network (CNN), the current state-of-the-art for most image classification and recognition problems. We also show how it is possible to simultaneously learn the optimal pooling radii for a subset of our features within the CNN framework. To evaluate our model and to help promote research in this area, we identify a set of location-sensitive concepts and annotate a subset of the Yahoo Flickr Creative Commons 100M dataset that has GPS coordinates with these concepts, which we make publicly available. By leveraging location context, we are able to achieve almost a 7% gain in mean average precision.

cs.CV↗

Learning Temporal Embeddings for Complex Video Analysis

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are sequences of temporally and semantically coherent images. We leverage this information to learn temporal embeddings for video frames by associating frames with the temporal context that they appear in. To do this, we propose a scheme for incorporating temporal context based on past and future frames in videos, and compare this to other contextual representations. In addition, we show how data augmentation using multi-resolution samples and hard negatives helps to significantly improve the quality of the learned embeddings. We evaluate various design decisions for learning temporal embeddings, and show that our embeddings can improve performance for multiple video tasks such as retrieval, classification, and temporal order recovery in unconstrained Internet video.

cs.CV↗