arXiv ScienceSearch

arXiv subjects

Andrew Finch

Publications and source records attributed to Andrew Finch.

10 recordsLinked to original sources

Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, prosody is rarely studied within the context of speech-to-text translation (S2TT) systems. In particular, end-to-end (E2E) systems have been proposed as well-suited for prosody-aware translation because they have direct access to the speech signal when making translation decisions, but the understanding of whether this is successful in practice is still limited. A main challenge is the difficulty of evaluating prosody awareness in translation. To address this challenge, we introduce an evaluation methodology and a focused benchmark (named ContraProST) aimed at capturing a wide range of prosodic phenomena. Our methodology uses large language models and controllable text-to-speech (TTS) to generate contrastive examples. Through experiments in translating English speech into German, Spanish, and Japanese, we find that (a) S2TT models possess some internal representation of prosody, but the prosody signal is often not strong enough to affect the translations, (b) E2E systems outperform cascades of speech recognition and text translation systems, confirming their theoretical advantage in this regard, and (c) certain cascaded systems also capture prosodic information in the translation, but only to a lesser extent that depends on the particulars of the transcript's surface form.

cs.CL

The $3+1$ Formalism in the Geometric Trinity of Gravity

The geometric trinity of gravity offers a platform in which gravity can be formulated in three analogous approaches, namely curvature, torsion and nonmetricity. In this vein, general relativity can be expressed in three dynamically equivalent ways which may offer insights into the different properties of these decompositions such as their Hamiltonian structure, the efficiency of numerical analyses, as well as the classification of gravitational field degrees of freedom. In this work, we take a $3+1$ decomposition of the teleparallel equivalent of general relativity and the symmetric teleparallel equivalent of general relativity which are both dynamically equivalent to curvature based general relativity. By splitting the spacetime metric and corresponding tetrad into their spatial and temporal parts as well as through finding the Gauss-like equations, it is possible to set up a general foundation for the different formulations of gravity. Based on these results, general $3$-tetrad and $3$-metric evolution equations are derived. Finally through the choice of the two respective connections, the metric $3+1$ formulation for general relativity is recovered as well as the tetrad $3+1$ formulation of the teleparallel equivalent of general relativity and the metric $3+1$ formulation of symmetric teleparallel equivalent of general relativity. The approach is capable, in principle, of resolving common features of the various formulations of general relativity at a fundamental level and pointing out characteristics that extensions and alternatives to the various formulations can present.

gr-qc

Scalable Multilingual Frontend for TTS

This paper describes progress towards making a Neural Text-to-Speech (TTS) Frontend that works for many languages and can be easily extended to new languages. We take a Machine Translation (MT) inspired approach to constructing the frontend, and model both text normalization and pronunciation on a sentence level by building and using sequence-to-sequence (S2S) models. We experimented with training normalization and pronunciation as separate S2S models and with training a single S2S model combining both functions. For our language-independent approach to pronunciation we do not use a lexicon. Instead all pronunciations, including context-based pronunciations, are captured in the S2S model. We also present a language-independent chunking and splicing technique that allows us to process arbitrary-length sentences. Models for 18 languages were trained and evaluated. Many of the accuracy measurements are above 99%. We also evaluated the models in the context of end-to-end synthesis against our current production system.

cs.CL

Gravitoelectromagnetism, Solar System Test and Weak-Field Solutions in $f(T,B)$ Gravity with Observational Constraints

Gravitomagnetism characterize phenomena in the weak field limit within the context of rotating systems. These are mainly manifested in the geodetic and Lense-Thirring effects. The geodetic effect describes the precession of the spin of a gyroscope in orbit about a massive static central object, while the Lense-Thirring effect expresses the analogous effect for the precession of the orbit about a rotating source. In this work, we explore these effects in the framework of Teleparallel Gravity and investigate how these effects may impact recent and future missions. We find that teleparallel theories of gravity may have an important impact on these effects which may constrain potential models within these theories.

gr-qc

Extraction of Templates from Phrases Using Sequence Binary Decision Diagrams

The extraction of templates such as ``regard X as Y'' from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-fly from only tagged text by using a novel relaxed variant of the Sequence Binary Decision Diagram (SeqBDD). A SeqBDD can compress a set of sequences into a graphical structure equivalent to a minimal DFA, but more compact and better suited to the task of template extraction. The main contribution of this paper is a relaxed form of the SeqBDD construction algorithm that enables it to form general representations from a small amount of data. The process of compression of shared structures in the text during Relaxed SeqBDD construction, naturally induces the templates we wish to extract. Experiments show that the method is capable of high-quality extraction on tasks based on verb+preposition templates from corpora and phrasal templates from short messages from social media.

cs.CL

Findings of the Third Workshop on Neural Generation and Translation

This document describes the findings of the Third Workshop on Neural Generation and Translation, held in concert with the annual conference of the Empirical Methods in Natural Language Processing (EMNLP 2019). First, we summarize the research trends of papers presented in the proceedings. Second, we describe the results of the two shared tasks 1) efficient neural machine translation (NMT) where participants were tasked with creating NMT systems that are both accurate and efficient, and 2) document-level generation and translation (DGT) where participants were tasked with developing systems that generate summaries from structured data, potentially with assistance from text in another language.

cs.CL

Galactic Rotation Dynamics in f(T) gravity

We investigate galactic rotation curves in $f(T)$ gravity, where $T$ represents a torsional quantity. Our study centers on the particular Lagrangian $f(T)=T+\alpha{T^n}$, where $|n|\neq 1$ and $\alpha$ is a small unknown constant. To do this we treat galactic rotation curves as being composed from two distinct features of galaxies, namely the disk and the bulge. This process is carried out for several values of the index $n$. The resulting curve is then compared with Milky Way profile data to constrain the value of the index $n$ while fitting for the parameter $\alpha$. These values are then further tested on three other galaxies with different morphologies. On the galactic scale we find that $f(T)$ gravity departs from standard Newtonian theory in an important way. For a small range of values of $n$ we find good agreement with data without the need for exotic matter components to be introduced.

astro-ph.GA

Findings of the Second Workshop on Neural Machine Translation and Generation

This document describes the findings of the Second Workshop on Neural Machine Translation and Generation, held in concert with the annual conference of the Association for Computational Linguistics (ACL 2018). First, we summarize the research trends of papers presented in the proceedings, and note that there is particular interest in linguistic structure, domain adaptation, data augmentation, handling inadequate resources, and analysis of models. Second, we describe the results of the workshop's shared task on efficient neural machine translation, where participants were tasked with creating MT systems that are both accurate and efficient.

cs.CL

Gravitomagnetic effects in quadratic gravity with a scalar field

The two gravitomagnetic effects which influence bodies orbiting around a gravitational source are the geodetic effect and the Lense-Thirring effect. The former describes the precession angle of the axis of a spinning gyroscope while in orbit around a nonrotating gravitational source whereas the latter provides a correction for this angle in the case of a spinning source. In this paper we derive the relevant equations in quadratic gravity and relate them to their equivalents in general relativity. Starting with an investigation into Kepler's third law in quadratic gravity with a scalar field, the effects of an axisymmetric and rotating gravitational source on an orbiting body in a circular, equatorial orbit are introduced.

gr-qc

Neural Machine Translation with Supervised Attention

The attention mechanisim is appealing for neural machine translation, since it is able to dynam- ically encode a source sentence by generating a alignment between a target word and source words. Unfortunately, it has been proved to be worse than conventional alignment models in aligment accuracy. In this paper, we analyze and explain this issue from the point view of re- ordering, and propose a supervised attention which is learned with guidance from conventional alignment models. Experiments on two Chinese-to-English translation tasks show that the super- vised attention mechanism yields better alignments leading to substantial gains over the standard attention based NMT.

cs.CL