arXiv ScienceSearch

arXiv subjects

Courtni Byun

Publications and source records attributed to Courtni Byun.

2 recordsLinked to original sources

Stepping into the Margins: How Readers Want AI to Generate Footnotes

Footnotes can be powerful tools to aid understanding, providing information that augments the reading experience. However, static footnotes cannot address every reader question. Current reading tools allow readers to view curated footnotes, allow personal and social annotation, and link dictionaries to reading material. Many other existing tools and natural language processing (NLP) techniques--such as generative AI, summarization and translation--could be used to address any reader question. However, no one has yet explored which of these features readers actually want. To bridge this gap, we conducted thirteen semi-structured interviews with readers from various backgrounds, followed by a thematic analysis of their responses. We develop themes describing the types of footnotes readers prefer and how to determine the quality of footnotes--specifically focusing on what sources of information a system considers, what the footnotes contain, and how the footnotes are presented to the reader.

cs.HC

Automatic Evaluation of Local Topic Quality

Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Even recent models, which aim to improve the quality of these token-level topic assignments, have been evaluated only with respect to global metrics. We propose a task designed to elicit human judgments of token-level topic assignments. We use a variety of topic model types and parameters and discover that global metrics agree poorly with human assignments. Since human evaluation is expensive we propose a variety of automated metrics to evaluate topic models at a local level. Finally, we correlate our proposed metrics with human judgments from the task on several datasets. We show that an evaluation based on the percent of topic switches correlates most strongly with human judgment of local topic quality. We suggest that this new metric, which we call consistency, be adopted alongside global metrics such as topic coherence when evaluating new topic models.

cs.IR