arXiv Science⌕ Search

arXiv · 2610.00007

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

Abstract

Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. We place nine systems across three paradigms and 13 M to 8 B parameters: a classical tagger (spaCy), bidirectional-encoder specialists (GLiNER, 166 to 460 M), and generative LLMs run locally (Qwen3-0.6B/1.7B/4B-Instruct, DeepSeek-R1-1.5B/8B), on three datasets of differing character, and report accuracy plus two axes the literature omits: latency and output validity. Because our corpus (RSS-News) had no gold, we built silver gold from a cross-family LLM judge panel, then measured its fidelity against benchmark gold and a full human re-validation of the corpus (strict F1 0.95, an upper bound since the human gold was silver-seeded); gold provenance flips the paradigm ranking, moving from LLM-authored silver to human gold raises every encoder and lowers every generative model. On accuracy alone a 4 B instruct LLM is competitive (it leads on clean newswire), so the encoder's case is deployability: it matches or slightly trails at one-ninth to one-twenty-fourth the size, at millisecond-to-second latency, with zero malformed output, while the smallest generative models emit up to 27% invalid output on long inputs, a failure fixed by scale, not output budget. We then characterize GLiNER's per-span confidence: it ranks correctness well (AUROC 0.76 to 0.86) but is overconfident (ECE 0.24 to 0.47, halved by temperature scaling); thresholding gives a small honest out-of-sample F1 gain; an all-local small-to-large cascade gives a modest, corpus-dependent gain over cost-matched random routing; and confidence tracks correctness but not novelty. Every number recomputes offline from per-span records.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vinay Kumar Chaganti. 2026-07-09. On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence. https://arxiv.org/abs/2610.00007

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.

cs.CL↗

PUMA: Learning a Mutation-Aware Vocabulary of Protein Units

Modeling protein sequences as a language has made language models a powerful tool in computational biology, yet the language itself remains poorly understood. A key step toward understanding it is identifying its constituent units. In natural languages, morphemes can occur in multiple forms; similarly, in proteins, mutations can give rise to variations of a unit that persist through evolution, forming families of related units. We introduce PUMA (Protein Units via Mutation-Aware Merging), an algorithm that learns protein units from sequence and explores their mutational variants using substitution matrices, forming a genealogy of unit families. Our results show that mutations remaining within a PUMA family are more often benign than the substitution matrix alone predicts, and that PUMA genealogy improves molecular function representations compared to treating units independently. A case study of a unit family demonstrates relatedness beyond homology. PUMA achieves competitive performance on downstream tasks when used as a protein language model tokenizer. Moreover, collapsing units into families results in a smaller embedding table and faster training. Together, these results support PUMA as a biologically grounded protein vocabulary that organizes protein units into plausible families of mutational variants. The source code is available at https://github.com/boun-tabi-lifelu/PUMA.

cs.CL↗

When Guessing is Rewarded: Rethinking Language Model Evaluation with Distributional Uncertainty Scoring

Standard language model evaluation assigns scores to single predicted answers, rewarding high-confidence responses regardless of how residual probability mass is distributed over alternative options. This creates a systematic pressure toward overconfident guessing: under accuracy-based schemes, a model maximises its expected score by always committing to an answer rather than abstaining, even when its uncertainty is high. While penalty-based approaches partially address this by raising the confidence threshold for strategic guessing, they still treat all sub-threshold responses identically, ignoring a fundamental distinction in how models can express uncertainty - for example between hedging toward incorrect answers versus hedging toward "I don't know" responses. This paper introduces a novel evaluation metric to solve this problem of not considering a model's entire probability distribution over answer choices. The metric naturally distinguishes between harmful overconfidence in wrong answers and uncertainty expressed through abstention, providing scores in an interpretable default range. Through theoretical analysis and illustrative examples, the metric is shown to offer a more nuanced and aligned evaluation paradigm that incentivises models to express genuine uncertainty rather than guessing. Adapting 12 existing evaluation benchmarks to the metric's variants and measuring performance on six language models shows that for half of the tested benchmarks scores are negative across all tested models, indicating significant tendencies towards hallucination.

cs.CL↗