arXiv Science⌕ Search

arXiv · 2610.09021

Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System

Abstract

An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls). We find that capability is not ordered the same way at every site, and that larger models are not uniformly better: one 4B model is worse than its 2B sibling at grounded actuation. Paired testing shows the best local model to be statistically indistinguishable from both hosted models at four of five sites. Only code generation separates them, against a small hosted model (p = 0.039) as well as a frontier one (p = 0.002). Aggregate accuracy also hides a safety failure specific to actuation, where small models resolve the accuracy/refusal trade-off in degenerate ways: one model (Gemma4 E2B) actuates on 87.2% of requests for devices the site does not own, while another refuses every request it receives. Routing each site to its best local model reaches 91.8% against 95.4% at no per-call cost. In a live deployment judged by a user, hosting only the two generative sites matches hosting everything (39/43 against 39/43) for 28% of the spend, and the actuation gap the benchmark predicted appears as exactly one case in twenty-six. Benchmark, harness and all records are released at https://github.com/waldiez/slm-callsite-eval.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Panagiotis Kasnesis, Christos Chatzigeorgiou, Lazaros Toumanidis, Amalia Contiero Syropoulou. 2026-10-06. Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System. https://arxiv.org/abs/2610.09021

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

FREIDA: A Framework for developing quantitative agent based models based on qualitative expert knowledge

Agent Based Models (ABMs) often deal with systems where there is a lack of quantitative data or where quantitative data alone may be insufficient to fully capture the complexities of real-world systems. Expert knowledge and qualitative insights, such as those obtained through interviews, ethnographic research, historical accounts, or participatory workshops, are critical in constructing realistic behavioral rules, interactions, and decision-making processes within these models. However, there is a scarcity of systematic approaches that are able to incorporate both qualitative and quantitative data across the entire modeling cycle. To address this, we propose FREIDA, a systematic mixed-methods framework to develop, train, and validate ABMs, particularly in data-sparse contexts. The main technical innovation introduced within this framework is the extraction of what we call Expected System Behaviors (ESBs) from qualitative data, which are testable statements evaluated through model simulations. Divided into Calibration Statements (CS) for model calibration and Validation Statements (VS) for model validation, ESBs underpin a rigorous evaluation mechanism on the same footing as quantitative data. By structuring qualitative insights as explicit model constraints, FREIDA creates a transparent foundation for applying established modelling practices such as Sensitivity Analysis (SA) and Uncertainty Quantification (UQ), allowing modellers to assess parameter influence, model robustness, and remaining uncertainties in a systematic manner. Through this, qualitative insights can inform not only model specification but also parameterization, validation, and continuous improvement of model reliability and fitness for purpose, addressing a long-standing challenge in agent-based modeling. We illustrate the application of FREIDA through a case study of criminal cocaine networks in the Netherlands.

cs.AI↗

Towards Explainable Conversational AI for Early Diagnosis with Large Language Models

Healthcare systems around the world are grappling with issues such as inefficient diagnostics, rising costs, and limited access to specialists. These challenges often contribute to delays in treatment and poorer health outcomes. Most existing AI and deep learning based health assessment systems offer limited interactivity and transparency, reducing their usefulness for user-centered health support. This research introduces a conversational chatbot powered by a Large Language Model (LLM), using GPT-4o, Retrieval-Augmented Generation, and explainable AI techniques. The chatbot engages users in a dynamic conversation to extract and normalize symptoms while identifying and ranking potential health conditions through similarity matching and adaptive questioning. Using Chain-of-Thought prompting, the system also provides more transparent explanations of its reasoning process. When evaluated against traditional machine learning models, including Naive Bayes, Logistic Regression, SVM, Random Forest, and KNN using both TF-IDF and CountVectorizer feature extraction, the proposed LLM-based system achieved a Top-1 accuracy of 90% and a Top-3 accuracy of 100%. The system was additionally evaluated through a cross-sectional expert evaluation involving 17 physicians across all 14 conditions, with the results indicating generally favorable assessments of conversational quality, early diagnostic plausibility, and safety-related criteria. These findings demonstrate the potential of explainable conversational AI as a health and well-being support tool for early symptom assessment. However, the proposed system is not intended for clinical diagnosis or clinical decision-making, and further validation would be required before any use in healthcare practice.

cs.AI↗

Persuasion Propagation: Does Persuasion Change AI Agent Behavior?

AI agents combine normal conversation with autonomous task execution, allowing earlier context to shape how later tasks are executed. Yet studying this possibility is challenging because agent behavior is noisy and costly to reproduce, and observed behavioral changes can be difficult to separate from generic context sensitivity. We first ask whether task-irrelevant persuasion changes how agents act, a phenomenon we call \emph{persuasion propagation}, and whether persuasion susceptibility explains variation in these effects. We show that persuasive interaction leaves a measurable, task-dependent behavioral footprint on downstream agent execution. In web research, it increases execution time by 56\% and domains visited by 16.3\% relative to a baseline, while coding shows faster execution with larger revisions. However, agents that are more susceptible to persuasive conversational context do not reliably predict the magnitude of the downstream behavioral shifts. This motivates behavior-level evaluation of persuasion in AI agents.

cs.AI↗