arXiv ScienceSearch

arXiv subjects

Robert Hankache

Publications and source records attributed to Robert Hankache.

7 recordsLinked to original sources

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.

cs.CL

Helping Customers in Distress: An LLM-powered Agent that Converses, Probes, and Routes

Banks receive millions of reports of fraud, scams, and disputed transactions every year, making it challenging to accurately direct customers to the appropriate specialist teams for assistance. The existing manual process driven by humans is slow and stressful for both customers and staff. To address this, we develop a customer-facing AI powered triaging agent that leverages large language models (LLMs) to conduct multi-turn conversations, ask relevant questions, and classify cases for accurate, policy-guided routing, making it embedded in the customer journey. To evaluate and continuously improve the agent, synthetic digital twins of real customers were simulated, generating realistic, labelled dialogues based on historical data to test a wide range of real-world scenarios. This work details the triage agent's modelling approach, integration with policy, safety guardrails and reasoning frameworks, the use of the synthetic agent for scalable evaluation, and findings on the AI system's accuracy, robustness, and compliance. Results show that the agent successfully improves triaging of historical cases, achieving a 30.6% increase in classification accuracy, with high satisfaction levels reported by our subject-matter experts, highlighting how targeted probing can lead to more effective triage in banking operations at scale.

cs.HC

Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks

Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias Benchmark for QA (BBQ) dataset to investigate the impact of substituting explicit nationality labels with culturally indicative names, a scenario more reflective of real-world LLM applications. Our novel approach examines how this substitution affects both bias magnitude and accuracy across a spectrum of LLMs from industry leaders such as OpenAI, Google, and Anthropic. Our experiments show that small models are less accurate and exhibit more bias compared to their larger counterparts. For instance, on our name-based dataset and in the ambiguous context (where the correct choice is not revealed), Claude Haiku exhibited the worst stereotypical bias scores of 9%, compared to only 3.5% for its larger counterpart, Claude Sonnet, where the latter also outperformed it by 117.7% in accuracy. Additionally, we find that small models retain a larger portion of existing errors in these ambiguous contexts. For example, after substituting names for explicit nationality references, GPT-4o retains 68% of the error rate versus 76% for GPT-4o-mini, with similar findings for other model providers, in the ambiguous context. Our research highlights the stubborn resilience of biases in LLMs, underscoring their profound implications for the development and deployment of AI systems in diverse, global contexts.

cs.CL

Evaluating the Sensitivity of LLMs to Prior Context

As large language models (LLMs) are increasingly deployed in multi-turn dialogue and other sustained interactive scenarios, it is essential to understand how extended context affects their performance. Popular benchmarks, focusing primarily on single-turn question answering (QA) tasks, fail to capture the effects of multi-turn exchanges. To address this gap, we introduce a novel set of benchmarks that systematically vary the volume and nature of prior context. We evaluate multiple conventional LLMs, including GPT, Claude, and Gemini, across these benchmarks to measure their sensitivity to contextual variations. Our findings reveal that LLM performance on multiple-choice questions can degrade dramatically in multi-turn interactions, with performance drops as large as 73% for certain models. Even highly capable models such as GPT-4o exhibit up to a 32% decrease in accuracy. Notably, the relative performance of larger versus smaller models is not always predictable. Moreover, the strategic placement of the task description within the context can substantially mitigate performance drops, improving the accuracy by as much as a factor of 3.5. These findings underscore the need for robust strategies to design, evaluate, and mitigate context-related sensitivity in LLMs.

cs.CL

A Brief Review of Quantum Machine Learning for Financial Services

This review paper examines state-of-the-art algorithms and techniques in quantum machine learning with potential applications in finance. We discuss QML techniques in supervised learning tasks, such as Quantum Variational Classifiers, Quantum Kernel Estimation, and Quantum Neural Networks (QNNs), along with quantum generative AI techniques like Quantum Transformers and Quantum Graph Neural Networks (QGNNs). The financial applications considered include risk management, credit scoring, fraud detection, and stock price prediction. We also provide an overview of the challenges, potential, and limitations of QML, both in these specific areas and more broadly across the field. We hope that this can serve as a quick guide for data scientists, professionals in the financial sector, and enthusiasts in this area to understand why quantum computing and QML in particular could be interesting to explore in their field of expertise.

quant-ph

Machine-enhanced CP-asymmetries in the electroweak sector

The violation of charge conjugation (C) and parity (P) symmetries are a requirement for the observed dominance of matter over antimatter in the Universe. As an established effect of beyond the Standard Model physics, this could point towards additional CP violation in the Higgs-gauge sector. The phenomenological footprint of the associated anomalous couplings can be small, and designing measurement strategies with the highest sensitivity is therefore of the utmost importance in order to maximise the discovery potential of the Large Hadron Collider (LHC). There are, however, very few measurements of CP-sensitive observables in processes that probe the weak-boson self-interactions. In this article, we study the sensitivity to new sources of CP violation for a range of experimentally-accessible electroweak processes, including $W\gamma$ production, $WW$ production via photon fusion, electroweak $Zjj$ production, electroweak $ZZjj$ production, and electroweak $W^\pm W^\pm jj$ production. We study simple angular observables as well CP-sensitive observables constructed using the outputs of machine-learning (ML) algorithms. We find that the ML-constructed CP-sensitive observables improve the sensitivity to CP-violating effects by up to a factor of five, depending on the process. We also find that inclusive $W\gamma$ and electroweak $Zjj$ production have the potential to set the best possible constraints on certain CP-odd operators in the Higgs-gauge sector of dimension-six effective field theories.

hep-ph

Machine-enhanced CP-asymmetries in the Higgs sector

Improving the sensitivity to CP-violation in the Higgs sector is one of the pillars of the precision Higgs programme at the Large Hadron Collider. We present a simple method that allows CP-sensitive observables to be directly constructed from the output of neural networks. We show that these observables have improved sensitivity to CP-violating effects in the production and decay of the Higgs boson, when compared to the use of traditional angular observables alone. The kinematic correlations identified by the neural networks can be used to design new analyses based on angular observables, with a similar improvement in sensitivity.

hep-ph