arXiv Science⌕ Search

arXiv · 2610.03232

Seeing through the Eyes of AI: Situated Explainability in Augmented Reality

Abstract

Explainable Artificial Intelligence (AI) enables humans to understand and interpret decisions of AI models. Instead of having a black box, explainability supports humans in understanding AI models' behavior. Existing explainable AI approaches often present explanations on 2D displays using pre-recorded data, requiring users to relate the displayed information back to the physical objects and real world locations involved in a model's decision. Users are forced to decouple data exploration and capture from AI model interpretation. For AI systems that work within physical environments, this separation can make explanations difficult to interpret in context. We propose using Augmented Reality (AR) to enhance the understanding of AI models by enabling spatial explainability information directly in a user's workspace, in real time, as they explore the world. We show how known explainability methods can be applied in AR and provide insights into user experiences with such an application.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ana Stanescu, Lucchas Ribeiro Skreinig, Tobias Langlotz, Stefanie Zollmann, Peter Mohr, Dieter Schmalstieg, Mark Billinghurst, Denis Kalkofen. 2026-10-02. Seeing through the Eyes of AI: Situated Explainability in Augmented Reality. https://arxiv.org/abs/2610.03232

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Characterizing Creativity in Data Visualization: Reflections and Future Directions

Creativity is valued in data visualization design and research, yet how it manifests remains unclear. Understanding creativity is essential for developing expressive visualization tools and representations. In this paper, we present two complementary studies to characterize creativity in data visualization. First, a systematic review of 104 papers yields a design space spanning four themes: design frameworks incorporating divergent and convergent thinking activities; creative representations developing unorthodox visualizations; visualization-enabled creativity support tools supporting creative tasks (e.g., writing); and understanding creativity, creative processes, and influencing factors. Second, qualitative interviews with 11 visualization practitioners and researchers explore practical challenges and contrast them with current academic framing through our design space. Findings indicate that artifacts or final products are often disproportionately considered as the primary indicator of creativity, whereas the design process remains undervalued in practical. We conclude by presenting a definition of creativity in data visualization and directions for future research.

cs.HC↗

Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI

A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6% of emergencies and concluded that consumer-facing AI triage poses safety risks. Its protocol, however, was an exam-style scaffold (forced A/B/C/D output, knowledge suppression, no clarifying questions) unlike how consumers use health chatbots. We ask whether the headline error rate is a property of the models or of the measurement. In a first, mechanistic study, five frontier LLMs on a 17-scenario bank scored 6.4 points higher under naturalistic patient-style messages than under the constrained scaffold (p=0.015), and on one vignette three models went from 0-24% with forced choice to 100% with free text. In a second, faithful replication we ran the authors' own 60 released vignettes through six frontier models under four matched formats, with clinician validation of the rewrites and a blinded clinician audit of the LLM adjudicators. Here the direction reversed: free-text rewrites scored slightly below the exact structured prompt (78.7% vs 81.8%, p=0.020), removing only the answer scaffold changed little (80.6% vs 81.4%, p=0.81), and naturalistic input with a forced categorical answer beat both the exact prompt (84.4%, p=0.023) and free text (p=0.0001). On the four vignettes defining the original emergency rate, under-triage was 17% with the exact scaffold, 50% in free text, and 21% when the same message was answered with a forced letter; every free-text "under-triage" was a same-day recommendation scored C rather than D, and 60% made escalation conditional on information the patient was asked to check, behavior a single-turn benchmark cannot score. The headline rate is therefore largely a property of output format and of mapping prose onto a four-point scale. Benchmark scaffolds are behaviorally active instruments; safety claims should report sensitivity to input wording, output format and adjudication.

cs.HC↗

HandMade: Spatial Prompting for Generative 3D Creation with Part-Labeled VR Sketches

Text-to-3D generation lowers the barrier to 3D content creation, but text alone is a weak interface for specifying spatial intent: where parts should be placed, how they relate, and how an object should be organized in 3D. We present HandMade, a workflow that combines VR 3D sketching and language for open-domain 3D asset generation. HandMade treats coarse, part-labeled 3D sketches not as incomplete geometry to reconstruct directly, but as spatial prompts for existing generative models. It converts segmented VR strokes into multi-view part guidance and structured prompts, allowing users to specify object layout and part relationships through 3D sketching while using language for identity, material, style, and local details. A technical evaluation shows that HandMade better preserves user-authored spatial scaffolds than text-only and sketch-based baselines on 20 varied examples. A user study with eight participants characterizes how users make use of 3D sketching for spatial layout and language for identity, materials, and details across initial authoring and subsequent revision. HandMade contributes an interaction paradigm and interface-to-generation pipeline for spatially guided 3D creation.

cs.HC↗