arXiv ScienceSearch

arXiv · 2609.00524

Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications

Abstract

Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.

Explore related subjects

Keep this discovery

BibTeXRIS

Satwik Ram Kodandaram, Monalika Padma Reddy, Xiaojun Bi, Jiawei Zhou, I. V. Ramakrishnan, Vikas Ashok. 2026-09-01. Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications. https://arxiv.org/abs/2609.00524

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression

This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.

cs.CY

An Agent Model Abstraction for Human-AI Teaming Cognitive Coupling

Industrial environments increasingly rely on collaboration between humans and AI-enabled agents. Effective teamwork requires aligning how agents perceive situations, plan actions to pursue goals, and adapt to changing conditions, yet existing systems lack mechanisms for cross-agent cognitive processes coupling. This paper presents a conceptual cognitive agent model that formalises cognitive coupling through eight components: Input, Process, Output, State, Value, Memory, World Model, and Goal. The model abstracts how agents coordinate and co-regulate their cognitive cycles, providing a basis for analysing distributed cognition and designing cognitively interoperable human-AI systems.

cs.HC

Disclosure and dissolution: explainability, AI power, and situated agency in understanding

This paper interrogates the political and philosophical stakes of AI through the lens of cyborg theory cosmotechnics and glitch feminism. It advocates for a tech-positive critically situated approach to AI as a collaborator and substrate for power relations rather than an autonomous agent of harm. By rejecting the naive pause stop narratives of the human AI binary and embracing an explainable AI XAI artistic practice within embodied intersectional and community rooted engagements AI can empower diverse voices and foster ethical creativity. It starts by examining a frontier cybersecurity AI situating its framing within AI public anxiety and arguing for a new critical stance on human navigation in an AI world.

cs.HC