arXiv Science⌕ Search

arXiv · 2609.32385

DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis

Abstract

Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failures arise from maintaining the analytical process, selecting actions, or grounding visual targets. We introduce DashAct, to our knowledge the first benchmark to diagnose these failures at a fine-grained level within the same dashboard task. DashAct contains 357 human-verified interaction trajectories with milestone dependencies and hierarchical target annotations. Its progressive diagnostic cascade evaluates end-to-end execution, restores verified context for next-action prediction, and provides target semantics and a local view for visual grounding. By progressively restoring the conditions for success, DashAct measures the minimum support an agent needs to recover rather than scoring isolated skills. Experiments show that current models struggle even as support is added. The cascade outcomes reveal bottlenecks hidden by end-to-end scores and provide actionable guidance for improving GUI agents.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chuhan Zhang, Qi Xie, Ziyue Wang, Jianing Yin, Yunfan Zhou, Dazhen Deng, Yingcai Wu. 2026-09-26. DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis. https://arxiv.org/abs/2609.32385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models

Smell's connections with food, memory, and social experience have motivated researchers to bring olfaction into interactive systems. Multimodal AI opens new possibilities for generating smells from natural language, yet it remains unclear whether pretrained models encode sufficient olfactory knowledge to translate language into perceptible compositions. We present AromaGen, an AI-powered 12-channel wearable olfactory system that maps free-form natural-language descriptions to 12 base odorants selected to cover a semantically derived olfactory space, while supporting iterative refinement through natural-language feedback. In a between-subjects study (N=60) across 50 real-world benchmark smells, participants distinguished AromaGen-generated target smells above chance in a three-alternative forced-choice task, with retrieval-augmented AI composition performing comparably to human-expert composition and zero-shot AI composition. Natural-language feedback significantly improved the perceived similarity of both AI- and human-expert compositions. Our findings demonstrate the feasibility of using pretrained multimodal AI, with human aroma composition data, to generate perceptible olfactory experiences from natural language.

cs.HC↗

Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing

Challenging behaviors including aggression, self-injury, and property destruction are observed in 68% of autistic youth and pose risks to youth and caregivers. These episodes are preceded by agitation, a rising state of distress expressed through movement, vocalization, and autonomic arousal. Its signs are subtle and individualized, and its autonomic components are invisible without instrumentation. We collected upper-body movement from inertial measurement units, physiology from a wrist-worn device, and vocalizations from lapel microphones across 30 clinician-led sessions with 15 autistic youth, paired with expert behavioral annotations. We adapt four pretrained foundation models, one per modality, project each to a shared 128-dimensional space, and fuse them into a single group model. The model detected agitation with an area under the ROC curve of 0.724 at the clinician-annotated onset (within-participant permutation p=0.0005), declining to 0.608 at 30,s before onset. Thirteen of fifteen participants were above chance. A from-scratch configuration reached only 0.58, while frozen and fine-tuned features performed comparably (0.71 and 0.72). Audio contributed most of the signal, and a watch-only configuration stayed near chance. Individualized agitation is therefore detectable, including in unannotated windows preceding the annotated onset, using foundation-model transfer with one shared model rather than one per child.

cs.HC↗

"If You're Not Doing It, Somebody Else Is": Active Negotiation and the Invisible Labor of Sustained LLM Use

Large language models (LLMs) have become fixtures of academic work even as their users describe them as degrading their writing, thinking, and skills. Dominant adoption frameworks read continued use as evidence of satisfaction, and cannot explain continued use of a distrusted tool. We interviewed 36 graduate student workers, balanced between English-as-a-foreign-language (EFL) and non-EFL speakers, and introduce the Active Negotiation framework: a model of sustained LLM use as a recurring cycle of risk, mitigation, and justification. A failure surfaces a risk, mitigation labor addresses it, and a justification renders the residual risk tolerable until the next failure reopens the cycle. The cycle runs across three dimensions: practical, auditing output; internal, auditing one's own cognition and identity; and social, managing how peers and institutions perceive use. EFL participants invoke linguistic parity as a further justification. We reframe continued adoption as compliance sustained by invisible labor.

cs.HC↗