Comparing Human Oversight Strategies for Computer-Use Agents
LLM-powered computer-use agents (CUAs) shift users from direct manipulation to supervising an unfolding action sequence, yet existing oversight research mostly evaluates static decisions rather than real-time oversight across an agent trajectory. We compared four oversight strategies with 48 participants across 192 live web sessions. Quantitative results show that oversight strategy affected exposure to problematic actions, but not users' ability to intervene. Behavioral analysis shows why: control often returned too early or too late, creating a timing mismatch within the action sequence. Even when control arrived at the right moment, users judged whether the agent acted correctly, not whether the action was safe. Front-loading oversight into a plan did not solve this either: plans constrained unplanned actions but left unlisted safeguards unaddressed, while upfront approval inhibited runtime scrutiny. These findings show that oversight depends less on maximizing control than on aligning authority, timing, and attention with decision-critical moments.