arXiv Science⌕ Search

arXiv · 2610.09608

When My Skill Becomes Agent Skill: How Knowledge Workers Share Their Expertise with AI Systems

Abstract

Organizations have long sought to make workers' expertise reusable by others. Agentic AI changes the nature of such reuse by enabling AI systems to act on workers' knowledge with limited human involvement. This raises the question of how this shift shapes workers' willingness to share their expertise. We conducted an experiment with knowledge workers who created materials incorporating domain knowledge and decided whether to authorize human or AI reuse. We find that AI and human reuse differ primarily in whether workers choose to share their knowledge, rather than in what they choose to share. Participants' reasoning shifts from prosocial considerations when sharing with humans toward concerns about loss of control, replacement, and downstream governance when sharing with AI. These findings suggest that AI reuse may intensify tensions between organizational knowledge reuse and contributors' interests. We discuss implications for workplace AI and knowledge management policies that preserve workers' rights and agency.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tianqi Song, Zicheng Zhu, Hancheng Cao, Yi-Chieh Lee. 2026-10-07. When My Skill Becomes Agent Skill: How Knowledge Workers Share Their Expertise with AI Systems. https://arxiv.org/abs/2610.09608

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Harnessing the Power of AI in Qualitative Research: Role Assignment, Engagement, and User Perceptions of AI-Generated Follow-Up Questions in Semi-Structured Interviews

Semi-structured interviews highly rely on the quality of follow-up questions, yet interviewers' knowledge and skills may limit their depth and potentially affect outcomes. While many studies have shown the usefulness of large language models (LLMs) for qualitative analysis, their possibility in the data collection process remains underexplored. We adopt an AI-driven "Wizard-of-Oz" setup to investigate how real-time LLM support in generating follow-up questions shapes semi-structured interviews. Through a study with 17 participants, we examine the value of LLM-generated follow-up questions, the evolving division of roles, relationships, collaborative behaviors, and responsibilities between interviewers and AI. Our findings (1) provide empirical evidence of the strengths and limitations of AI-generated follow-up questions (AGQs); (2) introduce a Human-AI collaboration framework in this interview context; and (3) propose human-centered design guidelines for AI-assisted interviewing. We position LLMs as complements, not replacements, to human judgment, and highlight pathways for integrating AI into qualitative data collection.

cs.HC↗

Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving

Partial driving automation creates a tension: drivers remain legally responsible while being less active in control. Meaningful human control (MHC), a normative framework that can potentially address this tension, proposes that automated systems are designed to track relevant human reasons and that humans should at all times remain in control and be responsible. However, empirical methods for evaluating whether systems are under MHC remain underdeveloped. In this driving simulator study, we investigated the extent to which 24 drivers experienced MHC when interacting with partially automated driving systems under two modes - haptic shared control and traded control. During overtaking manoeuvres on a two-lane, two-way road with fully automated longitudinal control and partially automated lateral control, drivers' actions were necessary to prevent crashes due to silent automation failures. Starting from hypotheses derived from the properties of systems under MHC, we used a mixed-methods approach that links behavioural metrics, subjective post-trial ratings, and qualitative feedback to assess drivers' perception of responsibility and control. A confirmatory analysis indicated a negative correlation between the perception of the automated vehicle understanding the driver and conflict in steering torques. Qualitative feedback revealed that mismatches in intentions between the driver and automation, lack of safety, and resistance to driver inputs reduced perceived MHC, while subtle haptic guidance aligned with driver intent had a positive effect. Thus, future designs should prioritise effortless driver interventions, transparent communication of automation intent through haptic, visual or auditory cues, and clear authority allocation to strengthen meaningful human control in partially automated driving.

cs.HC↗

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

Do large language models' (LLMs') answers to self-report questionnaires predict how they behave? Prior work finds they do not, but it uses human personality inventories, so the gap could reflect borrowed human constructs rather than LLM self-report itself. We test this with a self-report instrument built from LLM-specific behaviors (e.g., over-refusal, unsolicited disclaimers) whose structure is derived bottom-up. Administering 300 items 30 times to 25 LLMs from 17 developers yields five replicable, reliable factors (Tucker $ϕ\geq .957$, $α\geq .930$). We compare these self-reports with 2,500 open-ended behavioral samples rated by 151 humans and an LLM-judge ensemble. Humans and judges agree about model behavior ($\bar{r} = .51$), but self-report barely tracks human ratings ($\bar{r} = .09$, 95% CI $[-.07, .18]$) or rater-free text measures, and correcting for criterion unreliability leaves four of five factors near zero. Verbosity is the partial exception ($r = .40$, 71% of its reliability ceiling). On Responsiveness, self-report tracks LLM judges more than humans ($r = .53$ vs. $.18$; Steiger $p = .04$), and controlling for length and formatting does not remove this: agreement between LLM judges and LLM self-report is weak evidence that either tracks human judgment.

cs.HC↗