arXiv · 2609.15097
CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation
Abstract
Persona skill distillation can extract recurring patterns from personal information and encode them into reusable skills, enabling AI systems to closely replicate an individual's behavior. However, such replication also raises serious concerns regarding personal privacy and labor autonomy. Unlike existing perturbation-based defenses that require individuals to modify their data before collection, once historical records are collected by an attacker, they can no longer be altered, sanitized, or revoked. Therefore, such defenses are difficult to adapt to this append-only setting. To solve this challenge, we introduce CounterPersona, which constructs targeted counter-persona evidence, packs compatible behavioral states into compact realization units, and strengthens them through rationale-guided consistency rewriting. We conduct extensive experiments showing that CounterPersona achieves strong and consistent effectiveness across lexical, semantic, and LLM-based measures, while remaining robust across distillers. Our work establishes a skill anti-distillation paradigm for protecting personal privacy and labor autonomy against unauthorized skill distillation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pengwei Wang, Zihan Wang, Hangcheng Cao, Qingchuan Zhao, Hongwei Li, Guowen Xu. 2026-09-14. CounterPersona: Append-Only Defense Against Unauthorized Persona Skill Distillation. https://arxiv.org/abs/2609.15097
Cite the original work for its findings. Save a collection to share your selection of sources.