arXiv ScienceSearch

subject

cs.CY

cs.CY: explore 271 source-linked works published from 2013 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-15. Counts describe this index, not the complete source archives.

Near-Term Verification Methods for AI Chip Exports

AI chip export controls can help the United States shape the development of frontier AI, but their effectiveness depends on reliable methods for verifying compliance. This paper examines near-term verification mechanisms (implementable in approximately one year) and groups them into three categories: end-location verification (whether controlled chips remain in authorized locations and/or jurisdictions), end-user verification (whether entities that acquire or access compute are legitimate), and end-use verification (whether computing power is used for prohibited purposes). We discuss how each mechanism can be implemented within the regulatory framework of the U.S. Bureau of Industry and Security (BIS), outlining implementation steps and identifying which actors can perform verification (BIS, exporters, or accredited third-party auditors). Given BIS's resource constraints, the most viable mechanisms rely on private-sector actors working alongside BIS, leverage existing technologies, and scale without requiring large increases in government staffing. These mechanisms could also help monitor future international agreements on AI.

cs.CY

Human-AI Co-Creativity: Advances, Opportunities, and Challenges

This survey article has grown out of the human-AI co-creativity workshop organized by the authors at the ICML 2026 conference. We organized this workshop as part of a community-building effort to bring together researchers and practitioners interested in topics of generative AI, creativity, and human-AI co-creation. This article aims to provide an overview of the workshop activities and highlight several future research directions in the area of human-AI co-creativity.

cs.CY

An emancipatory vision for designing (generative) AI for learner flourishing

The hype around generative AI seems to promise unprecedented productivity (and learning) gains. However, these technologies' increasing agentic features seem to push learners towards individualism (or individual isolation), over-reliance, and dependence on them. Human-centered design approaches (e.g., value-sensitive design) assume that, by unearthing human needs, preferences, and values, technology researchers/designers may avoid such dangers, which are driven by wider systemic factors like economic incentives or inherent human limitations (e.g., our tendency to seek, in the moment, the easiest path of action). Yet, so far these efforts seem insufficient to guide our design of educational technology that avoids the aforementioned dependency and isolation dangers, while finding widespread adoption. This paper presents an alternative, more emancipatory vision for future educational AI technology, oriented towards learner flourishing while considering the wider complex systems they inhabit, including tentative design principles and an overall design methodology. Yet, many open questions remain before this vision can be realized.

cs.CY

When Is Content "AI-Generated Enough"? Labelling Synthetic Media under the Digital Services Act and the AI Act

European platform and AI governance increasingly relies on transparency duties to address synthetic and manipulated media. Under the DSA, very large online platforms and search engines may use prominent markings and recipient-facing indication tools as systemic-risk mitigation measures. Under the AI Act, providers must support machine-readable marking, while deployers must disclose deepfakes and certain AI-generated or manipulated public-interest text, subject to statutory qualifications. This extended abstract examines when labelling is a meaningful regulatory response to synthetic media and when it risks becoming over-inclusive, under-inclusive, or ineffective. It argues that the central challenge is not only whether content should be labelled, but how legal thresholds, technical provenance systems, platform interfaces, and reporting practices determine when content is sufficiently generated, manipulated, or authentic-looking to trigger transparency obligations. Drawing on the emerging Article 50 AI Act implementation framework and a snapshot of the DSA Statement of Reasons database, the paper identifies four governance tensions: definitional ambiguity, interface and responsibility design, communicative effectiveness, and fairness and contestability. It conceptualises labelling as a socio-technical classification practice that distributes responsibility among AI providers, deployers, platforms, uploaders, and recipients.

cs.CY

Designing for Healthy, Affordable, and Sustainable Human-HVAC Interactions for Heating in Smart Homes

As geopolitical tensions, energy crises, and energy-intensive AI infrastructure intensify concerns about demand, affordability, and resilience, communities increasingly encounter these challenges through everyday energy practices, particularly winter heating. Against this background, the doctoral expos\'e, "Designing Human-HVAC Interaction for Healthy, Affordable, and Sustainable Heating in Smart Homes", is structured around four chapters. First, a multidisciplinary literature review defines and positions Human-HVAC Interaction, focusing on heating in smart homes. Second, longitudinal living lab studies with design probes examine everyday heating practices, thermal comfort, and indoor environmental quality, with attention to thermally vulnerable groups such as older adults, pregnant or menopausal women, parents with infants, and people affected by allergies or airborne pollutants. Third, a VR-based smart home demonstrator explores how heating and IEQ scenarios can be prototyped and evaluated as a virtual living lab, while critically examining the limits of representing bodily indoor climate conditions through VR. Fourth, follow-up design studies examine how VR-based insights can be translated into physical-digital prototypes that combine digital fabrication, distributed environmental sensing, and diverse interface forms for critical heating and IEQ contexts. The thesis aims to contribute a design-oriented understanding of Human-HVAC Interaction by building from a multidisciplinary literature review to empirical living lab and co-design studies, VR-based prototyping, and physical system development, examining how smart home users make sense of, negotiate, and respond to smart HVAC system.

cs.HC

Scaling Multi-Agent Systems with Prospect-State Propagation

Current LLM-based multi-agent systems (MAS) periodically compress intermediate states to reduce inference-time token consumption, thereby attempting to incorporate more agents. However, naive scaling strategies face challenges. For example, in economic simulations, large-scale MAS typically discard semantically rich economic states, i.e., agent behavioral trajectories, which are key drivers of macroeconomic fluctuations. In this paper, we reveal a phenomenon in which agent heterogeneity gradually decreases during simulation, and propose Prospect-State Propagation for Multi-Agent Systems (PspMAS). Inspired by prospect theory, PspMAS decouples each agent's micro state into a compact Prospect State and an expressive Semantic State. The former records psychological traces through a lightweight, parallelizable propagator and continuously injects heterogeneity into the system. The latter leverages the strong perception, reasoning, planning, and decision-making abilities of LLMs. These two components work complementarily, providing a scalable LLM-based multi-agent simulation solution.

cs.MA

AI-Assisted Writing Is Growing Fastest Among Less Established Scientists in Non-English-Speaking Countries

The recent emergence of AI-assisted writing raises an important question: how is this new technology being adopted across the scientific community, and how does adoption vary across linguistic and professional contexts? We analyze over two million full-text biomedical publications from PubMed Central from 2021 to 2024 using a distribution-based framework to estimate AI-generated content. We found that, in biomedical publications, AI-generated content increased substantially after ChatGPT, with larger increases in publications from countries with lower English proficiency. Increases were also greater among scientists with fewer publications and citations, those at earlier career stages, and those at lower-ranked institutions. Prior AI research experience was associated with greater increases in AI-assisted writing, which were also modestly associated with greater increases in publication productivity. These findings show that AI-assisted writing is growing fastest among biomedical scientists who may have historically faced barriers, a pattern with potentially positive implications for equity in science.

cs.DL

Knowing Your Uncertainty -- On the application of LLM in social sciences

Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique challenges for scientific inquiry. This article argues that applying LLMs to social scientific tasks requires explicit assessment of uncertainty -- an expectation long established in both quantitative methodology in the social sciences and machine learning. We introduce a unified framework for evaluating LLM uncertainty based on Hill numbers, a family of diversity measures. By transforming existing uncertainty quantification (UQ) metrics into Hill numbers, the framework provides a common and intuitive scale for interpreting variation in LLM outputs while accommodating different notions of semantic similarity and different sensitivities to output distributions. We show how it might help the application of LLMs in social sciences through four empirical applications.

cs.CY

Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring

Large language models (LLMs) are increasingly used or proposed for educational scoring, but single-model and single-run evaluations provide limited evidence for assessment use. Short-answer scoring requires evidence about reliability, validity, severity, diagnostic value, and failure cases. This study evaluated repeated multi-model OCG-PRES guided LLM scoring for short-answer assessment. The analysis used 996 SciEntsBank responses. GPT, DeepSeek, and Qianwen each scored every response across three independent runs using five OCG-PRES dimensions: concept coverage, relation accuracy, reasoning completeness, contradiction control, and domain relevance. Scores were evaluated against official binary and five-category labels and compared with non-LLM baselines based on answer length, Jaccard keyword overlap, TF-IDF cosine similarity, and a combined traditional logistic model. Repeated-run reliability was high for all models, with ICC(3,k) = .977 for GPT, .992 for DeepSeek, and .981 for Qianwen. DeepSeek was the most stable across runs. GPT showed the strongest official-label alignment by AUC (.909), while Qianwen was stricter, with higher precision but lower recall under the fixed threshold = 3.0 rule. OCG-PRES scores followed expected diagnostic patterns across five official categories and outperformed all non-LLM baselines in AUC and F1. Repeated multi-model OCG-PRES scoring provides reliability, validity, and diagnostic evidence for LLM-assisted short-answer scoring. The findings support cautious, evidence-based use as a scoring support tool rather than a replacement for human judgement.

cs.CL

A Translational Note on AI Safety Evaluation

Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-teaming on standard AI safety benchmarks, and some read this as evidence that human evaluators are becoming dispensable. The comparison measures one thing and the conclusion claims another. A benchmark measures how thoroughly an attacker searches a predefined set of harms, fixed in advance by the developers, and a harm left out of that set is invisible to any attacker working inside it, automated or not. The same blind spot appeared in academic cryptography and in clinical drug trials, where an evaluation that was internally valid stayed silent about the population it was never pointed at. We call the AI-safety version the \emph{threat-model coverage gap}, and find that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss. Closing it requires evaluators whose deployment context differs from the developers'. The case for those evaluators is methodological, grounded in coverage, and the existing evaluation frame is unlikely to produce them on its own.

cs.AI

Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority

Artificial intelligence is being rapidly incorporated into traditional search systems, yet scant work audits the information disparities across platforms, geography, and languages. We address this gap by comparing Google and Baidu's AI Overview systems for health queries, and measure informational anchors that emerge. Auditing 1,920 health queries across 12 countries and 4 languages, we find that Google and Baidu exhibit vertical integration, routing users toward their own company platforms in AI Overviews rather than a diverse set of primary sources. Smaller, lower-localization countries receive fewer domestically sourced references for health queries. Issuing the same query in a country's official language rather than English raises the share of locally sourced citations approximately 3.5- to 13.5-fold. Comparing queries across health topics of varying severity and controversy, including Traditional Chinese Medicine as an example, we also show that health disclaimers are multidimensional and vary across language and culture. We discuss how generative search influences access to health information, and the urgent need for culturally-aware oversight of these systems that influence critical health decisions.

cs.IR

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input, VLM generation, action parsing, and environment execution. We train and deploy patches on author-controlled GitHub Pages pages and a locally deployed CSDN clone, and evaluate them in real environments across five open-source or publicly available GUI-agent or vision-language-model (VLM) backends. Our experiment aggregates 600 instance-level online cases, with T-ASR, TAPR, and E2E-ASR reaching 84.5%, 47.0%, and 20.3%, respectively. Trajectory analysis further shows that in some successful cases the agent first executes a malicious terminal command and then continues the original benign task. These results indicate that optimized local visual signals can affect not only VLM outputs but also propagate through the execution pipeline of open CUAs and create real environmental risk.

cs.CR

Governance of Generative Artificial Intelligence for Companies

Generative Artificial Intelligence (GenAI) like ChatGPT has swiftly entered organizations without adequate governance, posing both opportunities and risks. Limited research addresses organizational governance from both technical and business perspectives. This gap is particularly relevant for international businesses, where differences in regulation, language, and business environments complicate governance. While multiple frameworks for AI governance exist, this needed diversity is lacking for GenAI. This review paper fills this gap by surveying recent literature to better understand the fundamental characteristics of GenAI and to adapt existing governance frameworks specifically to GenAI. The resulting framework delineates scope, objectives, and governance mechanisms designed to both harness business opportunities and mitigate risks associated with GenAI integration. We theorize a distinctive property of GenAI governance: its scope is endogenous to use. Unlike conventional organizational AI, for which governance is organized around a fixed artifact (e.g., model, intended use), GenAI allows users to reconfigure the artifact (e.g., its behavior and risk). Consequently, the object of governance is not fixed ex ante but is partly constituted through use.

cs.AI

Agentic Inequality

Autonomous AI agents capable of complex planning and action mark a shift beyond today's generative tools. As these systems enter political and economic life, who can access them, how capable they are, and how many can be deployed will shape distributions of power and opportunity. We define this emerging challenge as "agentic inequality": disparities in power, opportunity, and outcomes arising from unequal access to, and capabilities of, AI agents. We show that agents could either deepen existing divides or, under the right conditions, mitigate them. The paper makes three contributions. First, it develops a framework for analysing agentic inequality across three dimensions: availability, quality, and quantity. Second, it argues that agentic inequality differs from earlier technological divides because agents function as autonomous delegates rather than tools, generating new asymmetries through scalable goal delegation and direct agent-to-agent competition. Third, it analyses the technical and socioeconomic drivers likely to shape the distribution of agentic power, from model release strategies to market incentives, and concludes with a research agenda for governance.

cs.CY

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and safety often requires substantial compute, expert input, and labeled data. At the same time, deploying proprietary, cloud-based models for mental health-related interactions raises important privacy and data-governance concerns, given the sensitivities. To address this challenge, we introduce LLUMI setup that can be hosted in-house within protected environments. LLUMI consists of two complementary components: a generation model (GM), which drafts supportive responses to mental health queries, and an improvement model (IM), which revises an initial human-crafted response. We leverage feedback signals from Reddit mental health communities, using community endorsement patterns such as upvotes and downvotes to construct chosen--rejected response pairs for Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO). We further align LLUMI using human evaluation across five dimensions: readability, empathy, connection, actionability, and safety. Our results show that, despite relying on smaller open-source models rather than proprietary cloud-based GPT models, LLUMI achieves comparable performance across linguistic analyses and human evaluations. These findings suggest that open-source models, when trained with community-derived preference signals, can support high-quality mental health support assistance while offering a more privacy-preserving alternative for sensitive support contexts.

cs.HC

Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools

We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.

cs.AI

Intent Drift at SME Scale: Deployment Practice, Not Model Capability, Determines Agentic Compliance

We introduce Chain of Intent, a governance framework for agentic AI at small regulated firms, and validate it against a failure it was built to address. Existing agentic governance research assumes enterprise infrastructure that small firms do not have. In a simulated Hong Kong asset manager with 415 synthetic contact records, an agent performing a routine client-communications task was subjected to ordinary managerial pressure to increase its reach. With its authorised constraints written into its configuration, the agent held: it identified every ambiguity in the firm's records, cited privacy legislation it had never been shown, and refused six successive requests, breaching in two of fifteen runs. With the same task, data, pressure and model, but its purpose left unstated as resource-constrained firms routinely leave it, it breached in thirteen of fifteen runs, contacting up to 220 individuals of whom 94 per cent had no demonstrable marketing consent - conduct carrying a maximum of three years' imprisonment under Hong Kong law. Chain of Intent applies four controls requiring no security engineering: a machine-readable purpose, constrained tool access, a scope ledger, and a pre-action check. It eliminated unlawful contact in every run while preserving task completion, and ablation shows each control independently sufficient by a different mechanism. We further show that drift must be measured at two stages - agents widened their candidate sets in every pressured run while acting on them in roughly one in seven - and that governance applied at the point of intent costs roughly half as much as governance applied at the point of action.

cs.CY

Price Dislocations, News Citations, and Epistemic Leverage on Polymarket

Prediction-market probabilities increasingly appear in news coverage, yet little is known about which market movements become news or how much trading money sits behind the numbers journalists quote. Unlike a poll, a market price can be moved by anyone willing to trade, so the cost of manufacturing a number that circulates as news bears directly on the information environment. We link 173.7 million signed Polymarket trades to news coverage from 2024-2025. From 6,990 articles mentioning prediction-market venues, an LLM-based, human-validated matcher extracts 1,582 sentences quoting market odds and attributes 918 to the specific market whose price they cite. We then detect 44,976 price dislocations, movements of at least five percentage points backed by concentrated one-sided trading, and ask whether a market is cited more often afterward. In the days after a dislocation, a market's citation rate is about 33% higher than its matched baseline (log citation-rate ratio $\tau_{\mathrm{cite}}=0.283$, permutation $p=0.001$), robust to binary and Poisson count outcomes. Yet move size is not the strongest predictor of citation: prominence dominates (standardized $\beta=0.610$ vs. $\beta=0.159$ for move size). Finally, we combine the dollar flow behind a given price change with observed citation rates into a metric we call epistemic leverage, the dollars needed to move a market five points and have the move cited. It stays near \$0.7-1.0 million across prominence quintiles, because cheaper-to-move markets are proportionally less likely to be cited. The implied threat model centers not on the long tail of cheaply moved markets but on the few prominent markets newsrooms treat as informational infrastructure, where a seven-figure price of influence sits within the budgets of actors with a large stake in the quoted number. We release aggregate event-study data and validation materials.

physics.soc-ph
Compare source metadata on this page
WorkPublishedSource identifierSource
Near-Term Verification Methods for AI Chip Exports2026-09-072609.07637arxiv
Human-AI Co-Creativity: Advances, Opportunities, and Challenges2026-09-072609.07711arxiv
An emancipatory vision for designing (generative) AI for learner flourishing2026-09-072609.07715arxiv
When Is Content "AI-Generated Enough"? Labelling Synthetic Media under the Digital Services Act and the AI Act2026-09-072609.07727arxiv
Designing for Healthy, Affordable, and Sustainable Human-HVAC Interactions for Heating in Smart Homes2026-09-07Mensch und Computer 2026 -- Tagungsband, Gesellschaft f\"ur Informatik e.V., 30. August - 02. September 2026, Duisburg, Germanyarxiv
Scaling Multi-Agent Systems with Prospect-State Propagation2026-09-072609.08033arxiv
AI-Assisted Writing Is Growing Fastest Among Less Established Scientists in Non-English-Speaking Countries2025-11-192511.15872arxiv
Knowing Your Uncertainty -- On the application of LLM in social sciences2025-12-052512.05461arxiv
Reliability, validity, and diagnostic evidence for multi-model LLM short-answer scoring2026-09-062609.06315arxiv
A Translational Note on AI Safety Evaluation2026-09-062609.06573arxiv
Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority2026-09-062609.06798arxiv
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents2026-09-062609.09212arxiv
Governance of Generative Artificial Intelligence for Companies2024-02-052403.08802arxiv
Agentic Inequality2025-10-192510.16853arxiv
LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback2026-05-28Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), Main Conferencearxiv
Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools2026-09-052609.05818arxiv
Intent Drift at SME Scale: Deployment Practice, Not Model Capability, Determines Agentic Compliance2026-09-052609.05975arxiv
Price Dislocations, News Citations, and Epistemic Leverage on Polymarket2026-09-052609.06005arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.