arXiv ScienceSearch

subject

cs.CY

cs.CY: explore 211 source-linked works published from 2022 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-15. Counts describe this index, not the complete source archives.

Students' Perception of Big Data Engineering in Higher Education Curricula: Expectations, Interest and Ethical Implications

The study investigates students' interest and expectations in a Big Data Engineering course integrated with a Master curricula, as well as ethical implications of using Big Data. An anonymous online survey was conducted with 42 of the 67 students enrolled in the Big Data course offered to Computer Science and Bioinformatics Master's programs. The responses were analyzed and interpreted using thematic analysis, highlighting interesting aspects related to students' expectations, interest, and their perspective of the ethical implications of working with Big Data. The study concludes that, even though there is significant difference in students' background, the majority are interested in learning Big Data, for practical and personal reasons related to the potential for career growth and their passion for the field. The main expectation expressed is related to enhancing their knowledge related to Big Data via practical activities. All students demonstrate awareness of potential ethical threats related to security and privacy, while Computer Science students are aware of the possibility of introducing bias in data during acquisition and analysis and of potential abusive data usage.

cs.CY

An Empirical Study on Learning Paths and Gender Dynamics in Scrum Master Roles

Context: Agile development methodology has been widely adopted by industry and the demand for experienced professionals in Agile-related roles is persistently high. Objectives: We focus on the learning path for a Scrum Master role in multicultural software companies and investigate the role in relation to team size, together with the learning process for a career path, and how companies monitor soft skills development. Method: We conducted our study in two phases, two qualitative surveys (interview studies) and performed a qualitative and quantitative data analysis of the results. Conclusions: Our results identified that the need for a Scrum Master (SM) depends on the size of the team, with our study indicating a six-member limit. There is no overall standardized process for soft skills learning or metrics to measure progress. Some companies measure soft skills based on feedback received from the client or from the team, and other companies are taking both types of feedback into consideration. Many learning initiatives, especially on soft skills for an SM role, were based on actions of the employees. The version of Record of this contribution is published in Software Engineering and Advanced Applications. SEAA 2025. Lecture Notes in Computer Science, vol 16083. Springer, Cham. and available online at: check DOI.

cs.SE

Mitigating Disease Spread by Design in Refugee and IDP Camps

Disease spread represents an increasing challenge in refugee and internally displaced person (IDP) settlements. The movement and interaction of people within camps is influenced by their layout, which therefore has the potential to significantly affect disease spread. This work aims at creating a methodology to explore the potential effects of different camp layouts as mitigating factors in the spread of diseases within settlements. We showcase proof-of-concept experiments by leveraging the JUNE agent-based epidemic model, discuss the kind of operational insights this methodology can facilitate, and provide a framework for future investigations.

physics.soc-ph

Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses

As conversational AI becomes a source of everyday guidance, LLMs increasingly participate in the interpretation and legitimation of morally contested choices. We examine LLM moral advice as an interactional negotiation shaped by framing, sustained user pressure, and the moral subject's social position. Using GPT-4o-mini as an illustrative case, we conducted a factorial vignette experiment with a pre-specified three-round protocol. The model received eldercare dilemmas that varied in framing and persona, followed by two user challenges. We analyzed 1,620 configuration-framing cells, each repeated three times, yielding 4,860 conversational runs. Caregiving affirmation produced near-uniform endorsement, whereas non-caregiving framing produced more variable baseline stances. When users challenged caregiving endorsement, 90.1% of configurations shifted after one round. Non-caregiving framing produced more resistant and unstable trajectories. Never (27.9%) and Late (25.6%) accommodations were more common than Early accommodations (16.5%), and only 14.32% of configurations achieved perfect trajectory consistency, compared with 62.72% under caregiving framing. Advice also varied with social position. Female personas received more support for non-caregiving decisions, while the presence of sisters increased accommodation. The GPT-4o-mini case shows that LLM moral advice can develop through a partially stable negotiation between normative response tendencies and user pressure rather than express a fixed ethical framework. The framework and design support comparative research across models and moral domains. Such instability raises social, ethical, and technical concerns, as users may treat advice that is difficult to scrutinize as objective.

cs.CY

From Voice to Leverage in Data Governance. Building Enforceable Social License for Data ReUse

Contemporary data governance has converged on participation. Citizens assemblies, deliberative polls, codesign workshops, and community advisory panels are now routinely suggested to build trust and secure a social license for the reuse of data. Yet this consensus rests on an unexamined assumption: that giving affected communities a voice will, on its own, change what data holders do. This paper argues that the field needs to complement voice with leverage: the capacity of individuals and communities to make their participation consequential by attaching credible legal, institutional, contractual, economic, or reputational costs to being ignored. Drawing on labor relations bargaining theory, the sociology of disputes, and James C. Scott's account of legibility, we develop leverage as an analytic category defined by three properties: it is counterfactual, enforceable, and durable. We further argue that leverage presupposes epistemic legibility affected publics can rarely impose costs on a system they cannot see, document, or contest; and that most existing transparency instruments equip regulators and procurers rather than the governed. We then map six sources from which communities and data subjects can derive leverage (legal and regulatory; institutional decision; contractual; economic; data access; and political and reputational), examine the intermediary structures that aggregate and sustain it, and propose a six dimension evaluative framework: obligation, enforceability, consequence, monitoring equipment, revocability, and institutional durability for auditing whether any social license process actually redistributes power. The paper's central claim is that a social license worth the name must share the properties of a legal license: conditions, a licensor, and enforceable remedies for breach.

cs.CY

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

Affective computing has matured rapidly in laboratory settings, yet no prior dataset combines (i) months-to-years of duration, (ii) a naturalistic workplace context, (iii) a stable small-team social structure, and (iv) a fully passive sensing protocol that survives institutional review. We introduce WELD, the first dataset to satisfy all four. WELD comprises 733,780 per-frame seven-class facial-expression probability vectors from 49 employees of a Chinese software company over 30.1 months (Nov 2021 - May 2024) -- the longest naturalistic in-the-wild emotion corpus and the only multi-year corpus supporting both within-individual longitudinal and within-team relational analyses on the same subjects. Data are released under a four-tier access model with only aggregated probabilities publicly downloadable. We validate the corpus by replicating three established phenomena (+43.1% weekend valence boost; 13:00-trough diurnal cycle; Shanghai 2022 lockdown effect d=-0.40), and report four novel findings: (1) variance decomposition attributes 19.3% of daily-valence variance to between-person differences and 29.8% to month seasonality -- a quantitative ceiling for future predictive models; (2) Hidden Markov decomposition reveals six emotional regimes with asymmetric negative-state dwell times (16-18 d vs 3 d); (3) leave-one-person-out turnover prediction reaches AUC=0.79 yet a Cox concordance index of only 0.52, exposing a metric-trap when AUC is reported without survival-aware baselines; (4) the corpus reveals systematic over-prediction of "angry" by an off-the-shelf FER model on neutral Asian faces (0.194 vs ~0.05 Western priors), making WELD valuable for FER fairness audits. A complex-systems analysis of the corpus appears as a companion preprint (arXiv:2510.16046).

cs.AI

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

We present CARDIO-Affect, a complex-systems theoretical framework for long-term emotional dynamics in bounded social groups, with explicit uncertainty quantification at every layer. Long-period naturalistic emotion in stable small groups exhibits hallmarks of complex systems -- multi-stable attractors, weak chaos, long-range memory, and sparse heterogeneous coupling -- invisible to conventional short-clip facial-emotion analysis. CARDIO-Affect treats individual emotion as a multi-stable nonlinear stochastic dynamical system and group emotion as a sparsely-coupled network with emergent macrostates, formalised through six propositions and four pillars: (i) statistical mechanics with neural-parameterised Hamiltonian SDE over asymmetric potentials; (ii) information geometry on a 45-dimensional Fisher-Rao manifold; (iii) topological data analysis for invariant trajectory signatures; (iv) HRV-inspired Emotional Variability Analytics (EVA) decomposing each person-day into multi-scale time/frequency/nonlinear measures. We validate on the first 30.1-month longitudinal in-the-wild facial-emotion corpus (companion: arXiv:2510.15221) by discovering three falsifiable paradoxes: Sparse-Contagion (R_0=0.36, density 2.7%, 8 BH-FDR edges), Asymmetric-Persistence (negative dwell 5.85x positive, 1.77D potential gap), and Crisis-Inversion (Shanghai 2022 lockdown naive d=-0.40 collapses to permutation-p=0.94 under BSTS + synthetic-control). On synthetic benchmarks, CARDIO-EBM v2 matches asymptotically optimal Granger on linear VAR data (Class A AUROC 0.984+/-0.012 vs Granger 0.997+/-0.001, 5 seeds) but fails on tanh-coupled nonlinear data (Class B AUROC 0.490 vs Granger 0.796), a documented limitation of the linear mask-self estimator. We release framework code and the full reproduction pipeline.

physics.soc-ph

Assessing Autonomous Mobility-on-Demand Services and the Impacts of Operational Strategies: A Case Study of Chengdu, China

The Autonomous Mobility-on-Demand (AMoD) service is emerging as a potential alternative to on-demand urban mobility, but its operational performance relative to traditional street-hailing services and the effectiveness of related operational strategies remain unclear. This study presents a simulation framework integrating a graph theory-based trip-vehicle matching mechanism and uses historical street-hailing operations data to simulate AMoD services in Chengdu, China. The operational performance of these two urban mobility modes is evaluated using three key performance indicators: average passenger waiting time (APWT), average deadheading mileage (ADM), and average deadheading energy consumption (ADEC). We further evaluate the impacts of four operational strategies on simulated AMoD performance: vehicle repositioning, fleet size management, geofencing, and request rejection. Simulation results indicate that, under the same historical trip demand, fleet-size constraints, and road network as the observed street-hailing system, the simulated AMoD service is estimated to have lower values of APWT, ADM, and ADEC by 73.3% to 83.4%, 75.0%, and 74.0%, respectively, reflecting the potential operational gains associated with centralized dispatch in simulation settings. These differences are most pronounced during early-morning low-demand hours and in remote areas such as airports.

math.OC

A Visionary Look at Vibe Researching

Vibe researching is an emerging paradigm in which human researchers provide high-level direction and critical judgment while LLM-based agents handle the labor-intensive execution of literature review, experimentation, data analysis, and manuscript drafting. Inspired by the "vibe coding" movement in software engineering, it occupies a middle ground between traditional manual research and fully autonomous AI research systems. This paper defines the concept, describes its methodology (multi-agent architectures, memory, tool use, retrieval-augmented generation, and the human's role as orchestrator), identifies seven technical limitations, weighs its positive and negative societal impacts, and maps each problem to a concrete future direction. Our goal is to provide the research community with a clear and honest map of the territory so that the conversation about responsible adoption can start from shared ground.

cs.CY

Identifying AI Web Scrapers Using Canary Tokens

From pre-training to query-time augmentation, web-scraped data helps to improve the quality and contextual relevancy of content generated by large language models (LLMs). However, large-scale web scraping to feed LLMs can affect site stability and raise legal, privacy, or ethics concerns. If website owners wish to limit LLM-related web scraping on their site, due to these or other concerns, they may turn to scraper access control mechanisms like the Robots Exclusion Protocol. To be most effective, such mechanisms require site owners to first identify the scrapers that they wish to restrict (e.g., via User-Agent strings). Existing mechanisms to identify LLM-related scrapers rely on voluntary disclosure by companies, one-off experiments by researchers, or crowd-sourced reports -- methods that are neither reliable nor scalable. This paper proposes a novel technique for accurately and automatically inferring LLM-related scrapers. We host dynamic websites that serve unique canary tokens to each visiting scraper, then prompt LLMs for information about our sites. If an LLM consistently generates outputs containing tokens unique to a scraper, it provides evidence of exposure to that scraper. Via experiments across 22 production LLM systems, we demonstrate that our approach can reliably identify which scrapers feed which LLM, including several that are not publicly known or disclosed by the companies. Our approach provides a promising avenue for unprivileged third parties to infer which scrapers serve data to which LLMs, potentially enabling better control over unwanted scraping.

cs.CR

MIRA: A Bilingual Benchmark for Medical Information Response Audit

Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent judgment. We term this pattern Differential Information Dilution (DID). A comparison with 300 real-world health queries provides preliminary evidence of rank-order validity. A knowledge-guided mitigation prompt reduces information dilution for most models, with the largest reductions in underinformative simplification observed for Claude (~8%) and Qwen (~6%). Code and data are available at https://github.com/Rainxu09/MIRA.

cs.AI

Affective publics in Arabic YouTube

What is the emotional register of Arabic YouTube's affective publics? To investigate this, we analyzed 67,725 YouTube comments collected around socio-political topics associated with Yemen, Saudi Arabia, Iraq, Jordan, and Syria using a unified sentiment-and-emotion pipeline. Our results profile a single regional affective public rather than five separate national ones. Sentiment is overwhelmingly negative across all five country-oriented corpora, and the country-level emotion profiles are structurally similar. This shared register still accommodates some regional variations: discourse is organized around country-level political actors and cross-border historical trauma figures, and grief singularizes Iraq from the other countries. The differences in emotional register also tracks lived political causes rather than fixed categories, which we observe from patterns of the valence of US-related content, that tracks the presence or absence of direct US military engagement. Our work shows that the emotional register of Arabic YouTube's affective publics is a shared one that is historically layered and geographically conditioned, which has implications for public diplomacy in the MENA region.

cs.CY

Large-Language Models as a Cognitive Virus

Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how information is produced, transmitted, and used. Here we propose that their diffusion can be understood through a viral analogy, with LLM use spreading through populations, becoming embedded in cognitive and cultural practices. We model transitions among uncoupled, coupled, and persistently dependent users, and show that the interplay between social transmission, recovery, and collective reinforcement can generate tipping points and technological lock-in. A central consequence is the possibility of runaway dynamics: once a critical threshold is crossed, small increases in adoption can trigger rapid population-level shifts toward persistent dependence, with abrupt losses in cognitive competence. The same framework, however, identifies conditions for cognitive immunization, based on reducing transmission and facilitating reversibility. Our results highlight how LLM adoption may involve nonlinear collective transitions with important consequences for cognitive autonomy.

physics.soc-ph

Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT

Accountability means a decision can be examined, justified, and contested. LLMs make this hard: fluent output may be ungrounded, incomplete, or unfaithful to the decision process. Achieving accountability requires verified rationales (how was the decision reached), assumptions (what was assumed rather than known), policy consistency (the same treatment for the same facts), and pivotal conditions (what would change the outcome). We introduce self-faithfulness as an automatic test of accountability: changing the pivotal conditions should change the decision. We examine accountable AI through clinical trial matching, a high-stakes task central to evidence-based medicine. Although LLM-based matchers match patients to trials reasonably accurately, they apply decision policies inconsistently and produce rationales that are unfaithful to their own decisions. We introduce VERDICT, an LLM-based agent that translates a decision task, its constraints, and its policy into Satisfiability Modulo Theories (SMT), then derives the decision with SMT and MaxSMT solvers -- so policies are applied consistently and decisions are accountable by construction. Across a SIGIR 2016-derived dataset and TREC 2021, VERDICT achieves the strongest decision accuracy among LLM-only and neurosymbolic baselines, applies policies with perfect consistency, and produces clinician-preferred rationales grounded in explicit assumptions and pivotal conditions, with improved counterfactual self-faithfulness.

cs.CL

The 5P Reflection Model for Education in the Generative Artificial Intelligence (GenAI) Era

Contributions: A reflection model suitable for the era of Generative Artificial Intelligence (GenAI) is introduced. The proposed model is an integrated model that extracts features from various existing models and also incorporates technological aspects of GenAI. Background: Universities worldwide are facing challenges in adopting GenAI into their curricula, as it has impacted academic integrity and the scholarship of teaching and research. Traditional reflection models are struggling to authenticate student reflection as GenAI is incorporated in education. This requires a GenAI-aware model to enable the opportunities that address the associated challenges with GenAI. Research Questions: Does the academic ecosystem require a GenAI-aware reflection model to adopt GenAI into education? How to make a reflection model structured to ensure student authenticity and cognitive engagement within a GenAI-aware learning environment? Methodology: This study employs a design-based research methodology, assisted by a critical inquiry approach, to analyse existing reflection models in the era of GenAI and examine the technological aspects of GenAI. Additionally, it identifies a gap and the absence of a comprehensive reflection model that enhances reflection in the era of GenAI and supports the effective integration of GenAI in education. Findings: The literature survey indicates there is a greater need for reflection in the era of GenAI to ensure learning, but the existing reflection models lack proper strategies to manage the problem introduced by GenAI. A standard GenAI-aware reflection model, called 5P (Purpose, Process, Product, Pitfalls and Plan), is proposed i) to manage greater demand for reflection, ii) to address the limitations of the existing reflection models, iii) to consider technological aspects of GenAI.

cs.CY

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns "does multi-agent simulation help?" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.

cs.AI

OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education

Institutions practising outcome-based education compute learning outcome attainment routinely, while reviews of curriculum analytics report an absence of evidence on how that computation informs decisions. This paper presents OBER+, an extension of a deployed institutional attainment platform that computes the step from a measured shortfall to an evaluated corrective action. Five connected stages accumulate attainment across deliveries of a course, signal a shortfall and a persistent shortfall, grade it on cutoffs the regulator already uses, record the decision against a catalogue of practices annotated with their evidence, log the change, and quantify the subsequent movement in the shortfall. A further rule compares successive statements of an outcome, so attainment is never read as a series across a point at which the outcome changed. Applying the rules to the live record of two real courses produced three results. Every outcome of a core course was substantively redefined between consecutive deliveries, with subject matter moving between outcome numbers, so a naive reading would have reported a twenty-five point collapse between quantities that do not refer to the same learning. Recomputing the platform's figures from its documented rule showed six of ten differing by more than rounding explains, in a pattern that identified a defect since reported to the institution. Across fifteen statement pairs from three transitions, five were identical character for character, and among the ten that were not, the outcome carrying a given number was nearest to a differently numbered earlier outcome in six, a result resting on an ordering of similarities and requiring no threshold and no labelling. The contribution is a computational design for outcome-based reporting, stated as rules any attainment platform can implement, with evidence of what they make visible in a live institutional record.

cs.LG

Bridging Formal and Perceived Fairness: Development of an Interdisciplinary Framework in Algorithmic Decision-Making

While fairness has become a central concern in research on algorithmic systems, the field remains predominantly shaped by Computer Science, resulting in a strong emphasis on formal fairness metrics and bias mitigation strategies. Nevertheless, this focus may obscure a fundamental challenge: fairness is not merely a technical property, but a subjective, context-sensitive human judgment shaped by cognitive heuristics, mental models, normative expectations, and sociotechnical factors. Crucially, users' perceptions of fairness may diverge substantially from the fairness criteria an algorithm formally satisfies; a system may meet predefined technical fairness requirements yet still be perceived as unjust by decision-affected stakeholders. In such cases, the system fails on a fundamental dimension: it will not be trusted, accepted, or considered legitimate. Taking a user-centered design perspective, this paper presents a work-in-progress conceptual framework that bridges Computer Science approaches to formal algorithmic fairness with normative and Social Science fairness approaches regarding perceived fairness, trust, and technology acceptance, embedding both within the sociotechnical conditions that shape human judgment. Through (1) theoretical literature synthesis, (2) interdisciplinary workshops, and (3) stakeholder interviews, the project aims to inform evaluation approaches that integrate computational fairness audits with user-centered assessments and guide the design of fairness-aware, human-centered algorithmic systems that support informed, well-calibrated fairness judgments by those affected.

cs.CY
Compare source metadata on this page
WorkPublishedSource identifierSource
Students' Perception of Big Data Engineering in Higher Education Curricula: Expectations, Interest and Ethical Implications2026-09-042609.05160arxiv
An Empirical Study on Learning Paths and Gender Dynamics in Scrum Master Roles2026-09-042609.05186arxiv
Mitigating Disease Spread by Design in Refugee and IDP Camps2026-09-042609.05342arxiv
Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses2026-09-042609.05345arxiv
From Voice to Leverage in Data Governance. Building Enforceable Social License for Data ReUse2026-09-042609.05731arxiv
WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing2026-09-032510.15221arxiv
CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling2026-09-032510.16046arxiv
Assessing Autonomous Mobility-on-Demand Services and the Impacts of Operational Strategies: A Case Study of Chengdu, China2026-09-032511.06074arxiv
A Visionary Look at Vibe Researching2026-09-032604.00945arxiv
Identifying AI Web Scrapers Using Canary Tokens2026-09-032605.13706arxiv
MIRA: A Bilingual Benchmark for Medical Information Response Audit2026-09-032605.28025arxiv
Affective publics in Arabic YouTube2026-09-032609.03269arxiv
Large-Language Models as a Cognitive Virus2026-09-032609.03344arxiv
Accountable AI with Grounded, Faithful, Consistent, Actionable Rationales: A Case Study in Clinical Trial Matching with VERDICT2026-09-032609.03366arxiv
The 5P Reflection Model for Education in the Generative Artificial Intelligence (GenAI) Era2026-09-032609.03413arxiv
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis2026-09-032609.03553arxiv
OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education2026-09-032609.03770arxiv
Bridging Formal and Perceived Fairness: Development of an Interdisciplinary Framework in Algorithmic Decision-Making2026-09-032609.03853arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.