arXiv ScienceSearch

arXiv subjects

Ekta Raj

Publications and source records attributed to Ekta Raj.

2 recordsLinked to original sources

Participant-Mediated Collection of Sensitive Digital Trace Data: The CANDOR Research Infrastructure

Digital trace data provide rich measures of behavior in everyday settings, but the research ecosystem supporting their collection is constrained by declining platform API access and a historical reliance on publicly observable data. Participant-mediated data donation offers a complementary approach in which individuals contribute selected portions of their own digital histories to research. Such data can include longitudinal and non-public behavior, span multiple platforms and modalities, and be linked to independently collected study measures, enabling study designs that are difficult to implement using public social media data alone. These opportunities also introduce methodological challenges around participant control, data minimization, heterogeneous platform exports, privacy, and governance, particularly when semantic or multimodal content is necessary to study the construct of interest. We present CANDOR (Collecting and Analyzing Networked Data for Open Research), an end-to-end infrastructure for participant-mediated collection and governance of sensitive digital trace data. CANDOR supports participant-directed selection of platforms, data types, and temporal ranges; modular platform- and modality-specific parsing and de-identification; linkage to independent study measures; and protected processing, storage, and access. We derive design requirements for this class of research and compare CANDOR with existing data donation infrastructures, identifying how different approaches support participant control, data minimization, scientifically necessary data richness, study-design flexibility, and governance. Together, this work provides a methodological and infrastructural framework for using participant-contributed digital traces in behavioral research, particularly when the data needed to address a scientific question are longitudinal, non-public, multimodal, or sensitive.

cs.HC

Do Large Language Models Align with Core Mental Health Counseling Competencies?

The rapid evolution of Large Language Models (LLMs) presents a promising solution to the global shortage of mental health professionals. However, their alignment with essential counseling competencies remains underexplored. We introduce CounselingBench, a novel NCMHCE-based benchmark evaluating 22 general-purpose and medical-finetuned LLMs across five key competencies. While frontier models surpass minimum aptitude thresholds, they fall short of expert-level performance, excelling in Intake, Assessment & Diagnosis but struggling with Core Counseling Attributes and Professional Practice & Ethics. Surprisingly, medical LLMs do not outperform generalist models in accuracy, though they provide slightly better justifications while making more context-related errors. These findings highlight the challenges of developing AI for mental health counseling, particularly in competencies requiring empathy and nuanced reasoning. Our results underscore the need for specialized, fine-tuned models aligned with core mental health counseling competencies and supported by human oversight before real-world deployment. Code and data associated with this manuscript can be found at: https://github.com/cuongnguyenx/CounselingBench

cs.CL