arXiv ScienceSearch

arXiv · 2506.03655

Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability

Abstract

The proliferation of misinformation necessitates scalable, automated fact-checking solutions. Yet, current benchmarks often overlook multilingual and topical diversity. This paper introduces a novel, dynamically extensible data set that includes 61,514 claims in multiple languages and topics, extending existing datasets up to 2024. Through a comprehensive evaluation of five prominent Large Language Models (LLMs), including GPT-4o, GPT-3.5 Turbo, LLaMA 3.1, and Mixtral 8x7B, we identify significant performance gaps between different languages and topics. While overall GPT-4o achieves the highest accuracy, it declines to classify 43% of claims. Across all models, factual-sounding claims are misclassified more often than opinions, revealing a key vulnerability. These findings underscore the need for caution and highlight challenges in deploying LLM-based fact-checking systems at scale.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lorraine Saju, Arnim Bleier, Jana Lasser, Claudia Wagner. 2025-10-21. Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability. https://arxiv.org/abs/2506.03655

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Mindspeller Neuroprofiling. How task performance, EEG, and association evidence support O*NET-based role guidance

Mindspeller produces a Neuroprofile from three sources: rational self-report, association-based semantic positioning, and performance recorded during cognitive tasks together with EEG. The task-and-EEG stream is the only source used to create occupational evidence. A result can enter role matching only after the participant's task performance supports the intended construct and the corresponding EEG data pass the required quality and evidence checks. The current pilot uses four EEG electrodes, twelve scored tasks, 23 O*NET abilities, and an internal bank of 376 occupations. Self-report and association evidence help explain motivation, preference, and alignment, but do not generate roles. The output is intended to support discussion about cognitive fit. It is not a hiring decision, a measure of practical job skill, or a prediction of job performance. Role confidence is currently capped at Moderate, and external psychometric and job-outcome validity have not yet been established.

cs.CY

From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development

Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation], Phase III [large-scale validation], and Phase IV [post-marketing surveillance]), highlighting the distinct characteristics and interconnections. Major challenges are identified, including ethical compliance, participant recruitment, and ensuring diversity and representativeness in trial populations, while proposing evidence-based mitigation strategies. To address these challenges, innovative technologies, such as artificial intelligence, big data analytics, and digital health tools, are transforming trial design and implementation, enhancing efficiency and data quality. Looking forward, the review explores how emerging therapies, including gene therapy and immunotherapy, are reshaping trial design requirements and emphasizes the growing importance of regulatory harmonization and global collaboration. Clinical trials remain central to advancing innovative drug development and improving patient outcomes.

cs.CY

Why we need an AI-resilient society- Profiling Large Language Models

Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit logic. In the second, neural networks learned programs from data. In the third, large language models turn natural language itself into a programming interface. These shifts reach far beyond computer science, reshaping how societies generate knowledge, make decisions, and govern themselves. While generative adversarial networks introduced the era of deepfakes and synthetic media, large language models have added a new class of systemic risks. This report performs "mindhunting" for LLMs by applying a forensic-psychology profiling methodology to characterize AI based on documented features, e.g., hallucinations, bias and toxicity, sycophancy, fabrication and confabulation, knowledge without understanding, discontinuity and the inability to learn from experience, jagged intelligence, shortcuts and fractured representations. The resulting profile reveals an "entity" that confabulates fluently, amplifies its users' biases, possesses encyclopedic recall without causal understanding, and erodes the competence of those who depend on it. The implications extend to institutional erosion across law, academia, journalism, and democratic governance. To address these challenges, this report proposes a four-pillar framework for AI resilience: (i) cognitive sovereignty, which preserves the capacity for independent judgment, (ii) measurable control, which translates ethical commitments into enforceable standards and red lines, (iii) partial autonomy, which maintains human agency at critical decision points, and (iv) openness to guarantee transparency and accessibility (open-source, open-access, and open-data). This report is an updated and extended version of arXiv:1912.08786v1.

cs.CY