arXiv ScienceSearch

arXiv subjects

Jianing Qiu

Publications and source records attributed to Jianing Qiu.

3 recordsLinked to original sources

GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space

Large Language Models (LLMs) show great potential as clinical agents, yet existing benchmarks reduce clinical workflows to static predictions or unconstrained Markov Decision Processes (MDPs) with coarse action sets. To address this, we introduce GPAgentBench-2K, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters. Our environment models a full spectrum of six foundational clinical actions, imposes a topological workflow prior over the action space, and operationalizes safety-informed abstention as a first-class outcome. Evaluating 16 state-of-the-art LLMs reveals a significant performance degradation as the action space scales. Crucially, we uncover a clinical quality-safety gap: even frontier models with the highest diagnosis accuracy violate safety constraints in over half of high-risk cases. Finally, we establish a reference point using Constrained Group Relative Policy Optimization (C-GRPO), and show that while explicitly modeling constraints improves performance over unconstrained RL methods, it remains far from clinically acceptable safety.

cs.CL

Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective

Stochastic sampling strategies are widely adopted in large language models (LLMs) to balance output coherence and diversity. These heuristics are often inherited in Multimodal LLMs (MLLMs) without task-specific justification. However, we contend that stochastic decoding can be suboptimal for Visual Question Answering (VQA). VQA is a closed-ended task with head-heavy answer distributions where uncertainty is usually epistemic, arising from missing or ambiguous visual evidence rather than plausible continuations. In this work, we provide a theoretical formalization of the relationship between model calibration and predictive accuracy, and derive the sufficient conditions for greedy decoding optimality. Extensive experiments provide empirical evidence for the superiority of greedy decoding over stochastic sampling across multiple benchmarks. Furthermore, we propose Greedy Decoding for Reasoning Models, which outperforms both stochastic sampling and standard greedy decoding in multimodal reasoning scenarios. Overall, our results caution against naively inheriting LLMs decoding heuristics in MLLMs and demonstrate that greedy decoding can be an efficient yet strong default for VQA.

cs.CL

CaseWeaver: A Multi-Agent Framework for Multimodal Virtual Clinical Case Generation

Clinical diagnosis relies on consistent multimodal data collected from the same patient throughout the disease course, yet such data are difficult to acquire at scale because of collection costs, missing modalities, fragmented systems, and longitudinal follow-ups. Existing synthetic-data approaches largely focus on individual modalities or vision-language dual modalities at report-level generation. Little work has been done to construct synthetic data with consistent patient backgrounds, coherent disease trajectories, and interrelated modality-specific evidence at a complete clinical case level. We introduce CaseWeaver, a multi-agent framework built around a timeline-anchored Latent Clinical Case Graph (LCCG). The LCCG organizes patient context, latent disease states, clinical events, and expected observations in a shared patient-level representation. Modality-agents use scoped observation subgraphs and clinical protocols to generate evidence including clinical records, laboratory results, physiological signals, and medical images. We evaluate clinical inferability using a calibrated AgentClinic protocol and case diversity using Virtual Case Diversity (VCD) score. CaseWeaver outperformed general-model and agentic-workflow baselines on both metrics, producing more diverse and coherent multimodal virtual clinical cases.

cs.MA