arXiv ScienceSearch

arXiv subjects

Ruth Etzioni

Publications and source records attributed to Ruth Etzioni.

7 recordsLinked to original sources

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection

Modeling patient trajectories from longitudinal electronic health records (EHRs) requires reasoning over sparse, noisy, and long-context multimodal sequences. Existing LLM-based multi-agent systems address context length but process patients in isolation, failing to mirror how clinicians leverage accumulated experience from similar prior cases. We present Traj-Evolve, a self-evolving multi-agent system with two complementary evolving mechanisms. First, an Experience Pool (ExPool) acts as a non-parametric memory, indexing rejection-sampled reasoning traces to retrieve similar patients as few-shot contexts. Second, multi-agent reinforcement learning (MARL) via reward-ranked fine-tuning parametrically optimizes inter-agent and agent-memory collaboration. A leave-one-out cross-retrieval strategy unifies the two, aligning training- and inference-time behavior under retrieval augmentation. On a lung cancer prediction task utilizing up to five years of multimodal EHRs, Traj-Evolve outperforms 9 strong baselines on the overall population and a challenging never-smoker population. Analysis of the evolving dynamics highlights three key findings: (1) expanding the ExPool shifts optimal retrieval from diverse to specific samples; (2) under MARL, the manager agent's prediction loss converges quickly while the worker agents' temporal reasoning continues to benefit from more verified patients; and (3) the two mechanisms are complementary on the predicted risk, where ExPool improves specificity while MARL improves sensitivity.

cs.AI

TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection

Accurate estimation of cancer risk from longitudinal electronic health records (EHRs) could support earlier detection and improved care, but modeling such complex patient trajectories remains challenging. We present TrajOnco, a training-free, multi-agent large language model (LLM) framework designed for scalable multi-cancer early detection. Using a chain-of-agents architecture with long-term memory, TrajOnco performs temporal reasoning over sequential clinical events to generate patient-level summaries, evidence-linked rationales, and predicted risk scores. We evaluated TrajOnco on de-identified Truveta EHR data across 15 cancer types using matched case-control cohorts, predicting risk of cancer diagnosis at 1 year. In zero-shot evaluation, TrajOnco achieved AUROCs of 0.64-0.80, performing comparably to supervised machine learning in a lung cancer benchmark while demonstrating better temporal reasoning than single-agent LLMs. The multi-agent design also enabled effective temporal reasoning with smaller-capacity models such as GPT-4.1-mini. The fidelity of TrajOnco's output was validated through human evaluation. Furthermore, TrajOnco's interpretable reasoning outputs can be aggregated to reveal population-level risk patterns that align with established clinical knowledge. These findings highlight the potential of multi-agent LLMs to execute interpretable temporal reasoning over longitudinal EHRs, advancing both scalable multi-cancer early detection and clinical insight generation.

cs.AI

Traj-CoA: Patient Trajectory Modeling via Chain-of-Agents for Lung Cancer Risk Prediction

Large language models (LLMs) offer a generalizable approach for modeling patient trajectories, but suffer from the long and noisy nature of electronic health records (EHR) data in temporal reasoning. To address these challenges, we introduce Traj-CoA, a multi-agent system involving chain-of-agents for patient trajectory modeling. Traj-CoA employs a chain of worker agents to process EHR data in manageable chunks sequentially, distilling critical events into a shared long-term memory module, EHRMem, to reduce noise and preserve a comprehensive timeline. A final manager agent synthesizes the worker agents' summary and the extracted timeline in EHRMem to make predictions. In a zero-shot one-year lung cancer risk prediction task based on five-year EHR data, Traj-CoA outperforms baselines of four categories. Analysis reveals that Traj-CoA exhibits clinically aligned temporal reasoning, establishing it as a promisingly robust and generalizable approach for modeling complex patient trajectories. Implementation of Traj-CoA is available on https://github.com/zengsihang/Traj-CoA.

cs.AI

TrajSurv: Learning Continuous Latent Trajectories from Electronic Health Records for Trustworthy Survival Prediction

Trustworthy survival prediction is essential for clinical decision making. Longitudinal electronic health records (EHRs) provide a uniquely powerful opportunity for the prediction. However, it is challenging to accurately model the continuous clinical progression of patients underlying the irregularly sampled clinical features and to transparently link the progression to survival outcomes. To address these challenges, we develop TrajSurv, a model that learns continuous latent trajectories from longitudinal EHR data for trustworthy survival prediction. TrajSurv employs a neural controlled differential equation (NCDE) to extract continuous-time latent states from the irregularly sampled data, forming continuous latent trajectories. To ensure the latent trajectories reflect the clinical progression, TrajSurv aligns the latent state space with patient state space through a time-aware contrastive learning approach. To transparently link clinical progression to the survival outcome, TrajSurv uses latent trajectories in a two-step divide-and-conquer interpretation process. First, it explains how the changes in clinical features translate into the latent trajectory's evolution using a learned vector field. Second, it clusters these latent trajectories to identify key clinical progression patterns associated with different survival outcomes. Evaluations on two real-world medical datasets, MIMIC-III and eICU, show TrajSurv's competitive accuracy and superior transparency over existing deep learning methods.

cs.LG

A general two-stage progressive model of cancer natural history to project downstaging due to multi-cancer screening tests

Multi-cancer early detection (MCED) tests offer to screen for multiple types of cancer with a single blood sample. Despite their promising diagnostic performance, evidence regarding their population benefit is not yet available. Expecting that benefit will derive from detecting cancer before it progresses to an advanced stage, we develop a general two-stage model to project the reduction in advanced-stage diagnoses given stage-specific test sensitivities and testing ages. The model can be estimated using cancer registry data and values for the mean overall and advanced-stage preclinical sojourn times. We first estimate the model for lung cancer and validate it against the stage shift observed in the National Lung Screening Trial. We then estimate the model for liver, pancreas, and bladder cancer, which have no recommended screening tests, and we project stage shifts under a shared MCED testing protocol. Our framework transparently integrates available data to project reductions in advanced-stage diagnoses due to MCED testing.

stat.AP

The price elasticity of Gleevec in patients with Chronic Myeloid Leukemia enrolled in Medicare Part D: Evidence from a regression discontinuity design

Objective To assess the price elasticity of branded imatinib in chronic myeloid leukemia (CML) patients on Medicare Part D to determine if high out-of-pocket payments (OOP) are driving the substantial levels of non-adherence observed in this population. Data sources and study setting We use data from the TriNetX Diamond Network (TDN) United States database for the period from first availability in 2011 through the end of patent exclusivity following the introduction of generic imatinib in early 2016. Study design We implement a fuzzy regression discontinuity design to separately estimate the effect of Medicare Part D enrollment at age 65 on adherence and OOP in newly-diagnosed CML patients initiating branded imatinib. The corresponding price elasticity of demand (PED) is estimated and results are assessed across a variety of specifications and robustness checks. Data collection/extraction methods Data from eligible patients following the application of inclusion and exclusion criteria were analyzed. Principal findings Our analysis suggests that there is a significant increase in initial OOP of $232 (95% Confidence interval (CI): $102 to $362) for individuals that enrolled in Part D due to expanded eligibility at age 65. The relatively smaller and non-significant decrease in adherence of only 6 percentage points (95% CI: -0.21 to 0.08) led to a PED of -0.02 (95% CI: -0.056, 0.015). Conclusion This study provides evidence regarding the financial impact of coinsurance-based benefit designs on Medicare-age patients with CML initiating branded imatinib. Results indicate that factors besides high OOP are driving the substantial non-adherence observed in this population and add to the growing literature on PED for specialty drugs.

econ.EM

Accounting for Uncertainty During a Pandemic

We discuss several issues of statistical design, data collection, analysis, communication, and decision making that have arisen in recent and ongoing coronavirus studies, focusing on tools for assessment and propagation of uncertainty. This paper does not purport to be a comprehensive survey of the research literature; rather, we use examples to illustrate statistical points that we think are important.

stat.AP