arXiv ScienceSearch

arXiv subjects

Brett Smith

Publications and source records attributed to Brett Smith.

4 recordsLinked to original sources

Biomedical knowledge graph-optimized prompt generation for large language models

Large Language Models (LLMs) are being adopted at an unprecedented rate, yet still face challenges in knowledge-intensive domains like biomedicine. Solutions such as pre-training and domain-specific fine-tuning add substantial computational overhead, requiring further domain expertise. Here, we introduce a token-optimized and robust Knowledge Graph-based Retrieval Augmented Generation (KG-RAG) framework by leveraging a massive biomedical KG (SPOKE) with LLMs such as Llama-2-13b, GPT-3.5-Turbo and GPT-4, to generate meaningful biomedical text rooted in established knowledge. Compared to the existing RAG technique for Knowledge Graphs, the proposed method utilizes minimal graph schema for context extraction and uses embedding methods for context pruning. This optimization in context extraction results in more than 50% reduction in token consumption without compromising the accuracy, making a cost-effective and robust RAG implementation on proprietary LLMs. KG-RAG consistently enhanced the performance of LLMs across diverse biomedical prompts by generating responses rooted in established knowledge, accompanied by accurate provenance and statistical evidence (if available) to substantiate the claims. Further benchmarking on human curated datasets, such as biomedical true/false and multiple-choice questions (MCQ), showed a remarkable 71% boost in the performance of the Llama-2 model on the challenging MCQ dataset, demonstrating the framework's capacity to empower open-source models with fewer parameters for domain specific questions. Furthermore, KG-RAG enhanced the performance of proprietary GPT models, such as GPT-3.5 and GPT-4. In summary, the proposed framework combines explicit and implicit knowledge of KG and LLM in a token optimized fashion, thus enhancing the adaptability of general-purpose LLMs to tackle domain-specific questions in a cost-effective fashion.

cs.CL

Staring at the Sun with the Keck Planet Finder: An Autonomous Solar Calibrator for High Signal-to-Noise Sun-as-a-Star Spectra

Extreme precision radial velocity (EPRV) measurements contend with internal noise (instrumental systematics) and external noise (intrinsic stellar variability) on the road to 10 cm/s "exo-Earth" sensitivity. Both of these noise sources are well-probed using "Sun-as-a-star" RVs and cross-instrument comparisons. We built the Solar Calibrator (SoCal), an autonomous system that feeds stable, disc-integrated sunlight to the recently commissioned Keck Planet Finder (KPF) at the W. M. Keck Observatory. With SoCal, KPF acquires signal-to-noise ~1200, R = ~98,000 optical (445--870 nm) spectra of the Sun in 5~sec exposures at unprecedented cadence for an EPRV facility using KPF's fast readout mode (<16 sec between exposures). Daily autonomous operation is achieved by defining an operations loop using state machine logic. Data affected by clouds are automatically flagged using a reliable quality control metric derived from simultaneous irradiance measurements. Comparing solar data across the growing global network of EPRV spectrographs with solar feeds will allow EPRV teams to disentangle internal and external noise sources and benchmark spectrograph performance. To facilitate this, all SoCal data products are immediately available to the public on the Keck Observatory Archive. We compared SoCal RVs to contemporaneous RVs from NEID, the only other immediately public EPRV solar dataset. We find agreement at the 30-40 cm/s level on timescales of several hours, which is comparable to the combined photon-limited precision. Data from SoCal were also used to assess a detector problem and wavelength calibration inaccuracies associated with KPF during early operations. Long-term SoCal operations will collect upwards of 1,000 solar spectra per six-hour day using KPF's fast readout mode, enabling stellar activity studies at high signal-to-noise on our nearest solar-type star.

astro-ph.IM

First faint dual-field phase-referenced observations on the Keck interferometer

Ground-based long baseline interferometers have long been limited in sensitivity by the short integration periods imposed by atmospheric turbulence. The first observation fainter than this limit was performed on January 22, 2011 when the Keck Interferometer observed a K=11.5 target, about one magnitude fainter than its K=10.3 limit. This observation was made possible by the Dual Field Phase Referencing instrument of the ASTRA project: simultaneously measuring the real-time effects of the atmosphere on a nearby bright guide star, and correcting for it on the faint target, integration time longer than the turbulence time scale are made possible. As a prelude to this demonstration, we first present the implementation of Dual Field Phase Referencing on the interferometer. We then detail its on-sky performance focusing on the accuracy of the turbulence correction, and on the resulting fringe contrast stability. We conclude with a presentation of early results obtained with Laser Guide Star AO and the interferometer.

astro-ph.IM

Astrometry with the Keck-Interferometer: the ASTRA project and its science

The sensitivity and astrometry upgrade ASTRA of the Keck Interferometer is introduced. After a brief overview of the underlying interferometric principles, the technology and concepts of the upgrade are presented. The interferometric dual-field technology of ASTRA will provide the KI with the means to observe two objects simultaneously, and measure the distance between them with a precision eventually better than 100 uas. This astrometric functionality of ASTRA will add a unique observing tool to fields of astrophysical research as diverse as exo-planetary kinematics, binary astrometry, and the investigation of stars accelerated by the massive black hole in the center of the Milky Way as discussed in this contribution.

astro-ph