arXiv ScienceSearch

arXiv subjects

Xintian Sun

Publications and source records attributed to Xintian Sun.

5 recordsLinked to original sources

From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development

Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation], Phase III [large-scale validation], and Phase IV [post-marketing surveillance]), highlighting the distinct characteristics and interconnections. Major challenges are identified, including ethical compliance, participant recruitment, and ensuring diversity and representativeness in trial populations, while proposing evidence-based mitigation strategies. To address these challenges, innovative technologies, such as artificial intelligence, big data analytics, and digital health tools, are transforming trial design and implementation, enhancing efficiency and data quality. Looking forward, the review explores how emerging therapies, including gene therapy and immunotherapy, are reshaping trial design requirements and emphasizes the growing importance of regulatory harmonization and global collaboration. Clinical trials remain central to advancing innovative drug development and improving patient outcomes.

cs.CY

From In Silico to In Vitro: A Comprehensive Guide to Validating Bioinformatics Findings

Translating computational predictions into experimentally validated biological knowledge remains one of the central challenges in modern bioinformatics. This review critically examines how in silico findings can be prioritized, tested, and interpreted through experimental validation. We organize the validation process around three recurring challenges: the specificity gap between genome-wide computational predictions and single-target experimental assays; the reproducibility-translatability tension, in which results validated in one model system may not generalize to another; and the scale-depth trade-off between high-throughput candidate discovery and the low-throughput nature of mechanistic validation. Rather than presenting an encyclopedic catalogue of techniques, we compare the strengths, limitations, and common failure modes of major validation approaches, including qPCR, RNA-seq, Western blotting, co-immunoprecipitation, luciferase reporter assays, CRISPR perturbation, and functional phenotypic assays. We also provide structured comparison tables for gene expression, protein-protein interaction, non-coding RNA, regulatory element, and pathway validation, together with decision-making frameworks to guide method selection according to prediction type, biological context, evidence stringency, throughput, and resource constraints. Case studies from cancer genomics, drug target discovery, miRNA regulation, and neurological disease illustrate how multi-step validation workflows can strengthen causal inference and reduce false-positive interpretation. Finally, we discuss how emerging technologies, including CRISPR screens, single-cell and spatial multi-omics, and AI-assisted experimental design, may reshape validation practice by improving scalability, context specificity, and reproducibility.

q-bio.GN

From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models

Word embeddings and language models have transformed natural language processing (NLP) by facilitating the representation of linguistic elements in continuous vector spaces. This review visits foundational concepts such as the distributional hypothesis and contextual similarity, tracing the evolution from sparse representations like one-hot encoding to dense embeddings including Word2Vec, GloVe, and fastText. We examine both static and contextualized embeddings, underscoring advancements in models such as ELMo, BERT, and GPT and their adaptations for cross-lingual and personalized applications. The discussion extends to sentence and document embeddings, covering aggregation methods and generative topic models, along with the application of embeddings in multimodal domains, including vision, robotics, and cognitive science. Advanced topics such as model compression, interpretability, numerical encoding, and bias mitigation are analyzed, addressing both technical challenges and ethical implications. Additionally, we identify future research directions, emphasizing the need for scalable training techniques, enhanced interpretability, and robust grounding in non-textual modalities. By synthesizing current methodologies and emerging trends, this survey offers researchers and practitioners an in-depth resource to push the boundaries of embedding-based language models.

cs.CL

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

Remote sensing has evolved from simple image acquisition to complex systems capable of integrating and processing visual and textual data. This review examines the development and application of multi-modal language models (MLLMs) in remote sensing, focusing on their ability to interpret and describe satellite imagery using natural language. We cover the technical underpinnings of MLLMs, including dual-encoder architectures, Transformer models, self-supervised and contrastive learning, and cross-modal integration. The unique challenges of remote sensing data--varying spatial resolutions, spectral richness, and temporal changes--are analyzed for their impact on MLLM performance. Key applications such as scene description, object detection, change detection, text-to-image retrieval, image-to-text generation, and visual question answering are discussed to demonstrate their relevance in environmental monitoring, urban planning, and disaster response. We review significant datasets and resources supporting the training and evaluation of these models. Challenges related to computational demands, scalability, data quality, and domain adaptation are highlighted. We conclude by proposing future research directions and technological advancements to further enhance MLLM utility in remote sensing.

cs.CV

From Aleatoric to Epistemic: Exploring Uncertainty Quantification Techniques in Artificial Intelligence

Uncertainty quantification (UQ) is a critical aspect of artificial intelligence (AI) systems, particularly in high-risk domains such as healthcare, autonomous systems, and financial technology, where decision-making processes must account for uncertainty. This review explores the evolution of uncertainty quantification techniques in AI, distinguishing between aleatoric and epistemic uncertainties, and discusses the mathematical foundations and methods used to quantify these uncertainties. We provide an overview of advanced techniques, including probabilistic methods, ensemble learning, sampling-based approaches, and generative models, while also highlighting hybrid approaches that integrate domain-specific knowledge. Furthermore, we examine the diverse applications of UQ across various fields, emphasizing its impact on decision-making, predictive accuracy, and system robustness. The review also addresses key challenges such as scalability, efficiency, and integration with explainable AI, and outlines future directions for research in this rapidly developing area. Through this comprehensive survey, we aim to provide a deeper understanding of UQ's role in enhancing the reliability, safety, and trustworthiness of AI systems.

cs.AI