arXiv ScienceSearch

arXiv subjects

Zheng Dou

Publications and source records attributed to Zheng Dou.

3 recordsLinked to original sources

Local-to-Global Sentence-Level Graph Reranking for Scientific Synthesis

Retrieval-augmented scientific synthesis aims to answer complex research questions by integrating information from multiple papers into comprehensive and well-grounded responses. Since the generator can only synthesize the information selected and organized by the reranker, the quality of the generated synthesis depends critically on the reranked results. However, most rerankers operate at the passage level, which leaves key methodological, empirical, and comparative information buried in long and flat contexts, weakening the grounding of generated claims. Moreover, existing rerankers mainly rely on independent query-candidate scoring which overlooks complementary, contextual, and contrasting relations across scientific candidates, limiting information coverage and the comprehensiveness of the resulting synthesis. To address these limitations, we propose LoG-Reranker, a local-to-global sentence-level graph reranking framework for scientific synthesis. LoG-Reranker performs role-aware local scoring to identify fine-grained, query-relevant sentences and then models their relations on a sentence graph across the candidate set to globally refine sentence rankings. Top-ranked sentences and their connected neighbors are organized into a structured input context for generator to produce more grounded and comprehensive synthesis. Extensive experiments on scientific synthesis and reranking benchmarks show that LoG-Reranker consistently outperforms competitive rerankers, yielding more reliable rankings and improving the quality of generated synthesis.

cs.IR

UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents

Scientific document retrieval (SDR) plays a critical role in modern scientific research, supporting knowledge discovery and evidence-based reasoning. It has evolved along two paradigms: document--document (doc-doc) retrieval driven by inter-document contrastive learning, and question--document (q-doc) retrieval emerging from LLMs and RAG for natural-language interaction. In practice, scientific workflows rely on both paradigms, requiring retrieval of related papers given a seed document and identifying relevant documents given a user question. However, existing methods typically treat these paradigms separately, hindering their complementary strengths. To address this, we propose UniFAR, a unified facet-aware retrieval framework that jointly supports doc-doc and q-doc retrieval within a shared representation space. UniFAR introduces a multi-granularity representation and aggregation module to unify the encoding of short questions and long documents, and a facet-level modeling mechanism with learnable anchors to capture structured semantic roles and complex user intents. It further adopts a facet-aware joint training strategy that integrates doc-doc and q-doc contrastive objectives with facet-level alignment, enabling unified learning from both inter-document relations and question-oriented supervision. Experiments on three benchmark datasets under both doc-doc and q-doc settings across multiple backbone models show that UniFAR consistently outperforms strong baselines and generalizes effectively across different backbones.

cs.IR

FLeW: Facet-Level and Adaptive Weighted Representation Learning of Scientific Documents

Scientific document representation learning provides powerful embeddings for various tasks, while current methods face challenges across three approaches. 1) Contrastive training with citation-structural signals underutilizes citation information and still generates single-vector representations. 2) Fine-grained representation learning, which generates multiple vectors at the sentence or aspect level, requires costly integration and lacks domain generalization. 3) Task-aware learning depends on manually predefined task categorization, overlooking nuanced task distinctions and requiring extra training data for task-specific modules. To address these problems, we propose a new method that unifies the three approaches for better representations, namely FLeW. Specifically, we introduce a novel triplet sampling method that leverages citation intent and frequency to enhance citation-structural signals for training. Citation intents (background, method, result), aligned with the general structure of scientific writing, facilitate a domain-generalized facet partition for fine-grained representation learning. Then, we adopt a simple weight search to adaptively integrate three facet-level embeddings into a task-specific document embedding without task-aware fine-tuning. Experiments show the applicability and robustness of FLeW across multiple scientific tasks and fields, compared to prior models.

cs.IR