arXiv ScienceSearch

arXiv subjects

Yuqing Fu

Publications and source records attributed to Yuqing Fu.

7 recordsLinked to original sources

SEARCH-R: Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator for Multi-hop Question Answering

Multi-hop Question Answering (MHQA) aims to answer questions that require multi-step reasoning. It presents two key challenges: generating correct reasoning paths in response to the complex user queries, and accurately retrieving essential knowledge in the face of potential limitations in large language models (LLMs). Existing approaches primarily rely on prompt-based methods to generate reasoning paths, which are further combined with traditional sparse or dense retrieval to produce the final answer. However, the generation of reasoning paths commonly lacks effective control over the generative process, thus leading the reasoning astray. Meanwhile, the retrieval methods over-rely on knowledge matching or similarity scores rather than evaluating the practical utility of the information, resulting in retrieving homogeneous or non-useful information. Therefore, we propose a Structured Entity-Aware Retrieval with Chain-of-Reasoning Navigator framework named SEARCH-R. Specifically, SEARCH-R trains an end-to-end reasoning path navigator, which is able to provide a powerful sub-question decomposer by fine-tuning the Llama3.1-8B model. Moreover, a novel dependency tree-based retrieval is designed to evaluate the informational contribution of the document quantitatively. Extensive experiments on three challenging multi-hop datasets validate the effectiveness of the proposed framework. The code and dataset are available at: https://github.com/Applied-Machine-Learning-Lab/ACL2026_SEARCH-R.

cs.CL

FMMD: A multimodal open peer review dataset based on F1000Research

Automated scholarly paper review (ASPR) has entered the coexistence phase with traditional peer review, where artificial intelligence (AI) systems are increasingly incorporated into real-world manuscript evaluation. In parallel, research on automated and AI-assisted peer review has proliferated. Despite this momentum, empirical progress remains constrained by several critical limitations in existing datasets. While reviewers routinely evaluate figures, tables, and complex layouts to assess scientific claims, most existing datasets remain overwhelmingly text-centric. This bias is reinforced by a narrow focus on data from computer science venues. Furthermore, these datasets lack precise alignment between reviewer comments and specific manuscript versions, obscuring the iterative relationship between peer review and manuscript evolution. In response, we introduce FMMD, a multimodal and multidisciplinary open peer review dataset curated from F1000Research. The dataset bridges the current gap by integrating manuscript-level visual and structural data with version-specific reviewer reports and editorial decisions. By providing explicit alignment between reviewer comments and the exact article iteration under review, FMMD enables fine-grained analysis of the peer review lifecycle across diverse scientific domains. FMMD supports tasks such as multimodal issue detection and multimodal review comment generation. It provides a comprehensive empirical resource for the development of peer review research.

cs.DL

Enhancing Conversational Agents via Task-Oriented Adversarial Memory Adaptation

Conversational agents struggle to handle long conversations due to context window limitations. Therefore, memory systems are developed to leverage essential historical information. Existing memory systems typically follow a pipeline of offline memory construction and update, and online retrieval. Despite the flexible online phase, the offline phase remains fixed and task-independent. In this phase, memory construction operates under a predefined workflow and fails to emphasize task relevant information. Meanwhile, memory updates are guided by generic metrics rather than task specific supervision. This leads to a misalignment between offline memory preparation and task requirements, which undermines downstream task performance. To this end, we propose an Adversarial Memory Adaptation mechanism (AMA) that aligns memory construction and update with task objectives by simulating task execution. Specifically, first, a challenger agent generates question answer pairs based on the original dialogues. The constructed memory is then used to answer these questions, simulating downstream inference. Subsequently, an evaluator agent assesses the responses and performs error analysis. Finally, an adapter agent analyzes the error cases and performs dual level updates on both the construction strategy and the content. Through this process, the memory system receives task aware supervision signals in advance during the offline phase, enhancing its adaptability to downstream tasks. AMA can be integrated into various existing memory systems, and extensive experiments on long dialogue benchmark LoCoMo demonstrate its effectiveness.

cs.CL

New upper bounds on the number of non-zero weights of constacyclic codes

For any simple-root constacyclic code $\mathcal{C}$ over a finite field $\mathbb{F}_q$, as far as we know, the group $\mathcal{G}$ generated by the multiplier, the constacyclic shift and the scalar multiplications is the largest subgroup of the automorphism group ${\rm Aut}(\mathcal{C})$ of $\mathcal{C}$. In this paper, by calculating the number of $\mathcal{G}$-orbits of $\mathcal{C}\backslash\{\bf 0\}$, we give an explicit upper bound on the number of non-zero weights of $\mathcal{C}$ and present a necessary and sufficient condition for $\mathcal{C}$ to meet the upper bound. Some examples in this paper show that our upper bound is tight and better than the upper bounds in [Zhang and Cao, FFA, 2024]. In particular, our main results provide a new method to construct few-weight constacyclic codes. Furthermore, for the constacyclic code $\mathcal{C}$ belonging to two special types, we obtain a smaller upper bound on the number of non-zero weights of $\mathcal{C}$ by substituting $\mathcal{G}$ with a larger subgroup of ${\rm Aut}(\mathcal{C})$. The results derived in this paper generalize the main results in [Chen, Fu and Liu, IEEE-TIT, 2024]}.

cs.IT

Two classes of LCD BCH codes over finite fields

BCH codes form an important subclass of cyclic codes, and are widely used in compact discs, digital audio tapes and other data storage systems to improve data reliability. As far as we know, there are few results on $q$-ary BCH codes of length $n=\frac{q^{m}+1}{q+1}$. This is because it is harder to deal with BCH codes of such length. In this paper, we study $q$-ary BCH codes with lengths $n=\frac{q^{m}+1}{q+1}$ and $n=q^m+1$. These two classes of BCH codes are always LCD codes. For $n=\frac{q^{m}+1}{q+1}$, the dimensions of narrow-sense BCH codes of length $n$ with designed distance $\delta=\ell q^{\frac{m-1}{2}}+1$ are determined, where $q>2$ and $2\leq \ell \leq q-1$. Moreover, the largest coset leader is given for $m=3$ and the first two largest coset leaders are given for $q=2$. The parameters of BCH codes related to the first few largest coset leaders are investigated. Some binary BCH codes of length $n=\frac{2^m+1}{3}$ have optimal parameters. For ternary narrow-sense BCH codes of length $n=3^m+1$, a lower bound on the minimum distance of their dual codes is developed, which is good in some cases.

cs.IT

The dual codes of two classes of LCD BCH codes

Cyclic BCH codes and negacyclic BCH codes form important subclasses of cyclic codes and negacyclic codes, respectively, and can produce optimal linear codes in many cases. To the best of our knowledge, there are few results on the dual codes of cyclic and negacyclic BCH codes. In this paper, we study the dual codes of narrow-sense cyclic BCH codes of length $q^m+1$ over a finite field $\mathbb{F}_q$, where $q$ is an odd prime power, and the dual codes of narrow-sense negacyclic BCH codes of length $\frac{q^{m}+1}{2}$ over $\mathbb{F}_q$, where $q$ is an odd prime power satisfying $q\equiv 3~({\rm mod}~4)$. Some lower bounds on the minimum distances of the dual codes are established, which are very close to the true minimum distances of the dual codes in many cases. Sufficient and necessary conditions for the even-like subcodes of narrow-sense cyclic BCH codes of length $q^{m}+1$ being cyclic dually-BCH codes are given in terms of designed distances, where $q$ is odd and $m$ is odd or $m\equiv 2~({\rm mod~}4)$. The concept of negacyclic dually-BCH codes is proposed, and sufficient and necessary conditions in terms of designed distances are presented to ensure that narrow-sense negacyclic BCH codes of length $\frac{q^{m}+1}{2}$ are dually-BCH codes, where $q\equiv 3~({\rm mod}~4)$.

cs.IT

Improved upper bounds on the number of non-zero weights of cyclic codes

Let C be an arbitrary simple-root cyclic code and let G be the subgroup of Aut(C) (the automorphism group of C) generated by the multiplier, the cyclic shift and the scalar multiplications. To the best of our knowledge, the subgroup G is the largest subgroup of Aut(C). In this paper, an explicit formula, in some cases an upper bound, for the number of orbits of G on C\{0} is established. An explicit upper bound on the number of non-zero weights of C is consequently derived and a necessary and sufficient condition for the code C meeting the bound is exhibited. Many examples are presented to show that our new upper bounds are tight and are strictly less than the upper bounds in [Chen and Zhang, IEEE-TIT, 2023]. In addition, for two special classes of cyclic codes, smaller upper bounds on the number of non-zero weights of such codes are obtained by replacing G with larger subgroups of the automorphism groups of these codes. As a byproduct, our main results suggest a new way to find few-weight cyclic codes.

cs.IT