arXiv ScienceSearch

arXiv subjects

Zhe Xue

Publications and source records attributed to Zhe Xue.

At least 19 recordsLinked to original sources

MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion

Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.

cs.AI

Research on Cross-media Science and Technology Information Data Retrieval

Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has become an integral part of people's daily life. Unlike news data and social data on the Internet, cross-media science and technology information data has different characteristics. This data has become an important basis for researchers and scholars to track current hot spots and explore future directions of technology development. As the volume of science and technology information data becomes richer, traditional science and technology information retrieval systems, which support only unimodal data retrieval and use outdated keyword-matching models, can no longer meet the daily retrieval needs of science and technology scholars. Therefore, in view of this research background, it is of profound practical significance to study cross-media science and technology information data retrieval systems based on deep semantic features, in line with domestic and international technology-development trends.

cs.IR

Research Team Identification Based on Representation Learning of Academic Heterogeneous Information Network

Academic networks in the real world can usually be described by heterogeneous information networks composed of multiple types of nodes and relationships. Existing representation-learning research for homogeneous information networks lacks the ability to explore the heterogeneity of such networks and therefore cannot be directly applied to heterogeneous information networks. To meet the practical need to identify and discover scientific research teams from academic heterogeneous information networks composed of massive and complex scientific and technological data, this paper proposes a research-team identification method based on representation learning. Node-level and meta-path-level attention mechanisms learn low-dimensional, dense, real-valued vector representations while retaining rich topological information and meta-path semantics. Scientific research teams and important team members are then identified by maximizing node influence. Experimental results show that the proposed method outperforms the comparison methods.

cs.IR

What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks

Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, using a human-validated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.

cs.CV

Mining and searching association relation of scientific papers based on deep learning

There is a complex correlation among the data of scientific papers. The phenomenon reveals the data characteristics, laws, and correlations contained in the data of scientific and technological papers in specific fields, which can realize the analysis of scientific and technological big data and help to design applications to serve scientific researchers. Therefore, the research on mining and searching the association relationship of scientific papers based on deep learning has far-reaching practical significance.

cs.DL

Entity Alignment Method of Science and Technology Patent based on Graph Convolution Network and Information Fusion

The entity alignment of science and technology patents aims to link the equivalent entities in the knowledge graph of different science and technology patent data sources. Most entity alignment methods only use graph neural network to obtain the embedding of graph structure or use attribute text description to obtain semantic representation, ignoring the process of multi-information fusion in science and technology patents. In order to make use of the graphic structure and auxiliary information such as the name, description and attribute of the patent entity, this paper proposes an entity alignment method based on the graph convolution network for science and technology patent information fusion. Through the graph convolution network and BERT model, the structure information and entity attribute information of the science and technology patent knowledge graph are embedded and represented to achieve multi-information fusion, thus improving the performance of entity alignment. Experiments on three benchmark data sets show that the proposed method has better Hits@$K$ evaluation indicators than existing methods.

cs.CL

Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging

Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.

eess.IV

Accurate Portraits of Scientific Resources and Knowledge Service Components

With the advent of the cloud computing era, the cost of creating, capturing, and managing information has gradually decreased. The amount of data on the Internet is showing explosive growth, and more scientific and technological resources are being uploaded to the network. Different from news and social media data, scientific and technological resources are mainly composed of academic-style resources or entities, such as papers, patents, authors, and research institutions. There is a rich relationship network between these resources, from which a large amount of cutting-edge scientific and technological information can be mined. Existing scientific resource management and classification standards are difficult to completely cover all entities and associations, and they cannot accurately extract the important information contained in scientific and technological resources. Therefore, how to construct a complete and accurate representation of scientific and technological resources from structured and unstructured reports and texts, and how to tap the potential value of scientific and technological resources, are urgent problems. A feasible solution is to construct accurate portraits of scientific and technological resources by combining knowledge graph technology, text representation learning, entity extraction, and knowledge service components.

cs.DL

Knowledge Graph and Accurate Portrait Construction of Scientific and Technological Academic Conferences

In recent years, with the continuous progress of science and technology, the number of scientific research achievements has increased rapidly. As an exchange platform and medium for scientific research achievements, scientific and technological academic conferences have become increasingly abundant. The convening of academic conferences brings large numbers of papers, researchers, institutions, projects, and research topics, but massive conference data also makes it difficult for researchers to obtain valuable information efficiently. It is therefore meaningful to use deep learning, knowledge graph technology, semantic similarity calculation, and portrait modeling to mine core information from conference data. This paper reviews the key technologies for constructing knowledge graphs and accurate portraits of scientific and technological academic conferences, including named entity recognition, semantic text similarity, trend prediction, graph storage, search engines, and visualization components. These techniques jointly support the construction of conference knowledge services that help researchers acquire scientific information more quickly.

cs.DL

Efficient Partitioning Method of Large-Scale Public Safety Spatio-Temporal Data based on Information Loss Constraints

The storage, management, and application of massive spatio-temporal data are widely used in practical scenarios, including public safety. However, due to the unique spatio-temporal distribution characteristics of real-world data, existing methods still face limitations in preserving spatio-temporal proximity and achieving load balancing in distributed storage. This paper proposes an efficient partitioning method for large-scale public safety spatio-temporal data based on information loss constraints, named IFL-LSTP. The model combines a spatio-temporal partitioning module (STPM) and a graph partitioning module (GPM). STPM reduces the scale of data under a predefined information-loss threshold, while GPM uses graph representation learning to obtain balanced graph partitions. Experiments on multiple real-world datasets show that IFL-LSTP can reduce data scale, shorten graph model training time, preserve spatio-temporal proximity, and improve load-balancing effectiveness.

cs.LG

Semantic Representation Learning of Scientific Literature based on Adaptive Feature and Graph Neural Network

Because most scientific literature data are unlabeled, semantic representation learning based on unsupervised graphs has become crucial. To enrich scientific-literature features, this paper proposes a semantic representation learning method based on adaptive features and graph neural networks. By introducing adaptive feature processing, scientific-literature features are considered globally and locally. The graph attention mechanism weights and aggregates features of scientific documents connected by citation relations, so that correlations among different documents can be expressed more effectively. In addition, an unsupervised graph neural network semantic representation learning method is proposed. By comparing the mutual information between positive and negative local semantic representations of scientific literature and the global graph semantic representation in the latent space, the graph neural network captures local and global information and improves semantic representation learning. Experimental results show that the proposed method is competitive for scientific literature classification.

cs.CL

Research on Domain Information Mining and Theme Evolution of Scientific Papers

In recent years, with the increase of social investment in scientific research, the number of research results in various fields has increased significantly. Cross-disciplinary research results have gradually become an emerging frontier research direction. There is a certain dependence between a large number of research results. It is difficult to effectively analyze today's scientific research results when looking at a single research field in isolation. How to effectively use the huge number of scientific papers to help researchers becomes a challenge. This paper introduces the research status at home and abroad in terms of domain information mining and topic evolution law of scientific and technological papers from three aspects: the semantic feature representation learning of scientific and technological papers, the field information mining of scientific and technological papers, and the mining and prediction of research topic evolution rules of scientific and technological papers.

cs.DL

Resolving the Blow-Up: A Time-Dilated Numerical Framework for Multiple Firing Events in Mean-Field Neuronal Networks

In large-scale excitatory neuronal networks, rapid synchronization manifests as {multiple firing events (MFEs)}, mathematically characterized by a finite-time blow-up of the neuronal firing rate in the mean-field Fokker-Planck equation. Standard numerical methods struggle to resolve this singularity due to the divergent boundary flux and the instantaneous nature of the population voltage reset. In this work, we propose a robust {multiscale numerical framework based on time dilation}. By transforming the governing equation into a dilated timescale proportional to the firing activity, we desingularize the blow-up, effectively stretching the instantaneous synchronization event into a resolved mesoscopic process. This approach is shown to be physically consistent with the {microscopic cascade mechanism} underlying MFEs and the system's inherent fragility. To implement this numerically, we develop a hybrid scheme that utilizes a {mesh-independent flux criterion} to switch between timescales and a semi-analytical ``moving Gaussian'' method to accurately evolve the post-blowup Dirac mass. Numerical benchmarks demonstrate that our solver not only captures steady states with high accuracy but also efficiently reproduces periodic MFEs, matching Monte Carlo simulations without the severe time-step restrictions associated with particle cascades.

math.NA

Time Sensitive Multiple POIs Route Planning on Bus Networks

This work addresses a route planning problem constrained by a bus road network that includes the schedules of all buses. Given a query with a starting bus stop and a set of Points of Interest (POIs) to visit, our goal is to find an optimal route on the bus network that allows the user to visit all specified POIs from the starting stop with minimal travel time, which includes both bus travel time and waiting time at bus stops. Although this problem resembles a variant of the Traveling Salesman Problem, it cannot be effectively solved using existing solutions due to the complex nature of bus networks, particularly the constantly changing bus travel times and user waiting times. In this paper, we first propose a modified graph structure to represent the bus network, accommodating the varying bus travel times and their arrival schedules at each stop. Initially, we suggest a brute-force exploration algorithm based on the Dijkstra principle to evaluate all potential routes and determine the best one; however, this approach is too costly for large bus networks. To address this, we introduce the EA-Star algorithm, which focuses on computing the shortest route for promising POI visit sequences. The algorithm includes a terminal condition that halts evaluation once the optimal route is identified, avoiding the need to evaluate all possible POI sequences. During the computation of the shortest route for each POI visiting sequence, it employs the A* algorithm on the modified graph structure, narrowing the search space toward the destination and improving search efficiency. Experiments using New York bus network datasets demonstrate the effectiveness of our approach.

cs.DB

Incorporating the Refractory Period into Spiking Neural Networks through Spike-Triggered Threshold Dynamics

As the third generation of neural networks, spiking neural networks (SNNs) have recently gained widespread attention for their biological plausibility, energy efficiency, and effectiveness in processing neuromorphic datasets. To better emulate biological neurons, various models such as Integrate-and-Fire (IF) and Leaky Integrate-and-Fire (LIF) have been widely adopted in SNNs. However, these neuron models overlook the refractory period, a fundamental characteristic of biological neurons. Research on excitable neurons reveal that after firing, neurons enter a refractory period during which they are temporarily unresponsive to subsequent stimuli. This mechanism is critical for preventing over-excitation and mitigating interference from aberrant signals. Therefore, we propose a simple yet effective method to incorporate the refractory period into spiking LIF neurons through spike-triggered threshold dynamics, termed RPLIF. Our method ensures that each spike accurately encodes neural information, effectively preventing neuron over-excitation under continuous inputs and interference from anomalous inputs. Incorporating the refractory period into LIF neurons is seamless and computationally efficient, enhancing robustness and efficiency while yielding better performance with negligible overhead. To the best of our knowledge, RPLIF achieves state-of-the-art performance on Cifar10-DVS(82.40%) and N-Caltech101(83.35%) with fewer timesteps and demonstrates superior performance on DVS128 Gesture(97.22%) at low latency.

cs.CV

DiffusionCom: Structure-Aware Multimodal Diffusion Model for Multimodal Knowledge Graph Completion

Most current MKGC approaches are predominantly based on discriminative models that maximize conditional likelihood. These approaches struggle to efficiently capture the complex connections in real-world knowledge graphs, thereby limiting their overall performance. To address this issue, we propose a structure-aware multimodal Diffusion model for multimodal knowledge graph Completion (DiffusionCom). DiffusionCom innovatively approaches the problem from the perspective of generative models, modeling the association between the $(head, relation)$ pair and candidate tail entities as their joint probability distribution $p((head, relation), (tail))$, and framing the MKGC task as a process of gradually generating the joint probability distribution from noise. Furthermore, to fully leverage the structural information in MKGs, we propose Structure-MKGformer, an adaptive and structure-aware multimodal knowledge representation learning method, as the encoder for DiffusionCom. Structure-MKGformer captures rich structural information through a multimodal graph attention network (MGAT) and adaptively fuses it with entity representations, thereby enhancing the structural awareness of these representations. This design effectively addresses the limitations of existing MKGC methods, particularly those based on multimodal pre-trained models, in utilizing structural information. DiffusionCom is trained using both generative and discriminative losses for the generator, while the feature extractor is optimized exclusively with discriminative loss. This dual approach allows DiffusionCom to harness the strengths of both generative and discriminative models. Extensive experiments on the FB15k-237-IMG and WN18-IMG datasets demonstrate that DiffusionCom outperforms state-of-the-art models.

cs.IR

Crossover from ballistic transport to normal diffusion: a kinetic view

The crossover between dispersion patterns has been frequently observed in various systems. Inspired by the pathway-based kinetic model for E. coli chemotaxis that accounts for the intracellular adaptation process and noise, we propose a kinetic model that can exhibit a crossover from ballistic transport to normal diffusion at the population level. At the particle level, this framework aligns with a stochastic individual-based model. Using numerical simulations and rigorous asymptotic analysis, we demonstrate this crossover both analytically and computationally. Notably, under suitable scaling, the model reveals two distinct limits in which the macroscopic densities exhibit either ballistic transport or normal diffusion.

math.AP

Scientific and Technological Information Oriented Semantics-adversarial and Media-adversarial Cross-media Retrieval

Cross-media retrieval of scientific and technological information is one of the important tasks in the cross-media study. Cross-media scientific and technological information retrieval obtain target information from massive multi-source and heterogeneous scientific and technological resources, which helps to design applications that meet users' needs, including scientific and technological information recommendation, personalized scientific and technological information retrieval, etc. The core of cross-media retrieval is to learn a common subspace, so that data from different media can be directly compared with each other after being mapped into this subspace. In subspace learning, existing methods often focus on modeling the discrimination of intra-media data and the invariance of inter-media data after mapping; however, they ignore the semantic consistency of inter-media data before and after mapping and media discrimination of intra-semantics data, which limit the result of cross-media retrieval. In light of this, we propose a scientific and technological information oriented Semantics-adversarial and Media-adversarial Cross-media Retrieval method (SMCR) to find an effective common subspace. Specifically, SMCR minimizes the loss of inter-media semantic consistency in addition to modeling intra-media semantic discrimination, to preserve semantic similarity before and after mapping. Furthermore, SMCR constructs a basic feature mapping network and a refined feature mapping network to jointly minimize the media discriminative loss within semantics, so as to enhance the feature mapping network's ability to confuse the media discriminant network. Experimental results on two datasets demonstrate that the proposed SMCR outperforms state-of-the-art methods in cross-media retrieval.

cs.IR