arXiv ScienceSearch

arXiv subjects

Wenlong Liu

Publications and source records attributed to Wenlong Liu.

18 recordsLinked to original sources

IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion

Accurate city-level IP Geolocation is an important enabler for the modern digital ecosystem, underpinning services ranging from local content delivery and targeting to digital rights enforcement. However, traditional heuristic and database-driven methods often struggle to resolve the complex, non-linear allocation patterns of modern network infrastructures, particularly within the exploding IPv6 address space and transient mobile networks. In this paper, we introduce IPGeoAI, a novel deep learning model architecture that reframes geolocation from a static lookup problem to a sequential modeling task. Our approach utilizes the Transformer Encoder to capture hierarchical dependencies inherent in IP subnet structures. We propose a method to resolve geographic ambiguity by integrating unstructured semantic context via a Zero-Shot LLM Feature Extraction pipeline. We utilize Large Language Models to transform raw, noisy Autonomous Systems (AS) descriptions into structured, domain-specific metadata (such as 'University' vs. 'ISP' or 'Global' vs. 'Local') via an offline pre-computation process. By fusing these semantic signals into the network via a Multi-Head Cross-Attention module, we bridge the gap between numerical network topology and real-world semantic identity. Extensive offline evaluation on a proprietary dataset spanning 200,000 cities demonstrates that IPGeoAI significantly outperforms a leading external vendor in city-level granularity. By adopting a hierarchical inference strategy that refines coarse-grained country signals, our model achieves a 6% improvement in city-level accuracy while extending coverage to 100% of the traffic. Furthermore, in large-scale online production tests, the model drove a statistically significant +0.35% improvement in our 1st-tier downstream use cases metric.

cs.AI

Quantum geometry induced anomalous chiral transport and hidden symmetry breaking in centrosymmetric 2M-WS2

Chirality, a widely existing material property in nature involving the breaking of the left-right symmetry, has profound influences in various fields of natural sciences. Nonlinear response, such as electronic magnetochiral anisotropy (eMChA), has been recognized as a sensitive probe for the effects of symmetry breaking and nontrivial quantum geometries in solids. So far, observations of eMChA have primarily been limited to inversion-symmetry broken materials. Here, we report a remarkable chiral transport in centrosymmetric candidate topological superconductor 2M-WS2 flakes observed via second-harmonic generation under an out-of-plane magnetic field. More importantly, the eMChA becomes significant around the crossover temperature TFL ~ 25 K from the Fermi liquid (FL) to strange metal (SM) in the normal state, which interestingly echoes with the anomalously large Nernst response at the same temperature in bulk 2M-WS2. These observations reveal a direct correspondence between the nonlinear response, Nernst response, and FL-SM transition in 2M-WS2. Theoretical analysis indicates that nontrivial quantum geometry is behind the simultaneous response of eMChA and Nernst effects in 2M-WS2 and the contribution from the orbital magnetic moment at the Fermi surface becomes significant during the FL-SM transition. Based on first-principles calculations, a thick-layer-sliding mechanism with minimal energy gain in 2M-WS2 provides one possibility for the generation of such nontrivial quantum geometry. The intertwined physics of remarkable eMChA, Nernst response, and FL-SM transition make 2M-WS2 a rare quantum platform to study the chiral transport and unexplored phenomena in strange metals, which may shed light on the trans-century, unresolved scientific issue in unconventional high-temperature superconductivity.

cond-mat.str-el

On the existence of meromorphic solutions of the complex Schrödinger equation with a q-shift

In this paper, we study the following complex Schrödinger equation with a $q$-difference term: \begin{align}\tag{†}\label{dagger} f'(z) = a(z)f(qz) + R(z, f(z)), \quad R(z, f(z)) = \frac{P(z, f(z))}{Q(z, f(z))}, \end{align} where $a(z) \not\equiv 0$ is a small meromorphic function with respect to $f(z)$, and all the coefficient functions of $R(z, f(z))$ are also small meromorphic functions with respect to $f(z)$. We assume that $q\in\mathbb{C}\setminus \left \{ 0,-1,1 \right \} $ and that $R(z, f(z))$ is an irreducible rational function in both $f(z)$ and $z$. We obtain some necessary conditions for \eqref{dagger} to have meromorphic solutions of zero order and non-constant entire solutions, respectively. In particular, if $R(z,f(z))$ reduces to a polynomial in $f(z)$ with degree at most 2 and all the coefficients are constant, then under this assumption and without imposing any restrictions on the growth order of $f(z),$ we prove the existence of entire solutions in many cases, study their number, and further investigate the local and global meromorphic solutions to \eqref{dagger}. Additionally, we consider the possible forms of the meromorphic solutions to \eqref{dagger} in certain conditions and examine exponential polynomials as possible solutions of \eqref{dagger}.

math.CV

Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model

Product matching aims to identify identical or similar products sold on different platforms. By building knowledge graphs (KGs), the product matching problem can be converted to the Entity Alignment (EA) task, which aims to discover the equivalent entities from diverse KGs. The existing EA methods inadequately utilize both attribute triples and relation triples simultaneously, especially the interactions between them. This paper introduces a two-stage pipeline consisting of rough filter and fine filter to match products from eBay and Amazon. For fine filtering, a new framework for Entity Alignment, Relation-aware and Attribute-aware Graph Attention Networks for Entity Alignment (RAEA), is employed. RAEA focuses on the interactions between attribute triples and relation triples, where the entity representation aggregates the alignment signals from attributes and relations with Attribute-aware Entity Encoder and Relation-aware Graph Attention Networks. The experimental results indicate that the RAEA model achieves significant improvements over 12 baselines on EA task in the cross-lingual dataset DBP15K (6.59% on average Hits@1) and delivers competitive results in the monolingual dataset DWY100K. The source code for experiments on DBP15K and DWY100K is available at github (https://github.com/Mockingjay-liu/RAEA-model-for-Entity-Alignment).

cs.AI

VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning

Tool-integrated visual reasoning (TiVR) has demonstrated great potential in enhancing multimodal problem-solving. However, existing TiVR paradigms mainly focus on integrating various visual tools through reinforcement learning, while neglecting to design effective response mechanisms for handling unreliable or erroneous tool outputs. This limitation is particularly pronounced in referring and grounding tasks, where inaccurate detection tool predictions often mislead TiVR models into generating hallucinated reasoning. To address this issue, we propose the VG-Refiner, the first framework aiming at the tool-refined referring grounded reasoning. Technically, we introduce a two-stage think-rethink mechanism that enables the model to explicitly analyze and respond to tool feedback, along with a refinement reward that encourages effective correction in response to poor tool results. In addition, we propose two new metrics and establish fair evaluation protocols to systematically measure the refinement ability of current models. We adopt a small amount of task-specific data to enhance the refinement capability of VG-Refiner, achieving a significant improvement in accuracy and correction ability on referring and reasoning grounding benchmarks while preserving the general capabilities of the pretrained model.

cs.CV

On the extension of analytic solutions of a class of first-order q-difference equations

In this paper, we use the Banach fixed point theorem to examine the existence of meromorphic solutions to the following first-order $q$-difference equation \begin{align}\tag{†}\label{dagger} y(qz)=\frac{a_1(z)y(z)+a_2(z)y(z)^2+\dots+a_p(z)y(z)^p}{1+b_1(z)y(z)+\cdots +b_t(z)y(z)^t}, \end{align} where $q\in \mathbb{C},$ $a_1(z), \dots, a_p(z); b_1(z), \dots, b_t(z)$ are all meromorphic functions. We establish sufficient conditions ensuring the existence and uniqueness of meromorphic solutions that can be extended to the entire complex plane $\mathbb{C}.$ More precisely, we have the following result. If $\left | q \right |\geq 3 $ and \[|a_1(z)| = \max_{1 \le j \le p} |a_j(z)| \le \frac{1}{|z|}, \quad \max_{1 \le k \le t} |b_k(z)| \le \frac{1}{|z|}, \quad z \in \{\, |\Re(z)| \ge ρ> 0 \,\}, \] and $y(0)\ne \infty,$ then we prove that~\eqref{dagger} admits a unique meromorphic solution in $D(ρ),$ which can be extended meromorphically to $\mathbb {C}.$ Moreover, if $a_1(z)\equiv 0,$ the conclusion still holds. Furthermore, if $\left | q \right |\geq 6$ and \begin{gather*} |a_1(z)| \le \frac{1}{|q|}, \quad |a_j(z)| \le |q|^{|z|} \quad (2 \le j \le p), \quad |b_k(z)| \le |q|^{|z|} \quad (1 \le k \le t), \\[4pt] z \in D(ρ,σ) = \{\, z : |\Re(z)| \le ρ,\; |\Im(z)| \le σ, \,\, ρ>0,\,\, σ>0 \,\}, \end{gather*} and $y(0)\ne \infty,$ then we prove that \eqref{dagger} admits a unique meromorphic solution in $D(ρ, σ),$ which can also be extended meromorphically to $\mathbb {C}.$ This conclusion remains valid in the case where $a_1(z)\equiv 0.$

math.CV

Transcendental meromorphic solutions and the complex Schrödinger equation with delay

In this article, we focus on studying the differential-difference equation \[ f'(z) = a(z)f(z+1) + R(z, f(z)), \quad R(z, f(z)) = \frac{P(z, f(z))}{Q(z, f(z))}, \] where the two nonzero polynomials \( P(z, f(z)) \) and \( Q(z, f(z)) \) in \( f(z) \), with small meromorphic coefficients, are coprime, and \( a(z) \) is a nonzero small meromorphic function of \( f(z) \). This equation includes the complex Schrodinger equation with delay as a special case. If \( f(z) \) is a transcendental meromorphic solution of the equation with subnormal growth, then we derive all possible forms of the equation. Additionally, under these assumptions, we classify these specific forms based on the degrees of \( P(z, f(z)) \) and \( Q(z, f(z)) \) to establish necessary conditions for the existence of transcendental meromorphic solutions. In particular, when the degree of \( P \) minus the degree of \( Q \) is 2, we demonstrate that the equation reduces to a Riccati differential equation. Finally, examples are provided to support our results.

math.CV

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encoder-decoder architecture as Grounding DINO 1.5 to pursue an object-level representation for open-world object understanding. To make long-tailed object detection easy, DINO-X extends its input options to support text prompt, visual prompt, and customized prompt. With such flexible prompt options, we develop a universal object prompt to support prompt-free open-world detection, making it possible to detect anything in an image without requiring users to provide any prompt. To enhance the model's core grounding capability, we have constructed a large-scale dataset with over 100 million high-quality grounding samples, referred to as Grounding-100M, for advancing the model's open-vocabulary detection performance. Pre-training on such a large-scale grounding dataset leads to a foundational object-level representation, which enables DINO-X to integrate multiple perception heads to simultaneously support multiple object perception and understanding tasks, including detection, segmentation, pose estimation, object captioning, object-based QA, etc. Experimental results demonstrate the superior performance of DINO-X. Specifically, the DINO-X Pro model achieves 56.0 AP, 59.8 AP, and 52.4 AP on the COCO, LVIS-minival, and LVIS-val zero-shot object detection benchmarks, respectively. Notably, it scores 63.3 AP and 56.5 AP on the rare classes of LVIS-minival and LVIS-val benchmarks, improving the previous SOTA performance by 5.8 AP and 5.0 AP. Such a result underscores its significantly improved capacity for recognizing long-tailed objects.

cs.CV

tSF: Transformer-based Semantic Filter for Few-Shot Learning

Few-Shot Learning (FSL) alleviates the data shortage challenge via embedding discriminative target-aware features among plenty seen (base) and few unseen (novel) labeled samples. Most feature embedding modules in recent FSL methods are specially designed for corresponding learning tasks (e.g., classification, segmentation, and object detection), which limits the utility of embedding features. To this end, we propose a light and universal module named transformer-based Semantic Filter (tSF), which can be applied for different FSL tasks. The proposed tSF redesigns the inputs of a transformer-based structure by a semantic filter, which not only embeds the knowledge from whole base set to novel set but also filters semantic features for target category. Furthermore, the parameters of tSF is equal to half of a standard transformer block (less than 1M). In the experiments, our tSF is able to boost the performances in different classic few-shot learning tasks (about 2% improvement), especially outperforms the state-of-the-arts on multiple benchmark datasets in few-shot classification task.

cs.CV

SymPoint Revolutionized: Boosting Panoptic Symbol Spotting with Layer Feature Enhancement

SymPoint is an initial attempt that utilizes point set representation to solve the panoptic symbol spotting task on CAD drawing. Despite its considerable success, it overlooks graphical layer information and suffers from prohibitively slow training convergence. To tackle this issue, we introduce SymPoint-V2, a robust and efficient solution featuring novel, streamlined designs that overcome these limitations. In particular, we first propose a Layer Feature-Enhanced module (LFE) to encode the graphical layer information into the primitive feature, which significantly boosts the performance. We also design a Position-Guided Training (PGT) method to make it easier to learn, which accelerates the convergence of the model in the early stages and further promotes performance. Extensive experiments show that our model achieves better performance and faster convergence than its predecessor SymPoint on the public benchmark. Our code and trained models are available at https://github.com/nicehuster/SymPointV2.

cs.CV

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

This paper introduces Grounding DINO 1.5, a suite of advanced open-set object detection models developed by IDEA Research, which aims to advance the "Edge" of open-set object detection. The suite encompasses two models: Grounding DINO 1.5 Pro, a high-performance model designed for stronger generalization capability across a wide range of scenarios, and Grounding DINO 1.5 Edge, an efficient model optimized for faster speed demanded in many applications requiring edge deployment. The Grounding DINO 1.5 Pro model advances its predecessor by scaling up the model architecture, integrating an enhanced vision backbone, and expanding the training dataset to over 20 million images with grounding annotations, thereby achieving a richer semantic understanding. The Grounding DINO 1.5 Edge model, while designed for efficiency with reduced feature scales, maintains robust detection capabilities by being trained on the same comprehensive dataset. Empirical results demonstrate the effectiveness of Grounding DINO 1.5, with the Grounding DINO 1.5 Pro model attaining a 54.3 AP on the COCO detection benchmark and a 55.7 AP on the LVIS-minival zero-shot transfer benchmark, setting new records for open-set object detection. Furthermore, the Grounding DINO 1.5 Edge model, when optimized with TensorRT, achieves a speed of 75.2 FPS while attaining a zero-shot performance of 36.2 AP on the LVIS-minival benchmark, making it more suitable for edge computing scenarios. Model examples and demos with API will be released at https://github.com/IDEA-Research/Grounding-DINO-1.5-API

cs.CV

Symbol as Points: Panoptic Symbol Spotting via Point-based Representation

This work studies the problem of panoptic symbol spotting, which is to spot and parse both countable object instances (windows, doors, tables, etc.) and uncountable stuff (wall, railing, etc.) from computer-aided design (CAD) drawings. Existing methods typically involve either rasterizing the vector graphics into images and using image-based methods for symbol spotting, or directly building graphs and using graph neural networks for symbol recognition. In this paper, we take a different approach, which treats graphic primitives as a set of 2D points that are locally connected and use point cloud segmentation methods to tackle it. Specifically, we utilize a point transformer to extract the primitive features and append a mask2former-like spotting head to predict the final output. To better use the local connection information of primitives and enhance their discriminability, we further propose the attention with connection module (ACM) and contrastive connection learning scheme (CCL). Finally, we propose a KNN interpolation mechanism for the mask attention module of the spotting head to better handle primitive mask downsampling, which is primitive-level in contrast to pixel-level for the image. Our approach, named SymPoint, is simple yet effective, outperforming recent state-of-the-art method GAT-CADNet by an absolute increase of 9.6% PQ and 10.4% RQ on the FloorPlanCAD dataset. The source code and models will be available at https://github.com/nicehuster/SymPoint.

cs.CV

Effects of Coronal Magnetic Field Configuration on Particle Acceleration and Release during the Ground Level Enhancement Events in Solar Cycle 24

Ground level enhancements (GLEs) are extreme solar energetic particle (SEP) events that are of particular importance in space weather. In solar cycle 24, two GLEs were recorded on 2012 May 17 (GLE 71) and 2017 September 10 (GLE 72), respectively, by a range of advanced modern instruments. Here we conduct a comparative analysis of the two events by focusing on the effects of large-scale magnetic field configuration near active regions on particle acceleration and release. Although the active regions both located near the western limb, temporal variations of SEP intensities and energy spectra measured in-situ display different behaviors at early stages. By combining a potential field model, we find the CME in GLE 71 originated below the streamer belt, while in GLE 72 near the edge of the streamer belt. We reconstruct the CME shock fronts with an ellipsoid model based on nearly simultaneous coronagraph images from multi-viewpoints, and further derive the 3D shock geometry at the GLE onset. The highest-energy particles are primarily accelerated in the shock-streamer interaction regions, i.e., likely at the nose of the shock in GLE 71 and the eastern flank in GLE 72, due to quasi-perpendicular shock geometry and confinement of closed fields. Subsequently, they are released to the field lines connecting to near-Earth spacecraft when the shocks move through the streamer cusp region. This suggests that magnetic structures in the corona, especially shock-streamer interactions, may have played an important role in the acceleration and release of the highest-energy particles in the two events.

astro-ph.SR

Why "solar tsunamis" rarely leave their imprints in the chromosphere

Solar coronal waves frequently appear as bright disturbances that propagate globally from the eruption center in the solar atmosphere, just like the tsunamis in the ocean on Earth. Theoretically, coronal waves can sweep over the underlying chromosphere and leave an imprint in the form of Moreton wave, due to the enhanced pressure beneath their coronal wavefront. Despite the frequent observations of coronal waves, their counterparts in the chromosphere are rarely detected. Why the chromosphere rarely bears the imprints of solar tsunamis remained a mystery since their discovery three decades ago. To resolve this question, all coronal waves and associated Moreton waves in the last decade have been initially surveyed, though the detection of Moreton waves could be hampered by utilising the low-quality H$α$ data from Global Oscillations Network Group. Here, we present 8 cases (including 5 in Appendix) of the coexistence of coronal and Moreton waves in inclined eruptions where it is argued that the extreme inclination is key to providing an answer to address the question. For all these events, the lowest part of the coronal wavefront near the solar surface appears very bright, and the simultaneous disturbances in the solar transition region and the chromosphere predominantly occur beneath the bright segment. Therefore, evidenced by observations, we propose a scenario for the excitation mechanism of the coronal-Moreton waves in highly inclined eruptions, in which the lowest part of a coronal wave can effectively disturb the chromosphere even for a weak (e.g., B-class) solar flare.

astro-ph.SR

CONSS: Contrastive Learning Approach for Semi-Supervised Seismic Facies Classification

Recently, seismic facies classification based on convolutional neural networks (CNN) has garnered significant research interest. However, existing CNN-based supervised learning approaches necessitate massive labeled data. Labeling is laborious and time-consuming, particularly for 3D seismic data volumes. To overcome this challenge, we propose a semi-supervised method based on pixel-level contrastive learning, termed CONSS, which can efficiently identify seismic facies using only 1% of the original annotations. Furthermore, the absence of a unified data division and standardized metrics hinders the fair comparison of various facies classification approaches. To this end, we develop an objective benchmark for the evaluation of semi-supervised methods, including self-training, consistency regularization, and the proposed CONSS. Our benchmark is publicly available to enable researchers to objectively compare different approaches. Experimental results demonstrate that our approach achieves state-of-the-art performance on the F3 survey.

cs.CV

nVFNet-RDC: Replay and Non-Local Distillation Collaboration for Continual Object Detection

Continual Learning (CL) focuses on developing algorithms with the ability to adapt to new environments and learn new skills. This very challenging task has generated a lot of interest in recent years, with new solutions appearing rapidly. In this paper, we propose a nVFNet-RDC approach for continual object detection. Our nVFNet-RDC consists of teacher-student models, and adopts replay and feature distillation strategies. As the 1st place solutions, we achieve 55.94% and 54.65% average mAP on the 3rd CLVision Challenge Track 2 and Track 3, respectively.

cs.CV

Double-power-law feature of energetic particles accelerated at coronal shocks

Recent observations have shown that in many large solar energetic particle (SEP) events the event-integrated differential spectra resemble double power laws. We perform numerical modeling of particle acceleration at coronal shocks propagating through a streamer-like magnetic field by solving the Parker transport equation, including protons and heavier ions. We find that for all ion species the energy spectra integrated over the simulation domain can be described by a double power law, and the break energy depends on the ion charge-to-mass ratio as $E_B \sim (Q/A)^α$, with $α$ varying from 0.16 to 1.2 by considering different turbulence spectral indices. We suggest that the double power law distribution may emerge as a result of the superposition of energetic particles from different source regions where the acceleration rates differ significantly due to particle diffusion. The diffusion and mixing of energetic particles could also provide an explanation for the increase of Fe/O at high energies as observed in some SEP events. Although further mixing processes may occur, our simulations indicate that either power-law break or rollover can occur near the Sun and predict that the spectral forms vary significantly along the shock front, which may be examined by upcoming near-Sun SEP measurements from Parker Solar Probe and Solar Orbiter.

astro-ph.SR

Yes-Net: An effective Detector Based on Global Information

This paper introduces a new real-time object detection approach named Yes-Net. It realizes the prediction of bounding boxes and class via single neural network like YOLOv2 and SSD, but owns more efficient and outstanding features. It combines local information with global information by adding the RNN architecture as a packed unit in CNN model to form the basic feature extractor. Independent anchor boxes coming from full-dimension k-means is also applied in Yes-Net, it brings better average IOU than grid anchor box. In addition, instead of NMS, Yes-Net uses RNN as a filter to get the final boxes, which is more efficient. For 416 x 416 input, Yes-Net achieves 79.2% mAP on VOC2007 test at 39 FPS on an Nvidia Titan X Pascal.

cs.CV