arXiv ScienceSearch

arXiv subjects

Jie Wen

Publications and source records attributed to Jie Wen.

At least 19 recordsLinked to original sources

One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?

Patch-based autoregressive time-series forecasting often ties input representation, learned transitions, and recursive execution to one patch length. We ask which of these roles can be adjusted separately. A supporting atomic-encoding study finds greater sensitivity to model width than to atom grouping on the evaluated grid. Our main finding is that a frozen parent's recursive trajectory is easier to fit than the observed future with lightweight parallel exits. Autoregressive Trajectory Distillation (ATD) turns this into selectable ATD-1/2/4/8 execution, with ATD-1 exactly recovering the parent. On a paired four-data-set comparison, ATD-8 reaches $5.54\times$ end-to-end speedup with stable quality across widths. Fewer calls do not automatically remove the parent's existing forecast error: ATD improves trajectory fidelity in all 21 seed runs but forecast accuracy in only 15 against matched clean-future supervision. We further find a correctable residual projection along a train-selected periodic history direction. Spectrum Tangent applies this correction without adding neural parameters or Transformer calls. At horizon 720, it reduces mean squared error (MSE) and mean absolute error (MAE) by 2.54% and 2.33% over seven data sets and two output widths, while remaining $3.24\times$ faster than recursive inference. Level and shape projections sometimes disagree. Trajectory compressibility, the fidelity-accuracy mismatch, and the correction recur across three public AR parents. Together these results separate representation, transition, and execution as AR design axes. Code is available at https://github.com/RowanFFF/ATD-Spectrum-Tangent.

cs.LG

Sharp stability for cross $t$-intersecting families of permutations in the linear range

Two families $\mathcal{F},\mathcal{G}\subseteq S_n$ are cross $t$-intersecting if every $σ\in\mathcal{F}$ and $τ\in\mathcal{G}$ agree on at least $t$ points. A $t$-coset is a coset of the stabilizer of $t$ points. A subset of $S_n$ is non-trivial if it is not contained in any $t$-coset. Let $d_m$ denote the $m$-th derangement number. We prove that, for all $t\geq1$ and $n\geq400t$, every pair of cross $t$-intersecting families $\mathcal{F},\mathcal{G}\subseteq S_n$ satisfies the following: (i) $|\mathcal{F}||\mathcal{G}|\leq((n-t)!-d_{n-t}-d_{n-t-1})((n-t)!+t)$ if $\mathcal{F}\cup\mathcal{G}$ is non-trivial. (ii) $|\mathcal{F}||\mathcal{G}|\leq((n-t)!-d_{n-t}-d_{n-t-1}+t)^2$ if both $\mathcal{F}$ and $\mathcal{G}$ are non-trivial. (iii) $\min\{|\mathcal{F}\setminus\mathcal{C}|,|\mathcal{G}\setminus\mathcal{C}|\}\leq t((n-t-1)!-(n-t-2)!)$ for some $t$-coset $\mathcal{C}$ if $t\geq2$. We also characterize all extremal configurations. The first result extends a theorem of Ellis (2011) to an exponentially wider range and sharpens the stability theorem of Keller, Lifshitz, Minzer and Sheinfeld (2024); the second gives a product version of the classical Hilton--Milner--Frankl theorem for permutations; and the third settles the remaining cases $t\geq2$ of a conjecture of Ellis (2011) in a stronger form. In all three results, the linear dependence on $t$ is essentially optimal. Our proofs are based on the spread approximation method introduced by Kupavskii and Zakharov and on an approach to cross $t$-intersection problems developed by the present authors, with several essential refinements. As an application of our approach, we prove a product version of the Hilton--Milner--Frankl theorem for the alternating group.

math.CO

Structure of large $t$-intersecting families I: Stability for the Hilton--Milner--Frankl theorem

We study the structure of large $t$-intersecting families. A family of $k$-subsets of an $n$-set is $t$-intersecting if every two of its members intersect in at least $t$ elements. A $t$-intersecting family is non-trivial if no $t$-subset is contained in all its members. We prove several stability results for the seminal Hilton--Milner--Frankl theorem. First, for any fixed $η,\varepsilon,θ\in(0,1)$, we prove that if $k/t\geq1+η$ and $n=Ω(tk^{1+\varepsilon})$, then every non-trivial $t$-intersecting family of size greater than $(1+θ)|\mathcal{K}|$ is a subfamily of one of the two extremal families in the theorem, where $\mathcal{K}$ is an explicit large non-trivial $t$-intersecting family. The key ingredient in the proof is a removal lemma. We also obtain a classification of all $t$-intersecting families with size bounded below by $|\mathcal{K}|$ minus an explicit lower-order term, provided that $k\geq t+4\geq6$ and $n\geq t+6\cdot\max\{(t+2)^2, k(k-t)\}$. This strengthens results of Cao--Lv--Wang (2021) and Frankl (2025) for a broad range of $k$ and $t$ (for example, when $k-t\geq2\sqrt{t}$). As an application of this classification, we determine the largest $t$-intersecting families for each prescribed lower bound on $t$-diversity not exceeding $t(n-k)$, thereby obtaining $t$-intersection versions of results of Han and Kohayakawa (2017) and Kupavskii (2025). To establish these results, we develop techniques based on the spread approximation method and the $t$-cover method, which may be useful for other intersection problems.

math.CO

When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning

Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observed views, but also label-aware exploitation of view-specific evidence. Existing approaches usually emphasize either shared representation learning or decision-level fusion. The former improves robustness against missing views, yet tends to compress label-discriminative view-specific cues into a single latent representation. The latter preserves individual view predictions, but often relies on fixed or globally learned fusion weights, ignoring that different labels of different instances may require different views. To address these limitations, this paper presents V2L, a unified representation-decision framework for incomplete multi-view multi-label classification. On the representation side, V2L constructs semantically consistent variational posteriors from incomplete views through a perturbation-aware encoding mechanism, which provides a stable shared semantic basis. On the decision side, V2L introduces an active view-label relevance modeling strategy that estimates instance-wise and label-wise view contributions, allowing each label prediction to adaptively select useful view-specific evidence. From the perspective of model architecture, these two important strategies are integrated into a unified framework through a hybrid fusion architecture, simultaneously meeting the requirements of cross-view semantic consistency and representational complementarity. Extensive experiments under both incomplete and complete settings show that V2L achieves leading performance on five benchmarks. Code is available at: https://github.com/justsmart/V2L.

cs.CV

On the Frankl--Tokushige conjecture and almost complete $r$-cross $t$-intersection theorems for vector spaces

Let $r\geq3$ and $k_1\geq k_2\geq\cdots\geq k_r\geq t$. Let $\mathcal{F}_1,\mathcal{F}_2,\ldots,\mathcal{F}_r$ be families of subspaces, of respective dimensions $k_1,k_2,\ldots,k_r$, in an $n$-dimensional vector space over the finite field $\mathbb{F}_q$. The $r$ families are called $r$-cross $t$-intersecting if $\dim \left(F_{1} \cap F_{2} \cap \cdots \cap F_{r}\right) \geq t$ for all $F_{i} \in \mathcal{F}_{i}, i = 1,2,\dots,r$. In 2016, Frankl and Tokushige conjectured that $\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-1\brack k_i-1}$ for $t=1$ and $n\geq rk_1/(r-1)$. The appealing conjecture suggests establishing intersection theorems for $n\sim ck_1$ with $c=c(r)\in(1,2)$, a direction that has long been challenging. In this paper, we overcome this barrier by proving that $$\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-t\brack k_i-t}\;\;\mbox{for all}\;\;t\geq1\;\mbox{and}\;n\geq rk_1/(r-1)+C(t,r),$$ where $C(t,r)=rt/(r-1)+1$. This proves the Frankl--Tokushige conjecture except for at most three values of $n$, and establishes an Erdős--Ko--Rado type theorem for almost all values of parameters. Furthermore, we characterize all extremal configurations. Our proof is purely combinatorial and based on the $t$-cover method, with several essential refinements. We also obtain almost complete intersection theorems for $r$-wise $t$-intersecting families and non-trivial $r$-cross $t$-intersecting families.

math.CO

A unified approach to cross-intersection problems with applications to Hilton--Milner type theorems and stability

We develop a new approach to cross-intersection problems in extremal set theory. The method builds on the iterative procedure introduced by Kupavskii and Zakharov (2024) and the $t$-cover method. It provides a flexible framework for deriving extremal and stability results for cross $t$-intersecting families. Our approach applies to a variety of combinatorial objects. As an application, we prove a product version of the seminal Erdős--Ko--Rado theorem for sufficiently spread set systems. Two families $\mathcal{F}$ and $\mathcal{G}$ of $k$-subsets of $[n]$ are called cross $t$-intersecting if $|F\cap G|\geq t$ for all $F\in\mathcal{F}$ and $G\in\mathcal{G}$. We determine the families maximizing $\min\{|\mathcal{F}|, |\mathcal{G}|\}$ for large $n$ and all $t\ge2$, generalizing results of Mörs (1985) and Füredi (1995) for cross $1$-intersecting families. We then determine the families maximizing $|\mathcal{F}||\mathcal{G}|$ under the condition $\max\{|\cap_{F\in\mathcal{F}}F|,|\cap_{G\in\mathcal{G}}G|\}<t$ for large $n$. This improves the bound obtained by Frankl and Wang (2024), and provides a characterization of extremal configurations. For a family $\mathcal{F}$ of subsets of $[n]$, we introduce its $t$-diversity $γ_t(\mathcal{F})$, defined as the minimum number of sets from $\mathcal{F}$ not containing a fixed $t$-subset. This serves as a natural generalization of the important notion of diversity for $t=1$. We obtain a stability result via $γ_t$, and determine the maximum of $\min\{γ_t(\mathcal{F}),γ_t(\mathcal{G})\}$ for cross $t$-intersecting families $\mathcal{F}$ and $\mathcal{G}$. These yield new results for $t$-intersecting families, including a stability theorem towards a conjecture of Ellis, Keller and Lifshitz (2019), which may also be regarded as a $t$-intersection version, for large $n$, of an influential theorem of Frankl (1987).

math.CO

A Semi-supervised Physics-Aware Triple-Stream Underwater Image Enhancement Network

Underwater images normally suffer from degradation due to the transmission medium of water bodies. Both traditional prior-based approaches and deep learning-based methods have been used to address this problem. However, the inflexible assumption of the former often impairs their effectiveness in handling diverse underwater scenes, while the generalization of the latter to unseen images is usually weakened by insufficient data. In this study, we leverage both the physics-based Image Formation Model (IFM) and deep learning techniques for Underwater Image Enhancement (UIE). To this end, we propose a novel Physics-Aware Triple-Stream Underwater Image Enhancement Network, i.e., PATS-UIENet, which comprises a Direct Signal Transmission Estimation Stream (D-Stream), a Backscatter Signal Transmission Estimation Stream (B-Stream) and an Ambient Light Estimation Stream (A-Stream). This network fulfills the UIE task by explicitly estimating the degradation parameters of a revised IFM. We also adopt an IFM-inspired semi-supervised learning framework, which exploits both the labeled and unlabeled images, to address the issue of insufficient data. To our knowledge, such a physics-aware deep network and the IFM-inspired semi-supervised learning framework have not been used for the UIE task before. Our method performs better than, or at least comparably to, sixteen baselines across four testing sets in the degradation estimation and UIE tasks. These promising results should be due to the fact that the proposed method can not only model the degradation but also learn the characteristics of diverse underwater scenes.

cs.CV

Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval

Domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, enabling effective retrieval while mitigating domain discrepancies. However, existing methods encounter several fundamental limitations: 1) neglecting class-level semantic alignment and excessively pursuing pair-wise sample alignment; 2) lacking either pseudo-label reliability consideration or geometric guidance for assessing label correctness; 3) directly quantizing original features affected by domain shift, undermining the quality of learned hash codes. In view of these limitations, we propose Prototype-Based Semantic Consistency Alignment (PSCA), a two-stage framework for effective domain adaptive retrieval. In the first stage, a set of orthogonal prototypes directly establishes class-level semantic connections, maximizing inter-class separability while gathering intra-class samples. During the prototype learning, geometric proximity provides a reliability indicator for semantic consistency alignment through adaptive weighting of pseudo-label confidences. The resulting membership matrix and prototypes facilitate feature reconstruction, ensuring quantization on reconstructed rather than original features, thereby improving subsequent hash coding quality and seamlessly connecting both stages. In the second stage, domain-specific quantization functions process the reconstructed features under mutual approximation constraints, generating unified binary hash codes across domains. Extensive experiments validate PSCA's superior performance across multiple datasets.

cs.LG

Expandable, Compressible, Mineable: Open-World Thermal Image Restoration

In open-world settings, thermal infrared (TIR) image degradations continuously emerge and evolve, while most existing all-in-one restoration methods are built on a closed-set assumption and struggle to continually adapt to novel degradations. To address this, we propose ECMRNet, an Expandable, Compressible, and Mineable Restoration Network for open-world TIR restoration from a continual learning perspective. Conceptually, ECMRNet unifies continual degradation learning as an "expand-compress-mine" closed-loop process, enabling sustained adaptation to new degradations with controllable evolution. Structurally, ECMRNet decomposes intermediate representations into group-isolated subspaces, and achieves strict parameter isolation and fast adaptation to new degradations by freezing historical groups and isomorphically expanding new ones. To curb model growth as tasks accumulate, we present Structural Entropy Pruning, which identifies and removes redundant channel groups via two-dimensional structural entropy minimization, achieving information contribution-driven adaptive compression. Moreover, we design a Sub-degradation Knowledge Mining Module that dynamically retrieves and recombines transferable components from historical representations to improve restoration under compound degradations. Experimental results demonstrate that ECMRNet achieves superior overall performance across diverse single and compound degradations while using fewer parameters and lower computational cost. The source code is available at https://github.com/Kust-lp/ECMRNet.

cs.CV

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses despite safety alignment. Existing defenses typically improve safety through external filtering, auxiliary guardrails, or decoding-time control. However, these interventions often reduce practical deployability because they may require additional model access, introduce extra inference cost, or affect benign-task utility. In this paper, we propose Safety-Aware Intent Defense (SAID), a training-free jailbreak defense framework based on intent-level safety probing. SAID first distills potentially obfuscated user inputs into concise core intents using the target model itself. It then applies a validated safety prefix to probe each distilled intent and elicit the model's safety-aware response. Finally, a conservative aggregation rule rejects the original request if any distilled intent is identified as unsafe. This design enables black-box-compatible defense without updating model parameters or modifying the decoding process. Experiments on four open-source LLMs under six representative jailbreak attacks show that SAID achieves state-of-the-art defense performance in reducing harmful responses while maintaining competitive utility on benign tasks. Further analyses on prefix variants, hierarchical distillation, and inference efficiency demonstrate that SAID provides a practical safety-utility trade-off for securing LLMs against jailbreak threats.

cs.CR

SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere

Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods universally depend on large-scale annotations or task-specific training procedures, severely limiting their rapid deployment to novel scenes. We observe that intermediate-layer features of pre-trained multimodal large language models (MLLMs) already encode rich anomaly semantics, yet existing approaches rely on the language output pathway and fail to exploit the geometric discriminability latent in these representations. Based on this finding, we propose SphereVAD, a fully training-free, zero-shot VAD framework that recasts anomaly discrimination as von Mises-Fisher (vMF) likelihood-ratio geodesic inference on the unit hypersphere, unleashing latent discriminability through principled geometric reasoning rather than learning new representations. Specifically, SphereVAD first applies Frechet mean centering to unfold feature distributions and eliminate domain biases, then employs Holistic Scene Attention (HSA) to reinforce feature consistency using cross-video priors, and finally performs vMF-guided Spherical Geodesic Pulling (SGP) to align ambiguous segments with directional prototypes on the spherical manifold. This training-free pipeline requires only minimal synthetic images for calibration. SphereVAD establishes new state-of-the-art results among training-free approaches on three major benchmarks and remains competitive with fully supervised baselines. Code will be available upon acceptance.

cs.CV

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills

Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering into cross-file security review. Existing guardrails often flag risk but recover malicious intent inconsistently under semantics-preserving rewrites. This paper formulates pre-load auditing for untrusted Agent Skills as a robust three-way classification task and introduces SkillGuard-Robust, which combines role-aware evidence extraction, selective semantic verification, and consistency-preserving adjudication. We evaluate SkillGuard-Robust on SkillGuardBench and two public-ecosystem extensions through five large evaluation views ranging from 254 to 404 packages. On the 404-package held-out aggregate, SkillGuard-Robust reaches 97.30% overall exact match, 98.33% malicious-risk recall, and 98.89% attack exact consistency. On the 254-package external-ecosystem view, it reaches 99.66%, 100.00%, and 100.00%, respectively. These results support a bounded conclusion: factorized package auditing materially improves frozen and public-ecosystem robustness, while harsher external-source transfer remains an open challenge.

cs.CR

Feature-Label Modal Alignment for Robust Partial Multi-Label Learning

In partial multi-label learning (PML), each instance is associated with a set of candidate labels containing both ground-truth and noisy labels. The presence of noisy labels disrupts the correspondence between features and labels, degrading classification performance. To address this challenge, we propose a novel PML method based on feature-label modal alignment (PML-MA), which treats features and labels as two complementary modalities and restores their consistency through systematic alignment. Specifically, PML-MA first employs low-rank orthogonal decomposition to generate pseudo-labels that approximate the true label distribution by filtering noisy labels. It then aligns features and pseudo-labels through both global projection into a common subspace and local preservation of neighborhood structures. Finally, a multi-peak class prototype learning mechanism leverages the multi-label nature where instances simultaneously belong to multiple categories, using pseudo-labels as soft membership weights to enhance discriminability. By integrating modal alignment with prototype-guided refinement, PML-MA ensures pseudo-labels better reflect the true distribution while maintaining robustness against label noise. Extensive experiments on both real-world and synthetic datasets demonstrate that PML-MA significantly outperforms state-of-the-art methods, achieving superior classification accuracy and noise robustness.

cs.LG

SEHFS: Structural Entropy-Guided High-Order Correlation Learning for Multi-View Multi-Label Feature Selection

In recent years, multi-view multi-label learning (MVML) has attracted extensive attention due to its close alignment to real-world scenarios. Information-theoretic methods have gained prominence for learning nonlinear correlations. However, two key challenges persist: first, features in real-world data commonly exhibit high-order structural correlations, but existing information-theoretic methods struggle to learn such correlations; second, commonly relying on heuristic optimization, information-theoretic methods are prone to converging to local optima. To address these two challenges, we propose a novel method called Structural Entropy Guided High-Order Correlation Learning for Multi-View Multi-Label Feature Selection (SEHFS). The core idea of SEHFS is to convert the feature graph into a structural-entropy-minimizing encoding tree, quantifying the information cost of high-order dependencies and thus learning high-order feature correlations beyond pairwise correlations. Specifically, features exhibiting strong high-order redundancy are grouped into a single cluster within the encoding tree, while inter-cluster feaeture correlations are minimized, thereby eliminating redundancy both within and across clusters. Furthermore, a new framework based on the fusion of information theory and matrix methods is adopted, which learns a shared semantic matrix and view-specific contribution matrices to reconstruct a global view matrix, thereby enhancing the information-theoretic method and balancing the global and local optimization. The ability of structural entropy to learn high-order correlations is theoretically established, and and both experiments on eight datasets from various domains and ablation studies demonstrate that SEHFS achieves superior performance in feature selection.

cs.LG

On $r$-cross $t$-intersecting families of partitions

In this paper, we address several intersection problems for $r$-cross $t$-intersecting families of partitions. A $k$-partition of an $n$-set $X$ is a set of $k$ pairwise disjoint non-empty subsets whose union is $X$. For $1\leq i\leq r$, let $\mathcal{F}_i$ be a family of $k_i$-partitions of $X$. We say that $\mathcal{F}_1,\mathcal{F}_2,\ldots,\mathcal{F}_r$ are $r$-cross $t$-intersecting if $|\cap_{i=1}^{r}F_i|\geq t$ for all $F_i\in\mathcal{F}_i$. The families are called non-trivial if $|\cap_{i=1}^r(\cap_{F\in\mathcal{F}_i}F)|<t$. Proving an Erdős-Ko-Rado type theorem, we determine the families maximizing $\prod_{i=1}^r|\mathcal{F}_i|$. We further determine non-trivial $r$-cross $t$-intersecting families with maximum product of sizes; this result also serves as a Hilton-Milner type theorem. In particular, for $r=2$ there are two potential structures for optimal families, and for $r\geq3$ exactly one remains.

math.CO

Prompt Tuning for CLIP on the Pretrained Manifold

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold toward directions that are unfavorable for transfer. This drift degrades generalization. To address this limitation, we propose ManiPT, a framework that performs prompt tuning on the pretrained manifold. ManiPT introduces cosine consistency constraints in both the text and image modalities to confine the learned representations within the pretrained geometric neighborhood. Furthermore, we introduce a structural bias that enforces incremental corrections, guiding the adaptation along transferable directions to mitigate reliance on shortcut learning. From a theoretical perspective, ManiPT alleviates overfitting tendencies under limited data. Our experiments cover four downstream settings: unseen-class generalization, few-shot classification, cross-dataset transfer, and domain generalization. Across these settings, ManiPT achieves higher average performance than baseline methods. Notably, ManiPT provides an explicit perspective on how prompt tuning overfits under limited supervision.

cs.CV

Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification

Few-shot image classification remains a critical challenge in the field of computer vision, particularly in data-scarce environments. Existing methods typically rely on pre-trained visual-language models, such as CLIP. However, due to the modality gap, which is the inconsistent distribution of image and text features in the joint embedding space, directly using these features as class prototypes often leads to suboptimal performance. To address this issue, we propose a novel Cross-Modal Mapping (CMM) method. This method globally aligns image features with the text feature space through linear transformation and optimizes their local spatial relationships using triplet loss, thereby significantly enhancing cross-modal consistency. Experimental results show that compared to other methods, CMM simplifies the training process and demonstrates higher efficiency. Furthermore, CMM improves the average Top-1 accuracy by 1.06% on 11 benchmark datasets compared to methods that partially fine-tune the backbone, and it performs excellently on 4 distribution shift datasets. Notably, CMM effectively mitigates the modality gap in pre-trained models, enabling text features to serve as effective class prototypes for image features, thus providing an efficient and highly generalizable solution for few-shot learning.

cs.CV

Single Image Reflection Separation via Dual Prior Interaction Transformer

Single image reflection separation aims to separate the transmission and reflection layers from a mixed image. Existing methods typically combine general priors from pre-trained models with task-specific priors such as text prompts and reflection detection. However, the transmission prior, as the most direct task-specific prior for the target transmission layer, has not been effectively modeled or fully utilized, limiting performance in complex scenarios. To address this issue, we propose a dual-prior interaction framework based on lightweight transmission prior generation and effective prior fusion. First, we design a Local Linear Correction Network (LLCN) that finetunes pre-trained models based on the physical constraint T=SI+B, where S and B represent pixel-wise and channel-wise scaling and bias transformations. LLCN efficiently generates high-quality transmission priors with minimal parameters. Second, we construct a Dual-Prior Interaction Transformer (DPIT) that employs a dual-stream channel reorganization attention mechanism. By reorganizing features from general and transmission priors for attention computation, DPIT achieves deep fusion of both priors, fully exploiting their complementary information. Experimental results on multiple benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance.

cs.CV