arXiv ScienceSearch

arXiv subjects

Shun Zhang

Publications and source records attributed to Shun Zhang.

At least 19 recordsLinked to original sources

Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection

The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.

cs.CV

Why the Kellogg Mesh Is Radial: A Mathematical Explanation of a Classical Computational Benchmark

Kellogg's checkerboard interface problem is a classical benchmark for robust adaptive finite element methods. Its successful adaptive meshes are radial: they refine strongly toward the interface crossing but show no angular structure, despite the large contrast and the asymmetric solution. We explain this by proving that the singular solution $u(r,\theta)=r^\gamma\mu(\theta)$ satisfies the exact identities $\kappa|\nabla u|^2=\Lambda r^{2\gamma-2}$ and $\kappa|\nabla^2 u|_F^2=2(1-\gamma)^2\Lambda r^{2\gamma-4}$, with $\Lambda=\gamma^2\cos^2(\pi\gamma/4)$ and the Hessian taken separately in each quadrant. The point is what has disappeared: the right-hand sides depend on $r$ alone, although $\kappa$ and $u$ each depend on the angle as well. Combined with equal discretization-error distribution, this shows that the target element density is radial, so a correct mesh should display nothing but refinement toward the center, and a Kellogg mesh that is not radial is visible evidence that the computation is not following the coefficient-weighted local difficulty. The reading is specific to this benchmark: on a second interface problem the same estimator correctly produces a strongly material-biased mesh, with a computed element-count ratio of $3.934$ against the predicted $4$. A byproduct gives the benchmark constants in closed form, so the problem data can be generated from $\gamma$ alone at any precision.

math.NA

Mixed Finite Element Methods for a Dirac Source: Divergence-Form Splitting and L^p Error Analysis

For a mixed finite element method, a Dirac source is first a failure of duality, not of regularity: the conservation equation is tested against a Lebesgue space, and a Dirac measure lies in the dual of none. We therefore remove the measure from the conservation law by a divergence-form splitting. An explicit field whose divergence is the Dirac measure is subtracted from the physical flux, and the modified flux is taken as the mixed unknown, so that only the regular part of the load remains in the conservation equation. Equivalently, and independently of any discretization, the Dirac problem is rewritten as an elliptic equation whose data are in divergence form, generated by a field of L^p. The subtracted field depends on the location of the source alone and not on the coefficient. No coefficient-dependent singular solution and no discrete delta is needed, only the load vector of the RT_0-P_0 system changes, and the source may sit anywhere relative to the mesh: at a vertex, on a coefficient interface, or inside an element. Unless the splitting is matched to the operator at the source, the modified flux lies in L^p for every p<2 but not in L^2, so the flux error analysis has to leave the Hilbert scale. We prove a quasi-best approximation bound for the flux, and with it that on a quasi-uniform family the flux error is exactly of order h^(2/p-1): the matching lower bound comes already from the single element carrying the source. Grading the mesh there restores first-order complexity, N^(-1/2) in the number of elements, and the adaptive computations attain it. The scalar variable is limited only by piecewise constant approximation of the solution, which it attains. We also prove a residual norm equivalence in the Lebesgue scale, yielding a computable L^p estimator, reliable and locally efficient for the mixed flux together with a recovered potential.

math.NA

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.

cs.CV

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code is available at https://github.com/eeNickTang/SPaTS.

cs.CV

Goal-Oriented Error Estimation for Least-Squares Finite Element Methods via Physically Meaningful Adjoint PDEs

We develop a goal-oriented error-estimation framework for first-order system least-squares (FOSLS) finite element methods based on the physical PDE adjoint rather than the adjoint induced by the least-squares formulation. Identified explicitly from the differential operator, its mixed boundary conditions, and the output, the physical adjoint admits its own first-order flux system and hence a native, built-in least-squares estimator, which the least-squares-induced adjoint does not. Our error identities rest on a primal-dual corrected functional: they follow from the continuous primal and adjoint equations alone, hold for arbitrary conforming approximations, and require no Galerkin orthogonality. For outputs containing a weighted Dirichlet-boundary flux, whose weight becomes the essential datum of the adjoint, we construct two corrected approximations, one from the potentials and one also using the fluxes, and prove product-type error estimates in the primal and adjoint errors. The least-squares functionals then yield computable a posteriori bounds and a balanced marking indicator whose element sum equals the product of the two estimators exactly. Numerical experiments confirm the predicted convergence and the effectiveness of the marking.

math.NA

Archetypometrics of 'Friends'

Storytelling inherently revolves around characters. Using the television sitcom `Friends' as a case study, we investigate how well archetype vectors capture both individual characterization and the relational structure of a specific ensemble. Our work is based on the archetypometrics framework, which locates 2,000 fictional characters from 341 stories in a continuous space derived from 464 bipolar traits. We proceed in three stages: interpreting each character's archetypal profile against narrative evidence, projecting the ensemble onto ousiograms of the six essential dimensions, and measuring pairwise similarity with vector inner products. We show that the six characters of `Friends' occupy distinct archetypal positions that accord with their established identities, while the projections expose ensemble structure invisible in individual profiles, including the collapse of the Angel--Demon dimension, a signature of the sitcom's uniformly sympathetic cast. Based on inner products, we construct a similarity matrix that resolves three main kinds of relational structure: alignment (e.g., Phoebe--Joey), contrast (e.g., Phoebe--Ross), and orthogonality (e.g., Rachel--Ross and Monica--Chandler). The orthogonality of the romantic pairings affords a detailed view of relationships built on complementary rather than overlapping character traits. Overall, our case study suggests that for ensemble-based stories the archetypometric geometry is fully interpretable in narrative terms, from individual identities to the structure of the group's relationships.

physics.soc-ph

Narrative Structure in Tropes: A Computational Analysis of `Friends'

Tropes are recurring narrative devices in television and film. We carry out a computational analysis of tropes in the sitcom Friends, using human-curated trope annotations from TVTropes, episode transcripts, and IMDb ratings. Because automatic trope detection remains challenging, we treat existing trope annotations as a curated analytical layer and focus on their downstream narrative and semantic functions. We first examine the relationship between episode-level trope frequency and audience reception. We find a statistically significant positive association between trope count and weighted IMDb ratings, although the modest explanatory power suggests that more than trope density alone explains audience evaluation. We then connect trope annotations to dialogue transcripts and represent trope-related dialogue using TF-IDF-based semantic features. Using PCA and k-means clustering, we group 1,954 distinct tropes into 15 semantically interpretable clusters. Chi-square analyses show that the six main characters are unevenly distributed across these clusters, with character-specific trope profiles that are broadly consistent with their established narrative identities. Finally, we project trope clusters into the ousiometric power-danger space to examine their semantic organization. The results show that "Physical and Sexual Comedy" occupies a region associated with relatively high danger, while "Revelation, Surprise, and Reaction" occupies a region associated with relatively high power. Overall, our work demonstrates a way to operationalize trope measurement and shows that identifiable trope clusters can provide holistic "distant reading" descriptions of characters and stories.

physics.soc-ph

NoRIN: Backbone-Adaptive Reversible Normalization for Time-Series Forecasting

Reversible instance normalization (RevIN) and its successors (Dish-TS, SAN, FAN) have become the de facto plug-in for time-series forecasting, yet the map they apply to each data point is strictly affine, $x \mapsto ax+b$, so they cannot reshape the underlying distribution -- heavy tails remain heavy and skewness remains uncorrected. We propose NoRIN, a non-linear reversible normalization based on the arcsinh-form Johnson $S_U$ transform with two shape parameters $(\delta,\varepsilon)$ that control tailedness and skewness; the linear $Z$-score used by RevIN is recovered only in the limit $\delta \to \infty$. Training $(\delta,\varepsilon)$ jointly with the backbone via gradient descent reliably pushes them toward this linear limit within a few epochs -- a phenomenon we name the degeneration problem: the forecasting loss is locally indifferent to shape, and the high-capacity backbone compensates for any monotone reparameterization of its input. NoRIN escapes the degeneration by decoupling shape selection from gradient training: $(\delta,\varepsilon)$ are initialized by a closed-form Slifker-Shapiro quantile fit and refined by Bayesian optimization on the validation objective, while the inner training loop is identical to standard RevIN-style training. Across six representative backbones x five real-world datasets x three prediction horizons (90 configurations), decoupled shape optimization recovers $(\delta^\star,\varepsilon^\star)$ that sit systematically far from the linear limit, with values that vary in a backbone-dependent way. This empirically supports the central thesis: different backbones genuinely require different normalization parameters to reach their best performance.

cs.LG

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context management, and state persistence. Yet this same architectural centrality makes the harness a high-value attack surface: a single compromise at the harness level can cascade through the entire execution pipeline. We observe that existing security approaches suffer from structural mismatch, leaving them blind to harness-internal state and unable to coordinate across the different phases of agent operation. In this paper, we introduce \safeharness{}, a security architecture in which four proposed defense layers are woven directly into the agent lifecycle to address above significant limitations: adversarial context filtering at input processing, tiered causal verification at decision making, privilege-separated tool control at action execution, and safe rollback with adaptive degradation at state update. The proposed cross-layer mechanisms tie these layers together, escalating verification rigor, triggering rollbacks, and tightening tool privileges whenever sustained anomalies are detected. We evaluate \safeharness{} on benchmark datasets across diverse harness configurations, comparing against four security baselines under five attack scenarios spanning six threat categories. Compared to the unprotected baseline, \safeharness{} achieves an average reduction of approximately 38\% in UBR and 42\% in ASR, substantially lowering both the unsafe behavior rate and the attack success rate while preserving core task utility.

cs.CR

Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion

Semi-supervised remote sensing (RS) image semantic segmentation offers a promising solution to alleviate the burden of exhaustive annotation, yet it fundamentally struggles with pseudo-label drift, a phenomenon where confirmation bias leads to the accumulation of errors during training. In this work, we propose Co2S, a stable semi-supervised RS segmentation framework that synergistically fuses priors from vision-language models and self-supervised models. Specifically, we construct a heterogeneous dual-student architecture comprising two distinct ViT-based vision foundation models initialized with pretrained CLIP and DINOv3 to mitigate error accumulation and pseudo-label drift. To effectively incorporate these distinct priors, an explicit-implicit semantic co-guidance mechanism is introduced that utilizes text embeddings and learnable queries to provide explicit and implicit class-level guidance, respectively, thereby jointly enhancing semantic consistency. Furthermore, a global-local feature collaborative fusion strategy is developed to effectively fuse the global contextual information captured by CLIP with the local details produced by DINOv3, enabling the model to generate highly precise segmentation results. Extensive experiments on six popular datasets demonstrate the superiority of the proposed method, which consistently achieves leading performance across various partition protocols and diverse scenarios. Project page is available at https://xavierjiezou.github.io/Co2S/.

cs.CV

Low-Complexity Channel Estimation for Internet of Vehicles AFDM Communications With Sparse Bayesian Learning

Affine frequency division multiplexing (AFDM) has been considered as a promising waveform to enable high-reliable connectivity in the internet of vehicles. However, accurate channel estimation is critical and challenging to achieve the expected performance of the AFDM systems in doubly-dispersive channels. In this paper, we propose a sparse Bayesian learning (SBL) framework for AFDM systems and develop a dynamic grid update strategy with two off-grid channel estimation methods, i.e., grid-refinement SBL (GR-SBL) and grid-evolution SBL (GE-SBL) estimators. Specifically, the GR-SBL employs a localized grid refinement method and dynamically updates grid for a high-precision estimation. The GE-SBL estimator approximates the off-grid components via first-order linear approximation and enables gradual grid evolution for estimation accuracy enhancement. Furthermore, we develop a distributed computing scheme to decompose the large-dimensional channel estimation model into multiple manageable small-dimensional sub-models for complexity reduction of GR-SBL and GE-SBL, denoted as distributed GR-SBL (D-GR-SBL) and distributed GE-SBL (D-GE-SBL) estimators, which also support parallel processing to reduce the computational latency. Finally, simulation results demonstrate that the proposed channel estimators outperform existing competitive schemes. The GR-SBL estimator achieves high-precision estimation with fine step sizes at the cost of high complexity, while the GE-SBL estimator provides a better trade-off between performance and complexity. The proposed D-GR-SBL and D-GE-SBL estimators effectively reduce complexity and maintain comparable performance to GR-SBL and GE-SBL estimators, respectively.

cs.IT

Possible quasi-periodic optical oscillations of ZTF blazars

Based on the Zwicky Transient Facility (ZTF), we selected 10 blazars as our sample sources. Among these, we found four blazars (J 0923.5+4125, J 1221.3+3010, J 1503.5+4759, and J 1652.7+4024) showing possible indications of quasi periodic oscillations (QPOs) modulation. We conducted a detailed analysis of their optical light curves (g- and r-bands) over the past five years using the root mean square (RMS)-Flux relation, flux distribution, and QPO detection methods to investigate their variability characteristics. A linear RMS-Flux relation is present in both bands, and their flux distributions follow a log-normal form. This suggests that optical variability may arise from multiplicative, nonlinear processes across different timescales and flux states. Further QPO analysis using the weighted wavelet Z-transform (WWZ), Lomb-Scargle periodogram (LSP), and autoregressive integrated moving average (ARIMA) methods identified candidate periodic signals in four blazars. J 0923.5+4125 (period $\sim$ 205 days) and J 1221.3+3010 ($\sim$ 630 days) show local significances of $\sim 3 \sigma$, whereas J 1503.5+4759 ($\sim$ 38.5 days) and J 1652.7+4024 ($\sim$ 48 days) reach $\sim 4 \sigma$. After accounting for the look-elsewhere effect, the global significances for J 1503.5+4759 in the g- and r-bands are $\sim 2.7 \sigma$, while for J 1652.7+4024 they are approximately $\sim 2.5 \sigma$ in both bands. These two blazars warrant further monitoring and investigation.

astro-ph.GA

CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions

We present CycliST, a novel benchmark dataset designed to evaluate Video Language Models (VLM) on their ability for textual reasoning over cyclical state transitions. CycliST captures fundamental aspects of real-world processes by generating synthetic, richly structured video sequences featuring periodic patterns in object motion and visual attributes. CycliST employs a tiered evaluation system that progressively increases difficulty through variations in the number of cyclic objects, scene clutter, and lighting conditions, challenging state-of-the-art models on their spatio-temporal cognition. We conduct extensive experiments with current state-of-the-art VLMs, both open-source and proprietary, and reveal their limitations in generalizing to cyclical dynamics such as linear and orbital motion, as well as time-dependent changes in visual attributes like color and scale. Our results demonstrate that present-day VLMs struggle to reliably detect and exploit cyclic patterns, lack a notion of temporal understanding, and are unable to extract quantitative insights from scenes, such as the number of objects in motion, highlighting a significant technical gap that needs to be addressed. More specifically, we find no single model consistently leads in performance: neither size nor architecture correlates strongly with outcomes, and no model succeeds equally well across all tasks. By providing a targeted challenge and a comprehensive evaluation framework, CycliST paves the way for visual reasoning models that surpass the state-of-the-art in understanding periodic patterns.

cs.CV

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling semantic controls from generation pipelines, we revisit the architectural potential of Diffusion Transformers (DiTs) through the lens of expert specialization. This paper introduces Face-MoGLE, a novel framework featuring: (1) Semantic-decoupled latent modeling through mask-conditioned space factorization, enabling precise attribute manipulation; (2) A mixture of global and local experts that captures holistic structure and region-level semantics for fine-grained controllability; (3) A dynamic gating network producing time-dependent coefficients that evolve with diffusion steps and spatial locations. Face-MoGLE provides a powerful and flexible solution for high-quality, controllable face generation, with strong potential in generative modeling and security applications. Extensive experiments demonstrate its effectiveness in multimodal and monomodal face generation settings and its robust zero-shot generalization capability. Project page is available at https://github.com/XavierJiezou/Face-MoGLE.

cs.CV

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment

Zero-shot voice conversion (VC) synthesizes speech in a target speaker's voice while preserving linguistic and paralinguistic content. However, timbre leakage-where source speaker traits persist-remains a challenge, especially in neural codec and LLM-based VC, where quantized representations entangle speaker identity with content. We introduce SemAlignVC, an architecture designed to prevent timbre leakage using SemAlign, a novel method that aligns text and audio representations to ensure speaker-independent semantic encoding. This disentangled representation conditions an autoregressive transformer for high-fidelity conversion without explicit speaker embeddings. Experiments show SemAlignVC significantly reduces timbre leakage, outperforming baselines in speaker timbre similarity, intelligibility, and naturalness, making it a robust, privacy-preserving, and generalizable VC solution. Audio samples can be accessed at https://shivammehta25.github.io/SemAlignVC/

eess.AS

AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models

The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks are primarily designed for text and fail to capture vulnerabilities introduced by the acoustic properties of audio. We find that significant trustworthiness risks in ALLMs arise from non-semantic acoustic cues, such as timbre, accent, and background noise, which can be exploited to manipulate model behavior. To address this gap, we propose AudioTrust, the first large-scale and systematic framework for evaluating ALLM trustworthiness under audio-specific risks. AudioTrust covers six key dimensions: fairness, hallucination, safety, privacy, robustness, and authenticition. It includes 26 sub-tasks and a curated dataset of more than 4,420 audio samples collected from real-world scenarios, including daily conversations, emergency calls, and voice assistant interactions, and is specifically designed to probe trustworthiness across multiple dimensions. Our comprehensive evaluation spans 18 experimental settings and uses human-validated automated pipelines to enable objective and scalable assessment of model outputs. Experimental results on 14 state-of-the-art open-source and closed-source ALLMs reveal important limitations and failure boundaries under diverse high-risk audio scenarios, providing critical insights for the secure and trustworthy deployment of future audio models. Our platform and benchmark are publicly available at https://github.com/JusperLee/AudioTrust.

cs.SD

Bipartite Randomized Response Mechanism for Local Differential Privacy

With the increasing importance of data privacy, Local Differential Privacy (LDP) has recently become a strong measure of privacy for protecting each user's privacy from data analysts without relying on a trusted third party. In this paper, we consider the problem of high-utility differentially private release. Given a domain of items and a distance-defined utility function, our goal is to design a differentially private mechanism that releases an item with the global expected error as small as possible. The most common LDP mechanism for this task is the Generalized Randomized Response (GRR) mechanism that treats all candidate items equally except for the true item. In this paper, we introduce Bipartite Randomized Response mechanism (BRR), which adaptively divides all candidate items into two parts by utility rankings. In the local search phase, we confirm how many high-utility candidates to be assigned with high release probability, which gives the locally optimal bipartite classification of all candidates. For preserving LDP, the global search phase uniformly selects the smallest number of dynamic high-utility candidates obtained locally. In particular, we give explicit formulas on the uniform number of dynamic high-utility candidates. The global expected error of our BRR can theoretically deliver a decrease with an asymptotically exact ratio, and when the privacy budget is set to $3$ the expected error can be reduced by $66.4\%$. Extensive experiments demonstrate that BRR outperforms the state-of-the-art methods across the standard metrics and datasets.

cs.CR