arXiv ScienceSearch

arXiv subjects

Shiyu Zhang

Publications and source records attributed to Shiyu Zhang.

At least 19 recordsLinked to original sources

Exploring K-12 Teachers' Perceptions of Students' Relationships with AI Companions: Boundaries, Intervention Strategies, and Design Implications

K-12 students increasingly form relationships with AI companions. Schools face growing expectations to teach AI literacy, yet existing frameworks treat AI as a tool rather than a relationship, and little is known about how teachers understand and act on students' relational use of AI. We conducted scenario-based interviews with 33 US K-12 teachers. Teachers welcomed academic companions but worried that intimate companions remove the developmental friction through which students learn to sustain human relationships. Teachers drew the boundaries of their jurisdiction by setting and observable wellbeing: within it they taught, talked, and watched; beyond it they positioned themselves as the adults best placed to notice and connect students with support. They envisioned AI companion literacy as shared work across the jurisdictions of counselors, parents, platforms, and policymakers, spiraling across grade levels. We introduce AI companion literacy as an extension of AI literacy and discuss implications for K-12 AI education.

cs.HC

Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent Annotation

Industrial query-to-agent matching fails when topical relevance is mistaken for executable capability, especially on long-tail and boundary-sensitive requests. We formulate annotation as \emph{capability-bound process supervision} and instantiate it with Debate-to-Skill, which uses reusable decision principles, structured deliberation, verifier-based verdict extraction, and disagreement-driven refinement. On an industrial Query2Agent benchmark, we compare Debate-to-Skill with direct-label supervision, reasoning-SFT, and structural ablations. The results test whether gains come from supervising the capability-critical decision process itself, especially on grey-zone cases where semantic relatedness and executable capability diverge.

cs.AI

CO Structures with Narrow Lines in Nearby Quiescent Regions

Using CO data from Phase I of the Milky Way Imaging Scroll Painting (MWISP) survey, we present a systematic study of molecular structures with narrow lines. We identify 57 CO structures, most of which exhibit low densities and subsonic/transonic turbulence. Among them, structures with large projected areas and diffuse, sheet-like geometries are identified as veil clouds. The low LSR velocities and the concentration of these CO structures toward both the Galactic center (e.g., Ophiuchus, Aquila) and anticenter (e.g., Cepheus, Taurus) regions suggest a local origin for the sample, as supported by distance measurements of about 200--300pc for a subset with relatively large angular extents. These nearby structures likely arise from large-scale compression driven by past supernova activity within the Local Bubble. The observed low-velocity-dispersion emission may trace quiescent regions where turbulence has decayed due to a lack of sustained energy injection. For diffuse veil clouds with an assumed magnetic field of ~10uG, ion-neutral friction may provide an additional mechanism for turbulent dissipation on sub-parsec scales corresponding to their thickness of 0.1--0.3pc. Tracing the atomic-to-molecular transition, veil clouds provide a unique window into the diffuse, quiescent precursor state of dense gas. They likely represent a widespread but previously overlooked component of the Galactic molecular gas reservoir, with significant implications for cloud formation and evolution, the total mass budget and spatial distribution of molecular gas, and the initial conditions of star formation as a related consequence.

astro-ph.GA

Unsupervised Graph Representation Learning with Complementary View Alignment

Unsupervised graph representation learning aims to derive meaningful node embeddings by capturing both structural and attribute information without relying on labeled data. Existing methods, such as GAEs, have demonstrated effectiveness but typically rely on message-passing mechanisms that assume homophily, leading to performance degradation on heterophilous graphs, where connected nodes exhibit dissimilar features. This homophily bias results in the loss of critical high-frequency components that are essential for identifying heterophilous patterns. To address these challenges, we propose \textsc{AlignGAE}, a novel extension of \textit{MaskGAE} that preserves the full frequency spectrum through complementary view alignment. Our framework introduces a dual-encoder architecture that separately processes structural and attribute information, incorporates node positional encoding to approximate Neighborhood Identity Distribution (NID), and employs dual reconstruction tasks for both edges and node attributes. We further propose theoretically grounded NID alignment strategies that ensure semantic consistency across views while preserving their distinct characteristics. Through comprehensive spectral analysis, we demonstrate that \textsc{AlignGAE} achieves optimal representation properties when the alignment loss converges. Extensive experiments across 12 benchmark datasets validate our approach, showing that \textsc{AlignGAE} outperforms state-of-the-art methods by up to 18.7\% on heterophilous graphs in node classification, while maintaining competitive performance on homophilous graphs. Our results establish a new paradigm for frequency-aware graph representation learning.

cs.LG

The Miyaoka-Yau inequality for singular varieties with big canonical or anticanonical divisors

We establish the Miyaoka-Yau inequality for $n$-dimensional projective klt varieties with big canonical divisor $K_X$: \[ (2(n+1)\widehat{c}_2(X) - n \widehat{c}_1(X)^2) \cdot \langle c_1(K_X)^{n-2} \rangle \ge 0. \] We also prove the Miyaoka-Yau inequality for K-semistable projective klt varieties with big anticanonical divisor $-K_X$. As part of our approach, we define the non-pluripolar product $\langle α_1 \cdots α_p \rangle$ on singular varieties, and establish the Bogomolov-Gieseker type inequality for $\langle α^{n-1} \rangle$-semistable Higgs sheaves with respect to a big class $α$.

math.AG

Semipositivity of the orbifold second Chern class in Fujiki's class

We study inequalities for orbifold second Chern classes of compact normal analytic varieties in Fujiki's class. We prove Miyaoka's inequality for singular varieties in Fujiki's class with nef canonical divisor, as well as the semipositivity of the orbifold second Chern class for varieties with nef anti-canonical divisor. To prove these results, we establish generic nefness theorems for tangent and cotangent sheaves and an orbifold Bogomolov--Gieseker inequality for mixed polarizations.

math.AG

Positive holomorphic sectional curvature on rational surfaces

In 1975, Hitchin proved that any compact complex surface admitting a Kähler metric with positive holomorphic sectional curvature $HSC>0$ is rational. Conversely, he constructed such metrics on all Hirzebruch surfaces $\mathbb{F}_k$, as a first step towards characterizing rational surfaces by the existence of a Kähler metric with suitable curvature positivity. In this paper, we prove that every projective manifold $X$ obtained from a projective toric manifold by a finite sequence of blow-ups at points admits a Kähler metric with $HSC>0$. This statement applies to all rational surfaces and therefore completes Hitchin's result, resolving the complex surface case of a problem of Yau listed in "Open Problems in Geometry". The proof has two main ingredients. First, we prove that the toric Kähler metric on a projective toric manifold arising from Delzant's construction has $HSC>0$. Second, via a one-parameter degeneration, we construct, for any such $X$, a smooth projective family $π:\mathcal X\to\mathbb C$ such that $\mathcal X_t\simeq X$ for $t\ne0$, while $\mathcal X_0$ is a projective toric manifold.

math.DG

A study of the Physical Properties and Star Formation Activity of a Large Sample of Molecular Clouds: I Distances

Accurate distances to molecular clouds are crucial for determining their physical properties, understanding star formation, and tracing Galactic spiral structure. A number of 103,517 molecular clouds has been identified by the DBSCAN algorithm in the MWISP Phase I CO survey (l = 9.75-229.75 deg, |b| <= 5.25 deg), most of which lack reliable distances. In this work, we propose three independent methods, all of which match the molecular cloud's velocity-integrated intensity maps of 12CO lines from the MWISP with the three-dimensional dust extinction maps derived from Gaia, Pan-STARRS 1, and 2MASS, to determine molecular cloud distances. We present a catalog of 1,573 molecular clouds with robust distances ranging from approximately 150 pc to 3000 pc, 90 percent of which are measured for the first time, with typical statistical and systematic uncertainties of approximately 20% and 10%, respectively. We also derive their physical properties, such as their mass and sizes. This publicly available catalog of molecular clouds with distances provides a foundation for testing molecular cloud scaling relations and probing how cloud conditions influence star formation across diverse Galactic environments.

astro-ph.GA

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Overestimation in evaluating large language models (LLMs) has become an increasing concern. Due to the contamination of public benchmarks or imbalanced model training, LLMs may achieve unreal evaluation results on public benchmarks, either intentionally or unintentionally, which leads to unfair comparisons among LLMs and undermines their realistic capability assessments. Existing benchmarks attempt to address these issues by keeping test cases permanently secret, mitigating contamination through human evaluation, or repeatedly collecting and constructing new samples. However, these approaches fail to ensure reproducibility, transparency, and high efficiency simultaneously. Moreover, the extent of overestimation in current LLMs remains unquantified. To address these issues, we propose ArxivRoll, a dynamic evaluation framework inspired by one-time pad encryption in cryptography. ArxivRoll comprises two key components: \emph{i) SCP (Sequencing, Cloze, and Prediction)}, an automated generator for private test cases, and \emph{ii) Rugged Scores (RS)}, metrics that measure the proportion of public benchmark contamination and training bias. Leveraging SCP, ArxivRoll constructs a new benchmark every six months using recent articles from ArXiv and employs them for one-time evaluations of LLM performance. Extensive experiments demonstrate the high quality of our benchmark, and we provide a systematic evaluation of current LLMs. The source code is available at https://github.com/liangzid/ArxivRoll/.

cs.CL

Value-Decomposed Reinforcement Learning Framework for Taxiway Routing with Hierarchical Conflict-Aware Observations

Taxiway routing and on-surface conflict avoidance are coupled safety-critical decision problems in airport surface operations. Existing planning and optimization methods are often limited by online computational cost, while reinforcement learning methods may struggle to represent downstream traffic conflicts and balance multiple objectives. This paper presents Conflict-aware Taxiway Routing (CaTR), a reinforcement learning framework for real-time multi-aircraft taxiway routing. CaTR constructs a grid-based airport surface environment with action masking, introduces a hierarchical foresight traffic representation to encode current and downstream conflict-related traffic conditions, and adopts a value-decomposed reinforcement learning strategy to prioritize sparse but safety-critical objectives. Experiments are conducted on a realistic environment based on Changsha Huanghua International Airport under multiple traffic density levels. Results show that CaTR achieves better safety--efficiency trade-offs than representative planning, optimization, and reinforcement learning baselines while maintaining practical runtime.

cs.AI

Repeated Deceptive Path Planning against Learnable Observer

We study the problem of deceptive path planning (DPP), where an agent aims to conceal its true destination from external observers. While existing work assumes static, non-learning observers, real-world adversaries-such as in critical goods transportation or military operations-can adapt by learning from historical trajectories. To address this gap, we introduce Repeated Deceptive Path Planning (RDPP), a new formulation that explicitly models learnable observers. We show that existing DPP methods fail under this setting, as they cannot adapt to evolving adversarial predictions. While incorporating observer previous predictions into updates enables some adaptation, such incremental updates cause accumulative lag that degrades deception. To this end, we propose Deceptive Meta Planning (DeMP), a two-level optimization framework that combines episode-level adaptation, which enables short-term policy adjustment to counter updated observer, and meta-level updates, which leverage cross-episode feedback to capture how observers update their models and accelerate adaptation in future episodes. In this way, DeMP mitigates the accumulation of adaptation lag, enabling sustained deception against a learning observer. Experiments across environments demonstrate that DeMP significantly outperforms existing approaches in RDPP while maintaining competitive path cost. Our results highlight the importance of modeling repeated interactions with learnable adversaries, providing new insights into deception and privacy in multi-agent systems.

cs.AI

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.

cs.AI

From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning

Although machine unlearning is essential for removing private, harmful, or copyrighted content from LLMs, current benchmarks often fail to faithfully represent the true ``forgetting scope'' learned by the model. We formalize two distinct unlearning granularities, domain-level and instance-level, and propose \BiForget, an automated framework for synthesizing high-quality forget sets. Unlike prior work relying on \emph{external} generators, \BiForget exploits the target model per se to elicit data that matches its internal knowledge distribution through seed-guided and adversarial prompting. Our experiments across diverse benchmarks show that it achieves a superior balance of relevance, diversity, and efficiency. Quantitatively, in the Harry Potter domain, it improves relevance by ${\sim}20$ and diversity by ${\sim}$0.05 while \emph{halving} the total data size compared to SOTAs. Ultimately, it facilitates more robust forgetting and better utility preservation, providing a more rigorous foundation for evaluating LLM unlearning.

cs.CL

Distances to molecular clouds in the Galactic longitude l=10-20 deg from the MWISP 12CO 1-0 survey

We present distances to 56 molecular clouds within $10\degr \leq l \leq 20\degr$ and $|b| \leq 5.25\degr$ from the Milky Way Imaging Scroll Painting (MWISP) $^{12}$CO survey, 47 of which are first-time determinations. The molecular clouds were identified using the DBSCAN algorithm, and their distances were measured with the model-calibrated color-distance method using $J-K{_s}$ colors and the distances provided by 2MASS and \textit{Gaia} EDR3. The distances range from $\sim$275 pc to $\sim$2118 pc. We also derived the physical properties of molecular clouds and found a moderate correlation between the dust extinction and the $^{12}$CO integrated intensity.

astro-ph.GA

Efficient Reasoning via Thought Compression for Language Segmentation

Chain-of-thought (CoT) reasoning has significantly improved the performance of large multimodal models in language-guided segmentation, yet its prohibitive computational cost, stemming from generating verbose rationales, limits real-world applicability. We introduce WISE (Wisdom from Internal Self-Exploration), a novel paradigm for efficient reasoning guided by the principle of \textit{thinking twice -- once for learning, once for speed}. WISE trains a model to generate a structured sequence: a concise rationale, the final answer, and then a detailed explanation. By placing the concise rationale first, our method leverages autoregressive conditioning to enforce that the concise rationale acts as a sufficient summary for generating the detailed explanation. This structure is reinforced by a self-distillation objective that jointly rewards semantic fidelity and conciseness, compelling the model to internalize its detailed reasoning into a compact form. At inference, the detailed explanation is omitted. To address the resulting conditional distribution shift, our inference strategy, WISE-S, employs a simple prompting technique that injects a brevity-focused instruction into the user's query. This final adjustment facilitates the robust activation of the learned concise policy, unlocking the full benefits of our framework. Extensive experiments show that WISE-S achieves state-of-the-art zero-shot performance on the ReasonSeg benchmark with 58.3 cIoU, while reducing the average reasoning length by nearly \textbf{5$\times$} -- from 112 to just 23 tokens. Code is available at \href{https://github.com/mrazhou/WISE}{WISE}.

cs.CV

PGR-Net: Prior-Guided ROI Reasoning Network for Brain Tumor MRI Segmentation

Brain tumor MRI segmentation is essential for clinical diagnosis and treatment planning, enabling accurate lesion detection and radiotherapy target delineation. However, tumor lesions occupy only a small fraction of the volumetric space, resulting in severe spatial sparsity, while existing segmentation networks often overlook clinically observed spatial priors of tumor occurrence, leading to redundant feature computation over extensive background regions. To address this issue, we propose PGR-Net (Prior-Guided ROI Reasoning Network) - an explicit ROI-aware framework that incorporates a data-driven spatial prior set to capture the distribution and scale characteristics of tumor lesions, providing global guidance for more stable segmentation. Leveraging these priors, PGR-Net introduces a hierarchical Top-K ROI decision mechanism that progressively selects the most confident lesion candidate regions across encoder layers to improve localization precision. We further develop the WinGS-ROI (Windowed Gaussian-Spatial Decay ROI) module, which uses multi-window Gaussian templates with a spatial decay function to produce center-enhanced guidance maps, thus directing feature learning throughout the network. With these ROI features, a windowed RetNet backbone is adopted to enhance localization reliability. Experiments on BraTS-2019/2023 and MSD Task01 show that PGR-Net consistently outperforms existing approaches while using only 8.64M Params, achieving Dice scores of 89.02%, 91.82%, and 89.67% on the Whole Tumor region. Code is available at https://github.com/CNU-MedAI-Lab/PGR-Net.

cs.CV

Compact Kähler manifolds with partially semi-positive curvature

In this paper, we study MRC fibrations of compact Kähler manifolds with partially semi-positive curvature. We first prove that a compact Kähler manifold is rationally connected if its tangent bundle is BC-$p$ positive for all $1\leq p\leq \dim X$. As applications, we confirm a conjecture that any compact Kähler manifold with positive orthogonal Ricci curvature must be rationally connected, and generalize a result of Heier-Wong and Yang to the conformally Kähler case. The second result concern structure theorems for two immediate curvature conditions. We prove that, a compact Kähler manifold with $k$-semi-positive Ricci curvature or semi-positive $k$-scalar curvature, either the rational dimension $\geq n-k+1$ or it admits a locally constant fibration $f: X\rightarrow Y$ such that the fibre is rationally connected and the image $Y$ is Ricci-flat.

math.DG