arXiv ScienceSearch

arXiv subjects

Tingting Li

Publications and source records attributed to Tingting Li.

At least 19 recordsLinked to original sources

Instability of naked singularities of perfect fluid under $C^{1,α}$ gravitational perturbations

We study the instability of the naked singularities arising in the spherically symmetric self-similar collapsing of the Einstein--Euler system under gravitational perturbations. We show that small $C^{1,α}$ perturbations (without any symmetries) of the initial conformal metric lead to trapped surface formation. One key point is that the speed of sound is strictly slower the the speed of light, so the gravitational radiation can become sufficiently concentrated to form a trapped surface before the fluid develops any singularity.

gr-qc

Interior gravitational perturbations to naked singularities of a scalar field

For the $k$-self-similar naked singularity solutions of Einstein--scalar field equations constructed by Christodoulou, we construct a family of interior gravitational perturbations leading to trapped surface formation that remain large but finite at the threshold Hölder regularity, while converging to zero in all regularities below the threshold. This extends the interior spherically symmetric result below the threshold in \cite{Li25} to non-spherically symmetric setting. We further construct a genuinely localized family of perturbations whose angular support shrinks to a single point, with smooth convergence to zero away from that point. This exploits the additional angular freedom available beyond spherical symmetry. Moreover, when measured instead in Sobolev regularity, such genuinely localized family of perturbations still converges to zero in all regularities below the same threshold.

gr-qc

Notes on Reversed Products of Two Elements in Rings

This paper investigates reversed product properties of two elements in rings. Motivated by Cline's formula, we characterize reversible and $\ast$-reversible rings in terms of group invertible elements, EP elements, and the transfer behaviour of generalized inverses for reversed products. We prove that a unital ring $R$ is reversible if and only if $ab\in R^{\sharp}$ yields $ba\in R^{\sharp}$. For an involutive ring $R$, $R$ is $\ast$-reversible precisely whenever $ab\in R^{\mathrm{EP}}$ implies $b^{\ast}a\in R^{\mathrm{EP}}$. Several counterexamples are constructed to differentiate these ring classes, and the mutual inclusion relations among them are also discussed.

math.RA

CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification

Microbial density is clinically important for tumor assessment and treatment decision-making, and recent advances in deep learning suggest that it can be non-invasively inferred from multimodal MRI. In this work, MRI-based Microbial Density Stratification (MRI-MDS) is first investigated as a patient-level representation learning task, and Center Heatmap-driven Macro-micro modeling Network (CHM-Net) is introduced for this task. CHM-Net first establishes the link between imaging phenotypes and microbial states through center heatmap-guided small-lesion response localization. Building upon this, it constructs patient-level macro-micro evidence from localized heatmap responses for microbial density prediction. Experiments on the novel GBNPC 2026 dataset constructed for MRI-MDS demonstrate the effectiveness of CHM-Net, achieving superior performance over representative baselines with a 12.06% absolute ACC gain over the strongest competing result. Additionally, auxiliary validation on two 3D medical image datasets further verifies its robustness across volumetric medical image classification scenarios. The project is available at https://anonymous.4open.science/r/CHM-Net-942E/.

eess.IV

Stochastic Liouville-transport theory of light-atom interaction noise in thermal atomic vapors

Atom-light interaction noise can limit thermal-vapor sensing. Existing theories often treat internal-state dynamics, finite-mode atomic motion, and stochastic renewal separately, obscuring their coupled contributions to measured noise. We develop a general stochastic Liouville-transport theory, tested against polarization-resolved resonant Cs D$_2$ spectra. Joint experiment-theory analysis identifies atom-light noise below approximately 100 kHz as transit-dominated. Ballistic motion through the finite Gaussian mode modulates both the coupling-weighted effective atom number and trajectory-dependent Rabi coupling, producing predominantly common-mode noise. Boundary renewal introduces atoms with independently sampled ground-state sublevels, generating differential population fluctuations with opposite effects on the circular channels. Under an applied longitudinal magnetic field, experiment and theory show the same qualitative nonmonotonic change in common-mode suppression, supporting Zeeman redistribution of the channel responses. The framework can analyze noise in other thermal-atom sensors, including Rydberg-atom electric-field measurements.

quant-ph

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and results. The challenge seeks to identify robust reconstruction pipelines that are robust under real-world adverse conditions, specifically extreme low-light and smoke-degraded environments, as captured by our RealX3D benchmark. A total of 279 participants registered for the competition, of whom 33 teams submitted valid results. We thoroughly evaluate the submitted approaches against state-of-the-art baselines, revealing significant progress in 3D reconstruction under adverse conditions. Our analysis highlights shared design principles among top-performing methods and provides insights into effective strategies for handling 3D scene degradation.

cs.CV

SciEval: A Benchmark for Automatic Evaluation of K-12 Science Instructional Materials

The need to evaluate instructional materials for K-12 science education has become increasingly important, as more educators use generative AI to create instructional materials. However, the review of instructional materials is time-consuming, expertise-intensive, and difficult to scale, motivating interest in automated evaluation approaches. While large language models (LLMs) have shown strong performance on general evaluation tasks, their performance and reliability on instructional materials remain unclear. To address this gap, we formulate Automatic Instructional Materials Evaluation (AIME) as a generative AI task that predicts scores and evidence using the rubric designed by the educator. We create a benchmark dataset and develop baseline models for AIME. First, we curate the first AIME dataset, SciEval, consisting of instructional materials annotated with pedagogy-aligned evaluation scores and evidence-based rationales. Expert annotations achieve high inter-rater reliability, resulting in a dataset of 273 lesson-level instructional materials evaluated across 13 criteria (N=3549) using the EQuIP rubric. Second, we test mainstream LLMs (GPT, Gemini, Llama, and Qwen) on SciEval and find that none achieve strong performance. Then we fine-tune Qwen3 on SciEval. Results on a held-out test set show that domain-aligned fine-tuning can achieve up to 11 percent performance gains, highlighting the importance of domain-specific fine-tuning for AIME and facilitating the use of LLMs in other educational tasks.

cs.AI

Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos

K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these interactions has made automated analysis elusive. Existing benchmarks for classroom discourse focus primarily on mathematics and rely solely on transcripts, overlooking the visual artifacts and model-based reasoning emphasized by the Next Generation Science Standards (NGSS). We address this gap with SciIBI, the first video benchmark for analyzing science classroom discourse, featuring 113 NGSS-aligned clips annotated with Core Instructional Practices (CIP) and sophistication levels. By evaluating eight state-of-the-art LLMs and Multimodal LLMs, we reveal fundamental limitations: current models struggle to distinguish pedagogically similar practices, suggesting that CIP coding requires instructional reasoning beyond surface pattern matching. Furthermore, adding video input yields inconsistent gains across architectures. Crucially, our evidence-based evaluation reveals that models often succeed through surface shortcuts rather than genuine pedagogical understanding. These findings establish science classroom discourse as a challenging frontier for multimodal AI and point toward human-AI collaboration, where models retrieve evidence to accelerate expert review rather than replace it.

cs.CY

Spatially Varying Coefficient Mallows Model Averaging

Model averaging, as an appealing ensemble technique, strategically integrates all valuable information from candidate models to construct fast and accurate prediction. Despite of having been widely practiced in many fields such as cross-sectional data, censored data and longitudinal data, its application to spatial data characterized by inherent spatial heterogeneity remains surprisingly limited. To mitigate risk of model misspecification and enhance the flexibility of prediction, we propose a combined estimator constructed by computing the weighted average of estimators derived from a set of spatially varying coefficient candidate models. Herein, the model weights are determined via a Mallows-type criterion, which dynamically calibrates the relative importance of individual candidate models in the ensemble. Theoretically, we establish desirable asymptotic properties under two practical scenarios. First, in the case where all candidate models are misspecified, the proposed model averaging estimator attains asymptotic optimality in the sense that it minimizes the squared error loss function asymptotically. Second, when the candidate model set encompasses at least one quasi-correct model, the weights assigned by the Mallows-type criterion asymptotically concentrate on the quasi-correct models, and the resulting model averaging estimator converges in probability to the true conditional mean. Both simulation studies and a real-world empirical example demonstrate that the proposed method generally outperforms alternative comparative approaches in terms of predictive accuracy and robustness.

stat.ME

DrawSim-PD: Simulating Student Science Drawings to Support NGSS-Aligned Teacher Diagnostic Reasoning

Developing expertise in diagnostic reasoning requires practice with diverse student artifacts, yet privacy regulations prohibit sharing authentic student work for teacher professional development (PD) at scale. We present DrawSim-PD, the first generative framework that simulates NGSS-aligned, student-like science drawings exhibiting controllable pedagogical imperfections to support teacher training. Central to our approach are apability profiles--structured cognitive states encoding what students at each performance level can and cannot yet demonstrate. These profiles ensure cross-modal coherence across generated outputs: (i) a student-like drawing, (ii) a first-person reasoning narrative, and (iii) a teacher-facing diagnostic concept map. Using 100 curated NGSS topics spanning K-12, we construct a corpus of 10,000 systematically structured artifacts. Through an expert-based feasibility evaluation, K--12 science educators verified the artifacts' alignment with NGSS expectations (>84% positive on core items) and utility for interpreting student thinking, while identifying refinement opportunities for grade-band extremes. We release this open infrastructure to overcome data scarcity barriers in visual assessment research.

cs.CY

Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials

Designing high-quality, standards-aligned instructional materials for K--12 science is time-consuming and expertise-intensive. This study examines what human experts notice when reviewing AI-generated evaluations of such materials, aiming to translate their insights into design principles for a future GenAI-based instructional material design agent. We intentionally selected 12 high-quality curriculum units across life, physical, and earth sciences from validated programs such as OpenSciEd and Multiple Literacies in Project-based Learning. Using the EQuIP rubric with 9 evaluation items, we prompted GPT-4o, Claude, and Gemini to produce numerical ratings and written rationales for each unit, generating 648 evaluation outputs. Two science education experts independently reviewed all outputs, marking agreement (1) or disagreement (0) for both scores and rationales, and offering qualitative reflections on AI reasoning. This process surfaces patterns in where LLM judgments align with or diverge from expert perspectives, revealing reasoning strengths, gaps, and contextual nuances. These insights will directly inform the development of a domain-specific GenAI agent to support the design of high-quality instructional materials in K--12 science education.

cs.CY

Adaptive Fidelity Estimation for Quantum Programs with Graph-Guided Noise Awareness

Fidelity estimation is a critical yet resource-intensive step in testing quantum programs on noisy intermediate-scale quantum (NISQ) devices, where the required number of measurements is difficult to predefine due to hardware noise, device heterogeneity, and transpilation-induced circuit transformations. We present QuFid, an adaptive and noise-aware framework that determines measurement budgets online by leveraging circuit structure and runtime statistical feedback. QuFid models a quantum program as a directed acyclic graph (DAG) and employs a control-flow-aware random walk to characterize noise propagation along gate dependencies. Backend-specific effects are captured via transpilation-induced structural deformation metrics, which are integrated into the random-walk formulation to induce a noise-propagation operator. Circuit complexity is then quantified through the spectral characteristics of this operator, providing a principled and lightweight basis for adaptive measurement planning. Experiments on 18 quantum benchmarks executed on IBM Quantum backends show that QuFid significantly reduces measurement cost compared to fixed-shot and learning-based baselines, while consistently maintaining acceptable fidelity bias.

quant-ph

Reactive near-field subwavelength microwave imaging with a non-invasive Rydberg probe

Non-invasive microwave field imaging--accurately mapping field distributions without perturbing them--is essential in areas such as aerospace engineering, biomedical imaging and integrated-circuit diagnostics. Conventional metal probes, however, inevitably perturb reactive near fields: they act as strong scatterers that drive induced currents and secondary radiation, remap evanescent components and thereby degrade both accuracy and spatial resolution, particularly in the reactive near-field regime that is most relevant to these applications. Here we demonstrate, to our knowledge for the first time, reactive near-field subwavelength imaging of microwave fields using the quantum non-demolition properties of Rydberg atoms, realized with a compact, non-invasive single-ended fibre-integrated Rydberg probe engineered to minimize field disturbance. The probe achieves an imaging resolution of {\unboldmath$λ/56$}, and the measured field distributions agree with full-wave simulations with structural similarity approaching unity, confirming both its subwavelength spatial resolution and its genuinely non-invasive character compared with conventional metal-based probes. Because the atomic sensor is intrinsically isotropic, the same device can faithfully image multi-dimensional field structures without orientation-dependent calibration. Our results therefore establish a general, non-invasive route to high-accuracy, subwavelength reactive near-field microwave imaging, with particular promise for applications such as chip-defect detection and integrated-circuit diagnostics, where even small perturbations by the probe can mask the underlying physics of interest.

quant-ph

Multi-Channel Amplitude-Phase Asymmetric-Encrypted Janus Acoustic Meta-Holograms

Encrypted optical and acoustic meta-holograms only focus on the encrypted hologram in a single channel, viz. modulating spatial amplitude to project a holographic image. In this research, the unique concept of multi-channel amplitude-phase asymmetric-encrypted Janus acoustic meta-holograms is proposed, demonstrating remarkable capabilities of generating, encrypting, and decrypting both amplitude and phase holographic images on both sides of a metascreen. The flexible and decoupled manipulation mechanism for the amplitude-phase of the bidirectional acoustic waves used in our concept offers multiple possibilities to apply various encryption methods. In this work, our system enables single-input, two-faced four-channel asymmetric encryption, which substantially increase the communication capacity of conventional acoustic holograms, and establish a security framework based on mathematical problem, proving its security. Our work can lead to concrete applications including, but not limited to, multi-channel acoustic field communications and acoustic illusion and cloaking in non-transparent media.

physics.app-ph

MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook

This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We hope it better allows researchers to follow the state-of-the-art in this very dynamic area. Meanwhile, a growing number of testbeds have boosted the evolution of general-purpose large language models. Thus, this year's MARS2 focuses on real-world and specialized scenarios to broaden the multimodal reasoning applications of MLLMs. Our organizing team released two tailored datasets Lens and AdsQA as test sets, which support general reasoning in 12 daily scenarios and domain-specific reasoning in advertisement videos, respectively. We evaluated 40+ baselines that include both generalist MLLMs and task-specific models, and opened up three competition tracks, i.e., Visual Grounding in Real-world Scenarios (VG-RS), Visual Question Answering with Spatial Awareness (VQA-SA), and Visual Reasoning in Creative Advertisement Videos (VR-Ads). Finally, 76 teams from the renowned academic and industrial institutions have registered and 40+ valid submissions (out of 1200+) have been included in our ranking lists. Our datasets, code sets (40+ baselines and 15+ participants' methods), and rankings are publicly available on the MARS2 workshop website and our GitHub organization page https://github.com/mars2workshop/, where our updates and announcements of upcoming events will be continuously provided.

cs.CV

Learning Progression-Guided AI Evaluation of Scientific Models To Support Diverse Multi-Modal Understanding in NGSS Classroom

Learning Progressions (LPs) can help adjust instruction to individual learners needs if the LPs reflect diverse ways of thinking about a construct being measured, and if the LP-aligned assessments meaningfully measure this diversity. The process of doing science is inherently multi-modal with scientists utilizing drawings, writing and other modalities to explain phenomena. Thus, fostering deep science understanding requires supporting students in using multiple modalities when explaining phenomena. We build on a validated NGSS-aligned multi-modal LP reflecting diverse ways of modeling and explaining electrostatic phenomena and associated assessments. We focus on students modeling, an essential practice for building a deep science understanding. Supporting culturally and linguistically diverse students in building modeling skills provides them with an alternative mode of communicating their understanding, essential for equitable science assessment. Machine learning (ML) has been used to score open-ended modeling tasks (e.g., drawings), and short text-based constructed scientific explanations, both of which are time-consuming to score. We use ML to evaluate LP-aligned scientific models and the accompanying short text-based explanations reflecting multi-modal understanding of electrical interactions in high school Physical Science. We show how LP guides the design of personalized ML-driven feedback grounded in the diversity of student thinking on both assessment modes.

cs.CY

Force sensing with a graphene nanomechanical resonator coupled to photonic crystal guided resonances

Achieving optimal force sensitivity with nanomechanical resonators requires the ability to resolve their thermal vibrations. In two-dimensional resonators, this can be done by measuring the energy they absorb while vibrating in an optical standing wave formed between a light source and a mirror. However, the responsivity of this method -- the change in optical energy per unit displacement of the resonator -- is modest, fundamentally limited by the physics of propagating plane waves. We present simulations showing that replacing the mirror with a photonic crystal supporting guided resonances increases the responsivity of graphene resonators by an order of magnitude. The steep optical energy gradients enable efficient transduction of flexural vibrations using low optical power, thereby reducing heating. Furthermore, the presence of two guided resonances at different wavelengths allows thermal vibrations to be resolved with a high signal-to-noise ratio across a wide range of membrane positions in free space. Our approach provides a simple optical method for implementing ultrasensitive force detection using a graphene nanomechanical resonator.

cond-mat.mes-hall

Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation

Short answer assessment is a vital component of science education, allowing evaluation of students' complex three-dimensional understanding. Large language models (LLMs) that possess human-like ability in linguistic tasks are increasingly popular in assisting human graders to reduce their workload. However, LLMs' limitations in domain knowledge restrict their understanding in task-specific requirements and hinder their ability to achieve satisfactory performance. Retrieval-augmented generation (RAG) emerges as a promising solution by enabling LLMs to access relevant domain-specific knowledge during assessment. In this work, we propose an adaptive RAG framework for automated grading that dynamically retrieves and incorporates domain-specific knowledge based on the question and student answer context. Our approach combines semantic search and curated educational sources to retrieve valuable reference materials. Experimental results in a science education dataset demonstrate that our system achieves an improvement in grading accuracy compared to baseline LLM approaches. The findings suggest that RAG-enhanced grading systems can serve as reliable support with efficient performance gains.

cs.CL