arXiv ScienceSearch

arXiv subjects

Yilin Zhu

Publications and source records attributed to Yilin Zhu.

At least 19 recordsLinked to original sources

Do LLMs Build Spatial World Models? Evidence from Grid-World Maze Tasks

Foundation models have shown remarkable performance across diverse tasks, yet their ability to construct internal spatial world models for reasoning and planning remains unclear. We systematically evaluate the spatial understanding of large language models through maze tasks, a controlled testing context requiring multi-step planning and spatial abstraction. Across comprehensive experiments with Gemini-2.5-Flash, GPT-5-mini, Claude-Haiku-4.5, and DeepSeek-Chat, we uncover significant discrepancies in spatial reasoning that challenge assumptions about LLM planning capabilities. Using chain-of-thought prompting, Gemini achieves 80-86% accuracy on smaller mazes (5x5 to 7x7 grids) with tokenized adjacency representations, but performance collapses to 16-34% with visual grid formats, which is a 2-5x difference, suggesting representation-dependent rather than format-invariant spatial reasoning. We further probe spatial understanding through sequential proximity questions and compositional distance comparisons. Despite achieving 96-99% semantic coverage in reasoning traces, models fail to leverage this understanding for consistent spatial computations, indicating that they treat each question independently rather than building cumulative spatial knowledge. Our findings based on the maze-solving tasks suggest that LLMs do not develop robust spatial world models, but rather exhibit representation-specific and prompting-dependent reasoning that succeeds only under narrow conditions. These results have critical implications for deploying foundation models in applications requiring spatial abstraction.

cs.AI

Charting Empirical Laws for LLM Fine-Tuning in Scientific Multi-Discipline Learning

While large language models (LLMs) have achieved strong performance through fine-tuning within individual scientific domains, their learning dynamics in multi-disciplinary contexts remains poorly understood, despite the promise of improved generalization and broader applicability through cross-domain knowledge synergy. In this work, we present the first systematic study of multi-disciplinary LLM fine-tuning, constructing a five-discipline corpus and analyzing learning patterns of full fine-tuning, LoRA, LoRA-MoE, and LoRA compositions. Particularly, our study shows that multi-disciplinary learning is substantially more variable than single-discipline training and distills four consistent empirical laws: (1) Balance-then-Diversity: low-resource disciplines degrade performance unless mitigated via diversity-aware upsampling; (2) Merge-then-Align: restoring instruction-following ability is critical for cross-discipline synergy; (3) Optimize-then-Scale: parameter scaling offers limited gains without prior design optimization; and (4) Share-then-Specialize: asymmetric LoRA-MoE yields robust gains with minimal trainable parameters via shared low-rank projection. Together, these laws form a practical recipe for principled multi-discipline fine-tuning and provide actionable guidance for developing generalizable scientific LLMs.

cs.LG

Optimization over Trained Neural Networks: Going Large with Gradient-Based Algorithms

When optimizing a nonlinear objective, one can employ a neural network as a surrogate for the nonlinear function. However, the resulting optimization model can be time-consuming to solve globally with exact methods. As a result, local search that exploits the neural-network structure has been employed to find good solutions within a reasonable time limit. For such methods, a lower per-iteration cost is advantageous when solving larger models. The contribution of this paper is two-fold. First, we propose a gradient-based algorithm with lower per-iteration cost than existing methods. Second, we further adapt this algorithm to exploit the piecewise-linear structure of neural networks that use Rectified Linear Units (ReLUs). In line with prior research, our methods become competitive with -- and then dominant over -- other local search methods as the optimization models become larger.

math.OC

An Extended Validity Domain for Constraint Learning

We consider embedding a predictive machine-learning model within a prescriptive optimization problem. In this setting, called constraint learning, we study the concept of a validity domain, i.e., a constraint added to the feasible set, which keeps the optimization close to the training data, thus helping to ensure that the computed optimal solution exhibits less prediction error. In particular, we propose a new validity domain which uses a standard convex-hull idea but in an extended space. We investigate its properties and compare it empirically with existing validity domains on a set of test problems for which the ground truth is known. Results show that our extended convex hull routinely outperforms existing validity domains, especially in terms of the function value error, that is, it exhibits closer agreement between the true function value and the predicted function value at the computed optimal solution. We also consider our approach within two stylized optimization models, which show that our method reduces feasibility error, as well as a real-world pricing case study.

math.OC

Multiscale Physics-Informed Neural Networks for the Inverse Design of Hyperuniform Optical Materials

In this article, we employ multiscale physics-informed neural networks (MscalePINNs) for the inverse design of finite-size photonic materials with stealthy hyperuniform (SHU) disordered geometries. Specifically, we show that MscalePINNs can capture the fast spatial variations of complex fields scattered by arrays of dielectric nanocylinders arranged according to isotropic SHU point patterns, thus enabling a systematic methodology to inversely retrieve their effective dielectric profiles. Our approach extends the recently developed high-frequency homogenization theory of hyperuniform media and retrieves more general permittivity profiles for applications-relevant finite-size SHU systems, unveiling unique features related to their isotropic nature. In particular, we numerically corroborate the existence of a transparency region beyond the long-wavelength approximation, enabling effective and isotropic homogenization even without disorder-averaging, in contrast to the case of uncorrelated Poisson random patterns. The flexible multiscale network approach introduced here enables the efficient inverse design of more general effective media and finite-size optical metamaterials with isotropic electromagnetic responses beyond the limitations of traditional homogenization theories.

physics.optics

Democratizing the Creation of Animatable Facial Avatars

In high-end visual effects pipelines, a customized (and expensive) light stage system is (typically) used to scan an actor in order to acquire both geometry and texture for various expressions. Aiming towards democratization, we propose a novel pipeline for obtaining geometry and texture as well as enough expression information to build a customized person-specific animation rig without using a light stage or any other high-end hardware (or manual cleanup). A key novel idea consists of warping real-world images to align with the geometry of a template avatar and subsequently projecting the warped image into the template avatar's texture; importantly, this allows us to leverage baked-in real-world lighting/texture information in order to create surrogate facial features (and bridge the domain gap) for the sake of geometry reconstruction. Not only can our method be used to obtain a neutral expression geometry and de-lit texture, but it can also be used to improve avatars after they have been imported into an animation system (noting that such imports tend to be lossy, while also hallucinating various features). Since a default animation rig will contain template expressions that do not correctly correspond to those of a particular individual, we use a Simon Says approach to capture various expressions and build a person-specific animation rig (that moves like they do). Our aforementioned warping/projection method has high enough efficacy to reconstruct geometry corresponding to each expressions.

cs.GR

Localization landscape of optical waves in multifractal photonic membranes

In this paper, we investigate the localization properties of optical waves in disordered systems with multifractal scattering potentials. In particular, we apply the localization landscape theory to the classical Helmholtz operator and, without solving the associated eigenproblem, show accurate predictions of localized eigenmodes for one- and two-dimensional multifractal structures. Finally, we design and fabricate nanoperforated photonic membranes in silicon nitride (SiN) and image directly their multifractal modes using leaky-mode spectroscopy in the visible spectral range. The measured data demonstrate optical resonances with multiscale intensity fluctuations in good qualitative agreement with numerical simulations. The proposed approach provides a convenient strategy to design multifractal photonic membranes, enabling rapid exploration of extended scattering structures with tailored disorder for enhanced light-matter interactions.

physics.optics

Revitalizing Sex Education for Chinese Children: A Formative Study

Sex education helps children obtain knowledge and awareness of sexuality, and protects them against sexually transmitted diseases, pregnancy, and sexual abuse. Sex education is not well taught to children in China -- both school-based education and parental communication on this topic are limited. To interrogate the status quo of sex education in China and explore suitable interventions, we conducted a series of formative studies including interviews and social media analysis. Multiple stakeholders such as children, parents, education practitioners, and the general public were engaged for an in-depth understanding of their unique needs regarding teaching and learning sex education. We found that school-based sex education for Chinese children was currently insufficient and restrictive. Involving parents in sex education posed several challenges, such as a lack of sexuality and pedagogy knowledge, and embarrassment in initiating sex education conversations. Culture and politics were major hurdles to effective sex education. Based on the findings, we reflect on the complex interactions between culture, politics, education policy, and pedagogy, and discuss situated design of sex education in broader cultural and social contexts.

cs.CY

High-throughput speckle spectrometers based on multifractal scattering media

We present compact integrated speckle spectrometers based on monofractal and multifractal scattering media in a silicon-on-insulator platform. Through both numerical and experimental studies we demonstrate enhanced optical throughput, and hence signal-to-noise ratio, for a number of random structures with tailored multifractal geometries without affecting the spectral decay of the speckle correlation functions. Moreover, we show that the developed multifractal media outperform traditional scattering spectrometers based on uniform random distributions of scattering centers. Our findings establish the potential of low-density random media with multifractal correlations for integrated on-chip applications beyond what is possible with uncorrelated random disorder.

physics.optics

Inverse design of functional photonic patches by adjoint optimization coupled to the generalized Mie theory

We propose a rigorous approach for the inverse design of functional photonic structures by coupling the adjoint optimization method and the two-dimensional generalized Mie theory (2D-GMT) for the multiple scattering problem of finite-size arrays of dielectric nanocylinders optimized to display desired functions. We refer to these functional scattering structures as "photonic patches". We briefly introduce the formalism of 2D-GMT and the critical steps necessary to implement the adjoint optimization algorithm to photonic patches with designed radiation properties. In particular, we showcase several examples of periodic and aperiodic photonic patches with optimal nanocylinder radii and arrangements for radiation shaping, wavefront focusing in the Fresnel zone, and for the enhancement of the local density of states (LDOS) at multiple wavelengths over micron-size areas. Moreover, we systematically compare the performances of periodic and aperiodic patches with different sizes and find that optimized aperiodic Vogel spiral geometries feature significant advantages in achromatic focusing compared to their periodic counterparts. Our results show that adjoint optimization coupled to 2D-GMT is a robust methodology for the inverse design of compact photonic devices that operate in the multiple scattering regime with optimal desired functionalities. Without the need of spatial meshing, our approach provides efficient solutions at strongly reduced computational burden compared to standard numerical optimization techniques and suggests compact device geometries for on-chip photonics and metamaterials technologies.

physics.optics

Design of ultracompact broadband focusing spectrometers based on deep diffractive neural networks

We propose the inverse design of ultracompact, broadband focusing spectrometers based on adaptive deep diffractive neural networks (a-D$^2$NNs). Specifically, we introduce and characterize two-layer diffractive devices with engineered angular dispersion that focus and steer broadband incident radiation along predefined focal trajectories with desired bandwidth and $5$ nm spectral resolution. Moreover, we systematically study the focusing efficiency of two-layer devices with side length $L=100~\mu\mathrm{m}$ and focal length $f=300~\,\mu\mathrm{m}$ across the visible spectrum and we demonstrate accurate reconstruction of the emission spectrum from a commercial superluminescent diode. The proposed a-D$^2$NNs design method extends the capabilities of efficient multi-focal diffractive optical devices to include single-shot focusing spectrometers with customized focal trajectories for applications to ultracompact multispectral imaging and lensless microscopy.

physics.optics

Phage family classification under Caudoviricetes: a review of current tools using the latest ICTV classification framework

Bacteriophages, which are viruses infecting bacteria, are the most ubiquitous and diverse entities in the biosphere. There is accumulating evidence revealing their important roles in shaping the structure of various microbiomes. Thanks to (viral) metagenomic sequencing, a large number of new bacteriophages have been discovered. However, lacking a standard and automatic virus classification pipeline, the taxonomic characterization of new viruses seriously lag behind the sequencing efforts. In particular, according to the latest version of ICTV, several large phage families in the previous classification system are removed. Therefore, a comprehensive review and comparison of taxonomic classification tools under the new standard are needed to establish the state-of-the-art. In this work, we retrained and tested four recently published tools on newly labeled databases. We demonstrated their utilities and tested them on multiple datasets, including the RefSeq, short contigs, simulated metagenomic datasets, and low-similarity datasets. This study provides a comprehensive review of phage family classification in different scenarios and a practical guidance for choosing appropriate taxonomic classification pipelines. To our best knowledge, this is the first review conducted under the new ICTV classification framework. The results show that the new family classification framework overall leads to better-conserved groups and thus makes family-level classification more feasible.

q-bio.GN

Wave localization in number-theoretic landscapes

We investigate the localization of waves in aperiodic structures that manifest the characteristic multiscale complexity of certain arithmetic functions with a central role in number theory. In particular, we study the eigenspectra and wave localization properties of tight-binding Schr\"{o}dinger equation models with on-site potentials distributed according to the Liouville function $\lambda(n)$, the M\"{o}bius function $\mu(n)$, and the Legendre sequence of quadratic residues modulo a prime (QRs). We employ Multifractal Detrended Fluctuation Analysis (MDFA) and establish the multifractal scaling properties of the energy spectra in these systems. Moreover, by systematically analyzing the spatial eigenmodes and their level spacing distributions, we show the absence of level repulsion with broadband localization across the entire energy spectra. Our study introduces deterministic aperiodic systems whose eigenmodes are all strongly localized in realistic finite one-dimensional systems and provides opportunities for novel quantum and classical devices of particular importance to cold-atom experiments in engineered speckle potentials and enhanced light-matter interactions.

cond-mat.dis-nn

Enhanced wave localization in multifractal scattering media

In this paper we study the structural, scattering, and wave localization properties of multifractal arrays of electric point dipoles generated from multiplicative random fields with different degrees of multiscale correlations. Specifically, using the rigorous Green's matrix method, we investigate the scattering resonances and wave localization behavior of systems with $N=10^{4}$ dipoles and demonstrate an enhanced localization behavior in highly inhomogeneous multifractal structures compared to homogeneous fractals, or monofractals. We show distinctive spectral properties, such as the absence of level repulsion in the strong multiple scattering regime and power-law statistics of level spacings, which indicate a clear localization transition enhanced in non-homogeneous multifractals. Our findings unveil the importance of multifractal structural correlations in the multiple scattering regime of electric dipole arrays and provide an efficient model for the design of multiscale nanophotonic systems with enhanced light-matter coupling and localization phenomena beyond what is possible with traditional fractal systems.

physics.optics

Leveraging Deepfakes to Close the Domain Gap between Real and Synthetic Images in Facial Capture Pipelines

We propose an end-to-end pipeline for both building and tracking 3D facial models from personalized in-the-wild (cellphone, webcam, youtube clips, etc.) video data. First, we present a method for automatic data curation and retrieval based on a hierarchical clustering framework typical of collision detection algorithms in traditional computer graphics pipelines. Subsequently, we utilize synthetic turntables and leverage deepfake technology in order to build a synthetic multi-view stereo pipeline for appearance capture that is robust to imperfect synthetic geometry and image misalignment. The resulting model is fit with an animation rig, which is then used to track facial performances. Notably, our novel use of deepfake technology enables us to perform robust tracking of in-the-wild data using differentiable renderers despite a significant synthetic-to-real domain gap. Finally, we outline how we train a motion capture regressor, leveraging the aforementioned techniques to avoid the need for real-world ground truth data and/or a high-end calibrated camera capture setup.

cs.CV

Inverse design of ultracompact multi-focal optical devices by diffractive neural networks

We propose an efficient inverse design approach for multifunctional optical elements based on adaptive deep diffractive neural networks (a-D$^2$NNs). Specifically, we introduce a-D$^2$NNs and design two-layer diffractive devices that can selectively focus incident radiation over two well-separated spectral bands at desired distances. We investigate focusing efficiencies at two wavelengths and achieve targeted spectral lineshapes and spatial point-spread functions (PSFs) with optimal focusing efficiency. In particular, we demonstrate control of the spectral bandwidths at separate focal positions beyond the theoretical limit of single-lens devices with the same aperture size. Finally, we demonstrate devices that produce super-oscillatory focal spots at desired wavelengths. The proposed method is compatible with current diffractive optics and doublet metasurface technology for ultracompact multispectral imaging and lensless microscopy applications.

physics.optics

Learning Topological Motion Primitives for Knot Planning

In this paper, we approach the challenging problem of motion planning for knot tying. We propose a hierarchical approach in which the top layer produces a topological plan and the bottom layer translates this plan into continuous robot motion. The top layer decomposes a knotting task into sequences of abstract topological actions based on knot theory. The bottom layer translates each of these abstract actions into robot motion trajectories through learned topological motion primitives. To adapt each topological action to the specific rope geometry, the motion primitives take the observed rope configuration as input. We train the motion primitives by imitating human demonstrations and reinforcement learning in simulation. To generalize human demonstrations of simple knots into more complex knots, we observe similarities in the motion strategies of different topological actions and design the neural network structure to exploit such similarities. We demonstrate that our learned motion primitives can be used to efficiently generate motion plans for tying the overhand knot. The motion plan can then be executed on a real robot using visual tracking and Model Predictive Control. We also demonstrate that our learned motion primitives can be composed to tie a more complex pentagram-like knot despite being only trained on human demonstrations of simpler knots.

cs.RO

Self-Supervised Learning of State Estimation for Manipulating Deformable Linear Objects

We demonstrate model-based, visual robot manipulation of linear deformable objects. Our approach is based on a state-space representation of the physical system that the robot aims to control. This choice has multiple advantages, including the ease of incorporating physics priors in the dynamics model and perception model, and the ease of planning manipulation actions. In addition, physical states can naturally represent object instances of different appearances. Therefore, dynamics in the state space can be learned in one setting and directly used in other visually different settings. This is in contrast to dynamics learned in pixel space or latent space, where generalization to visual differences are not guaranteed. Challenges in taking the state-space approach are the estimation of the high-dimensional state of a deformable object from raw images, where annotations are very expensive on real data, and finding a dynamics model that is both accurate, generalizable, and efficient to compute. We are the first to demonstrate self-supervised training of rope state estimation on real images, without requiring expensive annotations. This is achieved by our novel self-supervising learning objective, which is generalizable across a wide range of visual appearances. With estimated rope states, we train a fast and differentiable neural network dynamics model that encodes the physics of mass-spring systems. Our method has a higher accuracy in predicting future states compared to models that do not involve explicit state estimation and do not use any physics prior, while only using 3\% of training data. We also show that our approach achieves more efficient manipulation, both in simulation and on a real robot, when used within a model predictive controller.

cs.RO