arXiv ScienceSearch

SEARCH · arXiv Science

Results for “physics.optics”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

780 recordsLinked to original sources

Physical policy gradient theorem for in situ stochastic-adjoint training

In situ adjoint training extracts parameter gradients directly from measurement, but has so far been limited to reciprocal or restricted systems. Here, we introduce the physical counterpart of the policy gradient theorem: a stochastic-adjoint gradient estimator that lifts these constraints by trading reciprocity for nondegenerate diffusion. As validation, we train a nonlinear resonator network, whose own dynamics supply the policy, against antagonistic temporal modulations with gradients from measured stochastic trajectories alone, without finite differences or a separate adjoint experiment.

physics.optics

Towards a universal meta-optics solver via large language models

Metasurface design increasingly requires fast models that can operate across structurally distinct device families, rather than retraining a separate surrogate for every geometry class. Conventional neural network surrogates often depend on fixed-dimensional descriptors, family-specific output formats, and repeated architecture tuning, which limits their scalability across heterogeneous meta-atoms. Here, we present a unified large language model (LLM) workflow for multi-family metasurface modeling and inverse-design. Geometries, design parameters, and optical response channels were converted into a shared instruction-following text format and used to fine-tune Gemma-2-9B across 8 metasurface families. Compared with single-family baselines, the joint model simultaneously predicted the optical responses of all metasurface families while reducing the MSE for each family by an average of 56.5%. The same representation was also used for inverse design. These results show that a shared sequence-based LLM interface can provide a practical route to cross-family metasurface design while reducing the need for task-specific surrogate architectures.

physics.optics

Bubble2Heat: Optical to Thermal Inference in Pool Boiling Using Physics-encoded Generative AI

Phase change process plays a critical role in thermal management systems, yet quantitative characterization of multiphase heat transfer remains limited by the challenges of measuring temperature fields in chaotic, rapidly evolving flow regimes. While computational methods offer temperature data at a high spatiotemporal resolution in ideal cases, replicating complex experimental conditions remains prohibitively difficult. In this paper, we present a deep learning framework that can generate temperature field data at simulation resolution from segmented high-speed recordings and pointwise thermocouple readings which are typically available in a canonical pool boiling experimental configuration without requiring advanced techniques. This framework leverages a conditional generative adversarial network trained only on simulation data. To ensure direct applicability of the model to experimental data, our framework also introduces a preprocessing pipeline that aligns high resolution simulation data with experimental measurements through both conventional image processing and image segmentation with pretrained convolutional neural network. We further show that standard data augmentation strategies are effective in enhancing the physical plausibility of the inference when precise physical constraints are not applicable. Our results highlight the potential of deep generative models to bridge the gap between observable multiphase phenomena and underlying thermal transport, offering a powerful approach to augment and interpret experimental measurements in complex two-phase systems.

cs.LG

Design and Physical Constraints of Synthetic-Frequency Photonic Switching Fabrics

Electro-optic frequency conversion and synthetic-frequency coupling are established functions in integrated photonic devices. Their role within a multiport switching fabric, however, depends on how simultaneous optical connections share spatial paths, frequency channels, and device controls. Here, we investigate how coherent coupling among frequency modes can be incorporated into photonic switching fabrics and identify the corresponding architectural and physical constraints. We show that synthetic-frequency coupling does not increase the number of simultaneous orthogonal frequency channels when all channels are freely accessible, but can establish connections that are otherwise blocked by fixed input frequencies, channel-continuity requirements, or unavailable output channels. Under the tested conditions, coupling over the first three frequency spacings in an $8\times8$ fabric with eight frequency channels per port achieves 96.1% of the blocking reduction obtained with unrestricted inter-mode coupling. We further show that a separate frequency-only conversion stage cannot replace missing spatial connectivity. A nominal reduction in spatial switching elements instead requires a joint element whose spatial state can be programmed independently for each frequency channel. Finally, we evaluate a thin-film lithium niobate resonator model using reported electro-optic coupling and photon-decay scales within a multistage Mach-Zehnder interferometer switching fabric. These results clarify the architectural role of synthetic-frequency coupling and the device-level requirements for incorporating it into integrated photonic switching fabrics.

physics.optics

Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields

Geostationary atmospheric motion vectors (AMVs) provide the dense horizontal wind vectors (u,v) and heights ingested into data assimilation systems. Traditional AMVs track features using window-based cross-correlation and estimate heights via infrared brightness temperatures paired with numerical weather prediction (NWP) background states, creating a circular dependency that yields inaccurate heights, high computational cost, and sparse retrievals. Stereo winds from GEO-GEO and GEO-LEO geometrically resolve heights from parallax shifts across different poses, eliminating NWP dependence and improving accuracy, but they remain computationally heavy with limited coverage. In this work, we replace window-based tracking in stereo matching with deep optical flow for efficient, improved retrieval. Fine-tuning balances a self-supervised geometric residual loss with supervised radiosonde reconstruction. To eliminate multi-satellite overlap requirements, we distill the stereo teacher into a single-satellite student model. Chi-square and height uncertainties from the teacher are emulated by the student for quality assurance. The student generates winds across full-disk GEO imagery globally. Validation compares stereo and student models against radiosondes, operational AMVs, ERA5 reanalysis, and EarthCARE cloud profiles. Results through triple collocation show that stereo winds improve performance beyond operational AMVs for water vapor bands (6.2, 6.9, and 7.3 μm), wit degradation in the long-wave infrared (11.2 μm) band.

cs.LG

Collaborative On-Sensor Array Cameras

Modern nanofabrication techniques have enabled us to manipulate the wavefront of light with sub-wavelength-scale structures, offering the potential to replace bulky refractive surfaces in conventional optics with ultrathin metasurfaces. In theory, arrays of nanoposts provide unprecedented control over manipulating the wavefront in terms of phase, polarization, and amplitude at the nanometer resolution. A line of recent work successfully investigates flat computational cameras that replace compound lenses with a single metalens or an array of metasurfaces a few millimeters from the sensor. However, due to the inherent wavelength dependence of metalenses, in practice, these cameras do not match their refractive counterparts in image quality for broadband imaging, and may even suffer from hallucinations when relying on generative reconstruction methods. In this work, we investigate a collaborative array of metasurface elements that are jointly learned to perform broadband imaging. To this end, we learn a nanophotonics array with 100-million nanoposts that is end-to-end jointly optimized over the full visible spectrum--a design task that existing inverse design methods or learning approaches cannot support due to memory and compute limitations. We introduce a distributed meta-optics learning method to tackle this challenge. This allows us to optimize a large parameter array along with a learned meta-atom proxy and a non-generative reconstruction method that is parallax-aware and noise-aware. The proposed camera performs favorably in simulation and in all experimental tests irrespective of the scene illumination spectrum.

physics.optics

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven: long-tailed or unbalanced distributions hinder generalization, and the low number of samples in low density regions makes it hard to run evaluations. Synthetic data can fill these gaps, providing us with a way to sample the input space more continuously and improve data coverage for benchmarks. Focusing on the autonomous driving safety-critical case of pedestrian detection in the dark, we show how synthetic low-light samples can be used to better characterize the performance of a state-of-the-art object detection model as a function of the scene illumination. We use a synthetic RAW image augmentation technique to generate low-light samples that match the noise model of the camera sensor. Performance metrics on real and synthetic low-light data are similar, indicating that the AI model finds it hard to distinguish between them.

cs.CV

Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization

We present a direct optimization method for three-dimensional freeform reflectors that transform the light of a finite-étendue source into a prescribed far-field angular intensity distribution. The reflector profile is represented by a small neural network (a multilayer perceptron), which is trained end-to-end through a differentiable ray-tracing objective. We furthermore parameterize the emission directions in gnomonic coordinates, and show how we use this to ensure that every emitted ray intersects the reflector. At each iteration, the network is converted to a bicubic spline representation for ray-tracing efficiency, and intersections with this smooth surface are solved by a damped Newton solve, with gradients computed via the implicit function theorem. The traced output distribution is compared with the desired target on a 'soft' histogram, under an $H^{-1}$-type spectral weighting that emphasizes long-range transport of flux to improve convergence. Optimization is performed using a BFGS method with self-scaled Broyden updates and a plateau-perturbation rule to prevent stalling. The method converges reliably within seconds on a single GPU for all examples tested.

physics.optics

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

Metasurfaces have revolutionized the development of photonic devices by enabling unprecedented precision in light manipulation. However, their design processes are often constrained by computationally expensive simulations and complex high-dimensional design spaces. Although deep learning has accelerated the design process by serving as a surrogate model, it remains constrained by task-specific architectures and lacks universal reasoning capabilities. This review surveys how Large Language Models (LLMs) are adding semantic interfaces, code generation, and tool orchestration to established numerical nanophotonic workflows. We first outline the development from classical neural networks to transformer-based models and their applications in nanophotonic design. We then review the emergence of LLM-related methods in nanophotonics and organize them into two operational modes: surrogate models that treat structure-spectrum mapping as a language task, and agentic systems that have been demonstrated to generate code, orchestrate selected simulation steps, and support closed-loop optimization. Furthermore, to identify future cross-disciplinary opportunities, we briefly explore applications of LLMs in research fields such as materials science and wireless communications. This review concludes by looking ahead to the next generation of multimodal foundation models with physical perception capabilities. In this vision, artificial intelligence is evolving from passive tools into active collaborators, participating in autonomous scientific discovery.

physics.optics

Geometric Optics Approximation Sampling: A Reflector-Induced Transport Map Framework

In this paper, we propose Geometric Optics Approximation Sampling (GOAS), a reflector-induced transport-map framework for sampling from target measures. Once a reflecting surface is constructed, the associated transport map is explicitly determined by the physical law of reflection. As a concrete realization, we develop a supporting-hyperellipsoid construction that requires only a discrete approximation of the target measure and does not require gradient information of the target density. The formulation accommodates both density-based and sample-based target representations. A softmin smoothing technique is introduced to obtain a smooth approximate transport map from this piecewise hyperellipsoidal construction. We establish well-posedness and stability of the reflector-induced push-forward measure and derive quantitative error estimates in the maximum mean discrepancy metric, and convergence of continuous statistical observables, including fixed-order moments. Numerical experiments on an analytically tractable example, strongly non-Gaussian targets, sample-based target approximations, and Bayesian inverse problems demonstrate the accuracy and flexibility of GOAS.

math.NA

Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization

Training quantized neural networks remains fundamentally challenging due to non-convex loss landscapes and discrete parameter spaces. We introduce an exact Quadratic Constrained Binary Optimization (QCBO) framework with provable guarantees. We first characterize the stratified topology of network zero-loss level sets: generic interior strata are smooth, yet globally optimal components can remain disconnected even under overparameterization. To address this non-convex obstruction, we compile finite-depth architectures with parameter codebooks and Forward Interval Propagation (FIP)-bounded states into bounded QCBOs, yielding an exact completely positive convex formulation that preserves the global discrete optimum with zero relaxation gap. To overcome monolithic sample scaling, we formulate sample-wise Decomposed Lower-Bound Optimization (DLBO) to reduce each Ising call from dataset to single-sample scale. The DLBO moment hierarchy also forms a Hamiltonian-locality hierarchy, with order two giving an auxiliary-free pairwise QUBO oracle and higher orders trading interaction locality for tighter bounds. Strictly feasible discrete parameters are recovered via Spectral--ADMM and randomized rounding. Experiments on a coherent Ising machine achieve $94.95\%$ accuracy on binary Fashion-MNIST (coats vs. sandals) at 1.1-bit precision, demonstrating resilience against low-bit representational collapse. Multi-class DLBO evaluations on 3-class Fashion-MNIST, 3-class Wine, and 3-class Digits further validate scalable convergence.

cs.LG

LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering

The visual aesthetics of photographs are deeply influenced by lens characteristics such as aperture shape, optical vignetting and optical diffraction, which together define a camera's unique optical style. Existing lens effect rendering methods primarily focus on accurately simulating the blur transition from small to large apertures but overlook the stylistic aspects of lens effects. As a result, they fail to produce diverse bokeh effects under large apertures or capture distinctive photographic phenomena such as starbursts that emerge under small apertures. In this work, we introduce LensStyle, a unified framework for controllable stylized lens effect rendering that explicitly models lens aesthetics through joint continuous-discrete control. Our model incorporates a Dual-Path Controller that disentangles continuous optical parameter modulation (e.g., focus distance and blur strength) from discrete lens-style conditioning (e.g., circular, polygonal, donut, cat-eye, and starburst effects), enabling fine-grained, interpretable, and physically grounded lens manipulation within a single unified framework. To support model training, we curate a comprehensive MultiLens dataset containing multi-lens image pairs synthesized under real optical constraints. Extensive experiments demonstrate that LensStyle achieves superior realism, controllability, and aesthetic quality compared with existing lens effect rendering approaches and diffusion-based image editing models, advancing computational photography toward multiple-lens-style simulation.

cs.CV

Phase-field digital image correlation for integrated displacement and damage measurements

This work presents a novel digital image correlation (DIC) framework for full-field measurements of displacement, strain, and damage, based on a phase field (PF) approach. The idea is to take advantage of the ability of the PF method to track complex crack morphologies and to provide a natural way in DIC to perform damage and crack measurements from experimental speckle images, in addition to displacement and strain fields. Moreover, incorporating the damage variable into DIC can improve the displacement accuracy near the crack tip, and can avoid the need of user-defined masks when dealing with cracked samples, which is advantageous when cracks become complex and the manual application of masks becomes challenging. The theoretical formulation of the proposed framework, namely PF-DIC, was presented in detail in the paper, along with a finite element implementation. Numerical examples have demonstrated the capability of the proposed PF-DIC in terms of capturing different types of cracks while providing similar measurement accuracy to that of masked DIC. Additionally, it is shown that the PF-DIC can be easily adapted to selectively identify critical cracks under specific loading conditions or mechanisms for damage assessment and diagnostic purposes. The proposed DIC framework can be used to characterize material defects, support structural health monitoring, and enable a potential unification of PF simulations and experimental fracture measurements

math.NA

SpectraTac: A Compact Camera-Free Optical Tactile Sensor with Distributed Color Sensing

Tactile sensing is essential for physical interaction in robotics and human--machine systems. However, combining rich tactile information with compact hardware, low cost, and low computational overhead remains challenging. This work presents SpectraTac, a compact, camera-free optical tactile sensor that combines active red--green--blue (RGB) illumination with spatially distributed color sensing. Contact deforms a compliant transparent elastomer and modulates its internal light field, producing spatially differentiated changes in color and intensity. Three distributed color sensors capture these responses as low-dimensional spatio-spectral features, avoiding cameras, imaging optics, and high-dimensional image processing. The device measures 19.2 mm in diameter and 4 mm in height, with a material cost below USD~5. A data-driven decoding framework simultaneously estimates three-dimensional (3D) force and the contact region from the optical measurements. For 3D force prediction, the sensor achieved mean absolute errors (MAEs) of 0.161, 0.164, and 0.429 N along the x-, y-, and z-axes, respectively. The nine-region contact-classification accuracy was 99.9%. We further evaluated real-time 3D force tracking and contact-region-based human--machine interaction through an interactive control task. These results indicate that distributed color-resolved optical sensing offers a compact, low-cost alternative to camera-based tactile sensing for robotics, wearable sensing, and interactive systems.

cs.RO

Geometric Optics Approximation Sampling: A Far-Field Reflector-Induced Transport Framework

We develop a far-field geometric optics approximation sampling (GOAS) framework for constructing direct samplers from target measures. The method exploits the connection between the far-field reflector problem and optimal transport with logarithmic cost, leading to a natural primal--dual transport structure. The associated dual reflector provides a reciprocal backward transport and, in the invertible smooth setting, the inverse of the forward reflector map. For numerical realization, we adopt a supporting hyperparaboloid construction based on a discrete approximation of the target measure. This construction is gradient-free with respect to the target density and naturally accommodates both density-based and sample-based target representations. The resulting piecewise reflector admits two sampling realizations: a primal--dual consistency resampling strategy that operates directly on the nonsmooth reflector, and a softmin-regularized realization yielding an explicit smooth transport map through the physical law of reflection. We establish the well-posedness and stability of the reflector-induced sampling measure and derive Wasserstein error estimates. Numerical experiments on non-Gaussian targets and Bayesian inverse problems demonstrate the accuracy, stability, Wasserstein error behavior, and applicability of the proposed framework, and compare it with MCMC and polynomial transport-map methods.

math.NA

OptiXDE: A fast optical-inspired solver for differential equations

OptiXDE is a matrix-free spectral operator framework for differential equations on uniform grids and embedded domains. Inspired by angular-spectrum propagation in Fourier optics, it maps transform-diagonal spatial operators to analytical modal multipliers and composes them with physical-space operators for nonlinearities, geometry and boundary enforcement. A common transform--operator--inverse-transform backbone is demonstrated across transient diffusion, periodic and embedded-domain Poisson problems, the cubic nonlinear Schr"odinger equation, viscous Burgers dynamics, the two-dimensional Allen--Cahn equation and incompressible flows from the Taylor--Green vortex to embedded-cylinder vortex shedding. Transform-compatible linear problems are recovered near the floating-point limit, whereas errors on the singular L-shaped domain remain localized near the re-entrant corner and regularized interface. Nonlinear benchmarks recover second-order temporal convergence and the expected conservative or dissipative behavior, while incompressibility remains near round-off level during long-time vortex shedding. The matrix-free updates require \(\mathcal{O}(N\log N)\) work and \(\mathcal{O}(N)\) memory. Device-resident transform workloads reach \(94.9\times\) GPU acceleration, and the complete embedded-cylinder solver achieves a \(42.1\times\) CPU--GPU speedup under matched numerical settings. These results establish OptiXDE as a deterministic and extensible operator-centric framework for structured and embedded-domain differential equations.

math.NA

Reliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space Restoration

Indirect time-of-flight (iToF) cameras provide compact and cost-effective dense depth measurements, but their ranging accuracy is often degraded by sensor-intrinsic uncertainty under practical imaging conditions. Spatially uniform or range-only Gaussian perturbations cannot accurately reproduce the range-dependent and signal-dependent noise characteristics of real iToF measurements, leading to a synthetic-to-real gap for learning-based restoration. To address this problem, we propose a joint depth-uncertainty modeling and restoration framework for reliable iToF sensing. A sensor-intrinsic depth-uncertainty model is first developed from calibrated tap responses, returned-signal levels, and sensor noise statistics through a depth-oriented weighted least-squares formulation. The resulting pixel-wise uncertainty is used for heteroscedastic depth synthesis and uncertainty-aware restoration supervision. Based on this heteroscedastic data synthesis, we further develop a U-shaped restoration network with Depth Visual State Space (DVSS) blocks, which combine long-range state-space modeling with convolutional spatial-channel refinement for structure-preserving depth recovery. Experiments on synthetic data and measurements captured by an in-house iToF prototype validate the proposed uncertainty model under varying range and returned-signal conditions. Controlled comparisons with fixed and range-aware Gaussian noise, together with evaluations on U-Net, Restormer, and DVSS, further demonstrate that the proposed synthesis consistently benefits different restoration backbones. The complete framework achieves 40.85~dB PSNR and 2.54 mm MAE on the synthetic test set, and 35.42 dB PSNR and 4.87 mm MAE on real iToF measurements.

physics.optics