arXiv ScienceSearch

arXiv subjects

Lei Tian

Publications and source records attributed to Lei Tian.

At least 19 recordsLinked to original sources

Reassessing 3GPP NR CSI Codebook Structures in Near-Field Channels: Finite-Feedback Multilayer Precoding and Design Insights

3GPP TR 38.901 Rel-19 introduces antenna-element-level spherical-wave modeling, while NR Type-I and enhanced Type-II (eType-II) CSI codebooks continue to use plane-wave DFT beams. Whether this mismatch materially degrades finite-feedback multilayer precoding in standardized multipath channels remains unclear. To isolate its impact, we evaluate both codebooks over strictly paired far-field (FF) and near-field (NF) 3GPP channels that share user locations, multipath parameters, polarization, and link budgets and differ only in their wavefront models. Simulations cover Rank 1-4 transmission in 7-GHz UMi and 24-GHz InH-linear scenarios. We find no systematic FF/NF shift in singular-mode gains or equal-power SVD (SVD-EP) rates. Instead, spherical-wave phases reorder multipath projections onto plane-wave candidates and thereby alter codeword selection. For Rank-4 InH-linear users at 0.1 normalized Rayleigh distance, given FF-selected Type-I and eType-II codewords incur median direct mismatch losses of 2.28% and 5.34% on the NF channel, respectively; codebook reselection identifies better-matched codewords and reduces these losses to 0.770% and 3.28%. Nested candidate-set comparisons further show that relaxing Type-I interlayer constraints improves the SVD-EP-normalized rate by 20.7 percentage points, whereas finite-range sampling adds only 0.579 points. These results support prioritizing multilayer multibeam representation in large-aperture NR CSI codebooks, with range states providing complementary refinement.

eess.SP

Site-specific Channel Modeling Based on Remote-Sensing Maps for 6G Space--Air--Ground Digital Twins

Site-specific channel models are essential for wireless digital twins of 6G space--air--ground communication systems. However, 3D maps are difficult to obtain over wide areas, which limits large-area site-specific channel modeling. To address this issue, this paper proposes a remote-sensing-based augmented ray-tracing channel modeling framework. The framework comprises a deterministic RT branch, a measurement-statistical branch, and an RT augmentation branch. To overcome the difficulty of acquiring large-area 3D maps, the deterministic RT branch reconstructs a 3D RT scene from satellite remote-sensing imagery and calibrates its electromagnetic material parameters using measured path loss. To provide the statistical parameters required for RT augmentation, the measurement-statistical branch establishes the marginal distributions and interparameter dependence models of the channel parameters. Specifically, a wideband UAV channel measurement campaign is conducted at 4.60 GHz, and a proposed multipath estimation method estimates the complex amplitudes, delays, and Doppler shifts of the measured multipath. To bridge the gap between RT predictions and measurements, the RT augmentation branch organizes the RT multipath into LoS, LoS-tail, and NLoS components, generates additional short-delay LoS-tail paths, and reallocates the component and path powers according to the measurement-derived statistics while preserving the total RT received power. The validation results show that the proposed framework reduces the path loss RMSE from 5.45 to 4.35 dB and, relative to calibrated RT, decreases the RMS delay spread and normalized Doppler spread RMSEs by 53.03 and 26.48, respectively. The proposed framework provides a site-specific channel modeling approach for 6G space--air--ground digital-twin studies.

eess.SP

Point Spread Function Engineering Using Implicit Neural Representations

Point spread function (PSF) engineering through pupil plane modulation is a technique used in microscopy to achieve specific imaging properties, such as depth encoding or extended depth of field. Existing PSF design methods often rely on extensive domain knowledge and task-specific basis functions, making it difficult to generalize across different applications. We treat the PSF engineering task as a phase retrieval problem and propose a neural field pupil design method that optimizes a phase profile for any arbitrary, user-defined 3D PSF distribution. This provides a flexible framework for 3D PSF engineering for various applications with implicit regularization that proves robust to initialization compared to pixel-wise optimization methods

physics.optics

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination-resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision-making systems under partial observability, together with a three-level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self-evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self-improving agents, agent gyms, and test-time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this survey provides a roadmap toward trustworthy, self-improving medical imaging systems for real clinical practice.

cs.AI

WEKP-PLP: A Wireless Environment Knowledge Pool-Enhanced Path Loss Prediction Framework for Shore-to-Ship Communication

Accurate shore-to-ship path loss prediction is essential for maritime mobile communication systems, but it remains challenging because dynamic sea surface conditions can alter reflection paths and multipath effects, introducing uncertainty into the observed path loss. This paper proposes a wireless environment knowledge pool-enhanced path loss prediction (WEKP-PLP) method for point prediction and interval characterization of shore-to-ship path loss. The proposed method first constructs a wireless environment knowledge pool (WEKP) from ray tracing (RT) simulations under different wind speeds, temperatures, salinities, frequencies, antenna heights, and propagation distances. The WEKP learns the mapping from environmental and system parameters to path loss residual quantiles and provides a prior that contains both the median prediction and the associated prediction interval. To adapt this prior to real scenarios, a small number of measurement samples are used to learn the residual between the WEKP median prediction and the measured path loss through Gaussian process regression (GPR). The experimental results show that the WEKP-PLP provides accurate point prediction and compact intervals with reliable coverage. The proposed method achieves an MAE of 1.03 dB and an RMSE of 1.54 dB. The ablation results further confirm that both the WEKP prior and the residual correction using a small number of measurement samples are necessary for accurate and reliable shore-to-ship path loss prediction.

eess.SP

DeepFilters: Scattering-Aware Pupil Engineering with Learned Digital Filter Reconstruction for Extended Depth of Field Microscopy

Extended depth of field microscopy encodes axial information into a single acquisition through engineered point spread functions, but conventional and deep optics approaches are subject to degradation in scattering tissue. We introduce DeepFilters, a scattering-aware deep optics framework that jointly optimizes a parameterized pupil filter and a digital-filter-based reconstruction network through a calibrated differentiable forward model to achieve broad generalization without retraining. Incorporating empirical scattering kernels, physics-guided regularization, and a hybrid genetic-gradient initialization strategy, DeepFilters extends the PSF from 16 micron to >400 micron in clear media and enables signal recovery beyond 120 micron deep in biological tissues, validated across fixed brain slices and sea urchin embryos.

physics.optics

Transfer-Function Approach to Substrate-Enhanced Diffraction Tomography

Forward and backward scattering provide complementary volumetric and interfacial information, yet conventional three-dimensional (3D) imaging typically accesses only one. In this Letter, we present a substrate-enhanced diffraction tomography approach that simultaneously recovers both channels under multi-angle epi-illumination.This geometry captures one forward- and two backward-scattering bands in axially symmetric Fourier regions, where their complementary coverage enables phase-absorption separation in a non-Hermitian spectrum. Explicit 3D transfer functions are derived for both channels, and an axial Kramers-Kronig relation is established to incorporate substrate-induced boundary conditions in a unified framework. Our results establish a label-free, high-resolution 3D imaging modality that surpasses the limits of existing methods.

physics.optics

Reflection-mode Multi-slice Fourier Ptychographic Tomography

Diffraction tomography (DT) has been widely explored in transmission-mode configurations, enabling high-resolution, label-free 3D imaging. However, industrial metrology applications, such as semiconductor inspection, typically involve opaque or highly reflective substrates (e.g., silicon or metal), necessitating a reflection-mode imaging configuration. In this work, we introduce reflection-mode Multi-Slice Fourier Ptychographic Tomography (rMS-FPT) that achieves high-resolution, volumetric imaging of multi-layered, strongly scattering samples on reflective substrates. We develop a reflection-mode multi-slice beam propagation method (rMSBP) to model multiple scattering and substrate interactions, enabling precise 3D reconstruction. By incorporating darkfield measurements, rMS-FPT enhances resolution beyond the traditional brightfield limit and provides sub-micrometer lateral resolution while achieving optical sectioning. We validate rMS-FPT through numerical simulations on a four-layer resolution target and experimental demonstrations using a reflection-mode LED array microscope. Experiments on a two-layer resolution target and a multi-layer scattering sample confirm the method's effectiveness. Our optimized implementation enables rapid imaging, covering a 1.2 mm $\times$ 1.2 mm area in 1.6 seconds, reconstructing over $10^9$ voxels within a 0.4 mm$^3$ volume. This work represents a significant step in extending DT to reflection-mode configurations, providing a robust and scalable solution for 3D metrology and industrial inspection.

physics.optics

Coordinate-conditioned Deconvolution for Scalable Spatially Varying High-Throughput Imaging

Wide-field fluorescence microscopy with compact optics often suffers from spatially varying blur due to field-dependent aberrations, vignetting, and sensor truncation, while finite sensor sampling imposes an inherent trade-off between field of view (FOV) and resolution. Computational Miniaturized Mesoscope (CM2) alleviate the sampling limit by multiplexing multiple sub-views onto a single sensor, but introduce view crosstalk and a highly ill-conditioned inverse problem compounded by spatially variant point spread functions (PSFs). Prior learning-based spatially varying (SV) reconstruction methods typically rely on global SV operators with fixed input sizes, resulting in memory and training costs that scale poorly with image dimensions. We propose SV-CoDe (Spatially Varying Coordinate-conditioned Deconvolution), a scalable deep learning framework that achieves uniform, high-resolution reconstruction across a 6.5 mm FOV. Unlike conventional methods, SV-CoDe employs coordinate-conditioned convolutions to locally adapt reconstruction kernels; this enables patch-based training that decouples parameter count from FOV size. SV-CoDe achieves the best image quality in both simulated and experimental measurements while requiring 10x less model size and 10x less training data than prior baselines. Trained purely on physics-based simulations, the network robustly generalizes to bead phantoms, weakly scattering brain slices, and freely moving C. elegans. SV-CoDe offers a scalable, physics-aware solution for correcting SV blur in compact optical systems and is readily extendable to a broad range of biomedical imaging applications.

eess.IV

Mid-Infrared Photothermal Relaxation Intensity Diffraction Tomography for Video-rate Volumetric Chemical Imaging

Three-dimensional molecular imaging of living cells is essential for unraveling cellular metabolism and response to therapies. However, existing volumetric methods, including fluorescence microscopy and quantitative phase imaging, either require fluorescent labels or lack chemical specificity. Mid-infrared (mid-IR) photothermal microscopy provides label-free spectroscopic contrast with sub-micrometer resolution but is limited by slow acquisition rates, precluding 3D live-cell studies. Here, we present a photothermal relaxation intensity diffraction tomography (PRIDT) system that encodes mid-IR absorption induced refractive index change via a photothermal relaxation scheme and recovers it through intensity diffraction tomography. PRIDT achieves video-rate volumetric chemical imaging with up to 15 Hz per wavelength and offers lateral and axial resolutions of 264 nm and 1.12 um over a volumetric field of view of 50x50x10 um3. We showcase high-speed PRIDT imaging of protein and lipid metabolism in ovarian cancer cells and lipid-droplet dynamics in live cells. PRIDT opens new avenues for rapid, quantitative, three-dimensional molecular imaging in living systems.

physics.optics

Dual-wavelength Fourier Ptychographic Topography

We introduce a dual-wavelength Fourier ptychographic topography (FPT) method that extends the lambda/2 height-range limit of single-wavelength FPT. By reconstructing complex fields at two illumination wavelengths and exploiting their phase difference, the method achieves an effective synthetic wavelength lambda_s and an unambiguous range of lambda_s/2 without reducing lateral resolution. A noise-robust wrapped-number search is used to select per-pixel integer pairs (k1, k2), and a global refinement with circular TV regularization and soft bounds improves stability and preserves height discontinuities. The approach is validated through rigorous scattering-model-based simulations and experiments on structured silicon samples, demonstrating accurate height recovery in regimes where single-wavelength FPT exhibits phase wrapping. We analyze the limits of the FPT forward model and identify aspect ratio (AR) and phase modulation transfer function (ph-MTF) as key predictors of reconstruction fidelity. Simulations and experiments show that increasing AR beyond a practical threshold causes loss of high-frequency phase transfer and destabilizes dual-wavelength unwrapping. Within this AR range, dual-wavelength FPT provides robust, high-resolution topography suitable for semiconductor and industrial metrology.

physics.optics

A Comprehensive Survey of 3GPP Release 19 ISAC Channel Modeling: From Empirical Features to Unified Methodology and Standardized Simulator

Integrated Sensing and Communication (ISAC) has been identified as a key 6G application by ITU and 3GPP. Channel measurement and modeling is a prerequisite for ISAC system design and has attracted widespread attention from both academia and industry. 3GPP Release 19 initiated the ISAC channel study item in December 2023 and finalized its modeling specification in May 2025 after extensive technical discussions. However, a comprehensive survey that provides a systematic overview,from empirical channel features to modeling methodologies and standardized simulators,remains unavailable. In this paper, the key requirements and challenges in ISAC channel research are first analyzed, followed by a structured overview of the standardization workflow throughout the 3GPP Release 19 process. Then, critical aspects of ISAC channels, including physical objects, target channels, and background channels, are examined in depth, together with additional features such as spatial consistency, environment objects, Doppler characteristics, and shared clusters, supported by measurement-based analysis. To establish a unified ISAC channel modeling framework, an Extended Geometry-based Stochastic Model (E-GBSM) is proposed, incorporating all the aforementioned ISAC channel characteristics. Finally, a standardized simulator is developed based on E-GBSM, and a two-phase calibration procedure aligned with 3GPP Release 19 is conducted to validate both the model and the simulator, demonstrating close agreement with industrial reference results. Overall, this paper provides a systematic survey of 3GPP Release 19 ISAC channel standardization and offers insights into best practices for new feature characterization, unified modeling methodology, and standardized simulator implementation, which can effectively supporting ISAC technology evaluation and future 6G standardization.

eess.SP

DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging

High-dimensional neuroimaging analyses for clinical diagnosis are often constrained by compromises in spatiotemporal fidelity and by the limited adaptability of large-scale, general-purpose models. To address these challenges, we introduce Dynamic Curriculum Learning for Spatiotemporal Encoding (DCL-SE), an end-to-end framework centered on data-driven spatiotemporal encoding (DaSE). We leverage Approximate Rank Pooling (ARP) to efficiently encode three-dimensional volumetric brain data into information-rich, two-dimensional dynamic representations, and then employ a dynamic curriculum learning strategy, guided by a Dynamic Group Mechanism (DGM), to progressively train the decoder, refining feature extraction from global anatomical structures to fine pathological details. Evaluated across six publicly available datasets, including Alzheimer's disease and brain tumor classification, cerebral artery segmentation, and brain age prediction, DCL-SE consistently outperforms existing methods in accuracy, robustness, and interpretability. These findings underscore the critical importance of compact, task-specific architectures in the era of large-scale pretrained networks.

cs.CV

Whitened Score Diffusion: A Structured Prior for Imaging Inverse Problems

Conventional score-based diffusion models (DMs) may struggle with anisotropic Gaussian diffusion processes due to the required inversion of covariance matrices in the denoising score matching training objective \cite{vincent_connection_2011}. We propose Whitened Score (WS) diffusion models, a novel framework based on stochastic differential equations that learns the Whitened Score function instead of the standard score. This approach circumvents covariance inversion, extending score-based DMs by enabling stable training of DMs on arbitrary Gaussian forward noising processes. WS DMs establish equivalence with flow matching for arbitrary Gaussian noise, allow for tailored spectral inductive biases, and provide strong Bayesian priors for imaging inverse problems with structured noise. We experiment with a variety of computational imaging tasks using the CIFAR, CelebA ($64\times64$), and CelebA-HQ ($256\times256$) datasets and demonstrate that WS diffusion priors trained on anisotropic Gaussian noising processes consistently outperform conventional diffusion priors based on isotropic Gaussian noise. Our code is open-sourced at \href{https://github.com/jeffreyalido/wsdiffusion}{\texttt{github.com/jeffreyalido/wsdiffusion}}.

eess.IV

6G Channel Modeling: Requirement, Measurement, Methodology and Simulator

Sixth-generation (6G) mobile communications have attracted substantial attention in the global research community of information and communication technologies (ICTs). 6G systems are expected to support not only extended 5G usage scenarios but also new usage scenarios, such as integrated sensing and communication (ISAC), integrated artificial intelligence (AI) and communication, and communication and ubiquitous connectivity. To achieve this goal, channel characteristics must be comprehensively studied and properly exploited to promote the design, standardization, and optimization of 6G systems. In this paper, we first summarize the requirements and challenges in 6G channel research. Our focus is on channels for six promising technologies enabling 6G, including ISAC, extremely large-scale MIMO (XL-MIMO), mid-band and terahertz (THz) technologies, reconfigurable intelligent surfaces (RISs), and space-air-ground integrated networks (SAGINs). A survey of the progress in 6G channel research regarding the above six promising technologies is presented in terms of the latest measurement campaigns, new characteristics, modeling methods, and research prospects. To support testing, optimization and evaluation, existing 6G channel simulators are summarized. Then, BUPTCMCCCMG-IMT2030 is introduced as an example of a simulator that was developed on the basis of the ITU/3GPP 3D geometry-based stochastic model (GBSM) methodology. We also address open issues covering standardization activities, AI-enabled methods, and system performance analysis in the context of 6G channel research. This paper offers in-depth, hands-on insights into the best practices of channel measurements, modeling, and simulations for the evaluation of 6G technologies, the development of 6G standards, and the implementation and optimization of 6G systems.

eess.SP

CCL-LGS: Contrastive Codebook Learning for 3D Language Gaussian Splatting

Recent advances in 3D reconstruction techniques and vision-language models have fueled significant progress in 3D semantic understanding, a capability critical to robotics, autonomous driving, and virtual/augmented reality. However, methods that rely on 2D priors are prone to a critical challenge: cross-view semantic inconsistencies induced by occlusion, image blur, and view-dependent variations. These inconsistencies, when propagated via projection supervision, deteriorate the quality of 3D Gaussian semantic fields and introduce artifacts in the rendered outputs. To mitigate this limitation, we propose CCL-LGS, a novel framework that enforces view-consistent semantic supervision by integrating multi-view semantic cues. Specifically, our approach first employs a zero-shot tracker to align a set of SAM-generated 2D masks and reliably identify their corresponding categories. Next, we utilize CLIP to extract robust semantic encodings across views. Finally, our Contrastive Codebook Learning (CCL) module distills discriminative semantic features by enforcing intra-class compactness and inter-class distinctiveness. In contrast to previous methods that directly apply CLIP to imperfect masks, our framework explicitly resolves semantic conflicts while preserving category discriminability. Extensive experiments demonstrate that CCL-LGS outperforms previous state-of-the-art methods. Our project page is available at https://epsilontl.github.io/CCL-LGS/.

cs.CV

Theoretical Analysis of Near-Field MIMO Channel Capacity and Mid-Band Experimental Validation

With the increase of multiple-input-multiple-output (MIMO) array size and carrier frequency, near-field MIMO communications will become crucial in 6G wireless networks. Due to the increase of MIMO near-field range, the research of near-field MIMO capacity has aroused wide interest. In this paper, we focus on the theoretical analysis and empirical study of near-field MIMO capacity. First, the near-field channel model is characterized from the electromagnetic information perspective. Second, with the uniform planar array (UPA), the channel capacity based on effective degree of freedom (EDoF) is analyzed theoretically, and the closed-form analytical expressions are derived in detail. Finally, based on the numerical verification of near-field channel measurement experiment at 13 GHz band, we reveal that the channel capacity of UPA-type MIMO systems decreases continuously with the communication distance increasing. It can be observed that the near-field channel capacity gain is relatively obvious when large-scale MIMO is adopted at both receiving and transmitter ends, but the near-field channel capacity gain may be limited in the actual communication system with the small antenna array at receiving end. This work will give some reference to the near-field communication systems.

eess.SP

Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning

Modern robot navigation systems encounter difficulties in diverse and complex indoor environments. Traditional approaches rely on multiple modules with small models or rule-based systems and thus lack adaptability to new environments. To address this, we developed Astra, a comprehensive dual-model architecture, Astra-Global and Astra-Local, for mobile robot navigation. Astra-Global, a multimodal LLM, processes vision and language inputs to perform self and goal localization using a hybrid topological-semantic graph as the global map, and outperforms traditional visual place recognition methods. Astra-Local, a multitask network, handles local path planning and odometry estimation. Its 4D spatial-temporal encoder, trained through self-supervised learning, generates robust 4D features for downstream tasks. The planning head utilizes flow matching and a novel masked ESDF loss to minimize collision risks for generating local trajectories, and the odometry head integrates multi-sensor inputs via a transformer encoder to predict the relative pose of the robot. Deployed on real in-house mobile robots, Astra achieves high end-to-end mission success rate across diverse indoor environments.

cs.RO