arXiv Science⌕ Search

arXiv subjects

Zhixuan Zhao

Publications and source records attributed to Zhixuan Zhao.

3 recordsLinked to original sources

Perfect resonance and fragile localization suppression in correlated disordered chains

Spatial correlations can suppress scattering in disordered chains and produce perfectly transmitting resonances. The practical value of this protection, however, depends not on the resonance peak itself but on the width of the surrounding transmission window and its sensitivity to local ordering errors. We show that a resonance can remain exactly transparent at its center while arbitrarily rare adjacent-swap errors restore an inverse localization length proportional to the square of the energy detuning. For lossless single-channel chains assembled from independent blocks of fixed length and composition, this positive quadratic term is guaranteed by a local scattering invariant and holds for every arrangement and every fixed swap probability between zero and one. With exact tuning and matched contacts, the central transmission remains unity. A microscopic quantum chain exhibits both higher-order suppression of scattering in the ideal recursive arrangement and the predicted response to local exchanges. These results reveal a limitation of spatial ordering that is invisible to a measurement at the resonance alone: spatial ordering protects the resonance peak, not the transport around it. The effect can therefore be tested experimentally by measuring transmission spectra before and after exchanges, without identifying microscopic defects.

cond-mat.dis-nn↗

Neural Fields for NV-Center Inverse Sensing

Inverse problems in scientific sensing are often solved with either hand-designed regularizers or supervised networks trained on simulated labels, yet both can fail when the forward model is nonlinear, spectrally coupled, and physically delicate. We study this issue for noise sensing based on nitrogen-vacancy (NV) centers in diamond, where a quantum sensor measures magnetic-noise spectra generated by sparse spin sources. We show that replacing a common scalar/coherent forward approximation with a tensor power-summed dipolar operator changes the inverse landscape and exposes a center-collapse failure mode in free-density optimization. We propose NeTMY, an amortization-free coordinate neural field coupled to the differentiable NV forward model, with annealed positional encoding, multiscale optimization, sparsity/gating, and spectrum-fidelity losses. Across sparse synthetic reconstructions generated by the corrected operator, NeTMY achieves the best localization and distributional metrics in the tested benchmark. Mechanism experiments show that NeTMY does not directly execute the raw density-space gradient; its parameterization smooths and redistributes updates, mitigating the center-collapse pathology. These results position NV quantum sensing as a useful testbed for physics-faithful neural inverse problems.

cs.LG↗

PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

We introduce PerceptionComp, a manually annotated benchmark for complex, long-horizon, perception-centric video reasoning. PerceptionComp is designed so that no single moment is sufficient: answering each question requires multiple temporally separated pieces of visual evidence and compositional constraints under conjunctive and sequential logic, spanning perceptual subtasks such as objects, attributes, relations, locations, actions, and events, and requiring skills including semantic recognition, visual correspondence, temporal reasoning, and spatial reasoning. The benchmark contains 1,114 highly complex questions on 279 videos from diverse domains including city walk tours, indoor villa tours, video games, and extreme outdoor sports, with 100% manual annotation. Human studies show that PerceptionComp requires substantial test-time thinking and repeated perception steps: participants take much longer than on prior benchmarks, and accuracy drops to near chance (18.97%) when rewatching is disallowed. State-of-the-art MLLMs also perform substantially worse on PerceptionComp than on existing benchmarks: the best model in our evaluation, Gemini-3-Flash, reaches only 45.96% accuracy in the five-choice setting, while open-source models remain below 40%. These results suggest that perception-centric long-horizon video reasoning remains a major bottleneck, and we hope PerceptionComp will help drive progress in perceptual reasoning.

cs.CV↗