arXiv ScienceSearch

arXiv subjects

Feng Wang

Publications and source records attributed to Feng Wang.

At least 19 recordsLinked to original sources

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in which selected actions can be executed. Games have no such external substrate; the virtual world itself must be instantiated. Most playable games require a persistent and navigable space, while 3D games additionally require explicit geometry that supports movement and interaction. Action-conditioned video rollouts provide visual observations but not this spatial representation. We present \textsc{Valerant}, a training-free framework that transforms a pretrained action-conditioned world model into a WAM for exploring and constructing 3D game maps. By coupling predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection, \textsc{Valerant} progressively transforms a single image into a persistent 3D game map. This framework extends WAM-based interaction beyond 2D visual simulation and offers a new approach to reducing manual effort in 3D game-map creation.

cs.AI

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, and reaches the highest page throughput in our comparison at 2.57 pages per second. On a low-budget GPU such as the NVIDIA L4, FastMTP doubles decoding speed over greedy autoregressive decoding. The model is publicly available at https://huggingface.co/jinaai/jina-ocr-v1.

cs.CL

Optical Voltage Profiling of 2D Semiconductors via Proximal Exciton Sensing

High contact resistances in atomically thin semiconductors often mask intrinsic electrical transport properties, particularly at low carrier densities where exotic correlated states emerge. We introduce optical voltage profiling, a noninvasive wide-field technique that replaces local voltage probes with a proximal monolayer MoSe$_2$ exciton sensor. Isolated by thin hexagonal boron nitride, this sensor converts the target's local electrostatic potential into spatially resolved modulations of exciton reflectance. Through pixel-wise in situ calibration, these signals yield quantitative two-dimensional voltage maps of an actively biased semiconductor device. Using this method, we demonstrate the carrier-density-driven metal-insulator transition in bilayer MoSe$_2$ and obtain channel resistances below 1 k$\Omega$ despite M$\Omega$-scale two-terminal resistances in the metallic region. The optically derived resistance exhibits a metal-insulator crossover near the resistance quantum $h/e^2$, and the voltage maps and reconstructed local conductivity reveal pronounced spatial heterogeneity in both insulating and metallic regimes. Beyond resolving channel resistance under high contact-resistance conditions, the technique provides spatially resolved access to microscopic transport heterogeneity in functional van der Waals devices.

cond-mat.mes-hall

Best Reaction Target To Determine Proton Distribution Radii of Atomic Nuclei

We found that a heavy target such as Pb is most suitable for determining the proton distribution radii of unstable nuclei through charge-changing cross-section ($\sigma_\text{cc}$) measurements. As a heavy ion probe, low-$Z$ targets are routinely used to determine nucleon distribution radii of unstable isotopes. This approach has recently been extended to study proton distribution radii from $\sigma_\text{cc}$ measurements. However, empirical scaling factors have to be introduced to apply the Glauber models. In the present work, we systematically investigated the scaling factor using 39 new $\sigma_\text{cc}$ data of 18 $p$-shell nuclei on hydrogen, carbon, silver, and lead targets at around 240 MeV/nucleon. Together with the existing data, we reveal a universal dependence of the scaling factor on both the masses of target nuclei and the separation energies of projectile nuclei. The scaling factors decrease with increasing target-nucleus mass and converge to 1 for the highest-$Z$ target, making the scaling unnecessary. We conclude that instead of a low-$Z$ target, employing a heavy target such as Pb in $\sigma_\text{cc}$ measurements is the best option to determine the proton distribution radii of unstable nuclei.

nucl-ex

Evidence for Dynamical Filtering: High Binary Fraction, Hard-binary Excess, and Unresolved Triples in the Surviving Core of NGC 6791

We present a deep photometric analysis of the main-sequence (MS) population in the old, metal-rich open cluster (OC) NGC 6791 using Gaia Data Release 3 data. After correcting for differential reddening, we use the Bayesian model comparison to test whether stellar rotation can account for the observed MS broadening and find that a rotation-dominated interpretation is strongly disfavored. We therefore infer that unresolved multiplicity is the primary contributor to the photometric offsets. We derive a high-q companion fraction of $54.3\% \pm 2.8\%$ for systems with $q \gtrsim 0.5$, significantly higher than typical values reported for most OCs and the field. The inferred offset distribution is not consistent with a flat mass-ratio distribution but instead shows an excess toward high mass ratios ($q \sim 0.8$--$1.0$), suggestive of preferential survival of hard binaries in a dynamically evolved environment. We also identify a population of stars lying above the equal-mass binary limit ($\Delta G > 0.75$ mag), which is difficult to explain with ordinary MS binaries alone and is plausibly interpreted as candidate unresolved triple or higher-order multiple systems. A Kolmogorov--Smirnov test, together with Monte Carlo label-shuffling experiments, shows no statistically significant difference between the projected radial distributions of the single-star and binary/multiple populations within the observed field. Taken together, these results are consistent with the picture that NGC 6791 is the dynamically processed inner remnant of a once more massive cluster.

astro-ph.SR

A Pressure-Robust Nonconforming Immersed Finite Element Method for Stokes Interface Problems

It is well established that an appropriate modification of test functions may lead to pressure-robust mixed methods for Stokes problems. However, for immersed finite element approximations of Stokes interface problems on unfitted meshes, it remains unclear whether the velocity error is independent of the pressure, since the velocity and the pressure are coupled in one of the interface conditions. In this paper, we provide a positive answer through a novel decomposition of the discontinuous pressure into a continuous component and a velocity-dependent discontinuous component. We demonstrate that the immersed Crouzeix--Raviart/$P_0$ element method achieves pressure robustness via an $H(\operatorname{div})$-conforming reconstruction of the test functions on the right-hand side. The stability and optimal error estimates of the proposed method are established with constants independent of the interface position relative to the mesh. Numerical experiments are presented to validate the theoretical findings.

math.NA

Neutral Atom Quantum Computing: Principles, Routes, Progress, and Challenges

Neutral atom quantum computing utilizes laser-trapped neutral atoms as qubits and realizes quantum logic gate operations through Rydberg-state interactions. In recent years, it has become one of the most vibrant directions in quantum computing hardware. This paper systematically reviews the working principles of neutral-atom quantum computers, including qubit encoding, atom trapping and manipulation, Rydberg states and interactions, the Rydberg blockade quantum gate mechanism, and atom rearrangement with reconfigurable architectures. The mainstream technical routes are surveyed, represented by optical tweezer arrays combined with Rydberg interactions, optical lattice schemes, and dipole trap arrays. A panoramic review is provided of domestic and international research progress from theoretical foundations in 2000 to the latest achievements in 2026, including thousand-qubit-scale systems, logical qubits, and quantum error correction experiments. Key breakthroughs are highlighted, such as the 6100-atom qubit array, continuous operation of a 3000-qubit system, quantum simulation of the Kitaev honeycomb model, toric code error correction demonstrations, encoding rates exceeding 1/2, and fault-tolerant architectures. The core bottlenecks are analyzed in depth, including the scalability--fidelity trade-off, engineering implementation of quantum error correction, atom loss and mid-circuit replenishment, laser system industrialization, control electronics scalability, and long-distance quantum interconnection. This paper aims to provide a systematic reference for academic research and technological development in this field.

quant-ph

From Cloud to Crowd: Democratizing LLM Service with Decentralized Edge Collaboration for RAG

The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices. Cloud-hosted LLMs are powerful but expensive and difficult to scale due to vendor lock-in and high resource needs, resulting in high expenses and unstable performance under load. Recent efforts focus on deploying small language models (SLMs), distilled or pruned from LLMs, on resource-constrained edge devices to reduce costs and improve scalability. However, edge-based SLMs face limited knowledge coverage and notable accuracy gap compared to cloud-based LLMs. To address this, we present DEFRAG, a decentralized edge collaboration system for retrieval-augmented generation (RAG) that optimizes both retrieval and generation across heterogeneous edge devices. For retrieval, DEFRAG compresses and shares knowledge graphs, using hybrid retrieval to expand knowledge coverage. For generation, DEFRAG introduces an optimizer that adaptively selects SLMs and RAG parameters per query, balancing accuracy and cost. We implement DEFRAG on a heterogeneous edge testbed and evaluate it on benchmark QA datasets. We also test it under mobile route stress, non-uniform data placement, and a domain-specific QA workload. The results show that DEFRAG maintains stable service quality and cost efficiency under these broader settings. Results show that DEFRAG narrows the SLM-LLM accuracy gap, while reducing cost by up to 98.4% and increasing peak throughput by up to 97.8% over centralized services. These findings demonstrate the potential of DEFRAG for democratized LLM services at the edge.

cs.DC

Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation

We address the problem of reconstructing a high-fidelity, freely navigable 3D scene from a single $360^\circ$ panorama, without per-scene optimization or multi-view capture. Existing methods either lack metric trajectory control, which hinders reliable downstream 3D reconstruction, or struggle with large disocclusions under long-range camera motion while requiring high-end multi-GPU servers.We present Genie Sim PanoWorld, a two-stage feed-forward pipeline that bridges generation and reconstruction via an explicit, trajectory-controllable panoramic video. A NavMesh-planned $\mathrm{SE}(3)$ roaming trajectory is injected into a latent video diffusion model through dense geometry-warped conditioning; long--short trajectory mixed training and a self-consistency objective based on shortcut models together yield high-fidelity video in four CFG-free denoising steps. A feed-forward panoramic reconstructor then lifts the generated video into a high-fidelity 3D Gaussian scene that supports real-time, free-viewpoint roaming and can be directly used as a simulation-ready asset for embodied AI applications. Experiments show that Genie Sim PanoWorld outperforms geometry-conditioned baselines in both panoramic video generation and downstream 3D reconstruction, while generalizing zero-shot to unseen indoor scenes.

cs.CV

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present AgentOmnia, a framework coordinating task-space definition, data synthesis, post-training, evaluation, and improvement across To-Consumer (ToC), To-Business (ToB), and To-Employee (ToE) applications. An extensible Domain x Capability x Atomic Difficulty taxonomy aligns these stages and enables fine-grained diagnosis with OmniaBench. AgentOmnia combines bidirectional environment-task synthesis with tool-dependency, program-structured, and solver-based pipelines, constructing 5,018 stateful environments with 255,375 tools and 52,361 tasks. Programs, solvers, and verifiers provide correctness signals, while supervised fine-tuning, online agentic reinforcement learning, and a rollback curriculum support post-training. Evaluation failures translate into Product Requirement Documents (PRDs) for targeted self-evolution. Starting from Qwen3-30B-A3B-Thinking-2507, AgentOmnia raises the pass rate on the OmniaBench challenging subset from 9.16% to 37.11% and the macro-average across OmniaBench, $\tau^2$-Bench, DeepPlanning, and VitaBench from 22.86% to 41.69%. Under a unified protocol,it leads the evaluated agentic post-trained baselines on OmniaBench and retains the highest four-benchmark macro-average. It also surpasses Qwen3-235B-A22B-Thinking-2507 on all four benchmarks and exceeds Qwen3.5-35B-A3B on the macro-average. Gains span three application splits, ten capability dimensions, eight atomic-difficulty factors, and 76 of 90 level-1 domains, indicating broad rather than category-specific improvement. A one-round study provides initial evidence for PRD-guided self-evolution, motivating validation at larger scales and in industrial settings.

cs.AI

jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation

Listwise rerankers are the discriminative core of agentic retrieval pipelines, yet production deployment demands efficiency, domain robustness, and fluency on semi-structured data at the same time. We present jina-reranker-v3.5, a 0.6B-parameter listwise reranker that meets these demands together without sacrificing the cross-document comparison that makes its predecessor jina-reranker-v3 effective. jina-reranker-v3.5 keeps the last-but-not-late (LBNL) interaction of jina-reranker-v3 and reworks it along three axes. It replaces uniform global attention with a hybrid schedule of three sliding-window layers followed by two global layers, pinning the terminal layer to global as LBNL readout requires. It trains on a curated multi-domain mixture that spans legal, medical, financial, multilingual, and structured retrieval. It transfers quality through a three-stage self-distillation recipe in which a full-attention teacher sets an upper bound that a sparse-attention student then recovers under a staged adaptation protocol. jina-reranker-v3.5 reaches 63.20 nDCG@10 on BEIR, matching a 4B model at roughly 7x fewer parameters, and improves over jina-reranker-v3 on MIRACL and RTEB as well. Its largest gains come on semi-structured retrieval, where it lifts nDCG@10 by 9.6 points over jina-reranker-v3 and leads all rerankers of comparable size. The hybrid schedule further cuts listwise inference latency by up to 1.56x. We release the model weights on Hugging Face under a non-commercial license.

cs.IR

Let RGB Be the Language of Vision

This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted into a common RGB-to-RGB image editing problem. In this paradigm, different types of visual information internally share the same encoding and decoding architecture and parameters as natural images, enabling a single model to transfer across tasks through a unified visual interface, in a way analogous to how language models operate over text. We refer to this formulation as RGB In and RGB Out (RINO). Built upon a generic image editing backbone without task-specific fine-tuning, RINO demonstrates robust and competitive zero-shot performance on both dense understanding tasks such as segmentation and depth estimation (where we unify outputs as RGB), and dense-conditioned generation tasks such as pose-to-image generation (where we unify inputs as RGB). We hope this study provides useful insights toward general unified vision-language systems, where diverse visual tasks can be expressed, interpreted, and solved through a shared visual language. Code is available at https://github.com/yangtiming/RINO.

cs.CV

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all historical visual and textual inputs into a shared context window, limiting long-horizon multimodal dialogue due to visual token explosion and unreliable cross-turn referencing. We propose a Cognitive-structured Multimodal Agent that externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant episodes during reasoning. The agent consists of a Perceptual Abstraction Engine for structured visual abstraction, a Cognitive Retrieval Engine for cross-turn memory retrieval, and a Multimodal Executive Controller for autonomous task inference and action planning. To address the lack of turn-level retrieval supervision in existing datasets, we develop a Unified Scenario Engine that programmatically generates structured multi-turn conversations with fine-grained retrieval annotations, enabling reinforcement learning to optimize abstraction and retrieval policies. We also construct a long-horizon visual-dialogue benchmark stratified by difficulty to evaluate episodic visual recall. Our 8B agent achieves 91.4% retrieval accuracy over 20-turn sessions, surpassing 32B baselines by +8.2% while nearly halving per-turn inference time (23.1s -> 12.7s). We further present the Cognitive-structured Multimodal Agent Harness (CMA-Harness), a tool-augmented deployment of the same cognitive structure integrating persistent multimodal memory, web access, image generation/editing/composition tools, and OpenAI-compatible serving. Structured memory and modular decision-making offer a more scalable, efficient paradigm for long-horizon multimodal agents than monolithic parameter scaling. Code: https://github.com/caseclose/cma-harness ; Project page: https://caseclose.github.io/cma-harness/

cs.CV

Domain decomposition methods with Physics-informed neural networks for elliptic equations on manifolds

We propose two numerical domain decomposition methods (DDMs) for elliptic equations on compact Riemannian manifolds, based on physics-informed neural networks (PINNs). Our approach incorporates the DDM technique for manifolds with the advantages of neural networks in high-dimensional settings. The proposed methods are validated through numerical experiments on various manifolds, both with and without boundary, in dimensions ranging from $5$ to $10$.

math.NA

Infinite collisions of simple random walks on random recursive trees generated by Bernoulli sequences

In this paper, we study random recursive trees generated by Bernoulli sequences. Starting from a graph with two vertices and one edge, each new vertex is connected to the last vertex with probability $ p $, or to the second-last vertex with probability $ q = 1-p $, this recursive construction yields a random infinite recursive tree $T$. We prove that $T$ almost surely has exactly one topological end. Furthermore, we establish that $T$ has the infinite collision property: two independent simple random walks on $T$ collide infinitely often almost surely.

math.PR

An Inner-Outer Iteration Algorithm with Optimal Parameters for Stochastic Lyapunov Matrix Equation

This paper proposes an inner--outer (IO) iterative algorithm with optimal parameters for solving stochastic Lyapunov matrix equation associated with discrete-time stochastic linear system. First, under the assumption that the underlying stochastic linear system is asymptotically mean-square stable, the monotonicity and boundedness of the iterative sequence generated by the proposed algorithm are analyzed. On this basis, a sufficient convergence result is established for the zero initial condition. Second, by deriving the spectral radius of the corresponding iteration matrix, several necessary and sufficient convergence conditions are obtained for arbitrary initial conditions. In addition, the optimal parameter-selection strategies are developed to improve the convergence performance of the algorithm. Finally, numerical examples are presented to verify the theoretical results and demonstrate the advantages of the proposed algorithm over several existing iterative methods.

math.NA

Ai2-Kit: Streamlining AI-Accelerated Ab Initio Workflows for Complex Chemical Systems

Molecular simulations of complex chemical systems, such as catalysis, electrochemistry, and energy storage, often need to capture the interplay of effects such as electronic structure, finite-temperature fluctuations, and electric-field response. Such complexity is difficult to address with traditional ab initio calculations, which are limited by the time and length scales they can reach. AI-accelerated ab initio (AI2) methods use machine learning potentials trained on first-principles data to replace expensive electronic-structure calculations, extending ab initio accuracy to these regimes, but their routine application requires reliable workflows that connect first-principles calculations, model training, molecular dynamics, enhanced sampling, trajectory analysis, and HPC orchestration. Here we present ai2-kit, a software toolkit for developing accessible, reproducible, and extensible AI2 workflows. ai2-kit provides high-semantic-density command-line interfaces and Python APIs for structure and dataset conversion, batch task generation, active-learning screening, job orchestration, and workflow recovery. We demonstrate ai2-kit in four representative applications: active-learning-based machine learning potential construction, free-energy perturbation for redox and acid-base processes, electrochemical machine learning potentials for electrified interfaces, and spectroscopies from machine learning molecular dynamics. ai2-kit also provides AI-agent skills that help users adapt these use cases into customized workflows for their own chemical systems and computational software stacks. Together, ai2-kit helps turn AI2 methods from bespoke computational protocols into reusable and extensible workflows for complex chemical systems, from model construction to property prediction.

physics.chem-ph

Moir\'e Phonons and Emergent Exciton-Phonon Coupling in a Moir\'e Heterobilayer

Moir\'e superlattices have emerged as a new platform for engineering electronic and optical properties in van der Waals heterostructures, enabling control over correlated and excitonic phenomena. Yet the impact of moir\'e superlattices on exciton-phonon coupling remains largely unexplored. Here we demonstrate emergent, layer-selective coupling between moir\'e phonons and moir\'e excitons in angle-aligned WS2/WSe2 heterobilayers. Using a broadband terahertz phonon transducer, we coherently launch moir\'e phonons that resonantly perturb the excitonic states. We show that the exciton-phonon coupling is intrinsically modified by the moir\'e superlattice in a layer-selective manner. A driven oscillator model captures the dynamics, revealing three moir\'e phonon resonances with distinct coupling to the moir\'e excitons. First principles calculations show that many moir\'e phonon modes can arise with distinct strongly hybridized in-plane and out-of-plane vibrations in the moir\'e unit cells. The calculations further identify the three experimentally observed moir\'e phonons and their emergent characteristic coupling to the moir\'e excitons.

cond-mat.mtrl-sci