arXiv ScienceSearch

arXiv subjects

Yupeng Wu

Publications and source records attributed to Yupeng Wu.

12 recordsLinked to original sources

SolarBench: A global solar energy nowcasting benchmark

As the share of solar power grows, nowcasting weather-driven solar variability becomes critical for reliable energy system operation. State-of-the-art approaches increasingly apply deep learning to sky camera and geostationary satellite observations, but fragmented datasets and inconsistent evaluation make it difficult to determine whether reported improvements generalize across climates, cloud regimes, and photovoltaic (PV) systems. Here we introduce SolarBench, an open global benchmark for image-based solar nowcasting. SolarBench harmonizes more than six million sky and satellite images from 11 diverse sites spanning a decade, together with irradiance or PV output and auxiliary atmospheric data. An accompanying toolbox supports reproducible data access, processing, model development, and evaluation. Using SolarBench, we benchmark representative models and reveal a gap between average forecasting accuracy and the ability to capture rapid solar fluctuations. We further quantify predictability across cloud regimes and demonstrate data-efficient adaptation to new PV systems. SolarBench provides an extensible foundation for fair comparison and methodological innovation in solar nowcasting.

cs.CV

VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We study cross-episode ObjectNav, where each request is an independently initialized, single-goal episode and only self-acquired, scene-scoped memory persists across episodes. We ask whether an agent with fixed model parameters and navigation components can reuse such experience without retraining or oracle information. We introduce \method, a training-free framework with a persistent Hierarchical Visual-Topological Memory (VTM). VTM uses a coarse room topology to index room-owned visual memories, distinguishes in-room from remote-visible evidence, and retains successful approach cues. For each request, VTM-Nav re-localizes the agent in accumulated scene structure, retrieves target-relevant records from plausible rooms, and grounds memory guidance in candidates derived from the current observation. A conservative execution guard further handles local failures. Under matched 40-step comparisons, VTM-Nav exceeds the memory-reset WMNav control by 4.6, 2.0, and 0.8 SR points on HM3D v0.1, HM3D v0.2, and MP3D, respectively, with comparable or higher SPL. On HM3D, it also exceeds WMNav harnessed by textual memory by 3.1 and 5.5 SR points. These results demonstrate effective reuse of cross-episode scene experience through hierarchical visual-topological memory.

cs.CV

Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding

Multi-Pool Chemical Exchange Saturation Transfer (CEST) MRI provides valuable metabolic information but is clinically limited by long acquisition times. Although sparse sampling reduces scanning time, reconstructing high-resolution Z-spectra from limited data remains an ill-posed inverse problem. Conventional interpolation and generic Implicit Neural Rep-resentations (INRs) often lack physical constraints, leading to spectral artifacts and physically invalid signals. To address this, we propose Lorentz Encoding (LE), a physics-informed framework that formulates CEST reconstruction as a self-supervised reconstruction task via implicit continuous coordinate learning. Unlike generic positional encodings, LE regularizes the continuous spectral mapping by projecting sparse coordinates into a physically constrained space governed by a combination of parametric Lorentzian profiles with learnable basis functions. This mechanism effectively reduces noise and enforces consistency with physical models. Experiments on in vivo human brain data demonstrate that LE significantly outperforms state-of-the-art methods. Specifically, under a 39-point sampling strategy, LE achieves a PSNR of 57.58 dB and an SSIM of 0.9994. Furthermore, the learned physics-informed encodings form a continuous, geometrically ordered trajectory in the latent space, ensuring accurate quantitative metabo-lite mapping (APT, NOE, MT).

cs.LG

DOME: Learning Transferable Domain Variables from Sparse Supervision for Test-Time Adaptation

Test-time adaptation (TTA) aims to align a model to shifting test domains using only unlabeled streaming data. Most existing methods implicitly infer a single global domain distribution, ignoring the multidimensional and sample-specific nature of real-world domain shifts, leading to fragile adaptation. We propose DOME, an effective domain encoder that explicitly models each sample's domain in a zero-shot manner. DOME leverages vision-language pretraining to extract dense, continuous representations, parameterizes domains as distributional variables, and introduces a momentum-updated sparse domain bank for disentangled supervision. By injecting these explicit domain cues into downstream models, even a basic entropy-minimization TTA strategy achieves state-of-the-art performance across ImageNet-C, ImageNet-R, and ImageNet-Sketch, outperforming complex TTA approaches. Our results demonstrate that robust adaptation stems not from intricate adaptation algorithms, but from explicit, structured domain representation.

cs.CV

VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

This paper presents a summary of the VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models (LMMs), hosted as part of the ICCV 2025 Workshop on Visual Quality Assessment. The challenge aims to evaluate and enhance the ability of state-of-the-art LMMs to perform open-ended and detailed reasoning about visual quality differences across multiple images. To this end, the competition introduces a novel benchmark comprising thousands of coarse-to-fine grained visual quality comparison tasks, spanning single images, pairs, and multi-image groups. Each task requires models to provide accurate quality judgments. The competition emphasizes holistic evaluation protocols, including 2AFC-based binary preference and multi-choice questions (MCQs). Around 100 participants submitted entries, with five models demonstrating the emerging capabilities of instruction-tuned LMMs on quality assessment. This challenge marks a significant step toward open-domain visual quality reasoning and comparison and serves as a catalyst for future research on interpretable and human-aligned quality evaluation systems.

cs.CV

A framework for rapid, reproducible, and high-fidelity whole-brain multi-pool CEST imaging at 3T

Purpose: To develop and validate a framework for rapid, accurate, and reproducible whole-brain, multi-pool chemical exchange saturation transfer (CEST) imaging at 3T, addressing challenges of long acquisition times and confounding factors. Methods: A single-shot 3D true fast imaging with steady-state precession (True FISP) sequence was optimized for whole-brain multi-pool CEST. Rapid B0, B1, and T1 mapping was performed using a dual-echo modified four-angle method. A feed-forward neural network was developed for rapid B1 correction, trained against the conventional multi-power method. The apparent exchange-dependent relaxation (AREX) metric was used to correct for T1 and magnetization transfer (MT) effects. The framework was validated in phantoms and healthy human subjects (N=8), including a test-retest reproducibility assessment. Results: The True FISP sequence yielded high-quality, whole-brain images with minimal artifacts and distortion in a clinically feasible scan time (~9 minutes). Phantom studies confirmed the effectiveness of B1 correction (coefficient of variation [CV] for MT_MTRLD decreased from 22.49% to 4.61%) and AREX-based confounder correction (CV for APT_AREX reduced from 33.6% to 6.9%). The neural network B1 correction showed excellent agreement with the conventional multi-power method in vivo (ICC > 0.97). High test-retest reproducibility was demonstrated across 96 brain regions, with the average CV for APT_AREX under 10% for over 95% of regions. Conclusion: A rapid and robust framework for whole-brain quantitative multi-pool CEST imaging was successfully developed and validated. By integrating an efficient acquisition sequence with a streamlined correction pipeline, this approach overcomes key barriers to clinical translation, enabling reliable metabolic imaging for widespread brain pathologies.

physics.med-ph

3D Single-shot CEST imaging at 3T Based on True FISP Readout

To simultaneously fit multiple-pool effects, spectrally selective 3D CEST imaging typically requires single-shot readouts to save time. However, to date, FLASH and EPI have been the primary pulse sequences used for this purpose. They suffer from low SNR or image distortion related to B0 field inhomogeneity. In this work, we developed a 3D single-shot CEST sequence using true fast imaging with steady-state precession (True FISP) readout, also known as bSSFP, and optimized the scanning parameters through simulations. The performance of the CEST sequence was validated using an egg white phantom, ten healthy volunteers, and a patient with a brain tumor on a 3T human scanner. Subsequently, the proposed CEST sequence using True FISP was compared with the commonly used FLASH-based CEST sequence, focusing on SNR and image contrast, while maintaining identical pre-saturation modes, repetition time, echo time and scan time. In the simulation experiments, the maximum CEST signal obtained from the True FISP was significantly greater than that obtained from the FLASH sequence. In the egg white phantom, the SNRs of amide proton transfer (APT) and nuclear Overhauser enhancement (NOE) effect images obtained from the True FISP were 68.3% and 57.0% higher than those obtained from the FLASH sequence, respectively. In healthy volunteers, saturated images collected with the True FISP sequence at 3.5 ppm showed an approximate 84% increase in mean temporal SNR compared to those collected with the FLASH sequence. Compared to the FLASH sequence, the CEST images obtained from the True FISP sequence could display more detailed brain tissue structures of both normal individuals and the patient with a brain tumor. Therefore, due to the high SNR inherent in the sequence, True FISP has the potential to be used for fast and high-quality 3D image readout of CEST contrasts in clinical applications.

physics.med-ph

Instruct Large Language Models to Generate Scientific Literature Survey Step by Step

Abstract. Automatically generating scientific literature surveys is a valuable task that can significantly enhance research efficiency. However, the diverse and complex nature of information within a literature survey poses substantial challenges for generative models. In this paper, we design a series of prompts to systematically leverage large language models (LLMs), enabling the creation of comprehensive literature surveys through a step-by-step approach. Specifically, we design prompts to guide LLMs to sequentially generate the title, abstract, hierarchical headings, and the main content of the literature survey. We argue that this design enables the generation of the headings from a high-level perspective. During the content generation process, this design effectively harnesses relevant information while minimizing costs by restricting the length of both input and output content in LLM queries. Our implementation with Qwen-long achieved third place in the NLPCC 2024 Scientific Literature Survey Generation evaluation task, with an overall score only 0.03% lower than the second-place team. Additionally, our soft heading recall is 95.84%, the second best among the submissions. Thanks to the efficient prompt design and the low cost of the Qwen-long API, our method reduces the expense for generating each literature survey to 0.1 RMB, enhancing the practical value of our method.

cs.CL

Prototypical Contrastive Learning through Alignment and Uniformity for Recommendation

Graph Collaborative Filtering (GCF), one of the most widely adopted recommendation system methods, effectively captures intricate relationships between user and item interactions. Graph Contrastive Learning (GCL) based GCF has gained significant attention as it leverages self-supervised techniques to extract valuable signals from real-world scenarios. However, many methods usually learn the instances of discrimination tasks that involve the construction of contrastive pairs through random sampling. GCL approaches suffer from sampling bias issues, where the negatives might have a semantic structure similar to that of the positives, thus leading to a loss of effective feature representation. To address these problems, we present the \underline{Proto}typical contrastive learning through \underline{A}lignment and \underline{U}niformity for recommendation, which is called \textbf{ProtoAU}. Specifically, we first propose prototypes (cluster centroids) as a latent space to ensure consistency across different augmentations from the origin graph, aiming to eliminate the need for random sampling of contrastive pairs. Furthermore, the absence of explicit negatives means that directly optimizing the consistency loss between instance and prototype could easily result in dimensional collapse issues. Therefore, we propose aligning and maintaining uniformity in the prototypes of users and items as optimization objectives to prevent falling into trivial solutions. Finally, we conduct extensive experiments on four datasets and evaluate their performance on the task of link prediction. Experimental results demonstrate that the proposed ProtoAU outperforms other representative methods. The source codes of our proposed ProtoAU are available at \url{https://github.com/oceanlvr/ProtoAU}.

cs.IR

On the Opportunities of Green Computing: A Survey

Artificial Intelligence (AI) has achieved significant advancements in technology and research with the development over several decades, and is widely used in many areas including computing vision, natural language processing, time-series analysis, speech synthesis, etc. During the age of deep learning, especially with the arise of Large Language Models, a large majority of researchers' attention is paid on pursuing new state-of-the-art (SOTA) results, resulting in ever increasing of model size and computational complexity. The needs for high computing power brings higher carbon emission and undermines research fairness by preventing small or medium-sized research institutions and companies with limited funding in participating in research. To tackle the challenges of computing resources and environmental impact of AI, Green Computing has become a hot research topic. In this survey, we give a systematic overview of the technologies used in Green Computing. We propose the framework of Green Computing and devide it into four key components: (1) Measures of Greenness, (2) Energy-Efficient AI, (3) Energy-Efficient Computing Systems and (4) AI Use Cases for Sustainability. For each components, we discuss the research progress made and the commonly used techniques to optimize the AI efficiency. We conclude that this new research direction has the potential to address the conflicts between resource constraints and AI development. We encourage more researchers to put attention on this direction and make AI more environmental friendly.

cs.AI

DRL-ORA: Distributional Reinforcement Learning with Online Risk Adaption

One of the main challenges in reinforcement learning (RL) is that the agent has to make decisions that would influence the future performance without having complete knowledge of the environment. Dynamically adjusting the level of epistemic risk during the learning process can help to achieve reliable policies in safety-critical settings with better efficiency. In this work, we propose a new framework, Distributional RL with Online Risk Adaptation (DRL-ORA). This framework quantifies both epistemic and implicit aleatory uncertainties in a unified manner and dynamically adjusts the epistemic risk levels by solving a total variation minimization problem online. The framework unifies the existing variants of risk adaption approaches and offers better explainability and flexibility. The selection of risk levels is performed efficiently via a grid search using a Follow-The-Leader-type algorithm, where the offline oracle also corresponds to a ''satisficing measure'' under a specially modified loss function. We show that DRL-ORA outperforms existing methods that rely on fixed risk levels or manually designed risk level adaptation in multiple classes of tasks.

cs.LG

How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection

The introduction of ChatGPT has garnered widespread attention in both academic and industrial communities. ChatGPT is able to respond effectively to a wide range of human questions, providing fluent and comprehensive answers that significantly surpass previous public chatbots in terms of security and usefulness. On one hand, people are curious about how ChatGPT is able to achieve such strength and how far it is from human experts. On the other hand, people are starting to worry about the potential negative impacts that large language models (LLMs) like ChatGPT could have on society, such as fake news, plagiarism, and social security issues. In this work, we collected tens of thousands of comparison responses from both human experts and ChatGPT, with questions ranging from open-domain, financial, medical, legal, and psychological areas. We call the collected dataset the Human ChatGPT Comparison Corpus (HC3). Based on the HC3 dataset, we study the characteristics of ChatGPT's responses, the differences and gaps from human experts, and future directions for LLMs. We conducted comprehensive human evaluations and linguistic analyses of ChatGPT-generated content compared with that of humans, where many interesting results are revealed. After that, we conduct extensive experiments on how to effectively detect whether a certain text is generated by ChatGPT or humans. We build three different detection systems, explore several key factors that influence their effectiveness, and evaluate them in different scenarios. The dataset, code, and models are all publicly available at https://github.com/Hello-SimpleAI/chatgpt-comparison-detection.

cs.CL