arXiv ScienceSearch

arXiv subjects

Steven Lu

Publications and source records attributed to Steven Lu.

At least 19 recordsLinked to original sources

Fully Automatic Trace Gas Plume Detection

Future imaging spectrometers will expand contemporary data volumes by orders of magnitude, requiring automated methods to upscale labor-intensive detection of trace gas point sources. Here we present a fully-automated approach that achieves operational performance for plume detection and labelling without human participation. Our method combines machine learning (ML)-based morphological analysis with physics-based spectroscopic model fitting. We deploy it on data from the EMIT imaging spectrometer, operating in two modes. First, we present a "daily digest" that runs automatically on all downlinked data, flagging the largest events for immediate response. The daily digest demonstrates that a significant fraction of the largest plumes can be detected automatically with negligible false positives. This represents a significant new high-water mark in plume detection accuracy. Second, we use it for retrospective analysis to find plumes that were missed by the existing human review process. We observe that at least 25% of large plumes may have been passed over in the existing workflow due to confirmation bias and ambiguity in the visual cues used by human reviewers. Finally, we extend detection to three understudied trace gases: NH3, NO2 and the first observations of carbon monoxide (CO) plume in EMIT imagery.

cs.LG

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications

We introduce MOMO, the first multi-sensor foundation model for Mars remote sensing. MOMO uses model merge to integrate representations learned independently from three key Martian sensors (HiRISE, CTX, and THEMIS), spanning resolutions from 0.25 m/pixel to 100 m/pixel. Central to our method is our novel Equal Validation Loss (EVL) strategy, which aligns checkpoints across sensors based on validation loss similarity before fusion via task arithmetic. This ensures models are merged at compatible convergence stages, leading to improved stability and generalization. We train MOMO on a large-scale, high-quality corpus of $\sim 12$ million samples curated from Mars orbital data and evaluate it on 9 downstream tasks from Mars-Bench. MOMO achieves better overall performance compared to ImageNet pre-trained, earth observation foundation model, sensor-specific pre-training, and fully-supervised baselines. Particularly on segmentation tasks, MOMO shows consistent and significant performance improvement. Our results demonstrate that model merging through an optimal checkpoint selection strategy provides an effective approach for building foundation models for multi-resolution data. The model weights, pretraining code, pretraining data, and evaluation code are available at: https://github.com/kerner-lab/MOMO.

cs.CV

On anti-hyperbolicity for hyperk\"ahler varieties

By restricting to (a linear subspace of) an affine chart in projective space, a complex stably rational or unirational manifold of dimension $m$ is meromorphically dominable by $\mathbb C^m$, i.e., admits a meromorphic dominating map from $\mathbb C^m$. So are varieties that are birational to abelian varieties and Kummer K3 surfaces. G. Buzzard and the second author have shown that elliptic K3 surfaces are holomorphically dominable by $\mathbb C^2$, i.e. admitting a holomorphic map with nontrivial Jacobian. In this paper we explore various examples and criteria for meromorphic and holomorphic dominability by $\mathbb C^m$ of certain hyperk\"ahler manifolds, generalizing some known results about K3 surfaces. Anti-hyperbolicity has several interpretations in the sense of vanishing of the Kobayashi-Royden metrics, admitting dense entire holomorphic curves, or dominating holomorphic or meromorphic maps from the complex affine space of the same dimension.

math.CV

Mars-Bench: A Benchmark for Evaluating Foundation Models for Mars Science Tasks

Foundation models have enabled rapid progress across many specialized domains by leveraging large-scale pre-training on unlabeled data, demonstrating strong generalization to a variety of downstream tasks. While such models have gained significant attention in fields like Earth Observation, their application to Mars science remains limited. A key enabler of progress in other domains has been the availability of standardized benchmarks that support systematic evaluation. In contrast, Mars science lacks such benchmarks and standardized evaluation frameworks, which have limited progress toward developing foundation models for Martian tasks. To address this gap, we introduce Mars-Bench, the first benchmark designed to systematically evaluate models across a broad range of Mars-related tasks using both orbital and surface imagery. Mars-Bench comprises 20 datasets spanning classification, segmentation, and object detection, focused on key geologic features such as craters, cones, boulders, and frost. We provide standardized, ready-to-use datasets and baseline evaluations using models pre-trained on natural images, Earth satellite data, and state-of-the-art vision-language models. Results from all analyses suggest that Mars-specific foundation models may offer advantages over general-domain counterparts, motivating further exploration of domain-adapted pre-training. Mars-Bench aims to establish a standardized foundation for developing and comparing machine learning models for Mars science. Our data, models, and code are available at: https://mars-bench.github.io/.

cs.CV

Uncertainty Quantification for Surface Ozone Emulators using Deep Learning

Air pollution is a global hazard, and as of 2023, 94\% of the world's population is exposed to unsafe pollution levels. Surface Ozone (O3), an important pollutant, and the drivers of its trends are difficult to model, and traditional physics-based models fall short in their practical use for scales relevant to human-health impacts. Deep Learning-based emulators have shown promise in capturing complex climate patterns, but overall lack the interpretability necessary to support critical decision making for policy changes and public health measures. We implement an uncertainty-aware U-Net architecture to predict the Multi-mOdel Multi-cOnstituent Chemical data assimilation (MOMO-Chem) model's surface ozone residuals (bias) using Bayesian and quantile regression methods. We demonstrate the capability of our techniques in regional estimation of bias in North America and Europe for June 2019. We highlight the uncertainty quantification (UQ) scores between our two UQ methodologies and discern which ground stations are optimal and sub-optimal candidates for MOMO-Chem bias correction, and evaluate the impact of land-use information in surface ozone residual modeling.

cs.LG

Leveraging Deep Learning for Physical Model Bias of Global Air Quality Estimates

Air pollution is the world's largest environmental risk factor for human disease and premature death, resulting in more than 6 million permature deaths in 2019. Currently, there is still a challenge to model one of the most important air pollutants, surface ozone, particularly at scales relevant for human health impacts, with the drivers of global ozone trends at these scales largely unknown, limiting the practical use of physics-based models. We employ a 2D Convolutional Neural Network based architecture that estimate surface ozone MOMO-Chem model residuals, referred to as model bias. We demonstrate the potential of this technique in North America and Europe, highlighting its ability better to capture physical model residuals compared to a traditional machine learning method. We assess the impact of incorporating land use information from high-resolution satellite imagery to improve model estimates. Importantly, we discuss how our results can improve our scientific understanding of the factors impacting ozone bias at urban scales that can be used to improve environmental policy.

cs.LG

Evaluating the Retrieval Robustness of Large Language Models

Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG may also lead to performance degradation due to imperfect retrieval and the model's limited ability to leverage retrieved content. In this work, we evaluate the robustness of LLMs in practical RAG setups (henceforth retrieval robustness). We focus on three research questions: (1) whether RAG is always better than non-RAG; (2) whether more retrieved documents always lead to better performance; (3) and whether document orders impact results. To facilitate this study, we establish a benchmark of 1500 open-domain questions, each with retrieved documents from Wikipedia. We introduce three robustness metrics, each corresponds to one research question. Our comprehensive experiments, involving 11 LLMs and 3 prompting strategies, reveal that all of these LLMs exhibit surprisingly high retrieval robustness; nonetheless, different degrees of imperfect robustness hinders them from fully utilizing the benefits of RAG.

cs.CL

PASS: An Asynchronous Probabilistic Processor for Next Generation Intelligence

New computing paradigms are required to solve the most challenging computational problems where no exact polynomial time solution exists.Probabilistic Ising Accelerators has gained promise on these problems with the ability to model complex probability distributions and find ground states of intractable problems. In this context, we have demonstrated the Parallel Asynchronous Stochastic Sampler (PASS), the first fully on-chip integrated, asynchronous, probabilistic accelerator that takes advantage of the intrinsic fine-grained parallelism of the Ising Model and built in state of the art 14nm CMOS FinFET technology. We have demonstrated broad applicability of this accelerator on problems ranging from Combinatorial Optimization, Neural Simulation, to Machine Learning along with up to $23,000$x energy to solution improvement compared to CPUs on probabilistic problems.

cs.DC

Evaluating Terrain-Dependent Performance for Martian Frost Detection in Visible Satellite Observations

Seasonal frosting and defrosting on the surface of Mars is hypothesized to drive both climate processes and the formation and evolution of geomorphological features such as gullies. Past studies have focused on manually analyzing the behavior of the frost cycle in the northern mid-latitude region of Mars using high-resolution visible observations from orbit. Extending these studies globally requires automating the detection of frost using data science techniques such as convolutional neural networks. However, visible indications of frost presence can vary significantly depending on the geologic context on which the frost is superimposed. In this study, we (1) present a novel approach for spatially partitioning data to reduce biases in model performance estimation, (2) illustrate how geologic context affects automated frost detection, and (3) propose mitigations to observed biases in automated frost detection.

cs.CV

Interactive Mars Image Content-Based Search with Interpretable Machine Learning

The NASA Planetary Data System (PDS) hosts millions of images of planets, moons, and other bodies collected throughout many missions. The ever-expanding nature of data and user engagement demands an interpretable content classification system to support scientific discovery and individual curiosity. In this paper, we leverage a prototype-based architecture to enable users to understand and validate the evidence used by a classifier trained on images from the Mars Science Laboratory (MSL) Curiosity rover mission. In addition to providing explanations, we investigate the diversity and correctness of evidence used by the content-based classifier. The work presented in this paper will be deployed on the PDS Image Atlas, replacing its non-interpretable counterpart.

cs.CV

Finiteness of pointed maps to moduli spaces of polarized varieties

We establish a finiteness result for pointed maps to the base space $U$ of a smooth projective family of varieties with maximal variation in moduli. For its proof, we establish the rigidity of pointed maps to a (not necessarily compact) variety which is hyperbolic modulo a proper closed subset. Together with Viehweg's hyperbolicity conjecture on the bigness of log-canonical bundles of moduli spaces, resolved by Campana-Paun, we derive an optimal dimension bound on the Hom scheme from a curve to $U$ among other applications.

math.AG

MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies

Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P -- that is, minimizing the forward cross-entropy, which is equivalent to maximum likelihood estimation (MLE). We have observed that models trained in this way may "over-generalize", in the sense that they produce non-human-like text. Moreover, we believe that reverse cross-entropy, i.e., the cross-entropy of P relative to Q, is a better reflection of how a human would evaluate text generated by a model. Hence, we propose learning with MixCE, an objective that mixes the forward and reverse cross-entropies. We evaluate models trained with this objective on synthetic data settings (where P is known) and real data, and show that the resulting models yield better generated text without complex decoding strategies. Our code and models are publicly available at https://github.com/bloomberg/mixce-acl2023

cs.CL

BloombergGPT: A Large Language Model for Finance

The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering. Large Language Models (LLMs) have been shown to be effective on a variety of tasks; however, no LLM specialized for the financial domain has been reported in literature. In this work, we present BloombergGPT, a 50 billion parameter language model that is trained on a wide range of financial data. We construct a 363 billion token dataset based on Bloomberg's extensive data sources, perhaps the largest domain-specific dataset yet, augmented with 345 billion tokens from general purpose datasets. We validate BloombergGPT on standard LLM benchmarks, open financial benchmarks, and a suite of internal benchmarks that most accurately reflect our intended usage. Our mixed dataset training leads to a model that outperforms existing models on financial tasks by significant margins without sacrificing performance on general LLM benchmarks. Additionally, we explain our modeling choices, training process, and evaluation methodology. We release Training Chronicles (Appendix C) detailing our experience in training BloombergGPT.

cs.LG

Cosmic Microwave Background Recovery: A Graph-Based Bayesian Convolutional Network Approach

The cosmic microwave background (CMB) is a significant source of knowledge about the origin and evolution of our universe. However, observations of the CMB are contaminated by foreground emissions, obscuring the CMB signal and reducing its efficacy in constraining cosmological parameters. We employ deep learning as a data-driven approach to CMB cleaning from multi-frequency full-sky maps. In particular, we develop a graph-based Bayesian convolutional neural network based on the U-Net architecture that predicts cleaned CMB with pixel-wise uncertainty estimates. We demonstrate the potential of this technique on realistic simulated data based on the Planck mission. We show that our model accurately recovers the cleaned CMB sky map and resulting angular power spectrum while identifying regions of uncertainty. Finally, we discuss the current challenges and the path forward for deploying our model for CMB recovery on real observations.

cs.LG

Mars Image Content Classification: Three Years of NASA Deployment and Recent Advances

The NASA Planetary Data System hosts millions of images acquired from the planet Mars. To help users quickly find images of interest, we have developed and deployed content-based classification and search capabilities for Mars orbital and surface images. The deployed systems are publicly accessible using the PDS Image Atlas. We describe the process of training, evaluating, calibrating, and deploying updates to two CNN classifiers for images collected by Mars missions. We also report on three years of deployment including usage statistics, lessons learned, and plans for the future.

cs.LG

Holomorphic curves in Base Spaces of Families of Polarized Manifolds

For a smooth family $V \to U$ of polarized manifolds with semi-ample canonical sheaves, we show the following result: any entire curve must be contained in the fibers of the classifying map from the base space $U$ to the moduli space. This settles the Relative Isotriviality Conjecture, \cite[Conjecture 1.5]{DLSZ}.

math.AG

Picard theorems for moduli spaces of polarized varieties

As a result of our study of the hyperbolicity of the moduli space of polarized manifold, we give a general big Picard theorem for a holomorphic curve on a log-smooth pair $(X,D)$ such that $W=X\setminus D$ admits a Finsler pseudometric that is strongly negatively curved when pulled back to the curve. We show, by some refinements of the classical Viehweg-Zuo construction, that this latter condition holds for the base space $W$, if nonsingular, of any algebraic family of polarized complex projective manifolds with semi-ample canonical bundles whose induced moduli map $\phi$ to the moduli space of such manifolds is generically finite and any $\phi$-horizontal holomorphic curve in $W$. This yields the big Picard theorem for any holomorphic curves in the base space $U$ of such an algebraic family by allowing this base space to be singular but with generically finite moduli map. An immediate and useful corollary is that any holomorphic map from an algebraic variety to such a base space $U$ must be algebraic, i.e., the corresponding holomorphic family must be algebraic. We also show the related algebraic hyperbolicity property of such a base space $U$, which generalizes previous Arakelov inequalities and weak boundedness results for moduli stacks and offers, in addition to the Picard theorem above, another evidence in favor of the hyperbolic embeddability of such an $U$.

math.AG

Video Object Segmentation using Teacher-Student Adaptation in a Human Robot Interaction (HRI) Setting

Video object segmentation is an essential task in robot manipulation to facilitate grasping and learning affordances. Incremental learning is important for robotics in unstructured environments, since the total number of objects and their variations can be intractable. Inspired by the children learning process, human robot interaction (HRI) can be utilized to teach robots about the world guided by humans similar to how children learn from a parent or a teacher. A human teacher can show potential objects of interest to the robot, which is able to self adapt to the teaching signal without providing manual segmentation labels. We propose a novel teacher-student learning paradigm to teach robots about their surrounding environment. A two-stream motion and appearance "teacher" network provides pseudo-labels to adapt an appearance "student" network. The student network is able to segment the newly learned objects in other scenes, whether they are static or in motion. We also introduce a carefully designed dataset that serves the proposed HRI setup, denoted as (I)nteractive (V)ideo (O)bject (S)egmentation. Our IVOS dataset contains teaching videos of different objects, and manipulation tasks. Unlike previous datasets, IVOS provides manipulation tasks sequences with segmentation annotation along with the waypoints for the robot trajectories. It also provides segmentation annotation for the different transformations such as translation, scale, planar rotation, and out-of-plane rotation. Our proposed adaptation method outperforms the state-of-the-art on DAVIS and FBMS with 6.8% and 1.2% in F-measure respectively. It improves over the baseline on IVOS dataset with 46.1% and 25.9% in mIoU.

cs.CV