arXiv ScienceSearch

arXiv subjects

Youngjun Choe

Publications and source records attributed to Youngjun Choe.

16 recordsLinked to original sources

Open-source data pipeline for street-view images: a case study on community mobility during COVID-19 pandemic

Street View Images (SVI) are a common source of valuable data for researchers. Researchers have used SVI data for estimating pedestrian volumes, demographic surveillance, and to better understand built and natural environments in cityscapes. However, the most common source of publicly available SVI data is Google Street View. Google Street View images are collected infrequently, making temporal analysis challenging, especially in low population density areas. Our main contribution is the development of an open-source data pipeline for processing 360-degree video recorded from a car-mounted camera. The video data is used to generate SVIs, which then can be used as an input for temporal analysis. We demonstrate the use of the pipeline by collecting a SVI dataset over a 38-month longitudinal survey of Seattle, WA, USA during the COVID-19 pandemic. The output of our pipeline is validated through statistical analyses of pedestrian traffic in the images. We confirm known results in the literature and provide new insights into outdoor pedestrian traffic patterns. This study demonstrates the feasibility and value of collecting and using SVI for research purposes beyond what is possible with currently available SVI data. Limitations and future improvements on the data pipeline and case study are also discussed.

cs.CV

Street View Data Collection Design for Disaster Reconnaissance

Over the last decade, street-view type images have been used across disciplines to generate and understand various place-based metrics. However efforts to collect this data were often meant to support investigator-driven research without regard to the utility of the data for other researchers. To address this, we describe our methods for collecting and publishing longitudinal data of this type in the wake of the COVID-19 pandemic and discuss some of the challenges we encountered along the way. Our process included designing a route taking into account both broad area canvassing and community capitals transects. We also implemented procedures for uploading and publishing data from each survey. Our methods successfully generated the kind of longitudinal data that can be beneficial to a variety of research disciplines. However, there were some challenges with data collection consistency and the sheer magnitude of data produced. Overall, our approach demonstrates the feasibility of generating longitudinal street-view data in the wake of a disaster event. Based on our experience, we provide recommendations for future researchers attempting to create a similar data set.

cs.HC

U.S. Power Resilience for 2002--2019

Prolonged power outages debilitate the economy and threaten public health. Existing research is generally limited in its scope to a single event, an outage cause, or a region. Here, we provide one of the most comprehensive analyses of U.S. power outages for 2002--2019. We categorized all outage data collected under U.S. federal mandates into four outage causes and computed industry-standard reliability metrics. Our spatiotemporal analysis reveals six of the most resilient U.S. states since 2010, improvement of power resilience against natural hazards in the south and northeast regions, and a disproportionately large number of human attacks for its population in the Western Electricity Coordinating Council region. Our regression analysis identifies several statistically significant predictors and hypotheses for power resilience. Furthermore, we propose a novel framework for analyzing outage data using differential weighting and influential points to better understand power resilience. We share curated data and code as Supplementary Materials.

stat.AP

Data-driven sparse polynomial chaos expansion for models with dependent inputs

Polynomial chaos expansions (PCEs) have been used in many real-world engineering applications to quantify how the uncertainty of an output is propagated from inputs. PCEs for models with independent inputs have been extensively explored in the literature. Recently, different approaches have been proposed for models with dependent inputs to expand the use of PCEs to more real-world applications. Typical approaches include building PCEs based on the Gram-Schmidt algorithm or transforming the dependent inputs into independent inputs. However, the two approaches have their limitations regarding computational efficiency and additional assumptions about the input distributions, respectively. In this paper, we propose a data-driven approach to build sparse PCEs for models with dependent inputs. The proposed algorithm recursively constructs orthonormal polynomials using a set of monomials based on their correlations with the output. The proposed algorithm on building sparse PCEs not only reduces the number of minimally required observations but also improves the numerical stability and computational efficiency. Four numerical examples are implemented to validate the proposed algorithm.

eess.SY

Post-Hurricane Damage Assessment Using Satellite Imagery and Geolocation Features

Gaining timely and reliable situation awareness after hazard events such as a hurricane is crucial to emergency managers and first responders. One effective way to achieve that goal is through damage assessment. Recently, disaster researchers have been utilizing imagery captured through satellites or drones to quantify the number of flooded/damaged buildings. In this paper, we propose a mixed data approach, which leverages publicly available satellite imagery and geolocation features of the affected area to identify damaged buildings after a hurricane. The method demonstrated significant improvement from performing a similar task using only imagery features, based on a case study of Hurricane Harvey affecting Greater Houston area in 2017. This result opens door to a wide range of possibilities to unify the advancement in computer vision algorithms such as convolutional neural networks and traditional methods in damage assessment, for example, using flood depth or bare-earth topology. In this work, a creative choice of the geolocation features was made to provide extra information to the imagery features, but it is up to the users to decide which other features can be included to model the physical behavior of the events, depending on their domain knowledge and the type of disaster. The dataset curated in this work is made openly available (DOI: 10.17603/ds2-3cca-f398).

cs.CV

COVID-19 Economic Policy Effects on Consumer Spending and Foot Traffic in the U.S

To battle with economic challenges during the COVID-19 pandemic, the US government implemented various measures to mitigate economic loss. From issuance of stimulus checks to reopening businesses, consumers had to constantly alter their behavior in response to government policies. Using anonymized card transactions and mobile device-based location tracking data, we analyze the factors that contribute to these behavior changes, focusing on stimulus check issuance and state-wide reopening. Our finding suggests that stimulus payment has a significant immediate effect of boosting spending, but it typically does not reverse a downward trend. State-wide reopening had a small effect on spending. Foot traffic increased gradually after stimulus check issuance, but only increased slightly after reopening, which also coincided or preceded several policy changes and confounding events (e.g., protests) in the US. We also find differences in the reaction to these policies in different regions in the US. Our results may be used to inform future economic recovery policies and their potential consumer response.

stat.AP

Splitting Gaussian Process Regression for Streaming Data

Gaussian processes offer a flexible kernel method for regression. While Gaussian processes have many useful theoretical properties and have proven practically useful, they suffer from poor scaling in the number of observations. In particular, the cubic time complexity of updating standard Gaussian process models make them generally unsuitable for application to streaming data. We propose an algorithm for sequentially partitioning the input space and fitting a localized Gaussian process to each disjoint region. The algorithm is shown to have superior time and space complexity to existing methods, and its sequential nature permits application to streaming data. The algorithm constructs a model for which the time complexity of updating is tightly bounded above by a pre-specified parameter. To the best of our knowledge, the model is the first local Gaussian process regression model to achieve linear memory complexity. Theoretical continuity properties of the model are proven. We demonstrate the efficacy of the resulting model on multi-dimensional regression tasks for streaming data.

stat.ML

Infrastructure Recovery Curve Estimation Using Gaussian Process Regression on Expert Elicited Data

Infrastructure recovery time estimation is critical to disaster management and planning. Inspired by recent resilience planning initiatives, we consider a situation where experts are asked to estimate the time for different infrastructure systems to recover to certain functionality levels after a scenario hazard event. We propose a methodological framework to use expert-elicited data to estimate the expected recovery time curve of a particular infrastructure system. This framework uses the Gaussian process regression (GPR) to capture the experts' estimation-uncertainty and satisfy known physical constraints of recovery processes. The framework is designed to find a balance between the data collection cost of expert elicitation and the prediction accuracy of GPR. We evaluate the framework on realistically simulated expert-elicited data concerning the two case study events, the 1995 Great Hanshin-Awaji Earthquake and the 2011 Great East Japan Earthquake.

stat.ME

Modeling of Lifeline Infrastructure Restoration Using Empirical Quantitative Data

Disaster recovery is widely regarded as the least understood phase of the disaster cycle. In particular, the literature around lifeline infrastructure restoration modeling frequently mentions the lack of empirical quantitative data available. Despite limitations, there is a growing body of research on modeling lifeline infrastructure restoration, often developed using empirical quantitative data. This study reviews this body of literature and identifies the data collection and usage patterns present across modeling approaches to inform future efforts using empirical quantitative data. We classify the modeling approaches into simulation, optimization, and statistical modeling. The number of publications in this domain has increased over time with the most rapid growth of statistical modeling. Electricity infrastructure restoration is most frequently modeled, followed by the restoration of multiple infrastructures, water infrastructure, and transportation infrastructure. Interdependency between multiple infrastructures is increasingly considered in recent literature. Researchers gather the data from a variety of sources, including collaborations with utility companies, national databases, and post-event damage and restoration reports. This study provides discussion and recommendations around data usage practices within the lifeline restoration modeling field. Following the recommendations would facilitate the development of a community of practice around restoration modeling and provide greater opportunities for future data sharing.

stat.AP

Identifying the Influential Inputs for Network Output Variance Using Sparse Polynomial Chaos Expansion

Sensitivity analysis (SA) is an important aspect of process automation. It often aims to identify the process inputs that influence the process output's variance significantly. Existing SA approaches typically consider the input-output relationship as a black-box and conduct extensive random sampling from the actual process or its high-fidelity simulation model to identify the influential inputs. In this paper, an alternate, novel approach is proposed using a sparse polynomial chaos expansion-based model for a class of input-output relationships represented as directed acyclic networks. The model exploits the relationship structure by recursively relating a network node to its direct predecessors to trace the output variance back to the inputs. It, thereby, estimates the Sobol indices, which measure the influence of each input on the output variance, accurately and efficiently. Theoretical analysis establishes the validity of the model as the prediction of the network output converges in probability to the true output under certain regularity conditions. Empirical evaluation on two manufacturing processes shows that the model estimates the Sobol indices accurately with far fewer observations than a state-of-the-art Monte Carlo sampling method.

stat.ME

Benchmark Dataset for Automatic Damaged Building Detection from Post-Hurricane Remotely Sensed Imagery

Rapid damage assessment is of crucial importance to emergency responders during hurricane events, however, the evaluation process is often slow, labor-intensive, costly, and error-prone. New advances in computer vision and remote sensing open possibilities to observe the Earth at a different scale. However, substantial pre-processing work is still required in order to apply state-of-the-art methodology for emergency response. To enable the comparison of methods for automatic detection of damaged buildings from post-hurricane remote sensing imagery taken from both airborne and satellite sensors, this paper presents the development of benchmark datasets from publicly available data. The major contributions of this work include (1) a scalable framework for creating benchmark datasets of hurricane-damaged buildings and (2) public sharing of the resulting benchmark datasets for Greater Houston area after Hurricane Harvey in 2017. The proposed approach can be used to build other hurricane-damaged building datasets on which researchers can train and test object detection models to automatically identify damaged buildings.

cs.CV

Building Damage Annotation on Post-Hurricane Satellite Imagery Based on Convolutional Neural Networks

After a hurricane, damage assessment is critical to emergency managers for efficient response and resource allocation. One way to gauge the damage extent is to quantify the number of flooded/damaged buildings, which is traditionally done by ground survey. This process can be labor-intensive and time-consuming. In this paper, we propose to improve the efficiency of building damage assessment by applying image classification algorithms to post-hurricane satellite imagery. At the known building coordinates (available from public data), we extract square-sized images from the satellite imagery to create training, validation, and test datasets. Each square-sized image contains a building to be classified as either 'Flooded/Damaged' (labeled by volunteers in a crowd-sourcing project) or 'Undamaged'. We design and train a convolutional neural network from scratch and compare it with an existing neural network used widely for common object classification. We demonstrate the promise of our damage annotation model (over 97% accuracy) in the case study of building damage assessment in the Greater Houston area affected by 2017 Hurricane Harvey.

cs.CV

Cross-Entropy Based Importance Sampling for Stochastic Simulation Models

To efficiently evaluate system reliability based on Monte Carlo simulation, importance sampling is used widely. The optimal importance sampling density was derived in 1950s for the deterministic simulation model, which maps an input to an output deterministically, and is approximated in practice using various methods. For the stochastic simulation model whose output is random given an input, the optimal importance sampling density was derived only recently. In the existing literature, metamodel-based approaches have been used to approximate this optimal density. However, building a satisfactory metamodel is often difficult or time-consuming in practice. This paper proposes a cross-entropy based method, which is automatic and does not require specific domain knowledge. The proposed method uses an expectation-maximization algorithm to guide the choice of a mixture distribution model for approximating the optimal density. The method iteratively updates the approximated density to minimize its estimated discrepancy, measured by estimated cross-entropy, from the optimal density. The mixture model's complexity is controlled using the cross-entropy information criterion. The method is empirically validated using extensive numerical studies and applied to a case study of evaluating the reliability of wind turbine using a stochastic simulation model.

stat.ME

Data-Driven Sensitivity Indices for Models With Dependent Inputs Using the Polynomial Chaos Expansion

Uncertainties exist in both physics-based and data-driven models. Variance-based sensitivity analysis characterizes how the variance of a model output is propagated from the model inputs. The Sobol index is one of the most widely used sensitivity indices for models with independent inputs. For models with dependent inputs, different approaches have been explored to obtain sensitivity indices in the literature. Typical approaches are based on procedures of transforming the dependent inputs into independent inputs. However, such transformation requires additional information about the inputs, such as the dependency structure or the conditional probability density functions. In this paper, data-driven sensitivity indices are proposed for models with dependent inputs. We first construct ordered partitions of linearly independent polynomials of the inputs. The modified Gram-Schmidt algorithm is then applied to the ordered partitions to generate orthogonal polynomials with respect to the empirical measure based on observed data of model inputs and outputs. Using the polynomial chaos expansion with the orthogonal polynomials, we obtain the proposed data-driven sensitivity indices. The sensitivity indices provide intuitive interpretations of how the dependent inputs affect the variance of the output without a priori knowledge on the dependence structure of the inputs. Three numerical examples are used to validate the proposed approach.

stat.ME

Importance Sampling and its Optimality for Stochastic Simulation Models

We consider the problem of estimating an expected outcome from a stochastic simulation model. Our goal is to develop a theoretical framework on importance sampling for such estimation. By investigating the variance of an importance sampling estimator, we propose a two-stage procedure that involves a regression stage and a sampling stage to construct the final estimator. We introduce a parametric and a nonparametric regression estimator in the first stage and study how the allocation between the two stages affects the performance of the final estimator. We analyze the variance reduction rates and derive oracle properties of both methods. We evaluate the empirical performances of the methods using two numerical examples and a case study on wind turbine reliability evaluation.

stat.ME

Information Criterion for Boltzmann Approximation Problems

This paper considers the problem of approximating a density when it can be evaluated up to a normalizing constant at a limited number of points. We call this problem the Boltzmann approximation (BA) problem. The BA problem is ubiquitous in statistics, such as approximating a posterior density for Bayesian inference and estimating an optimal density for importance sampling. Approximating the density with a parametric model can be cast as a model selection problem. This problem cannot be addressed with traditional approaches that maximize the (marginal) likelihood of a model, for example, using the Akaike information criterion (AIC) or Bayesian information criterion (BIC). We instead aim to minimize the cross-entropy that gauges the deviation of a parametric model from the target density. We propose a novel information criterion called the cross-entropy information criterion (CIC) and prove that the CIC is an asymptotically unbiased estimator of the cross-entropy (up to a multiplicative constant) under some regularity conditions. We propose an iterative method to approximate the target density by minimizing the CIC. We demonstrate that the proposed method selects a parametric model that well approximates the target density.

stat.ME