arXiv ScienceSearch

arXiv subjects

Sungbin Lim

Publications and source records attributed to Sungbin Lim.

32 records · Page 2Linked to original sources

AutoCLINT: The Winning Method in AutoCV Challenge 2019

NeurIPS 2019 AutoDL challenge is a series of six automated machine learning competitions. Particularly, AutoCV challenges mainly focused on classification tasks on visual domain. In this paper, we introduce the winning method in the competition, AutoCLINT. The proposed method implements an autonomous training strategy, including efficient code optimization, and applies an automated data augmentation to achieve the fast adaptation of pretrained networks. We implement a light version of Fast AutoAugment to search for data augmentation policies efficiently for the arbitrarily given image domains. We also empirically analyze the components of the proposed method and provide ablation studies focusing on AutoCV datasets.

cs.LG

torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models

We design and implement a ready-to-use library in PyTorch for performing micro-batch pipeline parallelism with checkpointing proposed by GPipe (Huang et al., 2019). In particular, we develop a set of design components to enable pipeline-parallel gradient computation in PyTorch's define-by-run and eager execution environment. We show that each component is necessary to fully benefit from pipeline parallelism in such environment, and demonstrate the efficiency of the library by applying it to various network architectures including AmoebaNet-D and U-Net. Our library is available at https://github.com/kakaobrain/torchgpipe .

cs.DC

Scalable Neural Architecture Search for 3D Medical Image Segmentation

In this paper, a neural architecture search (NAS) framework is proposed for 3D medical image segmentation, to automatically optimize a neural architecture from a large design space. Our NAS framework searches the structure of each layer including neural connectivities and operation types in both of the encoder and decoder. Since optimizing over a large discrete architecture space is difficult due to high-resolution 3D medical images, a novel stochastic sampling algorithm based on a continuous relaxation is also proposed for scalable gradient based optimization. On the 3D medical image segmentation tasks with a benchmark dataset, an automatically designed architecture by the proposed NAS framework outperforms the human-designed 3D U-Net, and moreover this optimized architecture is well suited to be transferred for different tasks.

cs.LG

Fast AutoAugment

Data augmentation is an essential technique for improving generalization ability of deep learning models. Recently, AutoAugment has been proposed as an algorithm to automatically search for augmentation policies from a dataset and has significantly enhanced performances on many image recognition tasks. However, its search method requires thousands of GPU hours even for a relatively small dataset. In this paper, we propose an algorithm called Fast AutoAugment that finds effective augmentation policies via a more efficient search strategy based on density matching. In comparison to AutoAugment, the proposed algorithm speeds up the search time by orders of magnitude while achieves comparable performances on image recognition tasks with various models and datasets including CIFAR-10, CIFAR-100, SVHN, and ImageNet.

cs.LG

Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning

In this paper, we present a new class of Markov decision processes (MDPs), called Tsallis MDPs, with Tsallis entropy maximization, which generalizes existing maximum entropy reinforcement learning (RL). A Tsallis MDP provides a unified framework for the original RL problem and RL with various types of entropy, including the well-known standard Shannon-Gibbs (SG) entropy, using an additional real-valued parameter, called an entropic index. By controlling the entropic index, we can generate various types of entropy, including the SG entropy, and a different entropy results in a different class of the optimal policy in Tsallis MDPs. We also provide a full mathematical analysis of Tsallis MDPs, including the optimality condition, performance error bounds, and convergence. Our theoretical result enables us to use any positive entropic index in RL. To handle complex and large-scale problems, we propose a model-free actor-critic RL method using Tsallis entropy maximization. We evaluate the regularization effect of the Tsallis entropy with various values of entropic indices and show that the entropic index controls the exploration tendency of the proposed method. For a different type of RL problems, we find that a different value of the entropic index is desirable. The proposed method is evaluated using the MuJoCo simulator and achieves the state-of-the-art performance.

cs.LG

Task Agnostic Robust Learning on Corrupt Outputs by Correlation-Guided Mixture Density Networks

In this paper, we focus on weakly supervised learning with noisy training data for both classification and regression problems.We assume that the training outputs are collected from a mixture of a target and correlated noise distributions.Our proposed method simultaneously estimates the target distribution and the quality of each data which is defined as the correlation between the target and data generating distributions.The cornerstone of the proposed method is a Cholesky Block that enables modeling dependencies among mixture distributions in a differentiable manner where we maintain the distribution over the network weights.We first provide illustrative examples in both regression and classification tasks to show the effectiveness of the proposed method.Then, the proposed method is extensively evaluated in a number of experiments where we show that it constantly shows comparable or superior performances compared to existing baseline methods in the handling of noisy data.

cs.LG

Neural Stain-Style Transfer Learning using GAN for Histopathological Images

Performance of data-driven network for tumor classification varies with stain-style of histopathological images. This article proposes the stain-style transfer (SST) model based on conditional generative adversarial networks (GANs) which is to learn not only the certain color distribution but also the corresponding histopathological pattern. Our model considers feature-preserving loss in addition to well-known GAN loss. Consequently our model does not only transfers initial stain-styles to the desired one but also prevent the degradation of tumor classifier on transferred images. The model is examined using the CAMELYON16 dataset.

cs.CV

Uncertainty-Aware Learning from Demonstration using Mixture Density Networks with Sampling-Free Variance Modeling

In this paper, we propose an uncertainty-aware learning from demonstration method by presenting a novel uncertainty estimation method utilizing a mixture density network appropriate for modeling complex and noisy human behaviors. The proposed uncertainty acquisition can be done with a single forward path without Monte Carlo sampling and is suitable for real-time robotics applications. The properties of the proposed uncertainty measure are analyzed through three different synthetic examples, absence of data, heavy measurement noise, and composition of functions scenarios. We show that each case can be distinguished using the proposed uncertainty measure and presented an uncertainty-aware learn- ing from demonstration method of an autonomous driving using this property. The proposed uncertainty-aware learning from demonstration method outperforms other compared methods in terms of safety using a complex real-world driving dataset.

cs.CV

A Sobolev Space theory for stochastic partial differential equations with time-fractional derivatives

In this article we present an $L_p$-theory ($p\geq 2$) for the time-fractional quasi-linear stochastic partial differential equations (SPDEs) of type $$ \partial^{\alpha}_tu=L(\omega,t,x)u+f(u)+\partial^{\beta}_t \sum_{k=1}^{\infty}\int^t_0 ( \Lambda^k(\omega,t,x)u+g^k(u))dw^k_t, $$ where $\alpha\in (0,2)$, $\beta <\alpha+\frac{1}{2}$, and $\partial^{\alpha}_t$ and $\partial^{\beta}_t$ denote the Caputo derivative of order $\alpha$ and $\beta$ respectively. The processes $w^k_t$, $k\in \mathbb{N}=\{1,2,\cdots\}$, are independent one-dimensional Wiener processes defined on a probability space $\Omega$, $L$ is a second order operator of either divergence or non-divergence type, and $\Lambda^k$ are linear operators of order up to two. The coefficients of the equations depend on $\omega (\in \Omega), t,x$ and are allowed to be discontinuous. This class of SPDEs can be used to describe random effects on transport of particles in medium with thermal memory or particles subject to sticking and trapping.

math.PR

An $L_q(L_p)$-theory for the time fractional evolution equations with variable coefficients

We introduce an $L_q(L_p)$-theory for the quasi-linear fractional equations of the type $$ \partial^{\alpha}_t u(t,x)=a^{ij}(t,x)u_{x^i x^j}(t,x)+f(t,x,u), \quad t>0, \,x\in \mathbf{R}^d. $$ Here, $\alpha\in (0,2)$, $p,q>1$, and $\partial^{\alpha}_t$ is the Caupto fractional derivative of order $\alpha$. Uniqueness, existence, and $L_q(L_p)$-estimates of solutions are obtained. The leading coefficients $a^{ij}(t,x)$ are assumed to be piecewise continuous in $t$ and uniformly continuous in $x$. In particular $a^{ij}(t,x)$ are allowed to be discontinuous with respect to the time variable. Our approach is based on classical tools in PDE theories such as the Marcinkiewicz interpolation theorem, the Calderon-Zygmund theorem, and perturbation arguments.

math.AP

Asymptotic behaviors of fundamental solution and its derivatives related to space-time fractional differential equations

Let $p(t,x)$ be the fundamental solution to the problem $$ \partial_{t}^{\alpha}u=-(-\Delta)^{\beta}u, \quad \alpha\in (0,2), \, \beta\in (0,\infty). $$ In this paper we provide the asymptotic behaviors and sharp upper bounds of $p(t,x)$ and its space and time fractional derivatives $$ D_{x}^{n}(-\Delta_x)^{\gamma}D_{t}^{\sigma}I_{t}^{\delta}p(t,x), \quad \forall\,\, n\in\mathbb{Z}_{+}, \,\, \gamma\in[0,\beta],\,\, \sigma, \delta \in[0,\infty), $$ where $D_{x}^n$ is a partial derivative of order $n$ with respect to $x$, $(-\Delta_x)^{\gamma}$ is a fractional Laplace operator and $D_{t}^{\sigma}$ and $I_{t}^{\delta}$ are Riemann-Liouville fractional derivative and integral respectively.

math.AP

An $L_q(L_p)$-theory for parabolic pseudo-differential equations: Calder\'on-Zygmund approach

In this paper we present a Calder\'{o}n-Zygmund approach for a large class of parabolic equations with pseudo-differential operators $\mathcal{A}(t)$ of arbitrary order $\gamma\in(0,\infty)$. It is assumed that $\cA(t)$ is merely measurable with respect to the time variable. The unique solvability of the equation $$ \frac{\partial u}{\partial t}=\cA u-\lambda u+f, \quad (t,x)\in \fR^{d+1} $$ and the $L_{q}(\fR,L_{p})$-estimate $$ \|u_{t}\|_{L_{q}(\fR,L_{p})}+\|(-\Delta)^{\gamma/2}u\|_{L_{q}(\fR,L_{p})} +\lambda\|u\|_{L_{q}(\fR,L_{p})}\leq N\|f\|_{L_{q}(\fR,L_{p})} $$ are obtained for any $\lambda > 0$ and $p,q\in (1,\infty)$.

math.AP

Parabolic BMO estimates for pseudo-differential operators of arbitrary order

In this article we prove the BMO-$L_{\infty}$ estimate $$ \|(-\Delta)^{\gamma/2} u\|_{BMO(\mathbf{R}^{d+1})}\leq N \|\frac{\partial}{\partial t}u-A(t)u\|_{L_{\infty}(\mathbf{R}^{d+1})}, \quad \forall\, u\in C^{\infty}_c(\mathbf{R}^{d+1}) $$ for a wide class of pseudo-differential operators $A(t)$ of order $\gamma\in (0,\infty)$. The coefficients of $A(t)$ are assumed to be merely measurable in time variable. As an application to the equation $$ \frac{\partial}{\partial t}u=A(t)u+f,\quad t\in \mathbf{R} $$ we prove that for any $u\in C^{\infty}_c(\mathbf{R}^{d+1})$ $$ \|u_t\|_{L_p(\mathbf{R}^{d+1})}+\|(-\Delta)^{\gamma/2}u\|_{L_p(\mathbf{R}^{d+1})}\leq N\|u_t-A(t)u\|_{L_p(\mathbf{R}^{d+1})}, $$ where $p\ in (1,\infty)$ and the constant $N$ is independent of $u$.

math.AP