arXiv ScienceSearch

arXiv subjects

Simon Fischer

Publications and source records attributed to Simon Fischer.

11 recordsLinked to original sources

Non-maximizing policies that fulfill multi-criterion aspirations in expectation

In dynamic programming and reinforcement learning, the policy for the sequential decision making of an agent in a stochastic environment is usually determined by expressing the goal as a scalar reward function and seeking a policy that maximizes the expected total reward. However, many goals that humans care about naturally concern multiple aspects of the world, and it may not be obvious how to condense those into a single reward function. Furthermore, maximization suffers from specification gaming, where the obtained policy achieves a high expected total reward in an unintended way, often taking extreme or nonsensical actions. Here we consider finite acyclic Markov Decision Processes with multiple distinct evaluation metrics, which do not necessarily represent quantities that the user wants to be maximized. We assume the task of the agent is to ensure that the vector of expected totals of the evaluation metrics falls into some given convex set, called the aspiration set. Our algorithm guarantees that this task is fulfilled by using simplices to approximate feasibility sets and propagate aspirations forward while ensuring they remain feasible. It has complexity linear in the number of possible state-action-successor triples and polynomial in the number of evaluation metrics. Moreover, the explicitly non-maximizing nature of the chosen policy and goals yields additional degrees of freedom, which can be used to apply heuristic safety criteria to the choice of actions. We discuss several such safety criteria that aim to steer the agent towards more conservative behavior.

cs.AI

Uniqueness of First Passage Time Distributions via Fredholm Integral Equations

Let $W$ be a standard Brownian motion with $W_0 = 0$ and let $b: \mathbb{R}_+ \to \mathbb{R}$ be a continuous function with $b(0) > 0$. The first passage time (from below) is then defined as \begin{align*} \tau := \inf \{ t \geq 0 \vert W_t \geq b(t) \}. \end{align*} It is well-known that the distribution $F$ of $\tau$ satisfies a set of Fredholm equations of the first kind, which is used, for example, as a starting point for numerical approaches. For this, it is fundamental that the Fredholm equations have a unique solution. In this article, we prove this in a general setting using analytical methods.

math.PR

Low-energy electron microscopy intensity-voltage data -- factorization, sparse sampling, and classification

Low-energy electron microscopy (LEEM) taken as intensity-voltage (I-V) curves provides hyperspectral images of surfaces, which can be used to identify the surface type, but are difficult to analyze. Here, we demonstrate the use of an algorithm for factorizing the data into spectra and concentrations of characteristic components (FSC3) for identifying distinct physical surface phases. Importantly, FSC3 is an unsupervised and fast algorithm. As example data we use experiments on the growth of praseodymium oxide or ruthenium oxide on ruthenium single crystal substrates, both featuring a complex distribution of coexisting surface components, varying in both chemical composition and crystallographic structure. With the factorization result a sparse sampling method is demonstrated, reducing the measurement time by 1-2 orders of magnitude, relevant for dynamic surface studies. The FSC3 concentrations are providing the features for a support vector machine (SVM) based supervised classification of the types. Here, specific surface regions which have been identified structurally, via their diffraction pattern, as well as chemically by complementary spectro-microscopic techniques, are used as training sets. A reliable classification is demonstrated on both exemplary LEEM I-V datasets.

cond-mat.mtrl-sci

A new integral equation for Brownian stopping problems with finite time horizon

For classical finite time horizon stopping problems driven by a Brownian motion \[V(t,x) = \sup_{t\leq\tau\leq0}E_{(t,x)}[g(\tau,W_{\tau})],\] we derive a new class of Fredholm type integral equations for the stopping set. For large problem classes of interest, we show by analytical arguments that the equation uniquely characterizes the stopping boundary of the problem. Regardless of the uniqueness, we use the representation to rigorously find the limit behavior of the stopping boundary close to the terminal time. Interestingly, it turns out that the leading-order coefficient is universal for wide classes of problems. We also discuss how the representation can be used for numerical purposes.

math.PR

Massively strained VO2 thin film growth on RuO2

Strain engineering vanadium dioxide thin films is one way to alter this material's characteristic first order transition from semiconductor to metal. In this study we extend the exploitable strain regime by utilizing the very large lattice mismatch of 8.78 % occurring in the VO$_2$/RuO$_2$ system along the c axis of the rutile structure. We have grown VO$_2$ thin films on single domain RuO$_2$ islands of two distinct surface orientations by atomic oxygen-supported reactive MBE. These films were examined by spatially resolved photoelectron and x-ray absorption spectroscopy, confirming the correct stoichiometry. Low energy electron diffraction then reveals the VO$_2$ films to grow indeed fully strained on RuO$_2$(110), exhibiting a previously unreported ($2\times2$) reconstruction. On TiO$_2$(110) substrates, we reproduce this reconstruction and attribute it to an oxygen-rich termination caused by the high oxygen chemical potential. On RuO$_2$(100) on the other hand, the films grow fully relaxed. Hence, the presented growth method allows for simultaneous access to a remarkable strain window ranging from bulk-like structures to massively strained regions.

cond-mat.mtrl-sci

A Closer Look at Covering Number Bounds for Gaussian Kernels

We establish some new bounds on the log-covering numbers of (anisotropic) Gaussian reproducing kernel Hilbert spaces. Unlike previous results in this direction we focus on small explicit constants and their dependency on crucial parameters such as the kernel bandwidth and the size and dimension of the underlying space.

math.FA

Note on the (non-)smoothness of discrete time value functions

We consider the discrete time stopping problem \[ V(t,x) = \sup_{\tau}E_{(t,x)}[g(\tau, X_\tau)],\] where $X$ is a random walk. It is well known that the value function $V$ is in general not smooth on the boundary of the continuation set $\partial C$. We show that under some conditions $V$ is not smooth in the interior of $C$ either. More precisely we show that $V$ is not differentiable in the $x$ component on a dense subset of $C$. As an example we consider the Chow-Robbins game. We give evidence that as well $\partial C$ is not smooth and that $C$ is not convex, even if $g(t,\cdot)$ is for every $t$.

math.PR

On the Sn/n-Problem

The Chow-Robbins game is a classical still partly unsolved stopping problem introduced by Chow and Robbins in 1965. You repeatedly toss a fair coin. After each toss, you decide if you take the fraction of heads up to now as a payoff, otherwise you continue. As a more general stopping problem this reads \[V(n,x) = \sup_{\tau }\operatorname{E} \left [ \frac{x + S_\tau}{n+\tau}\right]\] where $S$ is a random walk. We give a tight upper bound for $V$ when $S$ has subgassian increments. We do this by usinf the analogous time continuous problem with a standard Brownian motion as the driving process. From this we derive an easy proof for the existence of optimal stopping times in the discrete case. For the Chow-Robbins game we as well give a tight lower bound and use these to calculate, on the integers, the complete continuation and the stopping set of the problem for $n\leq 10^{5}$.

math.PR

Some New Bounds on the Entropy Numbers of Diagonal Operators

Entropy numbers are an important tool for quantifying the compactness of operators. Besides establishing new upper bounds on the entropy numbers of diagonal operators $D_\sigma$ from $\ell_p$ to $\ell_q$, where $p\not=q$, we investigate the optimality of these bounds. In case of $p q$ we show optimality under weaker assumption than previously used in the literature. In addition, we illustrate the benefit of our results with examples not covered in the literature so far.

math.FA

Sobolev Norm Learning Rates for Regularized Least-Squares Algorithm

Learning rates for least-squares regression are typically expressed in terms of $L_2$-norms. In this paper we extend these rates to norms stronger than the $L_2$-norm without requiring the regression function to be contained in the hypothesis space. In the special case of Sobolev reproducing kernel Hilbert spaces used as hypotheses spaces, these stronger norms coincide with fractional Sobolev norms between the used Sobolev space and $L_2$. As a consequence, not only the target function but also some of its derivatives can be estimated without changing the algorithm. From a technical point of view, we combine the well-known integral operator techniques with an embedding property, which so far has only been used in combination with empirical process arguments. This combination results in new finite sample bounds with respect to the stronger norms. From these finite sample bounds our rates easily follow. Finally, we prove the asymptotic optimality of our results in many cases.

stat.ML

Concurrent Imitation Dynamics in Congestion Games

Imitating successful behavior is a natural and frequently applied approach to trust in when facing scenarios for which we have little or no experience upon which we can base our decision. In this paper, we consider such behavior in atomic congestion games. We propose to study concurrent imitation dynamics that emerge when each player samples another player and possibly imitates this agents' strategy if the anticipated latency gain is sufficiently large. Our main focus is on convergence properties. Using a potential function argument, we show that our dynamics converge in a monotonic fashion to stable states. In such a state none of the players can improve its latency by imitating somebody else. As our main result, we show rapid convergence to approximate equilibria. At an approximate equilibrium only a small fraction of agents sustains a latency significantly above or below average. In particular, imitation dynamics behave like fully polynomial time approximation schemes (FPTAS). Fixing all other parameters, the convergence time depends only in a logarithmic fashion on the number of agents. Since imitation processes are not innovative they cannot discover unused strategies. Furthermore, strategies may become extinct with non-zero probability. For the case of singleton games, we show that the probability of this event occurring is negligible. Additionally, we prove that the social cost of a stable state reached by our dynamics is not much worse than an optimal state in singleton congestion games with linear latency function. Finally, we discuss how the protocol can be extended such that, in the long run, dynamics converge to a Nash equilibrium.

cs.GT