arXiv ScienceSearch

arXiv subjects

Alex Gray

Publications and source records attributed to Alex Gray.

7 recordsLinked to original sources

Multilinear and Linear Programs for Partially Identifiable Queries in Quasi-Markovian Structural Causal Models

We investigate partially identifiable queries in a class of causal models. We focus on acyclic Structural Causal Models that are quasi-Markovian (that is, each endogenous variable is connected with at most one exogenous confounder). We look into scenarios where endogenous variables are observed (and a distribution over them is known), while exogenous variables are not fully specified. This leads to a representation that is in essence a Bayesian network where the distribution of root variables is not uniquely determined. In such circumstances, it may not be possible to precisely compute a probability value of interest. We thus study the computation of tight probability bounds, a problem that has been solved by multilinear programming in general, and by linear programming when a single confounded component is intervened upon. We present a new algorithm to simplify the construction of such programs by exploiting input probabilities over endogenous variables. For scenarios with a single intervention, we apply column generation to compute a probability bound through a sequence of auxiliary linear integer programs, thus showing that a representation with polynomial cardinality for exogenous variables is possible. Experiments show column generation techniques to be superior to existing methods.

cs.AI

A Comparison of Neuroelectrophysiology Databases

As data sharing has become more prevalent, three pillars - archives, standards, and analysis tools - have emerged as critical components in facilitating effective data sharing and collaboration. This paper compares four freely available intracranial neuroelectrophysiology data repositories: Data Archive for the BRAIN Initiative (DABI), Distributed Archives for Neurophysiology Data Integration (DANDI), OpenNeuro, and Brain-CODE. The aim of this review is to describe archives that provide researchers with tools to store, share, and reanalyze both human and non-human neurophysiology data based on criteria that are of interest to the neuroscientific community. The Brain Imaging Data Structure (BIDS) and Neurodata Without Borders (NWB) are utilized by these archives to make data more accessible to researchers by implementing a common standard. As the necessity for integrating large-scale analysis into data repository platforms continues to grow within the neuroscientific community, this article will highlight the various analytical and customizable tools developed within the chosen archives that may advance the field of neuroinformatics.

q-bio.QM

Towards a Unification of Logic and Information Theory

Today, the vast majority of the world's digital information is represented using the fundamental assumption, introduced by Claude Shannon in 1948, that ``...the semantic aspects of communication are irrelevant to the engineering problem (of the design of communication systems)...''. Consider, nonetheless, the observation that we often combine a message with other information in order to deduce new facts, thereby expanding the value of such a message. It is noteworthy that to-date, no rigorous theory of communication has been put forth which postulates the existence of deductive capabilities on the receiver's side. The purpose of this paper is to present such a theory. We formally model such deductive capabilities using logic reasoning, and present a rigorous theory which covers the following generic scenario: Alice and Bob each have knowledge of some logic sentence, and they wish to communicate as efficiently as possible with the shared goal that, following their communication, Bob should be able to deduce a particular logic sentence that Alice knows to be true, but that Bob currently cannot prove. Many variants of this general setup are considered in this article; in all cases we are able to provide sharp upper and lower bounds. Our contribution includes the identification of the most fundamental requirements that we place on a logic and associated logical language for all of our results to apply. Practical algorithms that are in some cases asymptotically optimal are provided, and we illustrate the potential practical value of the design of communication systems that incorporate the assumption of deductive capabilities at the receiver using experimental results that suggest significant possible gains compared to classical systems.

cs.IT

Introduction to astroML: Machine Learning for Astrophysics

Astronomy and astrophysics are witnessing dramatic increases in data volume as detectors, telescopes and computers become ever more powerful. During the last decade, sky surveys across the electromagnetic spectrum have collected hundreds of terabytes of astronomical data for hundreds of millions of sources. Over the next decade, the data volume will enter the petabyte domain, and provide accurate measurements for billions of sources. Astronomy and physics students are not traditionally trained to handle such voluminous and complex data sets. In this paper we describe astroML; an initiative, based on Python and scikit-learn, to develop a compendium of machine learning tools designed to address the statistical needs of the next generation of students and astronomical surveys. We introduce astroML and present a number of example applications that are enabled by this package.

astro-ph.IM

Detection of Cosmic Magnification with the Sloan Digital Sky Survey

We present an 8 sigma detection of cosmic magnification measured by the variation of quasar density due to gravitational lensing by foreground large scale structure. To make this measurement we used 3800 square degrees of photometric observations from the Sloan Digital Sky Survey (SDSS) containing \~200,000 quasars and 13 million galaxies. Our measurement of the galaxy-quasar cross-correlation function exhibits the amplitude, angular dependence and change in sign as a function of the slope of the observed quasar number counts that is expected from magnification bias due to weak gravitational lensing. We show that observational uncertainties (stellar contamination, Galactic dust extinction, seeing variations and errors in the photometric redshifts) are well controlled and do not significantly affect the lensing signal. By weighting the quasars with the number count slope, we combine the cross-correlation of quasars for our full magnitude range and detect the lensing signal at >4 sigma in all five SDSS filters. Our measurements of cosmic magnification probe scales ranging from 60 kpc/h to 10 Mpc/h and are in good agreement with theoretical predictions based on the WMAP concordance cosmology. As with galaxy-galaxy lensing, future measurements of cosmic magnification will provide useful constraints on the galaxy-mass power spectrum.

astro-ph

Galaxy ecology: groups and low-density environments in the SDSS and 2dFGRS

We analyse the observed correlation between galaxy environment and H-alpha emission line strength, using volume-limited samples and group catalogues of 24968 galaxies drawn from the 2dF Galaxy Redshift Survey (Mb<-19.5) and the Sloan Digital Sky Survey (Mr<-20.6). We characterise the environment by 1) Sigma_5, the surface number density of galaxies determined by the projected distance to the 5th nearest neighbour; and 2) rho1.1 and rho5.5, three-dimensional density estimates obtained by convolving the galaxy distribution with Gaussian kernels of dispersion 1.1 Mpc and 5.5 Mpc, respectively. We find that star-forming and quiescent galaxies form two distinct populations, as characterised by their H-alpha equivalent width, EW(Ha). The relative numbers of star-forming and quiescent galaxies varies strongly and continuously with local density. However, the distribution of EW(Ha) amongst the star-forming population is independent of environment. The fraction of star-forming galaxies shows strong sensitivity to the density on large scales, rho5.5, which is likely independent of the trend with local density, rho1.1. We use two differently-selected group catalogues to demonstrate that the correlation with galaxy density is approximately independent of group velocity dispersion, for sigma=200-1000 km/s. Even in the lowest density environments, no more than ~70 per cent of galaxies show significant H-alpha emission. Based on these results, we conclude that the present-day correlation between star formation rate and environment is a result of short-timescale mechanisms that take place preferentially at high redshift, such as starbursts induced by galaxy-galaxy interactions.

astro-ph

Fast Algorithms and Efficient Statistics: N-point Correlation Functions

We present here a new algorithm for the fast computation of N-point correlation functions in large astronomical data sets. The algorithm is based on kdtrees which are decorated with cached sufficient statistics thus allowing for orders of magnitude speed-ups over the naive non-tree-based implementation of correlation functions. We further discuss the use of controlled approximations within the computation which allows for further acceleration. In summary, our algorithm now makes it possible to compute exact, all-pairs, measurements of the 2, 3 and 4-point correlation functions for cosmological data sets like the Sloan Digital Sky Survey (SDSS; York et al. 2000) and the next generation of Cosmic Microwave Background experiments (see Szapudi et al. 2000).

astro-ph