arXiv Science⌕ Search

arXiv subjects

Ankit Gupta

Publications and source records attributed to Ankit Gupta.

At least 73 records · Page 4Linked to original sources

Break It Down: A Question Understanding Benchmark

Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. In this work, we introduce a Question Decomposition Meaning Representation (QDMR) for questions. QDMR constitutes the ordered list of steps, expressed through natural language, that are necessary for answering a question. We develop a crowdsourcing pipeline, showing that quality QDMRs can be annotated at scale, and release the Break dataset, containing over 83K pairs of questions and their QDMRs. We demonstrate the utility of QDMR by showing that (a) it can be used to improve open-domain question answering on the HotpotQA dataset, (b) it can be deterministically converted to a pseudo-SQL formal language, which can alleviate annotation in semantic parsing applications. Last, we use Break to train a sequence-to-sequence model with copying that parses questions into QDMR structures, and show that it substantially outperforms several natural baselines.

cs.CL↗

A study of topological structures on equi-continuous mappings

Function space topologies are developed for EC(Y,Z), the class of equi-continuous mappings from a topological space Y to a uniform space Z. Properties such as splittingness, admissibility etc. are defined for such spaces. The net theoretic investigations are carried out to provide characterizations of splittingness and admissibility of function spaces on EC(Y,Z). The open-entourage topology and point-transitive-entourage topology are shown to be admissible and splitting respectively. Dual topologies are defined. A topology on EC(Y,Z) is found to be admissible (resp. splitting) if and only if its dual is so.

math.GM↗

A hidden integral structure endows Absolute Concentration Robust systems with resilience to dynamical concentration disturbances

Biochemical systems that express certain chemical species of interest at the same level at any positive equilibrium are called "absolute concentration robust" (ACR). These species behave in a stable, predictable way, in the sense that their expression is robust with respect to sudden changes in the species concentration, regardless the new positive equilibrium reached by the system. Such a property has been proven to be fundamentally important in certain gene regulatory networks and signaling systems. In the present paper, we mathematically prove that a well-known class of ACR systems studied by Shinar and Feinberg in 2010 hides an internal integral structure. This structure confers these systems with a higher degree of robustness that what was previously unknown. In particular, disturbances much more general than sudden changes in the species concentrations can be rejected, and robust perfect adaptation is achieved. Significantly, we show that these properties are maintained when the system is interconnected with other chemical reaction networks. This key feature enables design of insulator devices that are able to buffer the loading effect from downstream systems - a crucial requirement for modular circuit design in synthetic biology.

q-bio.SC↗

CAMLPAD: Cybersecurity Autonomous Machine Learning Platform for Anomaly Detection

As machine learning and cybersecurity continue to explode in the context of the digital ecosystem, the complexity of cybersecurity data combined with complicated and evasive machine learning algorithms leads to vast difficulties in designing an end to end system for intelligent, automatic anomaly classification. On the other hand, traditional systems use elementary statistics techniques and are often inaccurate, leading to weak centralized data analysis platforms. In this paper, we propose a novel system that addresses these two problems, titled CAMLPAD, for Cybersecurity Autonomous Machine Learning Platform for Anomaly Detection. The CAMLPAD systems streamlined, holistic approach begins with retrieving a multitude of different species of cybersecurity data in real time using elasticsearch, then running several machine learning algorithms, namely Isolation Forest, Histogram Based Outlier Score (HBOS), Cluster Based Local Outlier Factor (CBLOF), and K Means Clustering, to process the data. Next, the calculated anomalies are visualized using Kibana and are assigned an outlier score, which serves as an indicator for whether an alert should be sent to the system administrator that there are potential anomalies in the network. After comprehensive testing of our platform in a simulated environment, the CAMLPAD system achieved an adjusted rand score of 95 percent, exhibiting the reliable accuracy and precision of the system. All in all, the CAMLPAD system provides an accurate, streamlined approach to real time cybersecurity anomaly detection, delivering a novel solution that has the potential to revolutionize the cybersecurity sector.

cs.CR↗

User-Interactive Machine Learning Model for Identifying Structural Relationships of Code Features

Traditional machine learning based intelligent systems assist users by learning patterns in data and making recommendations. However, these systems are limited in that the user has little means of understanding the rationale behind the systems suggestions, communicating their own understanding of patterns, or correcting system behavior. In this project, we outline a model for intelligent software based on a human computer feedback loop. The Machine Learning (ML) systems recommendations are reviewed by the user, and in turn, this information shapes the systems decision making. Our model was applied to developing an HTML editor that integrates ML with user interaction to ascertain structural relationships between HTML document features and apply them for code completion. The editor utilizes the ID3 algorithm to build decision trees, sequences of rules for predicting code the user will type. The editor displays the decision trees rules in the Interactive Rules Interface System (IRIS), which allows developers to prioritize, modify, or delete them. These interactions alter the data processed by ID3, providing the developer some control over the autocomplete system. Validation indicates that, absent user interaction, the ML model is able to predict tags with 78.4 percent accuracy, attributes with 62.9 percent accuracy, and values with 12.8 percent accuracy. Based off of the results of the user study, user interaction with the rules interface corrects feature relationships missed or mistaken by the automated process, enhancing autocomplete accuracy and developer productivity. Additionally, interaction is proven to help developers work with greater awareness of code patterns. Our research demonstrates the viability of a software integration of machine intelligence with human feedback.

cs.HC↗

AquaSight: Automatic Water Impurity Detection Utilizing Convolutional Neural Networks

According to the United Nations World Water Assessment Programme, every day, 2 million tons of sewage and industrial and agricultural waste are discharged into the worlds water. In order to address this pervasive issue of increasing water pollution, while ensuring that the global population has an efficient, accurate, and low cost method to assess whether the water they drink is contaminated, we propose AquaSight, a novel mobile application that utilizes deep learning methods, specifically Convolutional Neural Networks, for automated water impurity detection. After comprehensive training with a dataset of 105 images representing varying magnitudes of contamination, the deep learning algorithm achieved a 96 percent accuracy and loss of 0.108. Furthermore, the machine learning model uses efficient analysis of the turbidity and transparency levels of water to estimate a particular sample of waters level of contamination. When deployed, the AquaSight system will provide an efficient way for individuals to secure an estimation of water quality, alerting local and national government to take action and potentially saving millions of lives worldwide.

cs.LG↗

StrokeSave: A Novel, High-Performance Mobile Application for Stroke Diagnosis using Deep Learning and Computer Vision

According to the WHO, Cerebrovascular Stroke, or CS, is the second largest cause of death worldwide. Current diagnosis of CS relies on labor and cost intensive neuroimaging techniques, unsuitable for areas with inadequate access to quality medical facilities. Thus, there is a great need for an efficient diagnosis alternative. StrokeSave is a platform for users to self-diagnose for prevalence to stroke. The mobile app is continuously updated with heart rate, blood pressure, and blood oxygen data from sensors on the patient wrist. Once these measurements reach a threshold for possible stroke, the patient takes facial images and vocal recordings to screen for paralysis attributed to stroke. A custom designed lens attached to a phone's camera then takes retinal images for the deep learning model to classify based on presence of retinopathy and sends a comprehensive diagnosis. The deep learning model, which consists of a RNN trained on 100 voice slurred audio files, a SVM trained on 410 vascular data points, and a CNN trained on 520 retinopathy images, achieved a holistic accuracy of 95.0 percent when validated on 327 samples. This value exceeds that of clinical examination accuracy, which is around 40 to 89 percent, further demonstrating the vital utility of such a medical device. Through this automated platform, users receive efficient, highly accurate diagnosis without professional medical assistance, revolutionizing medical diagnosis of CS and potentially saving millions of lives.

cs.CV↗

A finite state projection method for steady-state sensitivity analysis of stochastic reaction networks

Consider the standard stochastic reaction network model where the dynamics is given by a continuous-time Markov chain over a discrete lattice. For such models, estimation of parameter sensitivities is an important problem, but the existing computational approaches to solve this problem usually require time-consuming Monte Carlo simulations of the reaction dynamics. Therefore these simulation-based approaches can only be expected to work over finite time-intervals, while it is often of interest in applications to examine the sensitivity values at the steady-state after the Markov chain has relaxed to its stationary distribution. The aim of this paper is to present a computational method for the estimation of steady-state parameter sensitivities, which instead of using simulations, relies on the recently developed stationary Finite State Projection (sFSP) algorithm [J. Chem. Phys. 147, 154101 (2017)] that provides an accurate estimate of the stationary distribution at a fixed set of parameters. We show that sensitivity values at these parameters can be estimated from the solution of a Poisson equation associated with the infinitesimal generator of the Markov chain. We develop an approach to numerically solve the Poisson equation and this yields an efficient estimator for steady-state parameter sensitivities. We illustrate this method using several examples.

q-bio.QM↗

Sensitivity analysis for multiscale stochastic reaction networks using hybrid approximations

We consider the problem of estimating parameter sensitivities for stochastic models of multiscale reaction networks. These sensitivity values are important for model analysis, and, the methods that currently exist for sensitivity estimation mostly rely on simulations of the stochastic dynamics. This is problematic because these simulations become computationally infeasible for multiscale networks due to reactions firing at several different timescales. However it is often possible to exploit the multiscale property to derive a "model reduction" and approximate the dynamics as a Piecewise Deterministic Markov process (PDMP), which is a hybrid process consisting of both discrete and continuous components. The aim of this paper is to show that such PDMP approximations can be used to accurately and efficiently estimate the parameter sensitivity for the original multiscale stochastic model. We prove the convergence of the original sensitivity to the corresponding PDMP sensitivity, in the limit where the PDMP approximation becomes exact. Moreover we establish a representation of the PDMP parameter sensitivity that separates the contributions of discrete and continuous components in the dynamics, and allows one to efficiently estimate both contributions.

math.PR↗

Computational identification of irreducible state-spaces for stochastic reaction networks

Stochastic models of reaction networks are becoming increasingly important in Systems Biology. In these models, the dynamics is generally represented by a continuous-time Markov chain whose states denote the copy-numbers of the constituent species. The state-space on which this process resides is a subset of non-negative integer lattice and for many examples of interest, this state-space is countably infinite. This causes numerous problems in analyzing the Markov chain and understanding its long-term behavior. These problems are further confounded by the presence of conservation relations among species which constrain the dynamics in complicated ways. In this paper we provide a linear-algebraic procedure to disentangle these conservation relations and represent the state-space in a special decomposed form, based on the copy-number ranges of various species and dependencies among them. This decomposed form is advantageous for analyzing the stochastic model and for a large class of networks we demonstrate how this form can be used for finding all the closed communication classes for the Markov chain within the infinite state-space. Such communication classes are irreducible state-spaces for the dynamics and they support all the extremal stationary distributions for the Markov chain. Hence our results provide important insights into the long-term behavior and stability properties of stochastic models of reaction networks. We discuss how the knowledge of these irreducible state-spaces can be used in many ways such as speeding-up stochastic simulations of multiscale networks or in identifying the stationary distributions of complex-balanced networks. We illustrate our results with several examples of gene-expression networks from Systems Biology.

math.PR↗

Estimation of parameter sensitivities for stochastic reaction networks using tau-leap simulations

We consider the important problem of estimating parameter sensitivities for stochastic models of reaction networks that describe the dynamics as a continuous-time Markov process over a discrete lattice. These sensitivity values are useful for understanding network properties, validating their design and identifying the pivotal model parameters. Many methods for sensitivity estimation have been developed, but their computational feasibility suffers from the critical bottleneck of requiring time-consuming Monte Carlo simulations of the exact reaction dynamics. To circumvent this problem one needs to devise methods that speed up the computations while suffering acceptable and quantifiable loss of accuracy. We develop such a method by first deriving a novel integral representation of parameter sensitivity and then demonstrating that this integral may be approximated by any convergent tau-leap method. Our method is easy to implement, works with any tau-leap simulation scheme and its accuracy is proved to be similar to that of the underlying tau-leap scheme. We demonstrate the efficiency of our methods through numerical examples. We also compare our method with the tau-leap versions of certain finite-difference schemes that are commonly used for sensitivity estimations.

math.PR↗

Variance reduction for antithetic integral control of stochastic reaction networks

The antithetic integral feedback motif recently introduced in Briat, Gupta & Khammash (Cell Systems, 2017) is known to ensure robust perfect adaptation for the mean dynamics of a given molecular species involved in a complex stochastic biomolecular reaction network. However, it was observed that it also leads to a higher variance in the controlled network than that obtained when using a constitutive (i.e. open-loop) control strategy. This was interpreted as the cost of the adaptation property and may be viewed as a performance deterioration for the overall controlled network. To decrease this variance and improve the performance, we propose to combine the antithetic integral feedback motif with a negative feedback strategy. Both theoretical and numerical results are obtained. The theoretical ones are based on a tailored moment closure method allowing one to obtain approximate expressions for the stationary variance for the controlled network and predict that the variance can indeed be decreased by increasing the strength of the negative feedback. Numerical results verify the accuracy of this approximation and show that the controlled species variance can indeed be decreased, sometimes below its constitutive level. Three molecular networks are considered in order to verify the wide applicability of two types of negative feedback strategies. The main conclusion is that there is a trade-off between the speed of the settling-time of the mean trajectories and the stationary variance of the controlled species; i.e. smaller variance is associated with larger settling-time.

math.OC↗

Dilated Convolutions for Modeling Long-Distance Genomic Dependencies

We consider the task of detecting regulatory elements in the human genome directly from raw DNA. Past work has focused on small snippets of DNA, making it difficult to model long-distance dependencies that arise from DNA's 3-dimensional conformation. In order to study long-distance dependencies, we develop and release a novel dataset for a larger-context modeling task. Using this new data set we model long-distance interactions using dilated convolutional neural networks, and compare them to standard convolutions and recurrent neural networks. We show that dilated convolutions are effective at modeling the locations of regulatory markers in the human genome, such as transcription factor binding sites, histone modifications, and DNAse hypersensitivity sites.

q-bio.GN↗

A finite state projection algorithm for the stationary solution of the chemical master equation

The chemical master equation (CME) is frequently used in systems biology to quantify the effects of stochastic fluctuations that arise due to biomolecular species with low copy numbers. The CME is a system of ordinary differential equations that describes the evolution of probability density for each population vector in the state-space of the stochastic reaction dynamics. For many examples of interest, this state-space is infinite, making it difficult to obtain exact solutions of the CME. To deal with this problem, the Finite State Projection (FSP) algorithm was developed by Munsky and Khammash (Jour. Chem. Phys. 2006), to provide approximate solutions to the CME by truncating the state-space. The FSP works well for finite time-periods but it cannot be used for estimating the stationary solutions of CMEs, which are often of interest in systems biology. The aim of this paper is to develop a version of FSP which we refer to as the stationary FSP (sFSP) that allows one to obtain accurate approximations of the stationary solutions of a CME by solving a finite linear-algebraic system that yields the stationary distribution of a continuous-time Markov chain over the truncated state-space. We derive bounds for the approximation error incurred by sFSP and we establish that under certain stability conditions, these errors can be made arbitrarily small by appropriately expanding the truncated state-space. We provide several examples to illustrate our sFSP method and demonstrate its efficiency in estimating the stationary distributions. In particular, we show that using a quantised tensor train (QTT) implementation of our sFSP method, problems admitting more than 100 million states can be efficiently solved.

q-bio.QM↗

Dynamic disorder in simple enzymatic reactions induces stochastic amplification of substrate

A growing amount of evidence points to the fact that many enzymes exhibit fluctuations in their catalytic activity, which are associated with conformational changes on a broad range of timescales. The experimental study of this phenomenon, termed dynamic disorder, has become possible due to advances in single-molecule enzymology measurement techniques, through which the catalytic activity of individual enzyme molecules can be tracked in time. The biological role and importance of these fluctuations in a system with a small number of enzymes such as a living cell have only recently started being explored. In this work, we examine a simple stochastic reaction system consisting of an inflowing substrate and an enzyme with a randomly fluctuating catalytic reaction rate that converts the substrate into an outflowing product. To describe analytically the effect of rate fluctuations on the average substrate abundance at steady-state, we derive an explicit formula that connects the relative speed of enzymatic fluctuations with the mean substrate level. We demonstrate that the relative speed of rate fluctuations can have a dramatic effect on the mean substrate, and lead to large positive deviations from predictions based on the assumption of deterministic enzyme activity. Our results also establish an interesting connection between the amplification effect and the mixing properties of the Markov process describing the enzymatic activity fluctuations, which can be used to easily predict the fluctuation speed above which such deviations become negligible. As the techniques of single-molecule enzymology continuously evolve, it may soon be possible to study the stochastic phenomena due to enzymatic activity fluctuations within living cells. Our work can be used to formulate experimentally testable hypotheses regarding the magnitude of these fluctuations, as well as their phenotypic consequences.

q-bio.QM↗

Grain boundary diffusion in severely deformed Al-based alloy

Grain boundary diffusion in severely deformed Al-based AA5024 alloy is investigated. Different states are prepared by combination of equal channel angular processing and heat treatments, with the radioisotope $^{57}$Co being employed as a sensitive probe of a given grain boundary state. Its diffusion rates near room temperature (320~K) are utilized to quantify the effects of severe plastic deformation and a presumed formation of a previously reported deformation-modified state of grain boundaries, solute segregation at the interfaces, increased dislocation content after deformation and of the precipitation behavior on the transport phenomena along grain boundaries. The dominant effect of nano-sized Al$_3$Sc-based precipitates is evaluated using density functional theory and the Eshelby model for the determination of elastic stresses around the precipitates.

cond-mat.mtrl-sci↗

Low temperature features in the heat capacity of unary metals and intermetallics for the example of bulk aluminum and Al$_3$Sc

We explore the competition and coupling of vibrational and electronic contributions to the heat capacity of Al and Al$_3$Sc at temperatures below 50 K combining experimental calorimetry with highly converged finite temperature density functional theory calculations. We find that semilocal exchange correlation functionals accurately describe the rich feature set observed for these temperatures, including electron-phonon coupling. Using different representations of the heat capacity, we are therefore able to identify and explain deviations from the Debye behaviour in the low-temperature limit and in the temperature regime 30 - 50 K as well as the reduction of these features due to the addition of Sc.

cond-mat.mtrl-sci↗

Antithetic Integral Feedback ensures robust perfect adaptation in noisy biomolecular networks

Homeostasis is a running theme in biology. Often achieved through feedback regulation strategies, homeostasis allows living cells to control their internal environment as a means for surviving changing and unfavourable environments. While many endogenous homeostatic motifs have been studied in living cells, some other motifs may remain under-explored or even undiscovered. At the same time, known regulatory motifs have been mostly analyzed at the deterministic level, and the effect of noise on their regulatory function has received low attention. Here we lay the foundation for a regulation theory at the molecular level that explicitly takes into account the noisy nature of biochemical reactions and provides novel tools for the analysis and design of robust homeostatic circuits. Using these ideas, we propose a new regulation motif, which we refer to as {\em antithetic integral feedback, and demonstrate its effectiveness as a strategy for generically regulating a wide class of reaction networks. By combining tools from probability and control theory, we show that the proposed motif preserves the stability of the overall network, steers the population of any regulated species to a desired set point, and achieves robust perfect adaptation -- all with low prior knowledge of reaction rates. Moreover, our proposed regulatory motif can be implemented using a very small number of molecules and hence has a negligible metabolic load. Strikingly, the regulatory motif exploits stochastic noise, leading to enhanced regulation in scenarios where noise-free implementations result in dysregulation. Finally, we discuss the possible manifestation of the proposed antithetic integral feedback motif in endogenous biological circuits and its realization in synthetic circuits.

math.OC↗