arXiv ScienceSearch

arXiv subjects

Archana Mathur

Publications and source records attributed to Archana Mathur.

12 recordsLinked to original sources

A Granger-Causal Perspective on Gradient Descent with Application to Pruning

Stochastic Gradient Descent (SGD) is the main approach to optimizing neural networks. Several generalization properties of deep networks, such as convergence to a flatter minima, are believed to arise from SGD. This article explores the causality aspect of gradient descent. Specifically, we show that the gradient descent procedure has an implicit granger-causal relationship between the reduction in loss and a change in parameters. By suitable modifications, we make this causal relationship explicit. A causal approach to gradient descent has many significant applications which allow greater control. In this article, we illustrate the significance of the causal approach using the application of Pruning. The causal approach to pruning has several interesting properties - (i) We observe a phase shift as the percentage of pruned parameters increase. Such phase shift is indicative of an optimal pruning strategy. (ii) After pruning, we see that minima becomes "flatter", explaining the increase in accuracy after pruning weights.

cs.LG

DeliverAI: Reinforcement Learning Based Distributed Path-Sharing Network for Food Deliveries

Delivery of items from the producer to the consumer has experienced significant growth over the past decade and has been greatly fueled by the recent pandemic. Amazon Fresh, Shopify, UberEats, InstaCart, and DoorDash are rapidly growing and are sharing the same business model of consumer items or food delivery. Existing food delivery methods are sub-optimal because each delivery is individually optimized to go directly from the producer to the consumer via the shortest time path. We observe a significant scope for reducing the costs associated with completing deliveries under the current model. We model our food delivery problem as a multi-objective optimization, where consumer satisfaction and delivery costs, both, need to be optimized. Taking inspiration from the success of ride-sharing in the taxi industry, we propose DeliverAI - a reinforcement learning-based path-sharing algorithm. Unlike previous attempts for path-sharing, DeliverAI can provide real-time, time-efficient decision-making using a Reinforcement learning-enabled agent system. Our novel agent interaction scheme leverages path-sharing among deliveries to reduce the total distance traveled while keeping the delivery completion time under check. We generate and test our methodology vigorously on a simulation setup using real data from the city of Chicago. Our results show that DeliverAI can reduce the delivery fleet size by 12\%, the distance traveled by 13%, and achieve 50% higher fleet utilization compared to the baselines.

cs.LG

A novel RNA pseudouridine site prediction model using Utility Kernel and data-driven parameters

RNA protein Interactions (RPIs) play an important role in biological systems. Recently, we have enumerated the RPIs at the residue level and have elucidated the minimum structural unit (MSU) in these interactions to be a stretch of five residues (Nucleotides/amino acids). Pseudouridine is the most frequent modification in RNA. The conversion of uridine to pseudouridine involves interactions between pseudouridine synthase and RNA. The existing models to predict the pseudouridine sites in a given RNA sequence mainly depend on user-defined features such as mono and dinucleotide composition/propensities of RNA sequences. Predicting pseudouridine sites is a non-linear classification problem with limited data points. Deep Learning models are efficient discriminators when the data set size is reasonably large and fail when there is a paucity of data ($<1000$ samples). To mitigate this problem, we propose a Support Vector Machine (SVM) Kernel based on utility theory from Economics, and using data-driven parameters (i.e. MSU) as features. For this purpose, we have used position-specific tri/quad/pentanucleotide composition/propensity (PSPC/PSPP) besides nucleotide and dineculeotide composition as features. SVMs are known to work well in small data regimes and kernels in SVM are designed to classify non-linear data. The proposed model outperforms the existing state-of-the-art models significantly (10%-15% on average).

q-bio.BM

To prune or not to prune : A chaos-causality approach to principled pruning of dense neural networks

Reducing the size of a neural network (pruning) by removing weights without impacting its performance is an important problem for resource-constrained devices. In the past, pruning was typically accomplished by ranking or penalizing weights based on criteria like magnitude and removing low-ranked weights before retraining the remaining ones. Pruning strategies may also involve removing neurons from the network in order to achieve the desired reduction in network size. We formulate pruning as an optimization problem with the objective of minimizing misclassifications by selecting specific weights. To accomplish this, we have introduced the concept of chaos in learning (Lyapunov exponents) via weight updates and exploiting causality to identify the causal weights responsible for misclassification. Such a pruned network maintains the original performance and retains feature explainability.

cs.LG

Quantifying the Classification of Exoplanets: in Search for the Right Habitability Metric

What is habitability? Can we quantify it? What do we mean under the term habitable or potentially habitable planet? With estimates of the number of planets in our Galaxy alone running into billions, possibly a number greater than the number of stars, it is high time to start characterizing them, sorting them into classes/types just like stars, to better understand their formation paths, their properties and, ultimately, their ability to beget or sustain life. After all, we do have life thriving on one of these billions of planets, why not on others? Which planets are better suited for life and which ones are definitely not worth spending expensive telescope time on? We need to find sort of quick assessment score, a metric, using which we can make a list of promising planets and dedicate our efforts to them. Exoplanetary habitability is a transdisciplinary subject integrating astrophysics, astrobiology, planetary science, even terrestrial environmental sciences. We review the existing metrics of habitability and the new classification schemes of extrasolar planets and provide an exposition of the use of computational intelligence techniques to evaluate habitability scores and to automate the process of classification of exoplanets. We examine how solving convex optimization techniques, as in computing new metrics such as CDHS and CEESA, cross-validates ML-based classification of exoplanets. Despite the recent criticism of exoplanetary habitability ranking, this field has to continue and evolve to use all available machinery of astroinformatics, artificial intelligence and machine learning. It might actually develop into a sort of same scale as stellar types in astronomy, to be used as a quick tool of screening exoplanets in important characteristics in search for potentially habitable planets for detailed follow-up targets.

astro-ph.EP

Evolution of Novel Activation Functions in Neural Network Training with Applications to Classification of Exoplanets

We present analytical exploration of novel activation functions as consequence of integration of several ideas leading to implementation and subsequent use in habitability classification of exoplanets. Neural networks, although a powerful engine in supervised methods, often require expensive tuning efforts for optimized performance. Habitability classes are hard to discriminate, especially when attributes used as hard markers of separation are removed from the data set. The solution is approached from the point of investigating analytical properties of the proposed activation functions. The theory of ordinary differential equations and fixed point are exploited to justify the "lack of tuning efforts" to achieve optimal performance compared to traditional activation functions. Additionally, the relationship between the proposed activation functions and the more popular ones is established through extensive analytical and empirical evidence. Finally, the activation functions have been implemented in plain vanilla feed-forward neural network to classify exoplanets.

astro-ph.IM

SBAF: A New Activation Function for Artificial Neural Net based Habitability Classification

We explore the efficacy of using a novel activation function in Artificial Neural Networks (ANN) in characterizing exoplanets into different classes. We call this Saha-Bora Activation Function (SBAF) as the motivation is derived from long standing understanding of using advanced calculus in modeling habitability score of Exoplanets. The function is demonstrated to possess nice analytical properties and doesn't seem to suffer from local oscillation problems. The manuscript presents the analytical properties of the activation function and the architecture implemented on the function. Keywords: Astroinformatics, Machine Learning, Exoplanets, ANN, Activation Function.

cs.LG

Time Reversed Delay Differential Equation Based Modeling Of Journal Influence In An Emerging Area

A recent independent study resulted in a ranking system which ranked Astronomy and Computing (ASCOM) much higher than most of the older journals highlighting its niche prominence. We investigate the notable ascendancy in reputation of ASCOM by proposing a novel differential equation based modeling. The modeling is a consequence of knowledge discovery from big data methods, namely L1-SVD. We propose a growth model by accounting for the behavior of parameters that contribute to the growth of a field. It is worthwhile to spend some time in analyzing the cause and control variables behind rapid rise in the reputation of a journal in a niche area. We intend to identify and probe the parameters responsible for its growing influence. Delay differential equations are used to model the change of influence on a journal's status by exploiting the effects of historical data. The manuscript justifies the use of implicit control variables and models those accordingly that demonstrate certain behavior in the journal influence.

cs.DL

Use of NoSQL database and visualization techniques to analyze massive scholarly article data from journals

Visualization of the massive data is a challenging endeavor. Extracting data and providing graphical representations can aid in its effective utilization in terms of interpretation and knowledge discovery. Publishing research articles has become a way of life for academicians. The scholarly publications can shape-up the professional growth of authors and also expand the research and technological growth of a country, continent and other demographic regions. Scholarly articles have grown in gigantic numbers that are published in different domains by various journals. Information related to articles, authors, their affiliations, number of citations, country, publisher, references and other information is like a gold mine for statisticians and data analysts. This data when used skillfully, via visual analysis tool, can provide valuable understanding and can aid in deeper exposition for researchers working in domains like scientometrics and bibliometrics. Since the data is not readily available, we used Google scholar, a comprehensive and free repository of scholarly articles, as data source for our study. Data was scraped from Google scholar and stored as a graph and later visualized in the form of nodes and its relationships, which offered discerning and concealed information of growing impact of articles, journals and authors in their domains. Not only this, evident domain shift of an author, various research domains spread for an author, predicting emerging domain and subdomains, detecting cartel behavior at Journal and author-level was also depicted by graphical analysis. Neo4j graph database was used in the background to help store the data in structured manner.

cs.DL

Model Visualization in understanding rapid growth of a journal in an emerging area

A recent independent study resulted in a ranking system which ranked Astronomy and Computing (ASCOM) much higher than most of the older journals highlighting the niche prominence of the particular journal. We investigate the remarkable ascendancy in reputation of ASCOM by proposing a novel differential equation based modeling. The Modeling is a consequence of knowledge discovery from big data-centric methods, namely L1-SVD. The inadequacy of the ranking method in explaining the reason behind the growth in reputation of ASCOM is reasonable to understand given that the study was post-facto. Thus, we propose a growth model by accounting for the behavior of parameters that contribute to the growth of a field. It is worthwhile to spend some time in analysing the cause and control variables behind rapid rise in reputation of a journal in a niche area. We intent to probe and bring out parameters responsible for its growing influence. Delay differential equations are used to model the change of influence on a journal's status by exploiting the effects of historical data.

cs.DL

ScientoBASE: A Framework and Model for Computing Scholastic Indicators of non-local influence of Journals via Native Data Acquisition algorithms

Defining and measuring internationality as a function of influence diffusion of scientific journals is an open problem. There exists no metric to rank journals based on the extent or scale of internationality. Measuring internationality is qualitative, vague, open to interpretation and is limited by vested interests. With the tremendous increase in the number of journals in various fields and the unflinching desire of academics across the globe to publish in "international" journals, it has become an absolute necessity to evaluate, rank and categorize journals based on internationality. Authors, in the current work have defined internationality as a measure of influence that transcends across geographic boundaries. There are concerns raised by the authors about unethical practices reflected in the process of journal publication whereby scholarly influence of a select few are artificially boosted, primarily by resorting to editorial maneuvres. To counter the impact of such tactics, authors have come up with a new method that defines and measures internationality by eliminating such local effects when computing the influence of journals. A new metric, Non-Local Influence Quotient(NLIQ) is proposed as one such parameter for internationality computation along with another novel metric, Other-Citation Quotient as the complement of the ratio of self-citation and total citation. In addition, SNIP and International Collaboration Ratio are used as two other parameters.

cs.DL

DSRS: Estimation and Forecasting of Journal Influence in the Science and Technology Domain via a Lightweight Quantitative Approach

The evaluation of journals based on their influence is of interest for numerous reasons. Various methods of computing a score have been proposed for measuring the scientific influence of scholarly journals. Typically the computation of any of these scores involves compiling the citation information pertaining to the journal under consideration. This involves significant overhead since the article citation information of not only the journal under consideration but also that of other journals for the recent few years need to be stored. Our work is motivated by the idea of developing a computationally lightweight approach that does not require any data storage, yet yields a score which is useful for measuring the importance of journals. In this paper, a regression analysis based method is proposed to calculate Journal Influence Score. Proposed model is validated using historical data from the SCImago portal. The results show that the error is small between rankings obtained using the proposed method and the SCImago Journal Rank, thus proving that the proposed approach is a feasible and effective method of calculating scientific impact of journals.

cs.DL