arXiv ScienceSearch

arXiv subjects

George Perry

Publications and source records attributed to George Perry.

2 recordsLinked to original sources

A Survey of Source Code Representations for Machine Learning-Based Cybersecurity Tasks

Machine learning techniques for cybersecurity-related software engineering tasks are becoming increasingly popular. The representation of source code is a key portion of the technique that can impact the way the model is able to learn the features of the source code. With an increasing number of these techniques being developed, it is valuable to see the current state of the field to better understand what exists and what is not there yet. This article presents a study of these existing machine learning based approaches and demonstrates what type of representations were used for different cybersecurity tasks and programming languages. Additionally, we study what types of models are used with different representations. We have found that graph-based representations are the most popular category of representation, and tokenizers and Abstract Syntax Trees (ASTs) are the two most popular representations overall (e.g., AST and tokenizers are the representations with the highest count of papers, whereas graph-based representations is the category with the highest count of papers). We also found that the most popular cybersecurity task is vulnerability detection, and the language that is covered by the most techniques is C. Finally, we found that sequence-based models are the most popular category of models, and Support Vector Machines are the most popular model overall.

cs.LG

Evolutionary mismatch and the role of GxE interactions in human disease

Globally, we are witnessing the rise of complex, non-communicable diseases (NCDs) related to changes in our daily environments. Obesity, asthma, cardiovascular disease, and type 2 diabetes are part of a long list of "lifestyle" diseases that were rare throughout human history but are now common. A key idea from anthropology and evolutionary biology--the evolutionary mismatch hypothesis--seeks to explain this phenomenon. It posits that humans evolved in environments that radically differ from the ones experienced by most people today, and thus traits that were advantageous in past environments may now be "mismatched" and disease-causing. This hypothesis is, at its core, a genetic one: it predicts that loci with a history of selection will exhibit "genotype by environment" (GxE) interactions and have differential health effects in ancestral versus modern environments. Here, we discuss how this concept could be leveraged to uncover the genetic architecture of NCDs in a principled way. Specifically, we advocate for partnering with small-scale, subsistence-level groups that are currently transitioning from environments that are arguably more "matched" with their recent evolutionary history to those that are more "mismatched". These populations provide diverse genetic backgrounds as well as the needed levels and types of environmental variation necessary for mapping GxE interactions in an explicit mismatch framework. Such work would make important contributions to our understanding of environmental and genetic risk factors for NCDs across diverse ancestries and sociocultural contexts.

q-bio.PE