arXiv ScienceSearch

arXiv subjects

Stephen Quake

Publications and source records attributed to Stephen Quake.

5 recordsLinked to original sources

Informational blueprints reveal condition-dependent gene regulatory architectures

While coding regions in the genome have a direct interpretation in terms of protein products, significant fractions are non-coding and yet control essential biological functions. Unlike the genetic code, there is no "lookup table" that identifies where regulatory proteins, known as transcription factors (TFs), bind. Here, we extract these binding sites by distilling sequences of nucleotide letters into collective coordinates (hyperletters) representing the binding sites that are active under specific environmental conditions. Going beyond local information footprints between individual bases and expression levels, our $\textit{information blueprint}$ algorithm compresses the global information by optimising filters that simultaneously scan an entire promoter sequence. Inspired by renormalisation-group techniques, we identify TF binding sites as coarse-grained variables combining groups of correlated mutations with the highest collective impact on gene expression. We validate our approach on experimental data for $\textit{E. coli}$ and discover novel regulatory elements illustrating its deployment at scale across growth conditions.

q-bio.GN

The Environment-Dependent Regulatory Landscape of the E. coli Genome

All cells respond to changes in both their internal milieu and the environment around them through the regulation of their genes. Despite decades of effort, there remain huge gaps in our knowledge of both the function of many genes (the so-called y-ome) and how they adapt to changing environments via regulation. Here we describe a joint experimental and theoretical dissection of the regulation of a broad array of over 100 biologically interesting genes in E. coli across 39 diverse environments, permitting us to discover the binding sites and transcription factors that mediate regulatory control. Using a combination of mutagenesis, massively parallel reporter assays, mass spectrometry and tools from information theory and statistical physics, we go from complete ignorance of a promoter's environment-dependent regulatory architecture to predictive models of its behavior. As a proof of principle of the biological insights to be gained from such a study, we chose a combination of genes from the y-ome, toxin-antitoxin pairs, and genes hypothesized to be part of regulatory modules; in all cases, we discovered a host of new insights into their underlying regulatory landscape and resulting biological function.

q-bio.GN

Corvo: Visualizing CellxGene Single-Cell Datasets in Virtual Reality

The CellxGene project has enabled access to single-cell data in the scientific community, providing tools for browsed-based no-code analysis of more than 500 annotated datasets. However, single-cell data requires dimensional reduction to visualize, and 2D embedding does not take full advantage of three-dimensional human spatial understanding and cognition. Compared to a 2D visualization that could potentially hide gene expression patterns, 3D Virtual Reality may enable researchers to make better use of the information contained within the datasets. For this purpose, we present \emph{Corvo}, a fully free and open-source software tool that takes the visualization and analysis of CellxGene single-cell datasets to 3D Virtual Reality. Similar to CellxGene, Corvo takes a no-code approach for the end user, but also offers multimodal user input to facilitate fast navigation and analysis, and is interoperable with the existing Python data science ecosystem. In this paper, we explain the design goals of Corvo, detail its approach to the Virtual Reality visualization and analysis of single-cell data, and briefly discuss limitations and future extensions.

cs.HC

Solving the Tyranny of Pipetting

Stephen Quake is the Lee Otterson Professor of Bioengineering and Applied Physics at Stanford University. Here he reviews the early history of microfluidics and discusses more recent developments, with a focus on applications in biology and biochemistry.

physics.flu-dyn

Designing the statistically optimal drug for cancer therapy

Cancer and healthy cells have distinct distributions of molecular properties and thus respond differently to drugs. Cancer drugs ideally kill cancer cells while limiting harm to healthy cells. However, the inherent variance among cells in both cancer and healthy cell populations increases the difficulty of selective drug action. Here we propose a classification framework based on the idea that an ideal cancer drug should maximally discriminate between cancer and healthy cells. We first explore how molecular markers can be used to discriminate cancer cells from healthy cells on a single cell basis, and then how the effects of drugs are statistically predicted by these molecular markers. We then combine these two ideas to show how to optimally match drugs to tumor cells. We find that expression levels of a handful of genes suffice to discriminate well between individual cells in cancer and healthy tissue. We also find that gene expression predicts the efficacy of cancer drugs, suggesting that the cancer drugs act as classifiers using gene profiles. In agreement with our first finding, a small number of genes predict drug efficacy well. Finally, we formulate a framework that defines an optimal drug, and predicts drug cocktails that may target cancer more accurately than the individual drugs alone. Conceptualizing cancer drugs as solving a discrimination problem in the high-dimensional space of molecular markers promises to inform the design of new cancer drugs and drug cocktails.

q-bio.OT