arXiv ScienceSearch

arXiv subjects

Daria Stepanova

Publications and source records attributed to Daria Stepanova.

13 recordsLinked to original sources

Stochastic modelling reveals that chromatin folding buffers epigenetic landscapes against sirtuin depletion during DNA damage

Epigenetic landscapes, represented by patterns of chemical modifications on histone tails, are essential for maintaining cell identity and tissue homeostasis. These landscapes are shaped by multiple factors, including local biochemical signals and the three-dimensional organisation of chromatin. However, their response to genomic stress, such as DNA double-strand breaks (DSBs), remains incompletely understood. Here, we use a stochastic model of histone modification dynamics integrated with chromatin architecture to investigate how local depletion of sirtuins, histone deacetylases involved in DSB repair, destabilises epigenetic patterns. Our simulations recapitulate experimental findings in which sirtuin relocalisation to DSB sites leads to the epigenetic erosion and suggest that the resulting landscape depends on enzyme levels and chromatin geometry. Importantly, chromatin regions with large domains of long-range contacts are more resilient to epigenetic destabilisation. These findings suggest that chromatin folding can buffer against relocation of histone-modifying enzymes, highlighting a structural mechanism for preserving epigenetic integrity under stress.

q-bio.QM

Hybrid methods in reaction-diffusion equations

Simulation of stochastic spatially-extended systems is a challenging problem. The fundamental quantities in these models are individual entities such as molecules, cells, or animals, which move and react in a random manner. In big systems, accounting for each individual is inefficient. If the number of entities is large enough, random effects are negligible, and often partial differential equations (PDEs) are used in which the fluctuations are neglected. When the system is heterogeneous, so that the number of individuals is large in certain regions and small in others, the PDE description becomes inaccurate in certain regions. To overcome this problem, the so-called hybrid schemes have been proposed that couple a stochastic description in parts of the domain with its mean field limit in the others. In this chapter, we review the different formulations of this approach and our recent contributions to overcome several of the limitations of previous schemes, including the extension of the concept to multiscale models of cell populations.

q-bio.QM

Learning Rules from KGs Guided by Language Models

Advances in information extraction have enabled the automatic construction of large knowledge graphs (e.g., Yago, Wikidata or Google KG), which are widely used in many applications like semantic search or data analytics. However, due to their semi-automatic construction, KGs are often incomplete. Rule learning methods, concerned with the extraction of frequent patterns from KGs and casting them into rules, can be applied to predict potentially missing facts. A crucial step in this process is rule ranking. Ranking of rules is especially challenging over highly incomplete or biased KGs (e.g., KGs predominantly storing facts about famous people), as in this case biased rules might fit the data best and be ranked at the top based on standard statistical metrics like rule confidence. To address this issue, prior works proposed to rank rules not only relying on the original KG but also facts predicted by a KG embedding model. At the same time, with the recent rise of Language Models (LMs), several works have claimed that LMs can be used as alternative means for KG completion. In this work, our goal is to verify to which extent the exploitation of LMs is helpful for improving the quality of rule learning systems.

cs.CL

Understanding how chromatin folding and enzyme competition affect rugged epigenetic landscapes

Epigenetics plays a key role in cellular differentiation and maintaining cell identity, enabling cells to regulate their genetic activity without altering the DNA sequence. Epigenetic regulation occurs within the context of hierarchically folded chromatin, yet the interplay between the dynamics of epigenetic modifications and chromatin architecture remains poorly understood. In addition, it remains unclear what mechanisms drive the formation of rugged epigenetic patterns, characterised by alternating genomic regions enriched in activating and repressive marks. In this study, we focus on post-translational modifications of histone H3 tails, particularly H3K27me3, H3K4me3, and H3K27ac. We introduce a mesoscopic stochastic model that incorporates chromatin architecture and competition of histone-modifying enzymes into the dynamics of epigenetic modifications in small genomic loci comprising several nucleosomes. Our approach enables us to investigate the mechanisms by which epigenetic patterns form on larger scales of chromatin organisation, such as loops and domains. Through bifurcation analysis and stochastic simulations, we demonstrate that the model can reproduce uniform chromatin states (open, closed, and bivalent) and generate previously unexplored rugged profiles. Our results suggest that enzyme competition and chromatin conformations with high-frequency interactions between distant genomic loci can drive the emergence of rugged epigenetic landscapes. Additionally, we hypothesise that bivalent chromatin can act as an intermediate state, facilitating transitions between uniform and rugged landscapes. This work offers a powerful mathematical framework for understanding the dynamic interactions between chromatin architecture and epigenetic regulation, providing new insights into the formation of complex epigenetic patterns.

q-bio.QM

Explaining Graph Neural Networks for Node Similarity on Graphs

Similarity search is a fundamental task for exploiting information in various applications dealing with graph data, such as citation networks or knowledge graphs. While this task has been intensively approached from heuristics to graph embeddings and graph neural networks (GNNs), providing explanations for similarity has received less attention. In this work we are concerned with explainable similarity search over graphs, by investigating how GNN-based methods for computing node similarities can be augmented with explanations. Specifically, we evaluate the performance of two prominent approaches towards explanations in GNNs, based on the concepts of mutual information (MI), and gradient-based explanations (GB). We discuss their suitability and empirically validate the properties of their explanations over different popular graph benchmarks. We find that unlike MI explanations, gradient-based explanations have three desirable properties. First, they are actionable: selecting inputs depending on them results in predictable changes in similarity scores. Second, they are consistent: the effect of selecting certain inputs overlaps very little with the effect of discarding them. Third, they can be pruned significantly to obtain sparse explanations that retain the effect on similarity scores.

cs.LG

Addressing the Scalability Bottleneck of Semantic Technologies at Bosch

At the heart of smart manufacturing is real-time semi-automatic decision-making. Such decisions are vital for optimizing production lines, e.g., reducing resource consumption, improving the quality of discrete manufacturing operations, and optimizing the actual products, e.g., optimizing the sampling rate for measuring product dimensions during production. Such decision-making relies on massive industrial data thus posing a real-time processing bottleneck.

cs.DC

Computational modelling of angiogenesis: The importance of cell rearrangements during vascular growth

Angiogenesis is the process wherein endothelial cells (ECs) form sprouts that elongate from the pre-existing vasculature to create new vascular networks. In addition to its essential role in normal development, angiogenesis plays a vital role in pathologies such as cancer, diabetes and atherosclerosis. Mathematical and computational modelling has contributed to unravelling its complexity. Many existing theoretical models of angiogenic sprouting are based on the 'snail-trail' hypothesis. This framework assumes that leading ECs positioned at sprout tips migrate towards low-oxygen regions while other ECs in the sprout passively follow the leaders' trails and proliferate to maintain sprout integrity. However, experimental results indicate that, contrary to the snail-trail assumption, ECs exchange positions within developing vessels, and the elongation of sprouts is primarily driven by directed migration of ECs. The functional role of cell rearrangements remains unclear. This review of the theoretical modelling of angiogenesis is the first to focus on the phenomenon of cell mixing during early sprouting. We start by describing the biological processes that occur during early angiogenesis, such as phenotype specification, cell rearrangements and cell interactions with the microenvironment. Next, we provide an overview of various theoretical approaches that have been employed to model angiogenesis, with particular emphasis on recent in silico models that account for the phenomenon of cell mixing. Finally, we discuss when cell mixing should be incorporated into theoretical models and what essential modelling components such models should include in order to investigate its functional role.

q-bio.TO

Contextual Reasoning for Scene Generation (Technical Report)

We present a continuation to our previous work, in which we developed the MR-CKR framework to reason with knowledge overriding across contexts organized in multi-relational hierarchies. Reasoning is realized via ASP with algebraic measures, allowing for flexible definitions of preferences. In this paper, we show how to apply our theoretical work to real autonomous-vehicle scene data. Goal of this work is to apply MR-CKR to the problem of generating challenging scenes for autonomous vehicle learning. In practice, most of the scene data for AV learning models common situations, thus it might be difficult to capture cases where a particular situation occurs (e.g. partial occlusions of a crossing pedestrian). The MR-CKR model allows for data organization exploiting the multi-dimensionality of such data (e.g., temporal and spatial). Reasoning over multiple contexts enables the verification and configuration of scenes, using the combination of different scene ontologies. We describe a framework for semantically guided data generation, based on a combination of MR-CKR and Algebraic Measures. The framework is implemented in a proof-of-concept prototype exemplifying some cases of scene generation.

cs.AI

Answer-Set Programming for Lexicographical Makespan Optimisation in Parallel Machine Scheduling

We deal with a challenging scheduling problem on parallel machines with sequence-dependent setup times and release dates from a real-world application of semiconductor work-shop production. There, jobs can only be processed by dedicated machines, thus few machines can determine the makespan almost regardless of how jobs are scheduled on the remaining ones. This causes problems when machines fail and jobs need to be rescheduled. Instead of optimising only the makespan, we put the individual machine spans in non-ascending order and lexicographically minimise the resulting tuples. This achieves that all machines complete as early as possible and increases the robustness of the schedule. We study the application of Answer-Set Programming (ASP) to solve this problem. While ASP eases modelling, the combination of timing constraints and the considered objective function challenges current solving technology. The former issue is addressed by using an extension of ASP by difference logic. For the latter, we devise different algorithms that use multi-shot solving. To tackle industrial-sized instances, we study different approximations and heuristics. Our experimental results show that ASP is indeed a promising KRR paradigm for this problem and is competitive with state-of-the-art CP and MIP solvers. Under consideration in Theory and Practice of Logic Programming (TPLP).

cs.AI

On Event-Driven Knowledge Graph Completion in Digital Factories

Smart factories are equipped with machines that can sense their manufacturing environments, interact with each other, and control production processes. Smooth operation of such factories requires that the machines and engineering personnel that conduct their monitoring and diagnostics share a detailed common industrial knowledge about the factory, e.g., in the form of knowledge graphs. Creation and maintenance of such knowledge is expensive and requires automation. In this work we show how machine learning that is specifically tailored towards industrial applications can help in knowledge graph completion. In particular, we show how knowledge completion can benefit from event logs that are common in smart factories. We evaluate this on the knowledge graph from a real world-inspired smart factory with encouraging results.

cs.LG

Combining Inductive and Deductive Reasoning for Query Answering over Incomplete Knowledge Graphs

Current methods for embedding-based query answering over incomplete Knowledge Graphs (KGs) only focus on inductive reasoning, i.e., predicting answers by learning patterns from the data, and lack the complementary ability to do deductive reasoning, which requires the application of domain knowledge to infer further information. To address this shortcoming, we investigate the problem of incorporating ontologies into embedding-based query answering models by defining the task of embedding-based ontology-mediated query answering. We propose various integration strategies into prominent representatives of embedding models that involve (1) different ontology-driven data augmentation techniques and (2) adaptation of the loss function to enforce the ontology axioms. We design novel benchmarks for the considered task based on the LUBM and the NELL KGs and evaluate our methods on them. The achieved improvements in the setting that requires both inductive and deductive reasoning are from 20% to 55% in HITS@3.

cs.AI

A method to coarse-grain multi-agent stochastic systems with regions of multistability

Hybrid multiscale modelling has emerged as a useful framework for modelling complex biological phenomena. However, when accounting for stochasticity in the internal dynamics of agents, these models frequently become computationally expensive. Traditional techniques to reduce the computational intensity of such models can lead to a reduction in the richness of the dynamics observed, compared to the original system. Here we use large deviation theory to decrease the computational cost of a spatially-extended multi-agent stochastic system with a region of multi-stability by coarse-graining it to a continuous time Markov chain on the state space of stable steady states of the original system. Our technique preserves the original description of the stable steady states of the system and accounts for noise-induced transitions between them. We apply the method to a bistable system modelling phenotype specification of cells driven by a lateral inhibition mechanism. For this system, we demonstrate how the method may be used to explore different pattern configurations and unveil robust patterns emerging on longer timescales. We then compare the full stochastic, coarse-grained and mean-field descriptions via pattern quantification metrics and in terms of the numerical cost of each method. Our results show that the coarse-grained system exhibits the lowest computational cost while preserving the rich dynamics of the stochastic system. The method has the potential to reduce the computational complexity of hybrid multiscale models, making them more tractable for analysis, simulation and hypothesis testing.

q-bio.QM

Hybrid ASP-based Approach to Pattern Mining

Detecting small sets of relevant patterns from a given dataset is a central challenge in data mining. The relevance of a pattern is based on user-provided criteria; typically, all patterns that satisfy certain criteria are considered relevant. Rule-based languages like Answer Set Programming (ASP) seem well-suited for specifying such criteria in a form of constraints. Although progress has been made, on the one hand, on solving individual mining problems and, on the other hand, developing generic mining systems, the existing methods either focus on scalability or on generality. In this paper we make steps towards combining local (frequency, size, cost) and global (various condensed representations like maximal, closed, skyline) constraints in a generic and efficient way. We present a hybrid approach for itemset, sequence and graph mining which exploits dedicated highly optimized mining systems to detect frequent patterns and then filters the results using declarative ASP. To further demonstrate the generic nature of our hybrid framework we apply it to a problem of approximately tiling a database. Experiments on real-world datasets show the effectiveness of the proposed method and computational gains for itemset, sequence and graph mining, as well as approximate tiling. Under consideration in Theory and Practice of Logic Programming (TPLP).

cs.AI