arXiv ScienceSearch

arXiv subjects

Alec Kirkley

Publications and source records attributed to Alec Kirkley.

At least 19 recordsLinked to original sources

Detectability limits of scaling laws

Power law scaling relations between size and output are central to quantitative theories of cities, organisms, and other complex systems. Competing theories predict scaling exponents that differ by small fractions, but there is no existing theory for verifying whether a given dataset can even distinguish exponents at the required resolution to address such discrepancies. Here we derive a resolution limit for scaling exponents, giving the smallest exponent difference that any method of analysis can detect. We find that the Hurst exponents governing the evolution of systems' sizes and deviations from the scaling law determine how long a record of growing systems must be before it can separate competing scaling theories. Empirical results suggest that many available data panels are insufficient for reliable scaling model selection.

physics.soc-ph

Heterogeneous Interaction Network Analysis (HINA): A New Learning Analytics Approach for Modelling, Analyzing, and Visualizing Complex Interactions in Learning Processes

Existing learning analytics approaches, which often model learning processes as sequences of learner actions or homogeneous relationships, are limited in capturing the distributed, multi-typed interactions in contemporary learning environments. To address this, we propose Heterogeneous Interaction Network Analysis (HINA), a multi-level learning analytics framework for modelling interactions and associations across diverse entities (e.g., learners, behaviours, and AI agents) in learning processes. Grounded in network science principles, HINA integrates a multi-level analytical framework that analyzes individual process metrics (node-level), prominent associations (dyad-level), and latent clusters (meso-level) to address questions about how different elements in a learning environment interact and co-influence each other. In this paper, we first detail the theoretical and mathematical foundations of HINA for individual, dyadic, and meso-level analysis. We then demonstrate HINA's utility through a case study on AI-assisted small-group collaborative learning, revealing students' interaction profiles with peers versus AI, distinct engagement patterns that emerge from these interactions, and specific types of learning behaviors (e.g., asking questions) directed to AI versus peers. HINA contributes a novel and unified analytical framework that enables researchers to quantify individual-level processes, identify significant associations, and uncover meaningful clusters within a single workflow. Accompanied by a dedicated web tool, HINA provides a complete workflow for supporting process-based assessment, and enables a new theoretical lens for modeling and understanding complex and mediated learning processes.

cs.SI

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

Geospatial Foundation Models (GFMs) are emerging as a powerful paradigm for learning semantically rich and geographically consistent visual and physical representations. However, their reliance on Earth-observation (EO) data leaves information about human activity largely underrepresented. Human mobility data reveals the functional and relational structure between regions that is missing from EO data, but is often limited only to the city where it is observed, making it challenging to use for transferable urban representation learning. We introduce MoRAX, a lightweight framework for augmenting geospatial embeddings with functional structure derived from human mobility. MoRAX preserves the coverage and consistency of a GFM while providing information about the functional connectivity among urban regions, permitting zero-shot deployment in unseen cities with or without available mobility data. Across four target cities spanning two countries, the MoRAX teacher model, which observes mobility, consistently outperforms GFMs and strong urban representation baselines in eight socioeconomic and environmental prediction tasks. Meanwhile, the student model, which never takes mobility data as input, approaches the teacher in performance on most tasks. Transfer results across countries further demonstrate that modulation conditioned on mobility flows provides a general mechanism for grounding geospatial foundations in the human dimension of cities.

cs.LG

Hypergraph backboning

Hypergraphs provide a natural framework for describing complex networked systems with higher-order, non-dyadic interactions. Due to their high dimensionality and often redundant structure, a key challenge is to develop methods that simplify hypergraph representations while preserving the essential structure of interactions. Here we present a principled, efficient, and non-parametric information-theoretic method for pruning nested and/or redundant structures in hypergraphs, enabling a minimal representation of higher-order interactions in the presence of local heterogeneity. Our approach naturally extends to weighted hypergraphs, where higher-order topology and hyperedge weights combine to identify the system's structural backbone. We validate the method on controlled synthetic hypergraphs and apply it to empirical datasets from diverse domains, demonstrating substantial sparsification without loss of core structural information.

cs.SI

Information theory for hypergraph similarity

Comparing networks is essential for a number of downstream tasks, from clustering to anomaly detection. Despite higher-order interactions being critical for understanding the dynamics of complex systems, traditional approaches for network comparison are limited to pairwise interactions only. Here we construct a general information theoretic framework for hypergraph similarity, capturing meaningful correspondence among higher-order interactions while correcting for spurious correlations. Our method operationalizes any notion of structural overlap among hypergraphs as a principled normalized mutual information measure, allowing us to derive a hierarchy of increasingly granular formulations of similarity among hypergraphs within and across orders of interactions, and at multiple scales. We validate these measures through extensive experiments on synthetic hypergraphs and apply the framework to reveal meaningful patterns in a variety of empirical higher-order networks. Our work provides foundational tools for the principled comparison of higher-order networks, shedding light on the structural organization of networked systems with non-dyadic interactions.

physics.soc-ph

Networks of amenities reveal universal homophily and heterophily across global cities

Agglomeration economies drive urban growth at different spatial scales by enabling productivity gains, knowledge spillovers, and shared inputs among proximate firms and amenities. To develop a unified science of cities it is thus important to understand how and to what extent different amenities cluster or mix across scales and regional contexts. By utilizing a novel Bayesian framework for nonparametrically quantifying the spectrum of possible mixing patterns of amenities in a city, we identify universal spatial scales of homophily (agglomeration) and heterophily (co-agglomeration) among different amenity types across roughly 800 cities worldwide. Through a detailed longitudinal case study, we also find that the changes in heterophilic mixing derived from our methodology more effectively predict changes in neighborhood rental values than the diversity of amenities present. These findings suggest that agglomeration economies exhibit universal spatial regularities that depend largely on the types of firms or amenities being considered, rather than their specifics or regional context, and highlight the benefit of heterophilic amenity mixing at walkable spatial scales.

physics.soc-ph

Scalable inference of spatial regions and temporal signatures from time series

Regionalization aims to partition a spatial domain into contiguous regions that share similar characteristics, enabling more effective spatial analysis, policy making, and resource management. Existing approaches for spatial regionalization typically rely on static spatial snapshots rather than evolving time series. Meanwhile, most time series clustering methods ignore spatial structure or enforce spatial continuity through ad hoc regularization, constraining the number of inferred regions a priori either explicitly or implicitly. Utilizing the minimum description length principle from information theory, here we propose an efficient and fully nonparametric framework for the regionalization of spatial time series. Our method jointly infers a spatial partition along with a set of representative time series archetypes ("drivers") that best compress a spatiotemporal dataset, with a runtime log-linear in the number of time series. We demonstrate that this method can accurately recover planted regional structure and drivers in synthetic time series, and can extract meaningful structural regularities in large-scale empirical air quality and vegetation index records. Our method provides a principled and scalable framework for spatially contiguous partitioning, allowing interpretable temporal patterns and homogeneous regions to emerge directly from the data itself.

stat.ML

Estimation of partial rankings from sparse, noisy comparisons

Ranking items based on pairwise comparisons is common, from using match outcomes to rank sports teams to using purchase or survey data to rank consumer products. Statistical inference-based methods such as the Bradley-Terry model, which extract rankings based on an underlying generative model, have emerged as flexible and powerful tools to tackle ranking in empirical data. In situations with limited and/or noisy comparisons, it is often challenging to confidently distinguish the performance of different items based on the evidence available in the data. However, most inference-based ranking methods choose to assign each item to a unique rank or score, suggesting a meaningful distinction when there is none. Here, we develop a principled nonparametric Bayesian method, adaptable to any statistical ranking method, for learning partial rankings (rankings with ties) that distinguishes among the ranks of different items only when there is sufficient evidence available in the data. We develop a fast agglomerative algorithm to perform Maximum A Posteriori (MAP) inference of partial rankings under our framework and examine the performance of our method on a variety of real and synthetic network datasets, finding that it frequently gives a more parsimonious summary of the data than traditional ranking, particularly when observations are sparse.

physics.soc-ph

Structural reducibility of hypergraphs

Higher-order interactions provide a nuanced understanding of the relational structure of complex systems beyond traditional pairwise interactions. However, higher-order network analyses also incur more cumbersome interpretations and greater computational demands than their pairwise counterparts. Here we present an information-theoretic framework for determining the extent to which a hypergraph representation of a networked system is structurally redundant, and for identifying its most critical higher orders of interaction that allow us to remove these redundancies while preserving essential higher-order structure.

physics.soc-ph

Normalized mutual information is a biased measure for classification and community detection

Normalized mutual information is widely used as a similarity measure for evaluating the performance of clustering and classification algorithms. In this paper, we argue that results returned by the normalized mutual information are biased for two reasons: first, because they ignore the information content of the contingency table and, second, because their symmetric normalization introduces spurious dependence on algorithm output. We introduce a modified version of the mutual information that remedies both of these shortcomings. As a practical demonstration of the importance of using an unbiased measure, we perform extensive numerical tests on a basket of popular algorithms for network community detection and show that one's conclusions about which algorithm is best are significantly affected by the biases in the traditional mutual information.

cs.SI

Transfer entropy for finite data

Transfer entropy is a widely used measure for quantifying directed information flows in complex systems. While the challenges of estimating transfer entropy for continuous data are well known, it has two major shortcomings for data of finite cardinality: it exhibits a substantial positive bias for sparse bin counts, and it has no clear means to assess statistical significance. By computing information content in finite data streams without explicitly considering symbols as instances of random variables, we derive a transfer entropy measure which is asymptotically equivalent to the standard plug-in estimator but remedies these issues for time series of small size and/or high cardinality, permitting a fully nonparametric assessment of statistical significance without simulation.

physics.data-an

Belief propagation for finite networks using a symmetry-breaking source node

Belief Propagation (BP) is an efficient message-passing algorithm widely used for inference in graphical models and for solving various problems in statistical physics. However, BP often yields inaccurate estimates of order parameters and their susceptibilities in finite systems, particularly in sparse networks with few loops. Here, we show for both percolation and Ising models that fixing the state of a single well-connected "source" node to break global symmetry substantially improves inference accuracy and captures finite-size effects across a broad range of networks, especially tree-like ones, at no additional computational cost.

cs.SI

Fast nonparametric inference of network backbones for weighted graph sparsification

Network backbones provide useful sparse representations of weighted networks by keeping only their most important links, permitting a range of computational speedups and simplifying network visualizations. A key limitation of existing network backboning methods is that they either require the specification of a free parameter (e.g. significance level) that determines the number of edges to keep in the backbone, or impose specific restrictions on the topology of the backbone (e.g. that it is a spanning tree). Here we develop a completely nonparametric framework for inferring the backbone of a weighted network that overcomes these limitations and automatically selects the optimal set of edges to retain using the Minimum Description Length (MDL) principle. We develop objective functions for global and local network backboning which evaluate the importance of an edge in the context of the whole network and individual node neighborhoods respectively and are generalizable to any weight distribution under Bayesian model specifications that fix the average edge weight either exactly or in expectation. We then construct an efficient and provably optimal greedy algorithm to identify the backbone minimizing our objectives, whose runtime complexity is log-linear in the number of edges. We demonstrate our methods by comparing them with existing methods in a range of tasks on real and synthetic networks, finding that both the global and local backboning methods can preserve network connectivity, weight heterogeneity, and spreading dynamics while removing a substantial fraction of edges.

cs.SI

Network compression with configuration models and the minimum description length

Random network models, constrained to reproduce specific statistical features, are often used to represent and analyze network data and their mathematical descriptions. Chief among them, the configuration model constrains random networks by their degree distribution and is foundational to many areas of network science. However, configuration models and their variants are often selected based on intuition or mathematical and computational simplicity rather than on statistical evidence. To evaluate the quality of a network representation, we need to consider both the amount of information required to specify a random network model and the probability of recovering the original data when using the model as a generative process. To this end, we calculate the approximate size of network ensembles generated by the popular configuration model and its generalizations, including versions accounting for degree correlations and centrality layers. We then apply the minimum description length principle as a model selection criterion over the resulting nested family of configuration models. Using a dataset of over 100 networks from various domains, we find that the classic Configuration Model is generally preferred on networks with an average degree above ten, while a Layered Configuration Model constrained by a centrality metric offers the most compact representation of the majority of sparse networks.

cs.SI

Network mutual information measures for graph similarity

A wide range of tasks in network analysis, such as clustering network populations or identifying anomalies in temporal graph streams, require a measure of the similarity between two graphs. To provide a meaningful data summary for downstream scientific analyses, the graph similarity measures used for these tasks must be principled, interpretable, and capable of distinguishing meaningful overlapping network structure from statistical noise at different scales of interest. Here we derive a family of graph mutual information measures that satisfy these criteria and are constructed using only fundamental information theoretic principles. Our measures capture the information shared among networks according to different encodings of their structural information, with our mesoscale mutual information measure allowing for network comparison under any specified network coarse-graining. We test our measures in a range of applications on real and synthetic network data, finding that they effectively highlight intuitive aspects of network similarity across scales in a variety of systems.

physics.soc-ph

Urban Boundary Delineation from Commuting Data with Bayesian Stochastic Blockmodeling: Scale, Contiguity, and Hierarchy

A common method for delineating urban and suburban boundaries is to identify clusters of spatial units that are highly interconnected in a network of commuting flows, each cluster signaling a cohesive economic submarket. It is critical that the clustering methods employed for this task are principled and free of unnecessary tunable parameters to avoid unwanted inductive biases while remaining scalable for high resolution mobility networks. Here we systematically assess the benefits and limitations of a wide array of Stochastic Block Models (SBMs)$\unicode{x2014}$a family of principled, nonparametric models for identifying clusters in networks$\unicode{x2014}$for delineating urban spatial boundaries with commuting data. We find that the data compression capability and relative performance of different SBM variants heavily depends on the spatial extent of the commuting network, its aggregation scale, and the method used for weighting network edges. We also construct a new measure to assess the degree to which community detection algorithms find spatially contiguous partitions, finding that traditional SBMs may produce substantial spatial discontiguities that make them challenging to use in general for urban boundary delineation. We propose a fast nonparametric regionalization algorithm that can alleviate this issue, achieving data compression close to that of unconstrained SBM models while ensuring spatial contiguity, benefiting from a deterministic optimization procedure, and being generalizable to a wide range of community detection objective functions.

physics.soc-ph

Mutual information and the encoding of contingency tables

Mutual information is commonly used as a measure of similarity between competing labelings of a given set of objects, for example to quantify performance in classification and community detection tasks. As argued recently, however, the mutual information as conventionally defined can return biased results because it neglects the information cost of the so-called contingency table, a crucial component of the similarity calculation. In principle the bias can be rectified by subtracting the appropriate information cost, leading to the modified measure known as the reduced mutual information, but in practice one can only ever compute an upper bound on this information cost, and the value of the reduced mutual information depends crucially on how good a bound is established. In this paper we describe an improved method for encoding contingency tables that gives a substantially better bound in typical use cases, and approaches the ideal value in the common case where the labelings are closely similar, as we demonstrate with extensive numerical results.

cs.SI

Identifying hubs in directed networks

Nodes in networks that exhibit high connectivity, also called ``hubs'', play a critical role in determining the structural and functional properties of networked systems. However, there is no clear definition of what constitutes a hub node in a network, and the classification of network hubs in existing work has either been purely qualitative or relies on ad hoc criteria for thresholding continuous data that do not generalize well to networks with certain degree sequences. Here we develop a set of efficient nonparametric methods that classify hub nodes in directed networks using the Minimum Description Length principle, effectively providing a clear and principled definition for network hubs. We adapt our methods to both unweighted and weighted networks and demonstrate them in a range of example applications using real and synthetic network data.

cs.SI