arXiv ScienceSearch

arXiv subjects

Max Hill

Publications and source records attributed to Max Hill.

6 recordsLinked to original sources

A Needs Assessment for Measuring Geographic - Legislative Associations in the U.S. House of Representatives

Political legislation affects the well-being and livelihoods of constituents. In the U.S. Congress a representative's voting record on bills and legislation is public. These bills have themes associated with them, such as veterans' affairs, coastal monitoring, agricultural appropriations, etc. A bill on veterans' affairs may affect a constituency differently if they have a high percentage of veterans. In this work, we demonstrate how congressional vote outcomes can be merged with typical geographic information systems (GIS) data to help compare a legislator's votes with the geographies of their constituencies to measure the association between district features and legislators' decisions. We retrieved and tagged bills from the 118th U.S. House of Representatives (Jan. 2023 - Jan. 2025) by manually assigning each bill a set of themes. We then retrieved spatial data at the congressional district level related to each theme. We built a backend database to connect bill information and spatial data, and an interactive map prototype displayed the spatial data related to each theme associated with each bill. Our work helps identify challenges to creating a more accessible and complete system that would facilitate the visualization and analysis of the nature of contemporary Congressional representation. This technical report is primarily written for members of the data science community, and specifically those interested in legislative and geographic data.

cs.HC

Semialgebraic Conditions for Identifying Triangles in Phylogenetic Networks

An important consideration for a model-based method of phylogenetic network inference is the identifiability of the network parameter of the model. A recurring theme in previous works exploring this issue is that it is often difficult to identify the orientation of edges in a triangle of the network. In fact, it has been shown that for some models it is impossible to determine the orientation of triangle edges utilizing the standard algebraic technique of phylogenetic invariants. In this work, we consider one such model with a Jukes-Cantor site-substitution process and no coalescence. We give a complete semialgebraic description of three, 3-leaf Jukes-Cantor phylogenetic network models with embedded triangles. By describing these base cases, we resolve several questions about the identifiability of networks with embedded triangles. We show that for any pair of models, the intersection and set differences of the models are full-dimensional regions of the space of site-pattern probability distributions. Thus, despite being algebraically indistinguishable, these network models are not identical, nor are they identifiable (or generically identifiable). Our results also yield a straightforward biological interpretation--that the signal from a hybridization event may be immediately detectable but decays over time until it is impossible to identify the orientation of edges in the triangle of a network.

q-bio.PE

Dynamical Systems Models for Market Evolution: A Mechanistic Alternative to Autoregressive Methods

We present a novel approach to modeling market dynamics using ordinary differential equations that explicitly incorporates product competitiveness and consumer behavior. Our framework treats market segments as interacting populations in a dynamical system analogous to predator-prey models, where competitive advantages drive market share transitions through mechanistic modeling of market flows including new product adoption, refresh cycles, and obsolescence dynamics.

math.DS

Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics

Recently, Sturma, Drton, and Leung proposed a general-purpose stochastic method for hypothesis testing in models defined by polynomial equality and inequality constraints. Notably, the method remains theoretically valid even near irregular points, such as singularities and boundaries, where traditional testing approaches often break down. In this paper, we evaluate its practical performance on a collection of biologically motivated models from phylogenetics. While the method performs remarkably well across different settings, we catalogue a number of issues that should be considered for effective application.

q-bio.PE

New directions in algebraic statistics: Three challenges from 2023

In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.

math.ST

Species tree estimation under joint modeling of coalescence and duplication: sample complexity of quartet methods

We consider species tree estimation under a standard stochastic model of gene tree evolution that incorporates incomplete lineage sorting (as modeled by a coalescent process) and gene duplication and loss (as modeled by a branching process). Through a probabilistic analysis of the model, we derive sample complexity bounds for widely used quartet-based inference methods that highlight the effect of the duplication and loss rates in both subcritical and supercritical regimes.

math.PR