arXiv ScienceSearch

arXiv subjects

Nicholas Smith

Publications and source records attributed to Nicholas Smith.

12 recordsLinked to original sources

The Critical Importance of Software for HEP

Particle physics has an ambitious and broad global experimental programme for the coming decades. Large investments in building new facilities are already underway or under consideration. Scaling the present processing power and data storage needs by the foreseen increase in data rates in the next decade for HL-LHC is not sustainable within the current budgets. As a result, a more efficient usage of computing resources is required in order to realise the physics potential of future experiments. Software and computing are an integral part of experimental design, trigger and data acquisition, simulation, reconstruction, and analysis, as well as related theoretical predictions. A significant investment in computing and software is therefore critical. Advances in software and computing, including artificial intelligence (AI) and machine learning (ML), will be key for solving these challenges. Making better use of new processing hardware such as graphical processing units (GPUs) or ARM chips is a growing trend. This forms part of a computing solution that makes efficient use of facilities and contributes to the reduction of the environmental footprint of HEP computing. The HEP community already provided a roadmap for software and computing for the last EPPSU, and this paper updates that, with a focus on the most resource critical parts of our data processing chain.

hep-ex

Indirect searches at CMS

In these proceedings, we present several new measurements of Standard Model (SM) processes, in the Higgs sector and beyond, that push the precision frontier forward at CMS. Results are presented in the context of a framework parameterizing deviations in Higgs boson couplings, as well as in the context of SM Effective Field Theory, where new analyses targeting Higgs, top, and multi-boson processes probe an increasingly diverse set of operators. Within these frameworks, CMS is efficiently exploring a large space of new physics models. No significant deviation from SM expectations is found.

hep-ex

Second Analysis Ecosystem Workshop Report

The second workshop on the HEP Analysis Ecosystem took place 23-25 May 2022 at IJCLab in Orsay, to look at progress and continuing challenges in scaling up HEP analysis to meet the needs of HL-LHC and DUNE, as well as the very pressing needs of LHC Run 3 analysis. The workshop was themed around six particular topics, which were felt to capture key questions, opportunities and challenges. Each topic arranged a plenary session introduction, often with speakers summarising the state-of-the art and the next steps for analysis. This was then followed by parallel sessions, which were much more discussion focused, and where attendees could grapple with the challenges and propose solutions that could be tried. Where there was significant overlap between topics, a joint discussion between them was arranged. In the weeks following the workshop the session conveners wrote this document, which is a summary of the main discussions, the key points raised and the conclusions and outcomes. The document was circulated amongst the participants for comments before being finalised here.

hep-ex

Collaborative Computing Support for Analysis Facilities Exploiting Software as Infrastructure Techniques

Prior to the public release of Kubernetes it was difficult to conduct joint development of elaborate analysis facilities due to the highly non-homogeneous nature of hardware and network topology across compute facilities. However, since the advent of systems like Kubernetes and OpenShift, which provide declarative interfaces for building fault-tolerant and self-healing deployments of networked software, it is possible for multiple institutes to collaborate more effectively since resource details are abstracted away through various forms of hardware and software virtualization. In this whitepaper we will outline the development of two analysis facilities: "Coffea-casa" at University of Nebraska Lincoln and the "Elastic Analysis Facility" at Fermilab, and how utilizing platform abstraction has improved the development of common software for each of these facilities, and future development plans made possible by this methodology.

physics.data-an

Cloud to Ground Secured Computing: User Experiences on the Transition from Cloud-Based to Locally-Sited Hardware

The application of high-performance computing (HPC) processes, tools, and technologies to Controlled Unclassified Information (CUI) creates both opportunities and challenges. Building on our experiences developing, deploying, and managing the Research Environment for Encumbered Data (REED) hosted by AWS GovCloud, Research Computing at Purdue University has recently deployed Weber, our locally-sited HPC solution for the storage and analysis of CUI data. Weber presents our customer base with advances in data access, portability, and usability at a low, stable cost while reducing administrative overhead for our information technology support team.

cs.DC

VLA Imaging of HI-bearing Ultra-Diffuse Galaxies from the ALFALFA Survey

Ultra-diffuse galaxies have generated significant interest due to their large optical extents and low optical surface brightnesses, which challenge galaxy formation models. Here we present resolved synthesis observations of 12 HI-bearing ultra-diffuse galaxies (HUDs) from the Karl G. Jansky Very Large Array (VLA), as well as deep optical imaging from the WIYN 3.5-meter telescope at Kitt Peak National Observatory. We present the data processing and images, including total intensity HI maps and HI velocity fields. The HUDs show ordered gas distributions and evidence of rotation, important prerequisites for the detailed kinematic models in Mancera Pi\~na et al. (2019b). We compare the HI and stellar alignment and extent, and find the HI extends beyond the already extended stellar component and that the HI disk is often misaligned with respect to the stellar one, emphasizing the importance of caution when approaching inclination measurements for these extreme sources. We explore the HI mass-diameter scaling relation, and find that although the HUDs have diffuse stellar populations, they fall along the relation, with typical global HI surface densities. This resolved sample forms an important basis for more detailed study of the HI distribution in this extreme extragalactic population.

astro-ph.GA

The Shooting Regressor; Randomized Gradient-Based Ensembles

An ensemble method is introduced that utilizes randomization and loss function gradients to compute a prediction. Multiple weakly-correlated estimators approximate the gradient at randomly sampled points on the error surface and are aggregated into a final solution. A scaling parameter is described that controls a trade-off between ensemble correlation and precision. Numerical methods for estimating optimal values of the parameter are described. Empirical results are computed over a popular dataset. Inferential statistics on these results show that the method is capable of outperforming existing techniques in terms of increased accuracy.

cs.LG

Coffea -- Columnar Object Framework For Effective Analysis

The coffea framework provides a new approach to High-Energy Physics analysis, via columnar operations, that improves time-to-insight, scalability, portability, and reproducibility of analysis. It is implemented with the Python programming language, the scientific python package ecosystem, and commodity big data technologies. To achieve this suite of improvements across many use cases, coffea takes a factorized approach, separating the analysis implementation and data delivery scheme. All analysis operations are implemented using the NumPy or awkward-array packages which are wrapped to yield user code whose purpose is quickly intuited. Various data delivery schemes are wrapped into a common front-end which accepts user inputs and code, and returns user defined outputs. We will discuss our experience in implementing analysis of CMS data using the coffea framework along with a discussion of the user experience and future directions.

cs.DC

A Cryogenic Silicon Interferometer for Gravitational-wave Detection

The detection of gravitational waves from compact binary mergers by LIGO has opened the era of gravitational wave astronomy, revealing a previously hidden side of the cosmos. To maximize the reach of the existing LIGO observatory facilities, we have designed a new instrument that will have 5 times the range of Advanced LIGO, or greater than 100 times the event rate. Observations with this new instrument will make possible dramatic steps toward understanding the physics of the nearby universe, as well as observing the universe out to cosmological distances by the detection of binary black hole coalescences. This article presents the instrument design and a quantitative analysis of the anticipated noise floor.

astro-ph.IM

Cell cycle time series gene expression data encoded as cyclic attractors in Hopfield systems

Modern time series gene expression and other omics data sets have enabled unprecedented resolution of the dynamics of cellular processes such as cell cycle and response to pharmaceutical compounds. In anticipation of the proliferation of time series data sets in the near future, we use the Hopfield model, a recurrent neural network based on spin glasses, to model the dynamics of cell cycle in HeLa (human cervical cancer) and S. cerevisiae cells. We study some of the rich dynamical properties of these cyclic Hopfield systems, including the ability of populations of simulated cells to recreate experimental expression data and the effects of noise on the dynamics. Next, we use a genetic algorithm to identify sets of genes which, when selectively inhibited by local external fields representing gene silencing compounds such as kinase inhibitors, disrupt the encoded cell cycle. We find, for example, that inhibiting the set of four kinases BRD4, MAPK1, NEK7, and YES1 in HeLa cells causes simulated cells to accumulate in the M phase. Finally, we suggest possible improvements and extensions to our model.

q-bio.MN

Evolutionary and topological properties of gene modules and driver mutations in a leukemia gene regulatory network

The diverse, specialized genes in today's lifeforms evolved from a common core of ancient, elementary genes. However, these genes did not evolve individually: gene expression is controlled by a complex network of interactions, and alterations in one gene may drive reciprocal changes in its proteins' binding partners. We show that the topology of a leukemia gene regulatory network is strongly coupled with evolutionary properties. Slowly-evolving ("cold"), old genes tend to interact with each other, as do rapidly-evolving ("hot"), young genes, causing genes to evolve in clusters. We argue that gene duplication placed old, cold genes at the center of the network, and young, hot genes on the periphery, and demonstrate this with single-node centrality measures and two new measures of efficiency. Integrating centrality measures with evolutionary information, we define a medically-relevant "cancer network core," strongly enriched for common cancer mutations ($p=2\times 10^{-14}$). This could aid in identifying driver mutations and therapeutic targets.

q-bio.MN

Multi-species network inference improves gene regulatory network reconstruction for early embryonic development in Drosophila

Gene regulatory network inference uses genome-wide transcriptome measurements in response to genetic, environmental or dynamic perturbations to predict causal regulatory influences between genes. We hypothesized that evolution also acts as a suitable network perturbation and that integration of data from multiple closely related species can lead to improved reconstruction of gene regulatory networks. To test this hypothesis, we predicted networks from temporal gene expression data for 3,610 genes measured during early embryonic development in six Drosophila species and compared predicted networks to gold standard networks of ChIP-chip and ChIP-seq interactions for developmental transcription factors in five species. We found that (i) the performance of single-species networks was independent of the species where the gold standard was measured; (ii) differences between predicted networks reflected the known phylogeny and differences in biology between the species; (iii) an integrative consensus network which minimized the total number of edge gains and losses with respect to all single-species networks performed better than any individual network. Our results show that in an evolutionarily conserved system, integration of data from comparable experiments in multiple species improves the inference of gene regulatory networks. They provide a basis for future studies on the numerous multi-species gene expression datasets for other biological processes available in the literature.

q-bio.GN