arXiv ScienceSearch

arXiv subjects

Sean P. Fillingham

Publications and source records attributed to Sean P. Fillingham.

16 recordsLinked to original sources

Open Problems in AI Risk Modeling: Insights from a Workshop on the Technical Foundations of AI Risk Modeling

We investigate the design of robust risk models to assess societal risks posed by advanced AI systems, an emerging area in AI governance. Many regulatory proposals increasingly require systemic risk assessment, but in the absence of rigorous quantitative methods, the question remains what state of the art risk modeling should look like in practice. We identify the key methodological and institutional challenges that currently limit the adoption of risk modeling. We review five research traditions that inform this problem: probabilistic risk assessment, catastrophic AI risk analysis, cybersecurity risk quantification, Bayesian causal inference, and threshold-based governance. We compare two leading proposals, scenario-based risk estimation and Bayesian network-based threshold setting. Drawing on a workshop with 22 experts and subsequent analysis, we identify a structured agenda of open questions concerning model structure, scope, evidence integration, validation, and governance. We close by outlining priorities for progress, arguing that it will depend on integrating quantitative modeling with independent evaluation, transparent and tiered disclosure, and institutions capable of maintaining and updating risk models over time.

cs.CY

Exploring Systems-Thinking Approaches to Loss of Control Risk

Internal deployment of agentic AI systems for coding and research creates a sociotechnical control problem that extends beyond model behaviour. We treat internal-deployment Loss of Control as the inability to reliably constrain, audit, reverse, or halt AI-mediated changes to code, infrastructure, evaluation, or deployment processes in time to prevent serious organisational or societal harms. We ask whether established systems-safety methods can identify risks that model-level evaluations may miss. Using a generic frontier-lab coding-agent scenario reconstructed from public materials, we apply STECA, STPA, and FRAM. The analyses surface complementary findings: published frameworks can leave governance responsibilities and feedback loops externally unverifiable; delays in monitoring and intervention can make otherwise valid control actions ineffective; and routine operational variability can gradually erode the calibration and independence of safeguards. We argue that frontier-AI risk management should pair model-focused evaluations with systems-level hazard analysis and operational assurance that tracks whether controls remain effective over time.

cs.CY

Lessons from External Review of DeepMind's Scheming Inability Safety Case

Safety cases for frontier AI systems should provide a convincing argument, supported by evidence, that the risk of harm is within an acceptable bound. When developers author their own safety cases, confirmation bias and conflicted incentives can affect the quality of argument. External review can help to address this. In this paper, we apply the Assurance 2.0 framework to perform an external review of Google DeepMind's public scheming inability safety case. We surface substantive new concerns that materially affect the scope of the safety case and its applicability for decision-making. Based on this experience, we provide concrete recommendations for how external review should be conducted and what information AI developers should provide to support it.

cs.CY

STAMP/STPA Informed Characterization of Factors Leading to Loss of Control in AI Systems

A major concern amongst AI safety practitioners is the possibility of loss of control, whereby humans lose the ability to exert control over increasingly advanced AI systems. The range of concerns is wide, spanning current day risks to future existential risks, and a range of loss of control pathways from rapid AI self-exfiltration scenarios to more gradual disempowerment scenarios. In this work we set out to firstly, provide a more structured framework for discussing and characterizing loss of control and secondly, to use this framework to assist those responsible for the safe operation of AI-containing socio-technical systems to identify causal factors leading to loss of control. We explore how these two needs can be better met by making use of a methodology developed within the safety-critical systems community known as STAMP and its associated hazard analysis technique of STPA. We select the STAMP methodology primarily because it is based around a world-view that socio-technical systems can be functionally modeled as control structures, and that safety issues arise when there is a loss of control in these structures.

cs.CY

SCALAR: Benchmarking SAE Interaction Sparsity in Toy LLMs

Mechanistic interpretability aims to decompose neural networks into interpretable features and map their connecting circuits. The standard approach trains sparse autoencoders (SAEs) on each layer's activations. However, SAEs trained in isolation don't encourage sparse cross-layer connections, inflating extracted circuits where upstream features needlessly affect multiple downstream features. Current evaluations focus on individual SAE performance, leaving interaction sparsity unexamined. We introduce SCALAR (Sparse Connectivity Assessment of Latent Activation Relationships), a benchmark measuring interaction sparsity between SAE features. We also propose "Staircase SAEs", using weight-sharing to limit upstream feature duplication across downstream features. Using SCALAR, we compare TopK SAEs, Jacobian SAEs (JSAEs), and Staircase SAEs. Staircase SAEs improve relative sparsity over TopK SAEs by $59.67\% \pm 1.83\%$ (feedforward) and $63.15\% \pm 1.35\%$ (transformer blocks). JSAEs provide $8.54\% \pm 0.38\%$ improvement over TopK for feedforward layers but cannot train effectively across transformer blocks, unlike Staircase and TopK SAEs which work anywhere in the residual stream. We validate on a $216$K-parameter toy model and GPT-$2$ Small ($124$M), where Staircase SAEs maintain interaction sparsity improvements while preserving feature interpretability. Our work highlights the importance of interaction sparsity in SAEs through benchmarking and comparing promising architectures.

cs.LG

The importance of gas starvation in driving satellite quenching in galaxy groups at $z\sim 0.8$

We present results from a Keck/DEIMOS survey to study satellite quenching in group environments at $z \sim 0.8$ within the Extended Groth Strip (EGS). We target $11$ groups in the EGS with extended X-ray emission. We obtain high-quality spectroscopic redshifts for group member candidates, extending to depths over an order of magnitude fainter than existing DEEP2/DEEP3 spectroscopy. This depth enables the first spectroscopic measurement of the satellite quiescent fraction down to stellar masses of $\sim 10^{9.5}~{\rm M}_{\odot}$ at this redshift. By combining an infall-based environmental quenching model, constrained by the observed quiescent fractions, with infall histories of simulated groups from the IllustrisTNG100-1-Dark simulation, we estimate environmental quenching timescales ($τ_{\mathrm{quench}}$) for the observed group population. At high stellar masses (${M}_{\star}=10^{10.5}~{\rm M}_{\odot}$) we find that $τ_{\mathrm{quench}} = 2.4\substack{+0.2 \\ -0.2}$ Gyr, which is consistent with previous estimates at this epoch. At lower stellar masses (${M}_{\star}=10^{9.5}~{\rm M}_{\odot}$), we find that $τ_{\mathrm{quench}}=3.1\substack{+0.5 \\ -0.4}$ Gyr, which is shorter than prior estimates from photometry-based investigations. These timescales are consistent with satellite quenching via starvation, provided the hot gas envelope of infalling satellites is not stripped away. We find that the evolution in the quenching timescale between $0 \lt z \lt 1$ aligns with the evolution in the dynamical time of the host halo and the total cold gas depletion time. This suggests that the doubling of the quenching timescale in groups since $z\sim1$ could be related to the dynamical evolution of groups or a decrease in quenching efficiency via starvation with decreasing redshift.

astro-ph.GA

A Machine Learning Approach to Measuring the Quenched Fraction of Low-Mass Satellites Beyond the Local Group

Observations suggest that satellite quenching plays a major role in the build-up of passive, low-mass galaxies at late cosmic times. Studies of low-mass satellites, however, are limited by the ability to robustly characterize the local environment and star-formation activity of faint systems. In an effort to overcome the limitations of existing data sets, we utilize deep photometry in Stripe 82 of the Sloan Digital Sky Survey, in conjunction with a neural network classification scheme, to study the suppression of star formation in low-mass satellite galaxies in the local Universe. Using a statistically-driven approach, we are able to push beyond the limits of existing spectroscopic data sets, measuring the satellite quenched fraction down to satellite stellar masses of ${\sim}10^7~{\rm M}_{\odot}$ in group environments (${M}_{\rm{halo}} = 10^{13-14}~h^{-1}~{\rm M}_{\odot}$). At high satellite stellar masses ($\gtrsim 10^{10}~{\rm M}_{\odot}$), our analysis successfully reproduces existing measurements of the quenched fraction based on spectroscopic samples. Pushing to lower masses, we find that the fraction of passive satellites increases, potentially signaling a change in the dominant quenching mechanism at ${M}_{\star} \sim 10^{9}~{\rm M}_{\odot}$. Similar to the results of previous studies of the Local Group, this increase in the quenched fraction at low satellite masses may correspond to an increase in the efficacy of ram-pressure stripping as a quenching mechanism in groups.

astro-ph.GA

The Lick AGN Monitoring Project 2016: Dynamical Modeling of Velocity-Resolved H\b{eta} Lags in Luminous Seyfert Galaxies

We have modeled the velocity-resolved reverberation response of the H\b{eta} broad emission line in nine Seyfert 1 galaxies from the Lick Active Galactic Nucleus (AGN) Monitioring Project 2016 sample, drawing inferences on the geometry and structure of the low-ionization broad-line region (BLR) and the mass of the central supermassive black hole. Overall, we find that the H\b{eta} BLR is generally a thick disk viewed at low to moderate inclination angles. We combine our sample with prior studies and investigate line-profile shape dependence, such as log10(FWHM/σ), on BLR structure and kinematics and search for any BLR luminosity-dependent trends. We find marginal evidence for an anticorrelation between the profile shape of the broad H\b{eta} emission line and the Eddington ratio, when using the root-mean-square spectrum. However, we do not find any luminosity-dependent trends, and conclude that AGNs have diverse BLR structure and kinematics, consistent with the hypothesis of transient AGN/BLR conditions rather than systematic trends.

astro-ph.GA

The Lick AGN Monitoring Project 2016: Velocity-Resolved Hβ Lags in Luminous Seyfert Galaxies

We carried out spectroscopic monitoring of 21 low-redshift Seyfert 1 galaxies using the Kast double spectrograph on the 3-m Shane telescope at Lick Observatory from April 2016 to May 2017. Targeting active galactic nuclei (AGN) with luminosities of λLλ (5100 Å) = 10^44 erg/s and predicted Hβ lags of 20-30 days or black hole masses of 10^7-10^8.5 Msun, our campaign probes luminosity-dependent trends in broad-line region (BLR) structure and dynamics as well as to improve calibrations for single-epoch estimates of quasar black hole masses. Here we present the first results from the campaign, including Hβ emission-line light curves, integrated Hβ lag times (8-30 days) measured against V-band continuum light curves, velocity-resolved reverberation lags, line widths of the broad Hβ components, and virial black hole mass estimates (10^7.1-10^8.1 Msun). Our results add significantly to the number of existing velocity-resolved lag measurements and reveal a diversity of BLR gas kinematics at moderately high AGN luminosities. AGN continuum luminosity appears not to be correlated with the type of kinematics that its BLR gas may exhibit. Follow-up direct modeling of this dataset will elucidate the detailed kinematics and provide robust dynamical black hole masses for several objects in this sample.

astro-ph.GA

APOGEE Chemical Abundance Patterns of the Massive Milky Way Satellites

The SDSS-IV Apache Point Observatory Galactic Evolution Experiment (APOGEE) survey has obtained high-resolution spectra for thousands of red giant stars distributed among the massive satellite galaxies of the Milky Way (MW): the Large and Small Magellanic Clouds (LMC/SMC), the Sagittarius Dwarf (Sgr), Fornax (Fnx), and the now fully disrupted \emph{Gaia} Sausage/Enceladus (GSE) system. We present and analyze the APOGEE chemical abundance patterns of each galaxy to draw robust conclusions about their star formation histories, by quantifying the relative abundance trends of multiple elements (C, N, O, Mg, Al, Si, Ca, Fe, Ni, and Ce), as well as by fitting chemical evolution models to the [$α$/Fe]-[Fe/H] abundance plane for each galaxy. Results show that the chemical signatures of the starburst in the MCs observed by Nidever et al. in the $α$-element abundances extend to C+N, Al, and Ni, with the major burst in the SMC occurring some 3-4 Gyr before the burst in the LMC. We find that Sgr and Fnx also exhibit chemical abundance patterns suggestive of secondary star formation epochs, but these events were weaker and earlier ($\sim$~5-7 Gyr ago) than those observed in the MCs. There is no chemical evidence of a second starburst in GSE, but this galaxy shows the strongest initial star formation as compared to the other four galaxies. All dwarf galaxies had greater relative contributions of AGB stars to their enrichment than the MW. Comparing and contrasting these chemical patterns highlight the importance of galaxy environment on its chemical evolution.

astro-ph.GA

The GOGREEN and GCLASS Surveys: First Data Release

We present the first public data release of the GOGREEN and GCLASS surveys of galaxies in dense environments, spanning a redshift range $0.8<z<1.5$. The surveys consist of deep, multiwavelength photometry and extensive Gemini GMOS spectroscopy of galaxies in 26 overdense systems ranging in halo mass from small groups to the most massive clusters. The objective of both projects was primarily to understand how the evolution of galaxies is affected by their environment, and to determine the physical processes that lead to the quenching of star formation. There was an emphasis on obtaining unbiased spectroscopy over a wide stellar mass range ($M\gtrsim 2\times 10^{10}~\mathrm{M}_\odot$), throughout and beyond the cluster virialized regions. The final spectroscopic sample includes 2771 unique objects, of which 2257 have reliable spectroscopic redshifts. Of these, 1704 have redshifts in the range $0.8<z<1.5$, and nearly 800 are confirmed cluster members. Imaging spans the full optical and near-infrared wavelength range, at depths comparable to the UltraVISTA survey, and includes \textit{HST}/WFC3 F160W (GOGREEN) and F140W (GCLASS). This data release includes fully reduced images and spectra, with catalogues of advanced data products including redshifts, line strengths, star formation rates, stellar masses and rest-frame colours. Here we present an overview of the data, including an analysis of the spectroscopic completeness and redshift quality.

astro-ph.GA

Characterizing the Infall Times and Quenching Timescales of Milky Way Satellites with $Gaia$ Proper Motions

Observations of low-mass satellite galaxies in the nearby Universe point towards a strong dichotomy in their star-forming properties relative to systems with similar mass in the field. Specifically, satellite galaxies are preferentially gas poor and no longer forming stars, while their field counterparts are largely gas rich and actively forming stars. Much of the recent work to understand this dichotomy has been statistical in nature, determining not just that environmental processes are most likely responsible for quenching these low-mass systems but also that they must operate very quickly after infall onto the host system, with quenching timescales $\lesssim 2~ {\rm Gyr}$ at ${M}_{\star} \lesssim 10^{8}~{\rm M}_{\odot}$. This work utilizes the newly-available $Gaia$ DR2 proper motion measurements along with the Phat ELVIS suite of high-resolution, cosmological, zoom-in simulations to study low-mass satellite quenching around the Milky Way on an object-by-object basis. We derive constraints on the infall times for $37$ of the known low-mass satellite galaxies of the Milky Way, finding that $\gtrsim~70\%$ of the `classical' satellites of the Milky Way are consistent with the very short quenching timescales inferred from the total population in previous works. The remaining classical Milky Way satellites have quenching timescales noticeably longer, with $τ_{\rm quench} \sim 6 - 8~{\rm Gyr}$, highlighting how detailed orbital modeling is likely necessary to understand the specifics of environmental quenching for individual satellite galaxies. Additionally, we find that the $6$ ultra-faint dwarf galaxies with publicly available $HST$-based star-formation histories are all consistent with having their star formation shut down prior to infall onto the Milky Way -- which, combined with their very early quenching times, strongly favors quenching driven by reionization.

astro-ph.GA

Environmental Quenching of Low-Mass Field Galaxies

In the local Universe, there is a strong division in the star-forming properties of low-mass galaxies, with star formation largely ubiquitous amongst the field population while satellite systems are predominantly quenched. This dichotomy implies that environmental processes play the dominant role in suppressing star formation within this low-mass regime (${M}_{\star} \sim 10^{5.5-8}~{\rm M}_{\odot}$). As shown by observations of the Local Volume, however, there is a non-negligible population of passive systems in the field, which challenges our understanding of quenching at low masses. By applying the satellite quenching models of Fillingham et al. (2015) to subhalo populations in the Exploring the Local Volume In Simulations (ELVIS) suite, we investigate the role of environmental processes in quenching star formation within the nearby field. Using model parameters that reproduce the satellite quenched fraction in the Local Group, we predict a quenched fraction -- due solely to environmental effects -- of $\sim 0.52 \pm 0.26$ within $1< R/R_{\rm vir} < 2$ of the Milky Way and M31. This is in good agreement with current observations of the Local Volume and suggests that the majority of the passive field systems observed at these distances are quenched via environmental mechanisms. Beyond $2~R_{\rm vir}$, however, dwarf galaxy quenching becomes difficult to explain through an interaction with either the Milky Way or M31, such that more isolated, field dwarfs may be self-quenched as a result of star-formation feedback.

astro-ph.GA

Discovery and Follow-up Observations of the Young Type Ia Supernova 2016coj

The Type~Ia supernova (SN~Ia) 2016coj in NGC 4125 (redshift $z=0.004523$) was discovered by the Lick Observatory Supernova Search 4.9 days after the fitted first-light time (FFLT; 11.1 days before $B$-band maximum). Our first detection (pre-discovery) is merely $0.6\pm0.5$ day after the FFLT, making SN 2016coj one of the earliest known detections of a SN Ia. A spectrum was taken only 3.7 hr after discovery (5.0 days after the FFLT) and classified as a normal SN Ia. We performed high-quality photometry, low- and high-resolution spectroscopy, and spectropolarimetry, finding that SN 2016coj is a spectroscopically normal SN Ia, but with a high velocity of \ion{Si}{2} $λ$6355 ($\sim 12,600$\,\kms\ around peak brightness). The \ion{Si}{2} $λ$6355 velocity evolution can be well fit by a broken-power-law function for up to a month after the FFLT. SN 2016coj has a normal peak luminosity ($M_B \approx -18.9 \pm 0.2$ mag), and it reaches a $B$-band maximum \about16.0~d after the FFLT. We estimate there to be low host-galaxy extinction based on the absence of Na~I~D absorption lines in our low- and high-resolution spectra. The spectropolarimetric data exhibit weak polarization in the continuum, but the \ion{Si}{2} line polarization is quite strong ($\sim 0.9\% \pm 0.1\%$) at peak brightness.

astro-ph.SR

Under Pressure: Quenching Star Formation in Low-Mass Satellite Galaxies via Stripping

Recent studies of galaxies in the local Universe, including those in the Local Group, find that the efficiency of environmental (or satellite) quenching increases dramatically at satellite stellar masses below ~ $10^8\ {\rm M}_{\odot}$. This suggests a physical scale where quenching transitions from a slow "starvation" mode to a rapid "stripping" mode at low masses. We investigate the plausibility of this scenario using observed HI surface density profiles for a sample of 66 nearby galaxies as inputs to analytic calculations of ram-pressure and viscous stripping. Across a broad range of host properties, we find that stripping becomes increasingly effective at $M_{*} < 10^{8-9}\ {\rm M}_{\odot}$, reproducing the critical mass scale observed. However, for canonical values of the circumgalactic medium density ($n_{\rm halo} < 10^{-3.5}$ ${\rm cm}^{-3}$), we find that stripping is not fully effective; infalling satellites are, on average, stripped of < 40 - 70% of their cold gas reservoir, which is insufficient to match observations. By including a host halo gas distribution that is clumpy and therefore contains regions of higher density, we are able to reproduce the observed HI gas fractions (and thus the high quenched fraction and short quenching timescale) of Local Group satellites, suggesting that a host halo with clumpy gas may be crucial for quenching low-mass systems in Local Group-like (and more massive) host halos.

astro-ph.GA

Taking Care of Business in a Flash: Constraining the Timescale for Low-Mass Satellite Quenching with ELVIS

The vast majority of dwarf satellites orbiting the Milky Way and M31 are quenched, while comparable galaxies in the field are gas-rich and star-forming. Assuming that this dichotomy is driven by environmental quenching, we use the ELVIS suite of N-body simulations to constrain the characteristic timescale upon which satellites must quench following infall into the virial volumes of their hosts. The high satellite quenched fraction observed in the Local Group demands an extremely short quenching timescale (~ 2 Gyr) for dwarf satellites in the mass range Mstar ~ 10^6-10^8 Msun. This quenching timescale is significantly shorter than that required to explain the quenched fraction of more massive satellites (~ 8 Gyr), both in the Local Group and in more massive host halos, suggesting a dramatic change in the dominant satellite quenching mechanism at Mstar < 10^8 Msun. Combining our work with the results of complementary analyses in the literature, we conclude that the suppression of star formation in massive satellites (Mstar ~ 10^8 - 10^11 Msun) is broadly consistent with being driven by starvation, such that the satellite quenching timescale corresponds to the cold gas depletion time. Below a critical stellar mass scale of ~ 10^8 Msun, however, the required quenching times are much shorter than the expected cold gas depletion times. Instead, quenching must act on a timescale comparable to the dynamical time of the host halo. We posit that ram-pressure stripping can naturally explain this behavior, with the critical mass (of Mstar ~ 10^8 Msun) corresponding to halos with gravitational restoring forces that are too weak to overcome the drag force encountered when moving through an extended, hot circumgalactic medium.

astro-ph.GA