arXiv ScienceSearch

arXiv subjects

Daniel Graham

Publications and source records attributed to Daniel Graham.

10 recordsLinked to original sources

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents

Large language model (LLM) agents are increasingly applied to penetration testing, but we still know little about what they can do or how they fail. We compare two PentestGPT-based systems: a legacy human-in-the-loop system running the open-weight Kimi K2.5, and a newer autonomous system running Claude Opus 4.8. Across three public targets, the autonomous system solves all three, including the two the legacy system never finishes. The legacy result is the more surprising of the two. Even on the machines the legacy system fails to solve, it completes about half the subtasks, while running on ordinary university GPUs with no provider guardrails. We can describe the trend but not explain it, since model, harness, autonomy, and memory architecture all change together. Its direction still points to the next question: what will limit these agents as they take on more complex tasks? The usual answer is long-horizon memory, the loss of access to earlier findings during long attack chains. We test it by adding a coverage-memory layer to both systems, and neither improves outcomes. In the legacy stalled runs we could review, the limiting factor appeared to be planning and commitment rather than lost memory: agents held the evidence for a route forward and never turned it into a concrete exploitation hypothesis, which may suggest that offensive capability will advance with agents' ability to plan rather than with better memory. The same subtask scoring that tracks this capability is available to defenders, who can measure it as it rises instead of waiting to meet it in the field.

cs.CR

The Normalization of Deviance in AI Development

Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the institutions building these systems are themselves predisposed to drift toward failure. This paper argues that they are. Regardless of how capable AI systems become, the organizations building them face the same structural dynamics that preceded past major technological disasters. Drawing on case studies of the Space Shuttle Challenger, the Three Mile Island accident, and the Boeing 737 MAX crashes, this paper identifies the common structural mechanisms preceding each failure and maps them onto contemporary AI development. The findings suggest that existing safety infrastructure may provide less protection than it appears, as organizations can complete safety processes in full compliance and still produce catastrophic outcomes. The pre-disaster period of AI development is still underway; the purpose of this paper is to make these dynamics legible while they can still be interrupted.

cs.AI

Tiny Enough to Break In: Agentic Remote Access Trojans Powered by Small Language Models

Agentic artificial intelligence raises a new security concern: cyber threats that reason, act, and adapt locally without continuous human direction. We examine this threat through an Agentic Remote Access Trojan (agentic RAT): a Remote Access Trojan augmented with a locally deployed Small Language Model (SLM). The SLM interprets host and network observations, selects actions, recovers from failed steps, and reduces reliance on an external operator. We implement the concept in a controlled, network-isolated lab built from Kali Linux, a Metasploitable2 target, LM Studio, and a local 8-billion-parameter Dolphin-family model. We then test whether a model this small can support autonomous cyber decision-making. This is architecturally feasible today. On commodity hardware, with no cloud service and no operator in the loop, the SLM closed the full observe-decide-act cycle: it interpreted ranked reconnaissance evidence supplied by the controller, selected actions, and obtained verified root-shell access on real vulnerable services. However, it is not yet operationally reliable. The same model hallucinated commands, misread output, and recovered from failure inconsistently, completing 10.9% of a deliberately strict checklist. That gap reflects the limits of today's small models, not a ceiling on the concept. As SLMs improve, agentic endpoint systems may become more practical, more autonomous, and harder to detect, straining existing monitoring, containment, and policy-enforcement mechanisms. Real-world incidents in 2025-2026 already show AI-driven intrusions moving from concept toward practice. That makes the local, self-contained variant we study a plausible near-term direction, not a hypothetical one.

cs.CR

Coherent Control of Quantum-Dot Spins with Cyclic Optical Transitions

Solid-state spins are promising as interfaces from stationary qubits to single photons for quantum communication technologies. Semiconductor quantum dots have excellent optical coherence, exhibit near unity collection efficiencies when coupled to photonic structures, and possess long-lived spins for quantum memory. However, the incompatibility of performing optical spin control and single-shot readout simultaneously has been a challenge faced by almost all solid-state emitters. To overcome this, we leverage light-hole mixing to realize a highly asymmetric lambda system in a negatively charged heavy hole exciton in Faraday configuration. By compensating GHz-scale differential Stark shifts, induced by unequal coupling to Raman control fields, and by performing nuclear-spin cooling, we achieve quantum control of an electron-spin qubit with a $\pi$-pulse contrast of 97.4% while preserving spin-selective optical transitions with a cyclicity of 471 (50). We demonstrate this scheme for both GaAs and InGaAs quantum dots, and show that it is compatible with the operation of a nuclear quantum memory. Our approach thus enables repeated emission of indistinguishable photons together with qubit control, as required for single-shot readout, photonic cluster-state generation, and quantum repeater technologies.

quant-ph

Earth's Alfv\'en wings driven by the April 2023 Coronal Mass Ejection

We report a rare regime of Earth's magnetosphere interaction with sub-Alfv\'enic solar wind in which the windsock-like magnetosphere transforms into one with Alfv\'en wings. In the magnetic cloud of a Coronal Mass Ejection (CME) on April 24, 2023, NASA's Magnetospheric Multiscale mission distinguishes the following features: (1) unshocked and accelerated cold CME plasma coming directly against Earth's dayside magnetosphere; (2) dynamical wing filaments representing new channels of magnetic connection between the magnetosphere and foot points of the Sun's erupted flux rope; (3) cold CME ions observed with energized counter-streaming electrons, evidence of CME plasma captured due to reconnection between magnetic-cloud and Alfv\'en-wing field lines. The reported measurements advance our knowledge of CME interaction with planetary magnetospheres, and open new opportunities to understand how sub-Alfv\'enic plasma flows impact astrophysical bodies such as Mercury, moons of Jupiter, and exoplanets close to their host stars.

physics.space-ph

Blankets, Heat, and Why Free Energy Has Not Illuminated the Workings of the Brain

What can we hope to learn about brains from the free energy principle? In adopting the "primordial soup" physical model, Bruineberg et al. perpetuate the unsupported notion that the free energy principle has a meaningful physical--and neuronal--interpretation. We examine how minimization of free energy arises in physical contexts, and what this can and cannot tell us about brains.

q-bio.NC

Observations of rapidly growing whistler waves in front of space plasma shock

Whistler mode wave is a fundamental perturbation of electromagnetic fields and plasmas in various environments including planetary space, laboratory and astrophysics. The origin and evolution of the waves are a long-standing question due to the limited instrumental capability in resolving highly variable plasma and electromagnetic fields. Here, we analyse data with the high time resolution from the multi-scale magnetospheric spacecraft in the weak magnetic environment (i.e., foreshock) enabling a relatively long gyro-period of whistler mode wave. Moreover, we develop a novel approach to separate the three-dimensional fluctuating electron velocity distributions from their background, and have successfully captured the coherent resonance between electrons and electromagnetic fields at high frequency, providing the resultant growth rate of unstable whistler waves. Regarding the energy origin for the waves, the ion distributions are found to also play crucial roles in determining the eigenmode disturbances of fields and electrons. The quantification of wave growth rate can significantly advance the understandings of the wave evolution and the energy conversion with particles.

physics.space-ph

Why is neural connection weight a weak predictor of correlated neural activity?

As the field of connectomics has matured, it has expanded from mapping the existence of connections between brain components to measuring the strength of connections. This information is increasingly accessible via methodologies such as pairing functional magnetic resonance (MR) imaging and MR tractography in the same human subject, as well as novel methods in non-human animals using optogenetics. Systems and network neuroscience have in recent years focused extensively on explaining correlation patterns of functional activity in the brain in terms of the degree of connectedness of brain components, the so-called functional connectivity-structural connectivity relationship (SC-FC). What has been surprising has been how low the SC-FC correlations are. Why is it that brain parts that are more well-connected appear not to engender more correlated activity between them? Several explanations have been proposed but one possibility has not been considered: perhaps more neural activity doesn't imply more functional involvement. This article examines this possibility and proposes a new general framework for understanding brain network dynamics based on the design constraints of large-scale network communication systems. With this new perspective, we may start to answer the article's main question, and perhaps others.

q-bio.NC

Fully quantum embedding with density functional theory for full configuration interaction quantum Monte Carlo

In common with many high-accuracy electronic structure methods, the initiator adaptation of full configuration interaction quantum Monte Carlo (i-FCIQMC) has difficulty treating realistic systems with large numbers of electrons. This barrier has prevented the application of i-FCIQMC to questions of catalysis that, even for the simplest of models, require high-accuracy modeling of several features of the electronic structure, such as strong and dynamic correlation, and localized vs. delocalized bonding. We here present a fully-quantum embedded version of i-FCIQMC , which we apply to calculate the bond dissociation energy of an ionic bond (LiH) and a covalent bond (HF) physisorbed to a benzene molecule. The embedding is performed using a recently-developed Huzinaga projection operator approach, which affords good synergy with i-FCIQMC by minimizing the number of orbitals in the calculation. We find that, without embedding, i-FCIQMC struggles to converge these calculations due to their substantial system sizes and a lack of error cancellation between reactants and products. With embedding, the i-FCIQMC calculation converges straightforwardly to CCSD(T) benchmarks. Our results suggest that embedded i-FCIQMC will be able treat system sizes well beyond our current reach (even though embedding introduces an error). We discuss how embedding might be improved (and thus the introduced error reduced) using i-FCIQMC energies as benchmarks.

physics.chem-ph

The Human Cell Atlas White Paper

The Human Cell Atlas (HCA) will be made up of comprehensive reference maps of all human cells - the fundamental units of life - as a basis for understanding fundamental human biological processes and diagnosing, monitoring, and treating disease. It will help scientists understand how genetic variants impact disease risk, define drug toxicities, discover better therapies, and advance regenerative medicine. A resource of such ambition and scale should be built in stages, increasing in size, breadth, and resolution as technologies develop and understanding deepens. We will therefore pursue Phase 1 as a suite of flagship projects in key tissues, systems, and organs. We will bring together experts in biology, medicine, genomics, technology development and computation (including data analysis, software engineering, and visualization). We will also need standardized experimental and computational methods that will allow us to compare diverse cell and tissue types - and samples across human communities - in consistent ways, ensuring that the resulting resource is truly global. This document, the first version of the HCA White Paper, was written by experts in the field with feedback and suggestions from the HCA community, gathered during recent international meetings. The White Paper, released at the close of this yearlong planning process, will be a living document that evolves as the HCA community provides additional feedback, as technological and computational advances are made, and as lessons are learned during the construction of the atlas.

q-bio.TO