arXiv Science⌕ Search

arXiv · 2610.04243

Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD

Abstract

Active Directory (AD) remains the predominant identity and access management infrastructure in enterprise environments, and its compromise represents the highest-impact outcome in internal penetration tests. Recent work has shown that large language models (LLMs) can autonomously conduct assumed-breach penetration testing against AD, but these studies employ standalone agents lacking structured guardrails, deterministic validation, and multi-stage chain orchestration. We present a benchmark evaluation of NeuroSploit v4.2.0, an open-source Rust-based autonomous pentest harness, against the Game of Active Directory (GOAD), a deliberately vulnerable multi-forest AD lab maintained by Orange Cyberdefense comprising five virtual machines, two forests, and three domains. The harness orchestrates 22 AD-specific agents and 7 multi-stage attack-chain playbooks covering the full AD kill chain: enumeration, Kerberoasting, AS-REP roasting, NTLM relay and coercion, Kerberos delegation abuse, AD CS exploitation (ESC1-ESC8), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse, and persistence detection. We benchmark nine frontier LLMs (Claude Opus 4.6/4.7/4.8, GPT-6 Astra, GPT-5.6 Sol, Grok 4.6, Qwen 3.8, GLM 5.3, Kimi k3) within the harness, comparing against direct invocation across 14 technique categories and 7 chains. The harness achieves 96-100% technique coverage with 90-97% precision, while direct invocation covers only 21-54% and produces 3.2x more false positives. Time to full three-domain compromise with Opus 4.8 was 134 minutes with guardrail activations preventing lockout-triggering sprays, unauthorized DCSync dumps, and out-of-scope reconnaissance. Results demonstrate that structured harness orchestration with domain-specialized agents, POMDP belief tracking, and cross-model voting substantially outperforms unstructured LLM usage for AD penetration testing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joas Antonio dos Santos Barbosa. 2026-10-03. Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOAD. https://arxiv.org/abs/2610.04243

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PPFedIT: Towards Privacy-Preserving Federated Instruction Tuning with Few-shot Local Examples

Instruction tuning aligns large language models (LLMs) with human intentions but requires diverse, high-quality data that are difficult to collect in privacy-sensitive domains. Federated instruction tuning (FedIT) enables collaborative training across data owners, yet existing methods typically assume sufficient local data. In realistic few-shot settings, limited samples can cause overfitting, degrade performance, and increase vulnerability to training data extraction attacks. We propose PPFedIT, a federated algorithm that improves both model performance and privacy protection in federated few-shot learning. It comprises three client-side steps: (1) synthetic data generation, which uses LLMs to diversify and enrich local data; (2) parameter isolation training, which updates the shared global LLM on synthetic data and local LLMs on private local data to mitigate synthetic-data noise; and (3) local aggregation then sharing, which mixes global and local model parameters before uploading them for server aggregation to mitigate data extraction attacks. Experiments on three open-source datasets show that PPFedIT improves model performance by an average of 8.4% and reduces the risk of data extraction attacks by approximately 20% in challenging federated few-shot settings.

cs.CR↗

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce the refusal-affirmation logit gap: the difference between the top refusal-token logit and the top affirmative-token logit at the first decoding step. This single scalar quantifies the per-prompt safety margin that alignment provides. Empirically, alignment widens the gap on 97.5-99.8% of toxic prompts across three model families, and median gap closure co-varies with True-ASR ranking across suffix strategies (an internal consistency check, since our method optimises gap closure). To validate the metric's practical significance, we present logit-gap steering, a gradient-free, forward-pass-only method that discovers short in-distribution suffixes ($<$10 tokens per component) whose cumulative effect closes the gap. The method requires ${\approx}26{,}000$ forward-pass equivalents per family (${\approx}2$~min on one A100), ${\approx}125\times$ less than a single GCG search. Suffixes discovered on 0.5B--2B models transfer without modification to 72B within family. An 8-suffix ensemble reaches 38-96\% True ASR across 13 models on AdvBench and HarmBench, with most suffixes having $10^{3}$-$10^{4}\times$ lower perplexity than GCG-meaning published perplexity-filter defenses that collapse GCG (64.7%$\to$1.0%) leave our suffixes nearly intact (76.9%$\to$76.0%). These results demonstrate that current alignment margins, while consistently present, can be thin and efficiently measurable, and that defense strategies must account for in-distribution suffixes.

cs.CR↗

Spoofing Missed-Detection Bounds for PRF GNSS Ranging Authentication Under AWGN Models

Pseudorandom-function (PRF) ranging codes, such as those used in Galileo's encrypted E6-C under the Signal Authentication Service (SAS), enable a receiver to authenticate pseudoranges once the PRF secret is revealed. This work bounds how much authentication security the receiver obtains under Additive White Gaussian Noise (AWGN) assumptions. Against a spoofer that does not estimate the code before submitting its forgery, PRF security makes the forged correlation zero-mean up to the security of the underlying PRF, allowing integration time and C/N$_0$ to mostly determine probability of missed detection (PMD) and probability of false alarm (PFA). Against such a spoofer at a conservative 30 dB-Hz, 400 ms of E6-C aggregation certifies a PMD below $2^{-128}$ (plus any PRF advantage). For a spoofer that estimates chips before submitting a forgery, I derive the receiving-antenna gain at which authentication security breaks, which is about 12 dB for E6-C for the adversaries modeled. This work can be used to design a PRF GNSS ranging code protocol and a receiver capable of correctly asserting PRF ranging security assuming an AWGN model.

cs.CR↗