arXiv Science⌕ Search

arXiv · 2610.02297

From Alert Floods to Precedence Forests: Zero-Prior-Knowledge Incident Triage with LOGOS

Abstract

Commercial observability platforms rely on domain artifacts like distributed traces, topology maps, and baseline metrics. However, when troubleshooting proprietary software, enterprise operators are left with only raw, unannotated text logs. We explore the extreme boundary of log-only diagnosis: To what extent can we isolate failure propagation using strictly raw text logs? We present LOGOS, an unsupervised system that exploits entity-event co-occurrence and temporal precedence to collapse millions of raw log lines into a compact precedence forest. Evaluated across 25 production enterprise outages and 12 open-source issues, LOGOS operates with zero prior knowledge---requiring no seed queries, observed symptoms, or pre-defined incident boundaries. In a median wall-time of 4.5 minutes, LOGOS eliminates a median 99.8% of background noise, achieves 0.76 mean recall, and detects failure cascades with a 16-hour median diagnosis-verified lead time---consolidating alert floods 124x to enable 80% enterprise (100% open-source) zero-shot LLM root-cause accuracy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Radhika Niranjan Mysore. 2026-10-01. From Alert Floods to Precedence Forests: Zero-Prior-Knowledge Incident Triage with LOGOS. https://arxiv.org/abs/2610.02297

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Grassroots Bonds: Financing by the People, for the People

Grassroots currencies turn mutual trust into liquidity: a grassroots coin is a unit of its issuer's debt, backed by the issuer's goods and services, which the issuer must redeem, 1-for-1, against any coin they hold; liquidity arises from mutual credit lines, formed by the voluntary exchange of coins among persons who trust each other. As coins are redeemable 1-for-1, the exchange must be 1-for-1 as well, lest prompt redemption after it leave one party with undue profit. Thus, grassroots coins are incongruent with interest-bearing credit. Here, we extend grassroots currencies to include also grassroots bonds, units of their issuer's debt due at a later date. Upon maturity, the bearer of the bond may redeem it against a coin of the issuer, so liquid coins can be lent against interest-bearing bonds. We specify bonds by adding Date and Escrow clauses to the social contract of grassroots currencies: the contract enforces the redemption of a bond when it is mature according to the date stated by the issuer, who undertakes to keep it current. We show that the voluntary swap of coins and bonds can realise the basic financial instruments: loans, sale of debt, and forward contracts, and, with an escrow agent, loans with payment schedules, options, collateral, guarantees, insurance, credit default swaps, letters of credit, and credit lines. We extend the liquidity ratios of corporate finance to bonds, with maturity as asset class, and prove that a community clears its debts by redemptions exactly when no member owes more than they hold, and that it does so without coordination once dates advance and persons act on their rights. Grassroots currencies that include coins and bonds are implemented in GLP, a concurrent logic programming language running on Dart, as a program derived from the contract and demonstrated by a village market of six agents and an escrow agent.

cs.DC↗

Characterization-Guided GPU Fault Resilience in NVIDIA MPS

NVIDIA Multi-Process Service (MPS) enables fine-grained GPU sharing by allowing multiple processes to execute concurrently on the same GPU, making it an important mechanism for improving GPU utilization. However, MPS has weak fault resilience: a fault in one process can terminate all co-running processes, limiting its adoption in resilience-critical settings such as multi-tenant GPU clusters. In this work, we design fault-resilient MPS to solve this problem. Our design is guided by insights from a systematic characterization of GPU faults and a deep analysis of their end-to-end processing pipeline. Based on these insights, we design two complementary mechanisms. First, we design a fault isolation mechanism for the dominant memory-related faults that can be fully isolated while preserving process-level fail-stop semantics by software intervention in the open GPU driver kernel module. For other faults whose process is within proprietary software, we design a fast-recovery substrate that combines virtual-memory-based GPU-resident state sharing with pre-initialized standbys. Our evaluation across GPUs and workloads demonstrates effective fault isolation and fast recovery with minimal overhead: in an end-to-end case study, isolation incurs no visible outage, while recovery restores pre-fault throughput in 355\,ms.

cs.DC↗

ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum

Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.

cs.DC↗