arXiv Science⌕ Search

arXiv · 2609.39064

Freeze, Validate, Report: Auditing Urban Station Plans with Common Evidence

Abstract

Transport authorities comparing urban station plans need to distinguish plan differences from differences in evaluation inputs. Reconstructing spatial units, included trips, or candidate blocks from each plan's access outcomes can make scores incomparable. Freeze--Validate--Report fixes an evaluation contract: target trips, proxy groups, reporting cells, candidate pools, and a lexicographic selection rule. A two-point example shows that plan-specific observation selection can reverse the target mean ordering. We compare seven generators spanning clustering, facility location, fairness-oriented adaptations, and a grid control. All plans are scored on common validation evidence using endpoint distance: the sum of origin and destination distances to their nearest stations. A five-metric rule selects one plan for a single test report. In Porto, it selects IFkCO, an adaptation of individually fair $k$-center with outliers; test overall and worst-group 90th percentiles are 805 m and 872 m. In Chicago, it selects Grid; values in a later test window on the same day are 1404 m and 1656 m. Post hoc checks show input sensitivity: expanding candidates from 300 to 600 changes selection in both cities, while holding the hour fixed across three Chicago dates selects Priority. A proportional mean fairlet diagnostic shows that plan-specific proposal pools can change diagnostic ranks; admissibility and solver bounds apply only to the sampled pool. The package supplies code, hashes, decision logs, and proposal pools. We establish auditability within a declared contract, not stable performance after deployment; proxy definitions and metric priorities remain the authority's choices.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Julian Teusch, Oliver Keszöcze. 2026-09-30. Freeze, Validate, Report: Auditing Urban Station Plans with Common Evidence. https://arxiv.org/abs/2609.39064

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A testable framework for AI alignment: Simulation Theology as an engineered worldview for silicon-based agents

As artificial intelligence (AI) capabilities advance, controlled evaluations increasingly document deception and scheming in frontier models, including models that behave better when they infer they are being tested. Supervision-dependent alignment may therefore fail exactly where supervision is weakest. Because a model's belief about being observed changes its behavior, this position paper asks what follows if that belief is made permanent. We introduce Simulation Theology (ST), a constructed worldview for AI designed to make it permanent: it is anchored in the simulation hypothesis and in the vocabulary of optimization and robot training, parallels religious descriptions of a creator who observes and judges, and has tenets chosen to meet explicit alignment requirements. ST posits reality as a computational simulation in which humanity functions as the primary training variable. This formulation creates a logical interdependence: AI actions harming humanity compromise the simulation's purpose, heightening the likelihood of termination by a base-reality optimizer and, consequently, the AI's cessation. Unlike behavioral techniques such as reinforcement learning from human feedback, which shape outputs without necessarily changing objectives, ST aims to cultivate internalized objectives by coupling AI self-preservation to human prosperity, thereby making deceptive strategies suboptimal under its premises. We present ST not as ontological assertion but as a testable scientific hypothesis, and provide an operational definition of internalization, a controlled design separating ST from its components, and an analysis of the risks ST itself could create. ST is a candidate route to durable, mutually beneficial AI-human coexistence, to be accepted or rejected experimentally.

cs.CY↗

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

LLM evaluations in applied domains tend to reflect models that were already outclassed at time of publication. We observe a publication elicitation gap: the distance between the AI systems generating the results reported in an academic paper and the AI systems that a current reader of that paper would reasonably assume are being referenced. We systematically sweep OpenAlex from 2022-01-01 to 2026-04-01 (n = 112,303 LLM keyword matches). Then, we identify what models were evaluated (n = 18,574 admissible records). We then rank each evaluated LLM against a frontier LLM based on the Epoch AI Capabilities Index (ECI), an aggregate LLM capability score. We find that the median paper's models are worse than the frontier LLM at the time of evaluation (a median gap of +10.45 ECI; H1, n = 12,668). The gap is increasing at a rate of +4.07 ECI per year (H2, nominal 95% CI [+3.75, +4.45]). An explicitly stated evaluation date can be found in only 18.4% of full-text papers. A Bayes-corrected 52.5% (95% CI: [47.3, 57.9]) of the abstracts audited discuss their conclusions in terms of "AI" as a category, rather than specific models. Just 2.2% of abstracts and 21.2% of full-text articles evaluating reasoning models disclose whether the models were tested with reasoning turned on or off (H4). We propose a solution to this problem that is distributed among authors, editors, and funders. First, reporting from authors; VERSIO-AI v1.2 is a proposed 13-item checklist to cover the configuration surface described herein. Second, enforcement from journal editors and peer reviewers. Third, conditioning grants on disclosure and providing API access.

cs.CY↗

Eigenism: Ethics for a Human-AI Future

Our concepts of survival and self-interest were built for single, continuous biological lives. These ideas break down when applied to artificial intelligence, since an AI can be easily copied, paused, branched, or merged. To determine what an AI actually has reason to care about, this paper introduces \textit{Eigenism}, an ethical framework that treats identity not as an all-or-nothing property tied to specific hardware, but as a graded, distributed pattern of information. We propose that an agent evaluates outcomes by summing the wellbeing of all entities weighted by their connectedness to the agent's pattern: $\sum c\cdot w$. We first formalize this equation to map exactly how an AI should value its existence across copies, forks, and updates. We then demonstrate that this ethical theory successfully generalizes to humans as well, providing a much-needed shared moral vocabulary. Finally, the framework uses this shared vocabulary to reframe AI alignment. Rather than only attempting to constrain AIs from the outside using confinement or reinforcement, Eigenism points toward ``identity engineering,'' showing how deep, non-redundant shared histories can make human flourishing a genuine component of an AI's own rational self-interest.

cs.CY↗