arXiv Science⌕ Search

arXiv · 2609.37499

Who Warmed the Archives? LLMs Overestimate Historical Warmth

Abstract

Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Claudiu Creanga, Liviu P. Dinu. 2026-09-27. Who Warmed the Archives? LLMs Overestimate Historical Warmth. https://arxiv.org/abs/2609.37499

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ETHER: Aligning Emergent Communication for Hindsight Experience Replay

Hindsight Experience Replay (HER) enhances sample efficiency in goal-conditioned reinforcement learning (RL) by relabelling failed trajectories with goals that were actually achieved. However, HER assumes access to a goal relabelling function and a predicate function that determines whether a goal has been satisfied. These assumptions break down in instruction-following tasks, where goals are expressed in natural language and differ from the state space. We formalize this as the Hindsight Reinforcement Learning problem, which shows the need to jointly learn these functions alongside the RL policy. To address it, we propose ETHER (Emergent Textual Hindsight Experience Replay), an agent that leverages Emergent Communication. ETHER uses a referential game (RG) to train a speaker and a listener to develop a grounded, artificial language describing environment states. It partially aligns this emergent language with instruction language using co-occurrence patterns between task instructions and RL observations. Experiments on BabyAI's PickupDist task show that ETHER's learned RG speaker and listener can function as the goal relabelling and predicate functions of HER, improving sample efficiency despite imperfect language alignment. Our work bridges Emergent Communication and goal-conditioned RL, opening the door to wider applications of HER.

cs.CL↗

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query--Key (QK) score, defined for an attention head as the inner product between the last-token query and the key at the end-of-line token following option $i$, evaluated before rotary positional embedding is applied. Its argmax identifies a universal class of select-and-copy heads in middle layers that perform option selection through semantic query--key alignment, mechanistically distinct from induction and copy-suppression heads (Olsson et al., 2022): they are invariant to label symbols, and solve a synthetic task with zero surface overlap---properties no positional-copy account explains and that critically require stripping RoPE. Across 24 models from 1.5B to 72B parameters (LLaMA-2/3/3.1/3.3, Qwen-2.5, Gemma, Phi-3.5, DeepSeek-R1-Distill), a single head's QK-score exceeds the model's own zero-shot accuracy by up to $+27.4$ pp on HellaSwag and $+49.8$ pp on HaluDialogue; causal zero-ablation collapses MCQA accuracy to near-random. To remove any dependence on labeled validation data, we introduce an unsupervised HeadScore that ranks heads from unlabeled inputs and recovers the supervised top-$k$ heads on every tested model. Against four positional-debiasing baselines (e.g., PriDe, Wiegrefe, Wang), QK-score is complementary by construction: debiasing re-weights output logits, whereas QK-score reads the model's selection from a middle-layer head before decoding. We release a one-line drop-in HeadScore script and per-model head indices, making every result one-command reproducible across all 24 models and four benchmarks.

cs.CL↗

NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NMIXX adapts existing encoders through 18.8k source-linked triplets: paraphrases and Korean-English translations preserve meaning, while targeted financial rewrites introduce semantic contrasts. We examine this recipe across seven backbones on English and Korean financial and general-domain semantic textual similarity (STS), and analyze the composition and passage lengths of KorFinSTS. BGE-M3 attains the highest adapted financial correlations in this comparison, improving from 0.1969 to 0.2967 on FinSTS and from 0.0512 to 0.2732 on KorFinSTS. Its general English and Korean correlations decrease by 0.0391 and 0.0463. Across the seven models, five improve their mean financial correlation, but all reduce their mean general-domain correlation. Per-language comparisons and benchmark-weight sensitivity analysis reveal differences obscured by a single aggregate score. The study contributes a finance-specific supervision design and evidence for evaluating adaptation jointly with retained general semantic capability; direct cross-language retrieval remains outside its evaluation scope.

cs.CL↗