arXiv ScienceSearch

arXiv · 2605.26431

Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English

Abstract

We show that LLMs encode syntactic distinctions not present in the Universal Dependencies (UD) tree distances that structural probes are trained to recover. On English wh-movement stimuli, we measure the probe distance between an embedded subject and its verb, whose UD tree distance is invariant across conditions. That distance is shorter than baseline when the embedded clause is finite and longer when it is infinitival -- a within-clause sign asymmetry present in all 13 models across four families we test, at a majority of layers. No account based on UD distance, linear order, or monotone structural complexity can produce a sign reversal, while Minimalist phase theory can. Holding the matrix verb fixed while varying only the complement type reproduces the same finite-infinitival ordering in every model, ruling out a lexical-semantic explanation. The cross-clause pair separately reproduces the phase-count ordering of earlier work, validating the probes. A single activation patch, interchanging the clause-selecting matrix verb, moves the two pairs in opposite directions in 8 of 13 models. Because the causal intervention leaves the target string unchanged, it also rules out the clause-length and surface-cue explanations. Together these results reveal syntactic structure in LLMs beyond the UD probe target, and a Minimalist phase account is consistent with it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuanhao Chen, Peter Chin. 2026-07-17. Probing LLMs for Syntactic Structure Beyond Universal Dependencies: A Minimalist Phase Account in English. https://arxiv.org/abs/2605.26431

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

Small models are made more capable through distillation from a larger one that shares their tokenization scheme. However, do distilled byte and token models behave similarly in terms of scaling trends as compute and data increases? To enable this comparison, we introduce two variants to efficiently convert token logits to Byte Logits: 1) approximate: Marginalize-It, and 2) exact: End-Of-Token. We then present the first large scale study of overtraining decoder-only dense transformer models varying two dimensions simultaneously: the tokenization scheme (Tokens, Bytes, Bytes w/ eot) and the training objective (Distillation vs. Cross-Entropy), sweeping layer-parameter-matched models with roughly 1 billion parameters up to 1 trillion bytes of data. Across eight benchmarks spanning three categories: Multiple Choice QA, Language Generation, and Machine Translation, we find that Token-1B models outperform byte models (End-Of-Token-1B and Bytes-1B) in the low-FLOP regime but eventually plateau; byte models start worse yet surpass Token-1B models with more compute, reaching a higher downstream task performance ceiling. Extrapolating the average top-1 error vs. validation BPB scaling laws predicts that, asymptotically, distilled End-Of-Token-1B outperforms distilled Token-1B by up to 4%. They are also far more data efficient, matching the performance of distilled Token-1B using only one-sixth of the training data. Moreover, by operating over a small vocabulary of 256 bytes instead of on the order of 100K tokens, they circumvent the need for top-k truncation during logit dumping, while also reducing logit storage costs to roughly one-fifth. Finally, our downstream performance scaling laws predict that our distilled End-Of-Token-1B models asymptotically surpass the Llama 3.2-1B, Gemma-3-1B-pt, and Gemma 2B models on averaged downstream tasks by up to 6.5%, 8.1%, and 2.1%, respectively.

cs.CL

A Short Survey of Viewing Large Language Models in Legal Aspect

Large language models (LLMs) have transformed many fields, including natural language processing, computer vision, and reinforcement learning. These models have also made a significant impact in the field of law, where they are being increasingly utilized to automate various legal tasks, such as legal judgement prediction, legal document analysis, and legal document writing. However, the integration of LLMs into the legal field has also raised several legal problems, including privacy concerns, bias, and explainability. In this survey, we explore the integration of LLMs into the field of law. We discuss the various applications of LLMs in legal tasks, examine the legal challenges that arise from their use, and explore the data resources that can be used to specialize LLMs in the legal domain. Finally, we discuss several promising directions and conclude this paper. By doing so, we hope to provide an overview of the current state of LLMs in law and highlight the potential benefits and challenges of their integration.

cs.CL

Customized large language models can outperform Community Notes in correcting misinformation

Addressing misinformation in real-world settings is challenging: content is often multimodal; factuality judgments are nuanced and context-dependent; new events emerge rapidly across domains; corrections must be timely, trustworthy, and politically impartial; and multidimensional, multistakeholder frameworks remain lacking. Crowdsourced fact-checking systems such as Community Notes have gained broad adoption, but timely, scalable coverage remains difficult. We introduce MUSE, which augments large language models (LLMs) with trust-aware retrieval of up-to-date evidence and task-specific multimodal reasoning. Given a piece of content, MUSE identifies whether and which parts may be false or misleading and provides explanations grounded in credible references. We also develop an evaluation framework that assesses expert-rated response quality---including identification accuracy, explanation factuality, and the relevance and credibility of supporting references---as well as user perceptions. Across social media posts spanning modalities, domains, political leanings, misinformation tactics, and popularity, MUSE consistently produces high-quality responses, including for content not previously fact-checked online, and outperforms even highly rated Community Notes by 29%. It also improves participants' recognition of misinformation by 10%. Our work establishes a general methodological and evaluative framework for timely, scalable, and trustworthy correction of misinformation.

cs.CL