arXiv ScienceSearch

arXiv subjects

Dingli Liang

Publications and source records attributed to Dingli Liang.

5 recordsLinked to original sources

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is close but not exact, together with scores for the other prediction tasks. Across systems, similar result accuracy can mask larger differences in detailed predictions. Four systems predicted champion Spain, and two of them also recovered the exact final pairing. Compared with betting-market and human-fan baselines, the best system shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline. New schedules can be added as they begin, allowing the benchmark to evaluate future models without using outcomes that are already known. Code, predictions and evaluation scripts will be publicly released.

cs.AI

Holder Policy Optimisation

Group Relative Policy Optimisation (GRPO) enhances large language models by estimating advantages across a group of sampled trajectories. However, mapping these trajectory-level advantages to policy updates requires aggregating token-level probabilities within each sequence. Relying on a fixed aggregation mechanism for this step fundamentally limits the algorithm's adaptability. Empirically, we observe a critical trade-off: certain fixed aggregations frequently suffer from training collapse, while others fail to yield satisfactory performance. To resolve this, we propose \textbf{HölderPO}, a generalised policy optimisation framework unifying token-level probability aggregation via the Hölder mean. By explicitly modulating the parameter $p$, our framework provides continuous control over the trade-off between gradient concentration and variance bounds. Theoretically, we prove that a larger $p$ concentrates the gradient to amplify sparse learning signals, whereas a smaller $p$ strictly bounds gradient variance. Because no static configuration can universally resolve this concentration-stability trade-off, we instantiate the framework with a dynamic annealing algorithm that progressively schedules $p$ across the training lifecycle. Extensive evaluations demonstrate superior stability and convergence over existing baselines. Specifically, our approach achieves a state-of-the-art average accuracy of $54.9\%$ across multiple mathematical benchmarks, yielding a substantial $7.2\%$ relative gain over standard GRPO and secures an exceptional $93.8\%$ success rate on ALFWorld.

cs.LG

On Non-Noetherian Iwasawa Theory

We prove a general structure theorem for finitely presented torsion modules over a class of commutative rings that need not be Noetherian. As a first application, we then use this result to study the Weil- étale cohomology groups of $\mathbb{G}_m$ for curves over finite fields.

math.NT

On the Iwasawa asymptotic class number formula for $\mathbb{Z}_p^r\rtimes\mathbb{Z}_p$-extensions

Let $p$ be an odd prime and $F_{\infty,\infty}$ a $p$-adic Lie extension of a number field $F$ with Galois group isomorphic to $\mathbb{Z}_p^r\rtimes\mathbb{Z}_p$, $r\geq 1$. Under certain assumptions, we prove an asymptotic formula for the growth of $p$-exponents of the class groups in the said $p$-adic Lie extension. This generalizes a previous result of Lei, where he establishes such a formula in the case $r=1$. An important and new ingredient towards extending Lei's result rests on an asymptotic formula for a finitely generated (not necessarily torsion) $\mathbb{Z}_p[[\mathbb{Z}_p^r]]$-module which we will also establish in this paper. We then continue studying the growth of $p$-exponents of the class groups under more restrictive assumptions and show that there is an asymptotic formula in our noncommutative $p$-adic Lie extension analogous to a refined formula of Monsky (which is for the commutative extension) in a special case.

math.NT