arXiv ScienceSearch

subject

cs.LG

cs.LG: explore 2000 source-linked works published from 2019 to 2026, with original documents and citations.

This collection is a preview while coverage and quality are evaluated.

Search within this collection

Coverage and selection

Includes records with this source-supplied label or an explicit phrase match in their metadata. Matches indicate a mention, not proof that a paper uses a method or tests a material. Source versions are consolidated by DOI.

Sources: arxiv. Collection updated 2026-09-15. Counts describe this index, not the complete source archives.

Marginal-Contribution Policy Gradients under Filtered Feedback for Multi-Agent LLMs

We develop a unified treatment of credit assignment for RL training in multi-agent LLM systems. We show that observed reward alone cannot distinguish an agent that determines it from one that never affects it, and that standard shared-reward training performs exact gradient ascent on each agent's private utility rather than system performance. Moreover, we prove no single scalar per agent can consistently account for joint performance once agents interact. We thus develop the unique background-dependent notion of marginal contribution satisfying natural consistency requirements. From it we derive gradient-correct marginal contribution training signals, identify them from filtered feedback, and optimally allocate a budget of exact counterfactual evaluations against learned-signal error. Instantiated in GRPO, our signal improves routed GSM8K accuracy over winner-take-all training at no extra generation cost.

cs.LG

Negative Ontology of True Target for Machine Learning: Towards Recognition, Evaluation and Learning under Democratic Supervision

This article philosophically examines how a shift in the assumed ontology of the true target (TT) can lead to a new paradigm for machine learning (ML)-based predictive modelling. By systematically analysing the existence assumption of the TT underlying mainstream ML paradigms, we adopt a negative ontology perspective, explicitly positing that the TT does not objectively exist in the real world as a universally accessible object. On this basis, we define Democratic Supervision as an alternative supervisory principle for ML, in which no single source is assumed to possess an objectively privileged target. We further introduce Multiple Inaccurate True Targets (MIATTs) as an instance-level realization of Democratic Supervision. Building upon MIATTs, we establish the logic-driven generation and assessment for MIATTs construction (recognition with MIATTs), formulate logical assessment formula for evaluation with MIATTs, and develop undefinable true target learning for learning with MIATTs. These components are integrated to formulate the Recognition, Evaluation, Learning with MIATTs (REL-MIATTs) framework. We further characterize REL-MIATTs from the perspective of Cognitive Machine Learning: a single cycle provides an elementary cognitive learning unit, MIATTs introduce complementary supervisory perspectives, and iterative cycles support the evolution of learning through accumulated experience, feedback, and state updating. This provides a basis for human-AI co-evolution. We examine how REL-MIATTs support human-AI co-evolution and adaptive knowledge discovery in a synthetic controlled environment. A real-world application further demonstrates the potential of the framework for supporting individual education and professional development, providing empirical evidence for the feasibility of Democratic Supervision, REL-MIATTs, and its broader implications for continuous human-AI co-evolution.

cs.LG

Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity Education

The escalating complexity of modern cyber threats demands innovative approaches to security education that transcend traditional pedagogical methods. Conventional training paradigms often fail to engage learners meaningfully or develop the intuitive reasoning necessary for effective threat recognition. This paper introduces an interactive educational framework that reimagines cybersecurity awareness through the lens of a structured guessing game. Our approach integrates explainable artificial intelligence (XAI) principles with reinforcement learning to create a dynamic learning environment where users discover cybersecurity concepts through guided inquiry. The proposed system employs a policy-based reinforcement learning agent that assumes the role of a knowledgeable questioner, systematically narrowing down user-described security scenarios until it can both identify the underlying threat and provide transparent reasoning for its conclusion. By framing security education as an interactive dialogue, we transform passive knowledge acquisition into active discovery. We present the complete system architecture, detail the underlying algorithmic foundations, and demonstrate practical application through comprehensive case studies examining diverse attack vectors including the Cyber Kill Chain, phishing campaigns, ransomware outbreaks, and web application vulnerabilities. This work represents a significant departure from static security training methodologies, offering a personalized and game-based approach to cybersecurity education.

cs.CY

Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits

While parameterized quantum computations have shown success in standard reinforcement learning (RL), whether these advantages adapt to hierarchical RL (HRL) remains a critical open question. This work demonstrates that variational quantum circuits (VQCs) can effectively enhance HRL agents based on the option-critic architecture. Evaluated in standard environments, a hybrid HRL agent with a quantum feature extractor outperforms classical baselines while using fewer parameters. We also identify an architectural bottleneck: using VQCs for option-value estimation severely degrades learning. Further ablations reveal how quantum circuit design affects performance. Our work establishes design principles for parameter-efficient hybrid HRL agents.

cs.LG

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher model should be used, and in the case of self-distillation, which specific context should serve as the supervisory signal? Does the optimal choice vary from one token to the next? At present, addressing these questions typically requires costly training runs whose aggregate performance metrics obscure the dynamics at the level of individual tokens. We introduce a training-free diagnostic framework that operates at the highest resolution: per token, per question, and per teacher. We derive an ideal per-node gradient defined as the parameter update that maximally increases the student's probability of success. We then develop a scalable targeted-rollout algorithm to estimate this gradient efficiently, even for long chains of intermediate thoughts. The gradient alignment score, defined as the cosine similarity between this ideal gradient and any given distillation gradient, quantifies the extent to which a particular configuration approximates the ideal signal. Across a range of self-distillation settings and external teacher models, we observe that distillation guidance exhibits substantially higher alignment with the ideal on incorrect rollouts than on correct ones, where the student already performs well and the teacher's signal tends to become noisy. Furthermore, we find that the optimal distillation context depends jointly on the student model's capacity and the target task, and that no single universally effective configuration emerges. These findings motivate the use of per-task, per-token diagnostic analyses for distillation.

cs.LG

Geometric Dictionary Learning of Dynamical Systems with Optimal Transport

Learning dynamical systems through operator-theoretic representations provides a powerful framework for analyzing complex dynamics, as spectral quantities such as eigenvalues and invariant structures encode characteristic time scales and long-term behavior. However, dynamical operators are typically estimated independently for each system, preventing the discovery of shared structure across related dynamics. To address this limitation, we posit that related dynamical systems lie near a low-dimensional manifold in spectral operator space. Based on this hypothesis, we introduce DOODL (Dynamical OperatOr Dictionary Learning), a framework that learns a dictionary of characteristic spectral dynamics whose combinations approximate this manifold and yield compact, interpretable embeddings of individual systems. Beyond representation learning, DOODL enables fast and interpretable operator estimation from short and partially observed trajectories by constraining the estimation to the learned operator manifold. Experiments on metastable Langevin dynamics and turbulent plasma simulations demonstrate that DOODL scales to highly complex multiscale regimes while capturing characteristic spectral structure governing the dynamics rather than merely fitting trajectories, achieving errors one to two orders of magnitude lower than independent operator estimation methods in challenging low-data regimes.

stat.ML

WaveGraphNet: Physics-Consistent Guided-Wave Damage Localization through Coupled Inverse-Forward Graph Learning

Guided-wave structural health monitoring enables damage localization in composite plates using sparse networks of bonded piezoelectric transducers. However, supervised localization remains weakly constrained when measurements are available from only a limited set of damage locations. Because exhaustive spatial coverage is impractical, models trained at observed locations may generalize poorly to unseen regions. We propose WaveGraphNet, an inverse-forward graph-learning framework for guided-wave damage localization in carbon-fiber-reinforced polymer (CFRP) plates. The sensing network is represented as a graph, with transducers as nodes and measured pitch-catch paths as edges. The inverse branch uses order-invariant message passing and a geometry-constrained decoder to map path-wise energy deviations to a damage coordinate. An independently trained forward branch predicts the energy-deviation pattern associated with a candidate coordinate. At test time, gradients propagated through the forward branch update only the inverse-predicted coordinate to reduce the discrepancy between the measured and predicted response patterns. The framework is evaluated on the OGW-1 benchmark using three spatial hold-out splits in which complete damage regions are excluded from training. Comparisons include non-graph and graph-learning baselines together with the training-free Reconstruction Algorithm for Probabilistic Inspection of Damage (RAPID). Under the validation-selected settings, refined WaveGraphNet achieves the lowest reported test mean absolute error on all three splits. Refinement reduces the standalone localization error by 23.6 percent-67.9 percent, with an average relative reduction of 50.4 percent. These results demonstrate that learned forward-response consistency can improve guided-wave localization under limited spatial training coverage.

cs.LG

EmoTrack: Clinical-Semantic Modeling for Text-Based Depression Severity Estimation

Text-based counseling provides a valuable source of information for assessing depression severity. We study prediction of the total score on the eight-item Patient Health Questionnaire (PHQ-8), a self-report measure of depression severity, from counseling transcripts. Clinical-based methods rely mainly on large language model (LLM) inference to obtain structured session-level assessments, but these assessments provide limited information about which utterances support each score. Training-based methods train predictors directly on sentences or their semantic embeddings and preserve local conversational detail, but must learn clinical structure from limited labeled transcripts. Integrating a small set of clinical feature scores with a long sequence of high-dimensional utterance representations is challenging, as simple fusion can overemphasize dialogue evidence and fail to link clinical features to supporting utterances. Combining structured clinical assessments with sentence-level semantics, we propose EmoTrack, which jointly encodes clinical feature scores and utterance representations to predict the total PHQ-8 score. For longitudinal assessment across successive sessions, the model can also incorporate a compressed representation of the preceding session as optional memory. On the real-world DAIC-WOZ benchmark, EmoTrack reduces mean absolute error from 2.8234 to 2.4708 relative to the strongest evaluated baseline. Across five settings covering single-session estimation, cross-dataset transfer without adaptation, and longitudinal assessment, it reduces normalized aggregate error by approximately 6.1% relative to the strongest baseline.

cs.LG

The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer

This book provides a compact, derivation-oriented introduction to the mathematical foundations of modern generative artificial intelligence. Rather than surveying every recent architecture or implementation detail, it develops a coherent route through the ideas connecting major families of generative models, from PCA, probabilistic PCA, variational autoencoders, and diffusion models to normalising flows, autoregressive factorisations, GANs, Wasserstein GANs, and energy-based models. The aim is to make the structure of generative modelling more accessible without removing the mathematical substance needed to understand how these models are derived and related. The book is intended as a foundation-building primer for mathematically curious researchers, practitioners, and students.

cs.LG

From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets

Standard machine learning pipelines often admit many near-optimal models. These "Rashomon sets" pose a range of challenges and opportunities for uncertainty-aware, robust decision making. They allow users to incorporate domain knowledge and preferences that would otherwise be difficult to specify directly in an objective, and they quantify diversity among valid models for a given training dataset and objective function. However, computation of Rashomon sets, even for simple, interpretable model classes such as sparse decision trees, continues to require immense memory and runtime resources. We present PRAXIS, an algorithm to approximate this Rashomon set with orders of magnitude improvement in runtime and memory usage. We validate that PRAXIS regularly recovers almost all of the full Rashomon set. PRAXIS allows researchers and practitioners to scalably model the Rashomon set for real-world datasets. Code for PRAXIS is available at https://github.com/zakk-h/PRAXIS

cs.LG

Subliminal Learning is a LoRA Artifact

Subliminal learning is a phenomenon where language models can transmit behavioral traits to other models through seemingly innocuous data (Cloud et al., 2025). In subliminal learning, a teacher model with a behavioral trait (e.g. obsession with cats) can transmit this cat obsession to a student model finetuned only on numerical sequences generated by the teacher. In this paper, we ask: how does this unexpected behavioral transmission occur? We show that subliminal learning is a LoRA artifact. When subliminal learning occurs, transmission has an inverted U-shaped relationship with LoRA rank; it also disappears with full finetuning. We show that subliminal learning is highly dependent on the context seen during finetuning and evaluation. For example, a Qwen model with the default system prompt during finetuning ("You are Qwen, created by Alibaba Cloud. You are a helpful assistant.") does not show subliminal learning during generation when no system prompt is included. We further demonstrate that subliminal behavior is localized to computation at tokens seen during both finetuning and evaluation (e.g. the model's default system prompt, the standard chat template tokens, etc.). Overall, subliminal learning seems to be a fragile artifact of LoRA hyperparameters and finetuning context, making it an unstable channel for behavioral transmission.

cs.AI

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning

Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require fragile trade-offs that sacrifice performance for stochasticity or rely on heuristic metrics that can misalign policy rankings. We argue that diversity is more naturally understood as the rational response to uncertainty in the reward. When the reward function is not perfectly known--as is the case with ambiguous preferences or imperfect reward models--committing to a single action can be sub-optimal. Building on this, we propose a fundamental reformulation of the RL objective by replacing the scalar reward with a distribution over reward functions, and applying a non-linear objective over sets of actions. The result is a framework in which calibrated behavioural diversity emerges naturally, remains controllable through the reward function distribution, and is obtained without sacrificing expected reward. Focusing on the contextual bandit setting as commonly used in large language model (LLM) post-training, we derive a principled gradient estimator for this objective and prove that our formulation naturally generalizes both vanilla policy gradient and more recently developed action-set approaches. We provide didactic experiments which complement our theoretical results, and our large-scale empirical results in LLM reasoning further demonstrate that this framework offers a robust and theoretically grounded alternative for complex RL tasks where the traditional formulation of the problem fails to induce the desired breadth of agent behaviour.

cs.LG

Rollout-Level Advantage-Prioritized Experience Replay for GRPO

Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single gradient update and then discarded. Naive replay is not well suited in this setting because LLM policies drift quickly per gradient step. Stored rollouts therefore become stale and can destabilize training. We propose a rollout-level replay buffer for GRPO that stores and samples individual rollouts rather than whole groups. The buffer bounds staleness through age eviction. Any rollout older than tau_max training steps is removed. The buffer also preserves on-policy data via fresh-anchored composition. Each batch keeps its fresh on-policy rollouts and then concatenates replay rollouts drawn separately from the buffer. We prioritize replay by per-rollout advantage magnitude and recycle individual rollouts whose advantages are large. Across three Qwen3-Base scales on five math benchmarks, our method outperforms GRPO and naive replay baselines. Gains are positive at every scale and reach +1.66 pp on the five-benchmark average at 4B. Under an AES metric that jointly measures accuracy and token efficiency, our method is the only condition with a positive margin over GRPO at every scale.

cs.LG

In-Context Multiple Instance Learning

Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied in fields ranging from computational pathology to satellite imagery. Nevertheless, existing algorithms struggle in the low-label regime that characterizes many real-world applications. Flexible models overfit and rigid ones fail to adapt to the task at hand. We show that pretraining an in-context learner with a Perceiver-style architecture on synthetic data yields a model that can solve new tasks from a handful of labeled bags. At inference time, classification happens in a single forward pass and requires no gradient updates. We propose and investigate different synthetic data generators for bag-structured data and find that they capture complementary inductive biases. A model pretrained on a mixture of these generators inherits their per-task strengths and achieves the best average performance across twelve MIL benchmarks, outperforming supervised baselines that require task-specific training.

cs.LG

HAARES Half-Split Residual Basis Routing for Deep Transformers

Block-level residual routing makes learned residual aggregation practical by routing over block summaries, but each summary compresses an ordered sequence of attention and MLP updates into one cumulative vector. We propose \method{}, a lightweight residual basis router that keeps the cumulative block source and adds one half-split detail basis, computed as the difference between first-half and second-half residual updates. The detail basis is RMS-matched and updated online, exposing coarse intra-block trajectory information without dense sublayer-level routing. Across OpenWebText, cross-domain character-level benchmarks, and BPE-tokenized OpenWebText, the empirical pattern is depth-dependent: gains are small or mixed at shallow depth and most reliable in 48-layer models. In the 201M 48-layer setting, \method{} improves over Block AttnRes across all three seeds, while a 453M two-seed probe shows the same direction. Ablations rule out source duplication, random signed details, fixed detail-source biases, or block-count changes alone. Cost analysis shows that the method is FLOP-light but not wall-clock-free: it adds memory and routing overhead, yet its relative arithmetic cost is amortized as width grows and earlier convergence can reduce time-to-target.

cs.LG

How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions

Merchant information extraction turns noisy financial transaction descriptors into structured fields at production scale. Our deployed LoRA-fine-tuned LLaMA~3.1-8B reaches 96.95\% F1, but its memory and throughput motivate smaller replacements. We evaluate 23 retained fine-tuning runs plus a separately trained production reference, spanning Gemma~3 (270M--4B), Qwen~3.5 (0.8B--4B), Aya~3.35B, and LLaMA~3.1-8B across LoRA ranks, prompts, training templates, and serving environments. A rank-8 LLaMA fine-tune reaches 96.75\% F1, only 0.20 points below the rank-32 production reference. Qwen~3.5~4B with JSON-Only prompting reaches 96.60\% F1 and strict record-level exact match of 91.67\%, with a $3.8\times$ lower inverse-throughput time estimate than the rank-8 8B model. Qwen~3.5~0.8B reaches 94.75\% F1, and Qwen Think and Nothink templates differ by less than 0.004 F1. Across 14 Databricks endpoints, mean F1 change from local evaluation is $-0.0081$; Aya is the only family with a 2.7--5.1 point decline. These results show that compact fine-tuned models can preserve most extraction accuracy, but model selection must account for prompt choice, throughput, and serving-stack behavior.

cs.AI

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds and when each fails --- a gap that leaves practitioners, especially in scientific domains with heterogeneous instruments and multiple levels of measurement, unable to diagnose why standard methods underperform the best single modality. We study both objectives under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, the ingredient that breaks the classical recovery guarantees, and derive separation ratios that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable through a one-sided whitening, with recovery governed by source-modality quality. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither --- refined by a recovery count that separates partial recovery from complete failure. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training, and identifying when no objective in the CA/CP family can improve on the stronger modality alone. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and two real scientific domains --- astronomy and single-cell multi-omics --- validate the predictions in the nonlinear regime, including both faces of the Neither regime. Code to reproduce the results is available at https://github.com/IlayMalinyak/mm_align_vs_pred.

cs.LG

Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs

We present Activation- and Influence-Aware Ranks (AIR), an SVD-based LLM compression framework that guides each weight matrix's low-rank approximation with a backward-signal influence metric. Starting from the activation-aware optimum of SVD-LLM(W), AIR runs a single closed-form alternating least squares (ALS) sweep that integrates influence element-wise under a monotone-descent guarantee. AIR is layer-local and composes orthogonally with end-to-end methods: alone it exceeds ACIP, and AIR+LoRA outperforms it further. AIR improves perplexity over SVD-LLM(W) by >18% at <=60% parameter retention, matches its quality with ~90% less calibration data, and turns parameter savings into FLOP, peak-memory, and per-token latency gains.

cs.LG
Compare source metadata on this page
WorkPublishedSource identifierSource
Marginal-Contribution Policy Gradients under Filtered Feedback for Multi-Agent LLMs2026-04-032604.22785arxiv
Negative Ontology of True Target for Machine Learning: Towards Recognition, Evaluation and Learning under Democratic Supervision2026-04-272604.24824arxiv
Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity Education2026-04-142604.26964arxiv
Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits2026-05-052605.03434arxiv
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why2026-05-112605.10889arxiv
Geometric Dictionary Learning of Dynamical Systems with Optimal Transport2026-05-182605.18276arxiv
WaveGraphNet: Physics-Consistent Guided-Wave Damage Localization through Coupled Inverse-Forward Graph Learning2026-05-192605.20311arxiv
EmoTrack: Clinical-Semantic Modeling for Text-Based Depression Severity Estimation2026-05-212605.22286arxiv
The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer2026-05-282605.29713arxiv
From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets2026-05-292606.00202arxiv
Subliminal Learning is a LoRA Artifact2026-05-302606.00831arxiv
Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning2026-06-022606.03962arxiv
Rollout-Level Advantage-Prioritized Experience Replay for GRPO2026-06-032606.04560arxiv
In-Context Multiple Instance Learning2026-06-042606.06458arxiv
HAARES Half-Split Residual Basis Routing for Deep Transformers2026-06-042606.06564arxiv
How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions2026-06-062606.08051arxiv
When to Align, When to Predict: A Phase Diagram for Multimodal Learning2026-06-092606.11190arxiv
Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs2026-06-182606.19993arxiv

These are bibliographic comparisons, not experimental rankings. Follow the original document for methods and conditions.