arXiv ScienceSearch

arXiv subjects

Sebastian Pokutta

Publications and source records attributed to Sebastian Pokutta.

At least 19 recordsLinked to original sources

Robot Programming with Augmented Reality: The Role of Spatial Ability

Programming a robot arm requires users to interpret coordinate frames, joint rotations, and trajectories that are not directly visible. Augmented reality (AR) can make these spatial relations visible, but its benefits may depend on users' spatial ability. We conducted a randomized between-subjects experiment ($N=71$) in which participants learned to program a physical UR5e robot using either conventional teach-pendant controls with PDF instructions or a head-mounted AR interface that displayed joints, coordinate frames, and waypoints. We measured users' spatial ability with the Mental Rotation Test and assessed subjective cognitive load and system usability. Overall, AR did not significantly improve cognitive load or usability compared with conventional instruction. However, exploratory analyses revealed a compensatory effect: spatial ability predicted higher usability and lower extraneous cognitive load in the control group, but not in the AR condition, suggesting AR mitigated the disadvantage typically faced by users with lower spatial ability. These findings point to a compensatory function of AR that should be explored to guide the design of personalized AR interfaces for human-robot interaction.

cs.RO

LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles

Large language model (LLM)-based student simulation offers a scalable alternative for educational research, teacher training, and learner practice. However, its validity depends on whether LLMs maintain stable personas across and within interactions. We test this using a dual-assessment framework measuring self-reported characteristics and observer-rated behavioral expressions. Across three ICD-11-informed ADHD-related intensity conditions and a default condition, five LLMs, and three prompt designs, we quantify between-conversation (Exp. I; N=$4,962$) and within-conversation stability (Exp. II; N=$3,952$). Experiment I shows that self-reports and observer ratings are more stable at high than moderate intensities across the tested models, prompts, and independent runs. Experiment II shows that self-reports remain stable throughout extended interactions, but observer-rated behavior drifts in unscripted dialog for high- and moderate-intensity personas. Scripted interactions with recurring task-relevant prompts eliminate this drift almost entirely (up to 97\% reduction). Structured interaction design helps simulated learners maintain behavioral stability in sustained, path-dependent interactions.

cs.HC

Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations

Simulating student personas with large language models (LLMs) enables scalable evaluation of educational systems. However, behavioral drift, a progressive decline in persona consistency, can emerge over extended conversations, limiting the validity of such simulations. We evaluate five prompt-level mechanisms using separate monitoring and intervention pipelines. Across 1,200 28-turn conversations spanning four LLMs and two ADHD persona intensities, we varied when to intervene (static vs. adaptive) and what to inject (reinjection vs. reflective reminder), plus a novel adaptive condition in which a monitor generates behavior-specific instructions. Relative to no intervention, reinjection reduced the modeled rate of LLM-rated drift by 35--38\%, reflective reminders by 22--27\%, and behavior-specific instruction by 87\%. None eliminated drift. We found no evidence that adaptive timing outperformed static scheduling. Monitoring therefore appears more useful for deciding \textit{what} to correct than \textit{when} to intervene, although behavior-specific instruction requires component-level testing.

cs.HC

Objective Coefficient Rounding and Almost Symmetries in Binary Programs

This article investigates the interplay of rounding objective coefficients in binary programs and almost symmetries. Empirically, reducing the number of significant bits through rounding often leads to instances that are easier to solve. One reason can be that the amount of symmetries increases, which enables solvers to be more effective when they are exploited. This can signify that the original instance contains 'almost symmetries'. Furthermore, solving the rounded problems provides approximations to the original objective values. We empirically investigate these relations on instances of the capacitated facility location problem, the knapsack problem and a diverse collection of additional instances, using the solvers SCIP and CP-SAT. For all investigated problem classes, we show empirically that this yields faster algorithms with guaranteed solution quality. The influence of symmetry depends on the instance type and solver.

math.OC

Neural Field Tokenizations with Hierarchy and Spatial Locality Priors

Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities. Existing approaches are dominated by per-sample meta-learning, which scales poorly due to memory-intensive inner-loop optimization. The natural alternative -- feed-forward encoding -- typically introduces modality-specific assumptions, sacrificing the generality that makes learning with neural fields attractive. We argue that locality and hierarchy are useful priors for learning field representations that can be injected without compromising modality-agnosticism. We propose LH-NeF, a framework to learn general-purpose tokenized representations of continuous signals. A locality-preserving hierarchical encoder maps raw coordinate-value field observations to structured tokens, from which the field is reconstructed during training. By replacing meta-learning's inner loop with a single forward pass, LH-NeF uses 42x less memory and supports 133x larger batches than the strongest modality-agnostic baseline. Across images, 3D shapes, and climate fields, our learned representations match or exceed performance of modality-agnostic, modality-specific, and specialized generative neural field baselines on both reconstruction and downstream tasks.

cs.LG

Discrete eigenvalue optimization from entropic smoothing and first-order methods

We study the maximization of the minimum eigenvalue under combinatorial and integrality constraints. We propose a new approach based on branch-and-bound combines entropic smoothing of the minimum eigenvalue function and Frank-Wolfe methods over concave relaxations of the constraints, thereby exploiting combinatorial structure through linear optimization oracles. We establish approximation and convergence guarantees, including for truncated gradients computed from partial eigendecompositions, and introduce rank- and eigenvalue-based pruning and duality-based variable fixing. We evaluate the method on E-optimal experimental design and maximum algebraic connectivity problems and compare it with SCIP-SDP. The results show that our approach is particularly effective for large-dimensional instances and problems with additional combinatorial structure, whereas SCIP-SDP performs better on moderately sized instances with simpler constraints.

math.OC

Optimal Gradient-Norm Minimization in Non-Euclidean Hölder-Smooth Convex Optimization

Minimizing gradients of a convex function is an important problem across optimization and learning tasks. The gradient provides a directly computable certificate of approximate stationarity, and its minimization usually implies stronger results than those for minimization of function values. In this work, we study gradient-norm minimization for convex functions that are $(L,κ)$-Hölder smooth with respect to the $\ell_p$-norms, $p \geq 1$. We develop algorithms that achieve near-optimal gradient-oracle complexity for this problem. In the smooth case, our results resolve the previously open setting $p>2$. For Hölder-smooth objectives, we close the complexity gap throughout the full $p$-range, including to the best of our knowledge, a gap in the Euclidean case. We provide two families of algorithms: the first one comes with a simple iteration and generalizes a phenomenon known as mirror duality, exploiting dual behaviours of algorithms with errors and inexact computations. The second makes use of accumulating regularizers centered at different approximate solutions, which we sequentially minimize in order to provide our near-optimal rates.

math.OC

Limits of combinatorial patchworking

It is shown that there are real plane algebraic curves of degree eight that cannot be realized as T-curves, i.e., via combinatorial patchworking. In fact, this holds for several real schemes (i.e., ambient isotopy types) with the maximal number of real components, called $M$-curves. On the other hand, each nonempty real scheme of lower degree, maximal or not, arises as a T-curve. By constructing one patchwork of the dilated triangle $d\cdotΔ_2$ for each nonempty real scheme of degree $d\leq 7$, we provide an explicit method for constructing polynomials realizing these real schemes. This resolves a question of Itenberg and Viro (1996).

math.AG

Frank-Wolfe Beyond 1/t Convergence

We consider smooth convex minimization over compact convex sets, i.e., $\min_{x \in C} f(x)$ with the (vanilla) Frank-Wolfe algorithm. Well-known lower bounds establish a worst-case $Ω(1/t)$ primal-gap barrier in the general smooth convex case, and faster convergence usually requires favorable function properties such as Hölder error bounds or strong convexity. We present a new Local Dual Sharpness (LDS) condition, essentially a property of the feasible region and its LMO, under which the Frank-Wolfe algorithm converges in $o(1/t)$ for any smooth convex function, ruling out an $Ω(1/t)$ lower bound under LDS. The condition is a generalization (and localization) of uniform convexity of sets and it is satisfied by any uniformly convex set. To our knowledge, this is the first unconditional $o(1/t)$ convergence result for uniformly convex sets. Combining LDS with stronger function properties, e.g., a local variant of Hölder error bounds, allows us to quantify the actual rates.

math.OC

Scalable Lindblad Noise Learning via Stochastic Tensor-Network Simulation

Learning dissipation rates in large-scale open quantum systems is a major obstacle for near-term quantum technologies, as existing Lindblad estimation methods are typically limited to small system sizes due to the computational complexity of repeatedly solving the Lindblad equation during optimization. Here, we propose a scalable noise-learning framework for Lindblad dissipation rates that combines a stochastic simulation method, the Tensor Jump Method (TJM), with gradient-free optimization of a least-squares cost-function defined on time series of local-observable expectation values. We demonstrate the approach on two noise models in the Ising model: a site-resolved (local) model, in which independent dissipation rates are learned for each site up to $N_{\mathrm{site}}=16$, and a spatially homogeneous (global) model with only seven parameters, scaled to $N_{\mathrm{site}}=160$ sites.We complement these numerical results with a series of exact, provable guarantees: the Frobenius variance of the TJM density-matrix estimator is shown to equal $(1-\mathrm{Tr}[ρ^2])/N_{\mathrm{traj}}$, an exact purity-based characterization of the stochastic estimation error; the corresponding purity evolution is proven to be monotonically non-increasing for Hermitian jump operators; and, under a finite covariance distance assumption, the standard deviation of the cost-function is shown to decrease with system size, so that fewer trajectories are needed to reach a fixed target accuracy as the system grows. Together, this combination of scalable numerics and rigorous theoretical guarantees positions TJM-based noise learning as a practical foundation for characterizing dissipation in large quantum devices and for guiding future work on error mitigation and quantum error correction.

quant-ph

Computational Algebra with Attention: Transformer Oracles for Border Basis Algorithms

Solving systems of polynomial equations, particularly those with finitely many solutions, is a crucial challenge across many scientific fields. Traditional methods like Gröbner and Border bases are fundamental but suffer from high computational costs, which have motivated recent Deep Learning approaches to improve efficiency, albeit at the expense of output correctness. In this work, we introduce the Oracle Border Basis Algorithm, the first Deep Learning approach that accelerates Border basis computation while maintaining output guarantees. To this end, we design and train a Transformer-based oracle that identifies and eliminates computationally expensive reduction steps, which we find to dominate the algorithm's runtime. By selectively invoking this oracle during critical phases of computation, we achieve substantial speedup factors of up to 3.5x compared to the base algorithm, without compromising the correctness of results. To generate the training data, we develop a sampling method and provide the first sampling theorem for border bases. We construct a tokenization and embedding scheme tailored to monomial-centered algebraic computations, resulting in a compact and expressive input representation, which reduces the number of tokens to encode an $n$-variate polynomial by a factor of $O(n)$. Our learning approach is data efficient, stable, and a practical enhancement to traditional computer algebra algorithms and symbolic computation.

cs.LG

Boscia.jl: A review and tutorial

Mixed-integer nonlinear optimization (MINLP) comprises a large class of problems that are challenging to solve and exhibit a wide range of structures. The Boscia framework (Hendrych et al., 2025b) focuses on convex MINLP where the nonlinearity appears in the objective only. This paper provides an overview of the framework, showcases extensions post-publication and practical examples to illustrate its use and customizability. One key aspect is the integration and exploitation of Frank-Wolfe methods as continuous solvers within a branch-and-bound framework, enabling inexact node processing, warm-starting and explicit use of combinatorial structure among others. Three examples illustrate its flexibility, the user control over the optimization process and the benefit of oracle-based access to the objective and its gradient. Additionally, ablation studies are performed on the three examples to investigate the performance impact of the different features and customizations. The aim of this tutorial is to provide readers with an understanding of the main principles of the framework.

math.OC

Bounded-Support Additive Latin Transversals

We consider the following additive Latin transversal problem. Given a multiset $A=(a_1,\dots,a_k)$ of elements of $\mathbb Z_m$ and a set $B\subseteq\mathbb Z_m$ of cardinality $k$, the task is to order $B$ as $b_1,\dots,b_k$ so that the sums $a_i+b_i$ are pairwise distinct. When $k=m$, Hall proved that a solution exists if and only if $\sum_{i=1}^m a_i\equiv 0 \pmod m$; moreover, his theorem yields a polynomial-time construction. Alon proved that a solution always exists when $m$ is prime and $k<m$, but no polynomial-time construction is known in general. Our main algorithmic contribution is a direct randomized algorithm for Color-Counted Matching: given an edge-colored graph and prescribed target counts for the colors, find a matching using exactly the prescribed number of edges of each color. If $q$ is the sum of the target counts and $h$ is the number of colors, our base-$(q+1)$ reduction to Exact Red Matching, combined with the algorithm of Mulmuley-Vazirani-Vazirani, gives a randomized algorithm with running time $\left(|V|^2+|E|(q+1)^{h-1}\right)^{O(1)} $ for an input graph $(V,E)$. Thus the dependence on the target matching size is $q^{O(h)}$, up to polynomial factors in the graph size. In contrast, applying the general matching-ILP theorem of Lassota and Ligthart as a black box yields a $q^{O(h^2)}$ dependence for the corresponding fixed-size color-counted instances. Applying this primitive to additive Latin transversals with $s=|\operatorname{supp}(A)|$, we obtain an algorithm in randomized time $(k+\log m)^{O(s)}$. In particular, additive Latin transversals are randomized polynomial-time constructible for every fixed support size.

cs.DS

Joint-Range Inequalities for Nonconvex QCQPs

We study cutting planes for nonconvex quadratically constrained quadratic programs (QCQPs) through a project-then-lift approach inspired by mixed-integer rounding (MIR) inequalities. Given two base valid inequalities for the extended QCQP formulation, we project the associated two-row relaxation into a two-dimensional set and analyze the joint range of quadratic functions in two base inequalities. For the nonconvex joint range, we give a closed-form convex hull description of the projected set; for the convex joint range, we give its semidefinite representation. This yields a new family of joint-range inequalities, which can be lifted back to the extended QCQP formulation. MIR inequalities can handle ``mixed'' terms: continuous variables or fractional linear combinations of integer variables. Similarly, we propose more flexible secant mixed-joint-range inequalities, which better expose and exploit the nonconvex joint range. The proposed approach preserves sparsity, since the support of each lifted inequality is controlled by that of the base inequalities. In preliminary geometric experiments, the joint-range inequalities yield substantial area reduction of the projected relaxation constructed via reformulation-linearization-technique.

math.OC

When Does LLM Orchestration Pay Off? A Controlled Evaluation of Accuracy, Cost, and Task Difficulty

LLM orchestration is often assumed to improve reasoning by allocating additional inference-time computation, yet its gains may not justify its cost. Existing comparisons also frequently overlook differences in optimization effort, making it difficult to isolate the value of orchestration itself. We conduct a controlled evaluation of Self-Refine, Best-of-$N$, and Debate against task-only and chain-of-thought (CoT) single-call baselines across five LLM backbones and three domains: competitive programming, chess puzzles, and mathematics. For comparability, we optimize each method with GEPA under the same optimization budget and evaluate all methods on the same difficulty-stratified benchmark items. Orchestration yields moderate but benchmark-dependent gains: averaged across backbones within each benchmark, the largest improvement is 4.6 percentage points over optimized CoT inference and 4.5 points over task-only inference, while requiring approximately 2 to 4 times the mean total tokens of task-only inference. Human-derived difficulty is associated with lower absolute accuracy in all three benchmarks, but within-benchmark analyses do not indicate that orchestration effects increase with task difficulty. By contrast, exploratory mixed-effects analyses reveal strong interactions between orchestration method and backbone model across all three benchmarks, showing that orchestration effectiveness depends substantially on the underlying model. Our results suggest that orchestration decisions should be model-specific and account for whether moderate accuracy gains justify the additional inference cost. More broadly, evaluations of LLM orchestrations should control optimization effort and report model-specific accuracy--cost trade-offs rather than treating additional inference-time structure as uniformly beneficial.

cs.AI

A Counterexample to Ziegler's Cross-Polytope Conjecture for Simplicial 0/1-Polytopes

Ziegler proved that every simplicial $d$-dimensional $0/1$-polytope has at most $2d$ vertices, and asked whether equality forces the polytope to be centrally symmetric and hence, equivalently, a $0/1$-realization of the $d$-dimensional cross polytope. In this note, we give a negative answer, exhibiting an explicit set of $14$ vertices in $\{0,1\}^7$ whose convex hull is a simplicial $7$-polytope and is not centrally symmetric. Moreover, via exhaustive enumeration we show that up to the symmetries of the cube, there are precisely five such polytopes in dimension $7$ (of two combinatorial types) that are not centrally symmetric.

math.CO

The Weight Distribution of the Third-Order Reed-Muller Code of Length 2048

We compute the weight distribution of the third-order Reed--Muller code RM(3,11) of length 2048. The weight enumerator is assembled from the coset weight enumerators of f+RM(2,10), evaluated for representatives of all 3691560 nonzero GL(10,2)-orbits of Boolean cubic forms in ten variables. The computation rests on a structural theorem: a nondegenerate Boolean cubic form admits a nondegenerate hyperplane restriction, except for a single orbit in each odd dimension. The same pass determines the second-order nonlinearity of every cubic form: the relative covering radius of RM(2,10) in RM(3,10) is 408, attained on 179 orbits. This raises the best known lower bound on the covering radius of RM(2,10) from 400 to 408. A complementary heuristic search shows that the relative covering radius of RM(6,10) in RM(7,10) is at most 32, improving the previous bound of 50.

cs.IT

When Does Sparsity Mitigate the Curse of Depth in LLMs

Recent work has demonstrated the curse of depth in large language models (LLMs), where later layers contribute less to learning and representation than earlier layers. Such under-utilization is linked to the accumulated growth of variance in Pre-Layer Normalization, which can push deep blocks toward near-identity behavior. In this paper, we provide evidence that sparsity-like mechanisms can dampen variance propagation and are associated with improved depth utilization Our investigation covers two sources of sparsity: (i) implicit sparsity, which emerges from training and data conditions, including weight sparsity induced by weight decay and attention sparsity induced by long-context inputs; and (ii) explicit sparsity, which is enforced by architectural design, including key/value-sharing in Grouped-Query Attention and expert-activation sparsity in Mixtureof-Experts. Our claim is thoroughly supported by controlled depth-scaling experiments and targeted layer effectiveness interventions. Across settings, we observe a consistent relationship: mechanisms with reduced effective interaction density tend to exhibit lower output variance and better layer differentiation. We eventually distill our findings into a practical rule-of-thumb recipe for training depth-effective LLMs, yielding a notable 4.6 accuracy improvement on downstream tasks. Our results suggest that sparsity-like design choices are an important and previously underemphasized factor in effective depth scaling for LLMs. Code is available at https://github. com/pUmpKin-Co/SparsityAndCoD.

cs.CL