arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

"You Can't Just Automate It": Negotiating and Sustaining a "Good" Family Life Through Energy Practices

This study examines how Taiwanese parent-child families negotiate a "good" family life through everyday energy use and imagine future smart homes that support it. We conducted in-home interviews and co-design sessions with 21 families, including 46 parents and children. We found that families pursued a good life through energy practices shaped by thrift, comfort, care, safety, and enjoyment. These arrangements were continually adapted and responded to changing bodies, schedules, people, and infrastructures. This adaptive work was unevenly distributed, which in turn shaped different smart-home imaginaries. Drawing on the lens of Nearby and adversarial design, we conceptualize adaptation as situated sociotechnical work through which families continually rework energy arrangements. We further distinguish collective goods from plural and contestable goods to show why family IoT must support shared values while preserving opportunities to question and revise household arrangements. We offer theoretical and design directions for more adaptive, participatory, and contestable family IoT.

cs.HC↗

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.

cs.AI↗

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.

cs.CV↗

pytest-gpu-proof: Enabling Cloud-CPU Continuous Integration for GPU Code with Local GPU Attestation

GPU acceleration is now routine across robotics, but cloud-hosted GPU continuous integration (CI) runners are expensive, resulting in severe under-testing of GPU-accelerated code. We present pytest-gpu-proof, an open-source pytest plugin offering a practical middle ground. Tests can be run on a local machine, signed with a receipt of exactly what ran and what it produced, and integrated into standard CPU CI workflows (e.g., GitHub Actions). The tool is open source and on PyPI, and we are actively integrating it across our lab's software stack.

cs.DC↗

Affine Pricing Models from Group Quantization and Holonomy

The analytic tractability of affine pricing models is usually expressed through two complementary formulations: a coordinate-space pricing operator and an exponential-affine transform representation governed by generalized Riccati equations. We develop \emph{Affine Holonomy Group Quantization} (AHGQ) as a geometric framework in which these two formulations arise from the same underlying structure. The construction separates the affine pricing symbol into a homogeneous quadratic sector and a complementary affine sector. The first generates a finite-dimensional symplectic transport and a centrally extended Lie group, while the second is represented by a multiplicative holonomy carried by a thin-path groupoid. Their combination determines an affine Poincaré--Cartan form. Its characteristic dynamics reduce in momentum variables to the generalized Riccati system and its scalar amplitude, whereas the coordinate representation recovers the standard affine pricing operator. Representative Gaussian and square-root models illustrate the construction. The contribution is structural: AHGQ gives a common geometric origin to the coordinate and transform representations of continuous-path, time-homogeneous affine pricing models.

q-fin.MF↗

Magnetostrictive properties in Fe$_{4-x}$Co$_{x}$N films: Insight from experiments and first-principles calculations

The ferromagnetic nitride Fe$_{4}$N has attracted attention for application in spintronic devices due to its high spin polarization and large magnetostriction. We present a combined experimental and theoretical study on the magnetostriction of Fe$_{4-x}$Co$_{x}$N films across a wide composition range. The Fe$_{4-x}$Co$_{x}$N films were grown on SrTiO$_{3}$(001) substrates using molecular beam epitaxy, and the magnetostriction constants along the [100] direction ($λ$$_{100}$) and [111] direction ($λ$$_{111}$) were precisely evaluated using an optical cantilever method. The experimental results reveal that $λ$$_{111}$ remains positive in the whole composition range and shows a maximum value of +82 ppm around x = 0.9. The variation in $λ$$_{111}$ with x is much smaller than that in $λ$$_{100}$, for which giant tunability and sign reversal are observed. First-principles calculations show reasonable agreement with the experimental $λ$$_{111}$ for x ${\ge}$ 1.6, but give negative values at lower x, and an exceptionally large negative $λ$$_{111}$ is obtained at x = 0.8, where the Fermi level coincides with a pronounced minority-spin peak in the density of states. The calculated $λ$$_{111}$ depends strongly on the smearing parameter, indicating that the rhombohedral magnetostriction is highly sensitive to the treatment of atomic disorder. The saturation magnetostriction constant ($λ$$_{s}$) derived from $λ$$_{100}$ and $λ$$_{111}$ is also compared with the $λ$$_{s}$ measured for the (001)-oriented polycrystalline Fe$_{4-x}$Co$_{x}$N films, and a possible scenario for deviation between them is discussed. Our findings clarify the basic features of magnetostriction in the Fe$_{4-x}$Co$_{x}$N system, providing essential magnetoelastic parameters for designing nitride-based spintronic devices.

cond-mat.mtrl-sci↗

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

cs.CV↗

Refutable Exclusion Restrictions in Competing Risks with Categorical Covariates

Competing-risks data do not identify latent marginal duration distributions or their dependence without additional restrictions. This paper asks whether restrictions introduced to restore identification themselves restrict the observable law. For a two-risk Archimedean model with categorical exclusion restrictions, we derive a necessary-and-sufficient observable characterization. A discrete single-crossing argument identifies the scalar copula parameter from cell-specific overall survival probabilities, while cause indicators recover the remaining allocation and generate additional specification restrictions. We construct an identification-robust, self-normalized quadratic statistic, invert it to obtain confidence sets, and use empty inverted sets as a conservative specification test. Simulations show that the cause indicator can be decisive: under the weakest contrast considered it converts a frequently uninformative survival-based confidence set into an informative joint set without loss of coverage, whereas excessive categorical contrast can eliminate causes from individual cells and make cause-specific recovery inadmissible. The results turn latent exclusion restrictions into refutable restrictions without estimating covariate derivatives.

stat.ME↗

Bondal--Orlov reconstruction for tame stacks with trivial generic stabilizer

We generalize the Bondal--Orlov Reconstruction Theorem to smooth, proper, tame algebraic stacks with generically trivial stabilizer, whose canonical bundles are nowhere torsion: They do not become trivial after pullback along any finite morphism from a projective curve. Our main tools are coherent Tannaka duality as proved by Lurie and by Hall and Rydh, and Grothendieck duality for proper tame stacks as developed by Hall and Priver. We additionally use the notion of nowhere torsion line bundle on a scheme to prove a version of the Bondal--Orlov Reconstruction Theorem for finite type, separated, Gorenstein schemes which are not necessarily proper, which generalizes work of Ballard, Favero, Ito, and Matsui.

math.AG↗

Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography

Neural fields for scientific tomography are optimized from 2D images, but the actual quantity of interest is often a latent 3D physical field. Because the forward map is many-to-one, low 2D image error need not certify a correct 3D field. Moreover, the latent field is not directly supervised during training, and its error cannot be evaluated against truth at deployment. We develop CoroNeRF to jointly optimize 3D electron density and temperature fields directly from multiview, multiline intensities through a differentiable atomic-emission renderer. Using solar coronal tomography as a controlled testbed, we evaluate physical-field recovery and test whether cross-seed instability provides a ground-truth-free-at-inference indicator of local physical-field error. We underscore the following two observations. (i) Image fidelity is not field fidelity: spectral ablations show that limited-channel reconstructions can fit their available observations well while recovering substantially worse fields, whereas evaluation on a common richer probe exposes the discrepancy. (ii) Cross-seed instability ranks local physical-field error across tested matched-model conditions, supported by sparsification and physical signal-strength controls. Seed-deviation projections provide complementary directional validation, but shared forward-model mismatch can still produce incorrect cross-seed consensus. These results characterize joint thermodynamic recovery and the usefulness and limits of seed-based error localization in a controlled, single-scene solar tomography testbed.

cs.LG↗

Computations of Cohomology of Arithmetic Groups, Part 1

The key part of the current paper is the computation of boundary and Eisenstein cohomology of $GL_4({\mathbb Z})$ with coefficient in any highest weight representations. The method we develop let us compute in an alternative way the cohomology of $SL_3({\mathbb Z})$ and of $GL_3({\mathbb Z})$ with coefficients in any highest weight representation. This is done in a simpler, faster and in a more structured way compared to \cite{BHHM}. We state a duality for the boundary cohomology of $GL_m({\mathbb Z})$ of the type of Serre's duality, where the dualizing sheaf is a power of the determinant representation. We refine this duality to a duality on the level of the spectral sequence for the boundary cohomology $E_\infty^{p,q}$. We state it as a conjecture. However, all the computations, 30 different families of representations, satisfy this conjecture. We compute the Eisenstein cohomology of $GL_4({\mathbb Z})$ with coefficients in the symmetric powers and their twist by the determinant representation. For several other representations, we compute the Eisenstein cohomology, based a few conjectures. Based on those conjectured, one can compute the Eisenstein cohomology in most of the cases. They will be included in the next version of the paper.

math.NT↗

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.

cs.DC↗

Panel Conditioning in Fixed-Effects Models: Identification and Bias Propagation

Panel conditioning, the causal effect of prior survey participation on responses, can vary with tenure. Under an additive model of cell means in period, entry cohort, and tenure, we characterize which features of the conditioning path a staggered panel identifies on its observed support, and how the unidentified component affects common panel estimators. The identified set of the path is an affine translate of the tenure projection of the cell design's kernel, and a linear functional of the path is identified exactly when it annihilates that projection. It always contains an affine direction and, when the entry cohorts share a stride, periodic directions, which exhaust it under a connectivity condition on observed increments; second differences at that stride are then identified, and ordinary ones generally are not when the stride exceeds one. Under a recruitment condition, an interrupted schedule such as the four-eight-four rotation of the Current Population Survey (CPS) distinguishes a constant increment per interview from one per calendar month, which no equally spaced schedule can. We give support conditions for recovery under a plateau, entry-wave negative controls, or bounded cohort drift. A second set of results links identification to regression: two-way fixed effects absorb every unidentified direction, so the remaining conditioning bias is normalization-invariant and itself identified, and a two-way regression with tenure indicators corrects it under a residual-rank condition. When event time is aligned with tenure, conditioning shifts event-study coefficients by a known linear functional of the path, producing pre-trends without anticipation; bounds on identified curvature bound those shifts. Simulations verify the identities, a 19-wave Japanese panel illustrates the support calculations, and published CPS month-in-sample indices give a descriptive, not identifying, example.

stat.ME↗

Least-false Cox coefficients under affine follow-up contamination: exact continuous- and grouped-time benchmarks

Covariates summarized over a subject's completed follow-up are sometimes entered into Cox regression as though observed at baseline. This practice incorporates future event or censoring information and changes both the estimand and its sampling behavior. We analyze an affine class in which a genuine baseline covariate is contaminated by realized follow-up time. Treating partial likelihood as an observed-data estimation criterion, we derive the population score and characterize its unique least-false coefficient. Exact continuous-time benchmarks show scale reduction and saturation under strong contamination, while administrative censoring destroys the reduction and may produce overshoot. Grouping exit times with Breslow ties changes the geometry: the coefficient has a single hump and eventually returns to zero even though the induced association diverges. We also derive an observed-data influence function and show why model-based variance can be either too small or too large. A subject-level sandwich consistently estimates uncertainty around the least-false target under the stated conditions, but it does not correct the target itself.

stat.ME↗

Improved Revenue Guarantees for Selling Separately and Bundling

We study how much revenue a seller can lose by restricting attention to selling separately or grand bundling, in the setting of a single additive buyer with independent item values. Although revenue-optimal mechanisms can require lotteries and infinite menus, Babaioff, Immorlica, Lucier, and Weinberg showed that the better of these two simple formats always achieves a constant fraction of optimal revenue. We prove that $\mathrm{OPT} \le 3.52 \max\{\mathrm{SREV}, \mathrm{BREV}\}$, where $\mathrm{SREV}$ and $\mathrm{BREV}$ are the optimal revenues from selling separately and grand bundling, respectively. This improves the previous best-known approximation factor of $5.2$ due to Ma and Simchi-Levi and narrows the gap to the known lower bound of $2$.

cs.GT↗

A Task-Based Framework for Evaluating Raman Spectral Quality Measures

Raman spectral preprocessing and enhancement are often evaluated by comparing output spectra with a reference. Interpreting these comparisons requires evidence that spectral quality measures reflect downstream task performance. We present a controlled-perturbation framework for testing this relationship. Five perturbation types (baseline distortion, independent noise, correlated noise, a global wavenumber shift, and nonlinear axis warping) generate paired changes in a spectral measure (metric harm) and in downstream performance (task harm). An alignment gap (AG) quantifies how much the relationship between metric harm and task harm changes with perturbation type. Ordering concordance (OC) measures how often a metric correctly ranks two conditions by their task harm. The framework evaluates thirteen outputs (MSE, RMSE, MAE, NMSE, spectral angle, Pearson correlation, Wasserstein distance, a structure-to-noise ratio, peak precision, recall, F1, artifact ratio, and missing ratio). Three public datasets provide bacterial classification, sugar-mixture quantification, and mineral identification tasks. PCA with logistic regression, partial least squares regression, and cosine library matching supply the task outcomes. Classifiers and calibrations are fitted either to unperturbed training spectra or to each perturbed training condition, then evaluated on the same perturbed test spectra. Mineral queries are compared with an unchanged or correspondingly perturbed library. The resulting comparisons identify task-specific strengths and limitations, including cases where better ordering does not accompany a smaller AG. Removing axis perturbations and comparing spectra on a common physical grid test how these findings depend on the evaluation design. The framework provides a reproducible procedure for assessing existing measures and testing new candidates against downstream task performance.

physics.chem-ph↗

Concentration of bounded sparse chaoses and sparse Khatri-Rao embeddings

We establish moment and mixed-tail inequalities for fixed-order decoupled homogeneous chaoses generated by independent, centered, sparse bounded random variables. Our bounds apply to arbitrary real rectangular coefficient tensors and describe the fluctuation scales through weighted slice and partition norms, with a Bennett-type logarithmic improvement in the largest-entry term. As an application, we derive guarantees for sparse Khatri--Rao embeddings that explicitly account for sparsity and input geometry.

math.PR↗

Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.

cs.AI↗