arXiv ScienceSearch

arXiv subjects

Ioannis Patras

Publications and source records attributed to Ioannis Patras.

3 recordsLinked to original sources

Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images

Multimodal Large Language Models (MLLMs) show strong visual perception, yet remain limited in reasoning about space under changing viewpoints. We study this challenge as Perspective-Conditioned Spatial Reasoning (PCSR) in 360 degree omnidirectional images, where broad scene coverage reduces ambiguity from partial observations without eliminating the need for viewpoint-dependent inference. To assess this capability, we introduce PCSR-Bench, a diagnostic benchmark of 84,373 QA pairs from 2,600 omnidirectional images across 26 indoor environments, organized into eight tasks under three cognitive groups--- Perception, Spatial, and advanced PCSR. We evaluate 14 representative MLLMs and observe a substantial perception--reasoning gap: accuracy reaches 57.59% on Limited Field-of-View Reasoning (T7) but drops to 13.49%, 7.13%, and 0.64% on Relative Direction (T2), Egocentric Rotation (T4), and open-ended Compositional Directional Chains (T3), respectively. To probe the plasticity of this gap, we conduct an RL-based diagnostic study on a 7B-scale model. Reward shaping improves a matched 7B baseline from 31.10% to 60.06% under a controlled setting, suggesting that PCSR exhibits partial plasticity rather than being fully immutable. Still, these gains are task-selective, sensitive to reward design, and partially dependent on the evaluation protocol. These results position PCSR as a key bottleneck in current MLLMs and highlight meaningful yet bounded room for recovery under targeted optimization. Details and access are available at https://github.com/Caleb-ychen/PCSR-Benchmark.

cs.CV

TPSO: Training-Free Diverse Image Generation via Semantic Prompt Embedding Optimization

Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive outputs, increasing sampling redundancy and hindering both creative exploration and downstream applications. A key factor is the tendency of diffusion models to collapse toward strong modes in the learned distribution. Existing attempts to improve diversity, such as steering-based guidance, often introduce distortions that degrade image quality. To address this issue, we propose Token-Prompt Embedding Space Optimization (TPSO), a training-free and model-agnostic module. TPSO introduces learnable parameters to explore underrepresented regions of the token embedding space, reducing the tendency to repeatedly sample from strong modes of the distribution. Meanwhile, a prompt-level semantic constraint regulates distribution shifts, preventing quality degradation while preserving semantic fidelity. Extensive experiments on MS-COCO across three representative diffusion backbones demonstrate that TPSO substantially improves diversity, boosting performance from 1.10 to 4.18, while maintaining image quality with only a modest inference-time overhead of 3.6% to 8.9%. Code is available at: https://github.com/Open-Debin/TPSO.

cs.CV

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

Autoregressive video diffusion has recently emerged as a promising paradigm for long-video generation, enabling causal synthesis beyond the temporal limits of bidirectional models. Existing forcing-based training strategies reduce exposure bias by conditioning models on their own predictions during rollout, yet minute-scale generation remains challenging due to progressive temporal degradation and constrained motion evolution. In this work, we study the role of temporal KV memory during long-horizon autoregressive inference. Our analysis shows that simply retaining more historical frames does not consistently improve generation quality; instead, both the quantity and temporal placement of memory strongly affect motion dynamics. These findings suggest that temporal memory should be treated as structured context rather than a homogeneous chronological buffer. Motivated by this observation, we introduce Relax Forcing, a training-free memory mechanism for autoregressive video diffusion. Relax Forcing decomposes temporal context into three functional components: Sink frames that provide global stability, Tail frames that preserve short-term continuity, and dynamically selected History frames that supply mid-range motion structure. History frames are selected using a relaxation-based criterion that encourages alignment with global anchors while suppressing redundancy with recent context. This role-aware sparse memory design mitigates error accumulation during long-horizon rollout while preserving motion evolution and reducing attention overhead. Experiments on VBench-Long show that Relax Forcing improves long-video generation, achieving stronger motion dynamics and higher overall scores than existing autoregressive baselines. These results indicate that structured temporal memory is an effective and complementary direction for scalable long-video generation.

cs.CV