arXiv ScienceSearch

arXiv subjects

Sunghwan Kim

Publications and source records attributed to Sunghwan Kim.

At least 19 recordsLinked to original sources

Self-Evolving Search Index

Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

cs.IR

Nonlinear Diamagnetic Interactions in Ultrastrongly Coupled 2D Electrons

The quantum Hopfield model is widely used to describe ultrastrong light--matter coupling between cavity photons and collective bosonic excitations in solids, where the diamagnetic interaction is conventionally assumed to be a constant. We experimentally demonstrate that the diamagnetic response of Landau polaritons is reduced under strong terahertz field excitation. We show that this behavior originates from field-driven redistribution of electrons into the nonparabolic regime of the conduction band of GaAs, which reduces the plasma frequency and consequently the diamagnetic interaction strength. A microscopic hot-electron model reproduces the observed nonlinear response. Motivated by this microscopic picture, we propose a nonlinear extension of the Hopfield model with a Kerr-like interaction. Our results establish a route toward nonlinear cavity quantum electrodynamics and driven ultrastrong light--matter coupling beyond the conventional linear Hopfield description, which is capable of creating uniquely quantum optical effects such as squeezed light generation.

quant-ph

Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SBP), an end-to-end policy learning approach that operates directly on a 3D map of latent features. In SBP, the map extends perception beyond the robot's current field of view and aggregates observations over long horizons. Our mapping approach incrementally fuses multiview observations into a grid of scene-specific latent features. A pre-trained, scene-agnostic decoder reconstructs target embeddings from these features and enables online optimization of the map features during task execution. A policy, trainable with behavior cloning or reinforcement learning, treats the latent map as a state variable and uses global context from the map obtained via a 3D feature aggregator. We evaluate SBP on scene-level mobile manipulation and sequential tabletop manipulation tasks. Our experiments demonstrate that SBP (i) reasons globally over the scene, (ii) leverages the map as long-horizon memory, and (iii) outperforms image-based policies in both in-distribution and novel scenes, e.g., improving the success rate by 15% for the sequential manipulation task.

cs.RO

SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization

Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generative capabilities to deliver synthesized answers. This shift has fundamentally reshaped how web content gains exposure online, giving rise to Search-Augmented Generative Engine Optimization (SAGEO), the practice of optimizing web documents to improve their visibility in AI-generated responses. Despite growing interest, no evaluation environment currently supports comprehensive investigation of SAGEO. Specifically, existing benchmarks lack end-to-end visibility evaluation of optimization strategies, operating on pre-determined candidate documents that abstract away retrieval and reranking preceding generation. Moreover, existing benchmarks discard structural information (e.g., schema markup) present in real web documents, overlooking the rich signals that search systems actively leverage in practice. Motivated by these gaps, we introduce SAGEO Arena, a realistic and reproducible environment for stage-level SAGEO analysis. Our objective is to jointly target search-oriented optimization (SEO) and generation-centric optimization (GEO). To achieve this, we integrate a full generative search pipeline over a large-scale corpus of web documents with rich structural information. Our findings reveal that existing approaches remain largely impractical under realistic conditions and often degrade performance in retrieval and reranking. We also find that structural information helps mitigate these limitations, and that effective SAGEO requires tailoring optimization to each pipeline stage. Overall, our benchmark paves the way for realistic SAGEO evaluation and optimization beyond simplified settings.

cs.IR

SERF: Spatiotemporal Environment and Robot Feature Map for Long-Horizon Mobile Manipulation

Long-horizon robot mobile manipulation requires continual reasoning about localization, environment changes, and task progress, all of which are challenging to infer from image observations alone. In this paper, we show that conditioning a mobile manipulation policy on a spatiotemporal feature map improves reasoning over long horizons. The map represents the environment and the articulated robot body as neural points in a shared latent space and is updated online from egocentric observations and proprioceptive state. We update the environment neural points using object-level rigid tracking and the robot neural points using forward kinematics. We use our spatiotemporal environment and robot feature (SERF) map as a state input to a vision-language-action (VLA) model by extracting map tokens from multiple reference frames and spatial scales, providing the policy with both local and global context. We demonstrate SERF on BEHAVIOR-1K, a benchmark for long-horizon mobile manipulation in household environments. Experiments show that the SERF VLA policy outperforms image-only baselines, reaches subgoals faster by following more direct trajectories, improves robustness to scene-configuration shifts, and recovers from object-drop failures.

cs.RO

Towards Direct Evaluation of Harness Optimizers via Priority Ranking

Harness optimization enables automated agent creation by having an optimizer agent iteratively update the harness of target agents. Despite its success, current studies evaluate optimizers solely by observing target agents' performance gains. This indirect end-improvement evaluation neglects optimizers' actions at intermediate steps, which are often erroneous and hinder agent performance. Therefore, it is unclear whether harness optimization is driven by optimizers' informed update actions or simply trial-and-error. This necessitates direct evaluation of harness optimizers. However, evaluating harness optimizers directly is non-trivial and costly due to the lack of oracle harnesses. To address this, we present a simple, low-cost design to directly evaluate them, namely priority ranking. By asking harness optimizers to rank components (e.g., tools) in a given harness by their potential to improve/hinder agent performance when updated, our design quantifies optimizer ability at the step level without expensive rollouts or manual examination. More importantly, optimizers' ranking performance correlates with their ability to improve agents in actual multi-step harness optimization, establishing priority ranking as a reliable predictor of optimization ability. Priority ranking is enabled by Shor, a collection of 182 human-verified optimization scenarios spanning across domains, designs, and time stages. Codes and data can be found at https://github.com/k59118/Harness_Optimizer_Evaluation.

cs.AI

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length

Large language models (LLMs) have shown promise as interactive agents that solve tasks through extended sequences of environment interactions. While prior work has primarily focused on system-level optimizations or algorithmic improvements, the role of task horizon length in shaping training dynamics remains poorly understood. In this work, we present a systematic empirical study that examines horizon length through controlled task constructions. Specifically, we construct controlled tasks in which agents face identical decision rules and reasoning structures, but differ only in the length of action sequences required for successful completion. Our results reveal that increasing horizon length alone constitutes a training bottleneck, inducing severe training instability driven by exploration difficulties and credit assignment challenges. We demonstrate that horizon reduction is a key principle to address this limitation, stabilizing training and achieving better performance in long-horizon tasks. Moreover, we find that horizon reduction is related to stronger generalization across horizon lengths: models trained under reduced horizons generalize more effectively to longer-horizon variants at inference time, a phenomenon we refer to as horizon generalization.

cs.AI

Development and characterization of the efficient portable X-ray imaging device based on Raspberry Pi camera

This study reports the development and characterization of an efficient portable X-ray imaging device built from Raspberry Pi components, including a high-quality 12.3-megapixel camera configured for indirect detection with a Gd2O2S:Tb scintillation screen. The device was evaluated under both ambient light and X-ray exposure conditions. Initial characterization under ambient light ensured proper optical focusing; subsequently, camera settings (ISO and exposure time) were evaluated and optimized for X-ray imaging performance. Spatial resolution of the developed device was quantified using the Slanted-Edge method to derive the Modulation Transfer Function (MTF). Besides the low-noise feature, the device achieves MTF20 values of 68 lp/mm under ambient light and 25 lp/mm under X-ray irradiation (50 and 70 kV). Moreover, the modularity of the developed device was confirmed by conducting the tests with LYSO:Ce and GAGG:Ce screens. The results demonstrate that this efficient, scientific-grade, compact platform achieves spatial resolution comparable to that of clinical radiography systems, highlighting its potential for applications in scientific, educational, and medical contexts where efficient and portability are critical considerations.

physics.ins-det

Integration of TinyML and LargeML: A Survey of 6G and Beyond

The evolution from fifth-generation (5G) to sixth-generation (6G) networks is driving an unprecedented demand for advanced machine learning (ML) solutions. Deep learning has already demonstrated significant impact across mobile networking and communication systems, enabling intelligent services such as smart healthcare, smart grids, autonomous vehicles, aerial platforms, digital twins, and the metaverse. At the same time, the rapid proliferation of resource-constrained Internet-of-Things (IoT) devices has accelerated the adoption of tiny machine learning (TinyML) for efficient on-device intelligence, while large machine learning (LargeML) models continue to require substantial computational resources to support large-scale IoT services and ML-generated content. These trends highlight the need for a unified framework that integrates TinyML and LargeML to achieve seamless connectivity, scalable intelligence, and efficient resource management in future 6G systems. This survey provides a comprehensive review of recent advances enabling the integration of TinyML and LargeML in next-generation wireless networks. In particular, we (i) provide an overview of TinyML and LargeML, (ii) analyze the motivations and requirements for unifying these paradigms within the 6G context, (iii) examine efficient bidirectional integration approaches, (iv) review state-of-the-art solutions and their applicability to emerging 6G services, and (v) identify key challenges related to performance optimization, deployment feasibility, resource orchestration, and security. Finally, we outline promising research directions to guide the holistic integration of TinyML and LargeML for intelligent, scalable, and energy-efficient 6G networks and beyond.

cs.NI

Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization

LLM-powered embodied agents have shown success on conventional object-rearrangement tasks, but providing personalized assistance that leverages user-specific knowledge from past interactions presents new challenges. We investigate these challenges through the lens of agents' memory utilization along two critical dimensions: object semantics (identifying objects based on personal meaning) and user patterns (recalling sequences from behavioral routines). To assess these capabilities, we construct MEMENTO, an end-to-end two-stage evaluation framework comprising single-memory and joint-memory tasks. Our experiments reveal that current agents can recall simple object semantics but struggle to apply sequential user patterns to planning. Through in-depth analysis, we identify two critical bottlenecks: information overload and coordination failures when handling multiple memories. Based on these findings, we explore memory architectural approaches to address these challenges. Given our observation that episodic memory provides both personalized knowledge and in-context learning benefits, we design a hierarchical knowledge graph-based user-profile memory module that separately manages personalized knowledge, achieving substantial improvements on both single and joint-memory tasks. Project website: https://connoriginal.github.io/MEMENTO

cs.CL

AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping

The proliferation of e-commerce has made web shopping platforms key gateways for customers navigating the vast digital marketplace. Yet this rapid expansion has led to a noisy and fragmented information environment, increasing cognitive burden as shoppers explore and purchase products online. With promising potential to alleviate this challenge, agentic systems have garnered growing attention for automating user-side tasks in web shopping. Despite significant advancements, existing benchmarks fail to comprehensively evaluate how well agentic systems can curate products in open-web settings. Specifically, they have limited coverage of shopping scenarios, focusing only on simplified single-platform lookups rather than exploratory search. Moreover, they overlook personalization in evaluation, leaving unclear whether agents can adapt to diverse user preferences in realistic shopping contexts. To address this gap, we present AgenticShop, the first benchmark for evaluating agentic systems on personalized product curation in open-web environment. Crucially, our approach features realistic shopping scenarios, diverse user profiles, and a verifiable, checklist-driven personalization evaluation framework. Through extensive experiments, we demonstrate that current agentic systems remain largely insufficient, emphasizing the need for user-side systems that effectively curate tailored products across the modern web.

cs.IR

Web-Shepherd: Advancing PRMs for Reinforcing Web Agents

Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks. Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench. Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10 less cost compared to using GPT-4o-mini as the verifier. Our model, dataset, and code are publicly available at LINK.

cs.CL

ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions

Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capabilities in long-term interactions. Each test instance in ToolHaystack includes multiple tasks execution contexts and realistic noise within a continuous conversation, enabling assessment of how well models maintain context and handle various disruptions. By applying this benchmark to 14 state-of-the-art LLMs, we find that while current models perform well in standard multi-turn settings, they often significantly struggle in ToolHaystack, highlighting critical gaps in their long-term robustness not revealed by previous tool benchmarks.

cs.CL

Symmetry-Controlled Ultrastrong Phonon-Photon Coupling in a Terahertz Cavity

Optical cavities provide a powerful means to engineer light-matter hybrid states by coupling confined electromagnetic fields with matter excitations. Achieving in situ control of the coupling strength is essential for investigating how such hybridization evolves with the coupling strength. In this work, we use a symmetry-changing structural phase transition in lead halide perovskites to reversibly tune the phonon-photon coupling strength, leveraging the fact that their phonon frequencies and oscillator strengths are dictated by lattice symmetry. Terahertz time-domain spectroscopy of MAPbI3 embedded in nanoslot cavities reveals three polariton branches above the critical temperature Tc = 162.5 K, and the emergence of an additional branch below Tc, activated by a new phonon mode in the low-temperature phase. The full dispersion is accurately reproduced using a multimode Hopfield model, confirming that all normalized coupling strengths remain in the ultrastrong coupling regime. These results demonstrate symmetry-controlled tuning of ultrastrong coupling via phonon engineering in optical cavities.

quant-ph

Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems

Recent approaches in Conversational Recommender Systems (CRSs) have tried to simulate real-world users engaging in conversations with CRSs to create more realistic testing environments that reflect the complexity of human-agent dialogue. Despite the significant advancements, reliably evaluating the capability of CRSs to elicit user preferences still faces a significant challenge. Existing evaluation metrics often rely on target-biased user simulators that assume users have predefined preferences, leading to interactions that devolve into simplistic guessing game. These simulators typically guide the CRS toward specific target items based on fixed attributes, limiting the dynamic exploration of user preferences and struggling to capture the evolving nature of real-user interactions. Additionally, current evaluation metrics are predominantly focused on single-turn recall of target items, neglecting the intermediate processes of preference elicitation. To address this, we introduce PEPPER, a novel CRS evaluation protocol with target-free user simulators constructed from real-user interaction histories and reviews. PEPPER enables realistic user-CRS dialogues without falling into simplistic guessing games, allowing users to gradually discover their preferences through enriched interactions, thereby providing a more accurate and reliable assessment of the CRS's ability to elicit personal preferences. Furthermore, PEPPER presents detailed measures for comprehensively evaluating the preference elicitation capabilities of CRSs, encompassing both quantitative and qualitative measures that capture four distinct aspects of the preference elicitation process. Through extensive experiments, we demonstrate the validity of PEPPER as a simulation environment and conduct a thorough analysis of how effectively existing CRSs perform in preference elicitation and recommendation.

cs.IR

Multimode Phonon-Polaritons in Lead-Halide Perovskites in the Ultrastrong Coupling Regime

Phonons play a central role in fundamental solid-state phenomena, including superconductivity, Raman scattering, and symmetry-breaking phases. Harnessing phonons to control these effects and enable quantum technologies is therefore of great interest. However, most existing phonon control strategies rely on external driving fields or anharmonic interactions, limiting their applicability. Here, we realize multimode ultrastrong light--matter coupling and theoretically show the modulation of phonon emission. This regime is realized by coupling two optical phonon modes in lead halide perovskites to a nanoslot array functioning as a single-mode cavity. The small mode volume of the nanoslots enables high coupling strengths in the phonon-polariton system. We show theoretically that the nanoslot resonator mediates an effective interaction between phonon modes, leading to superthermal phonon bunching in thermal equilibrium between distinct modes. Our findings are well described by a multimode Hopfield model. This work establishes a pathway for engineering phononic properties for light-harvesting and light-emitting technologies.

quant-ph

Cavity-Mediated Coupling between Local and Nonlocal Modes in Landau Polaritons

The multimode ultrastrong coupling (USC) regime has emerged as a novel platform for accessing previously inaccessible phenomena in cavity quantum electrodynamics. Of particular interest are cavity-mediated correlations between local and nonlocal excitations, or equivalently, between modes at zero and finite in-plane momentum modes, which offer new opportunities for controlling light-matter interactions across space. However, direct experimental evidence of such interactions has remained elusive. Here, we demonstrate nonlocal multimode coupling in a Landau polariton system, where cavity photons simultaneously interact with the zero-momentum cyclotron resonance and finite-momentum magnetoplasmons of a two-dimensional electron gas in a GaAs quantum well. Our slot cavities, with their subwavelength mode volumes, supply in-plane momentum components that enable the excitation of finite-momentum matter modes. Terahertz time-domain magnetospectroscopy measurements reveal a clear splitting of the upper-polariton branch, arising from hybridization between magnetoplasmon modes and the cavity--cyclotron-resonance hybrids. Extracted coupling strengths confirm USC of the cyclotron resonance and strong coupling of the magnetoplasmon modes to the cavity field, respectively. The experimental results are well captured by the multimode Hopfield model and finite-element simulations. These findings establish a pathway for engineering multimode light-matter interactions involving zero- and finite-momentum matter modes in the USC regime.

quant-ph

Towards Personalized Conversational Sales Agents: Contextual User Profiling for Strategic Action

Conversational Recommender Systems (CRSs)aim to engage users in dialogue to provide tailored recommendations. While traditional CRSs focus on eliciting preferences and retrieving items, real-world e-commerce interactions involve more complex decision-making, where users consider multiple factors beyond simple attributes. To capture this complexity, we introduce Conversational Sales (CSALES), a novel task that integrates preference elicitation, recommendation, and persuasion within a unified conversational framework. To support realistic and systematic evaluation, we present CSUSER, an evaluation protocol with LLM-based user simulator grounded in real-world behavioral data by modeling fine-grained user profiles for personalized interaction. We also propose CSI, a conversational sales agent that proactively infers contextual user profiles and strategically selects actions through conversation. Comprehensive experiments show that CSI significantly improves both recommendation success and persuasive effectiveness across diverse user profiles.

cs.IR