arXiv ScienceSearch

arXiv subjects

Nimesh Sinha

Publications and source records attributed to Nimesh Sinha.

4 recordsLinked to original sources

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

Multi-merchant e-commerce catalogs contain equivalent and related products under different merchant-scoped identifiers, fragmenting behavioral evidence across merchants. Expert-defined taxonomies, meanwhile, are often too coarse for fine-grained discovery. We investigate whether a single hierarchical Semantic ID (\sid{}) representation can support personalized ranking and query reformulation. Learned once from product-content embeddings, the hierarchy defines product concepts at multiple granularities that each application combines with its own behavioral and serving context. For ranking, we aggregate consumer affinity and product performance over \sid{} prefixes and derive sequence features for candidate products and consumer histories. Controlled ablations show improved offline relevance, while online evaluation of the full ranking treatment shows stronger top-slot add-to-cart engagement and broader exposure for less-popular products. For query reformulation, we ground queries and session transitions in \sid{} concepts, use the hierarchy for navigation and refinement, and filter suggestions against the merchant's assortment. Offline evaluation shows finer intent preservation than taxonomy and higher-quality suggestions than raw query-string transitions; online evaluation shows reduced search effort and earlier access to purchasable products. These results show that a shared semantic product hierarchy can support both recommendation and search while preserving the task-specific context required by each application.

cs.IR

Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations

In multi-vertical e-commerce platforms like DoorDash, relatively newer product verticals such as grocery and retail present a significant opportunity for personalization innovation. A key challenge lies in solving the "cold start" problem for users. This paper introduces a novel framework for enhancing recommendation quality by transferring knowledge from data-rich verticals (e.g., restaurants at DoorDash) to data-sparse ones. We leverage Large Language Models (LLMs) to perform generative inference, synthesizing sparse, high-dimensional features that encapsulate latent user affinities. Specifically, we employ a hierarchical Retrieval-Augmented Generation (RAG) pipeline to derive multi-level taxonomic features from user restaurant order histories and search queries. These generated features, encoding both long-term cross-vertical preferences and short-term intent, are integrated into a production Multi-Task Learning (MTL) ranking model. We demonstrate through extensive offline and online evaluation that this approach significantly improves personalization and engagement in emerging business verticals, effectively bridging the behavioral data gap.

cs.IR

Event-based Product Carousel Recommendation with Query-Click Graph

Many current recommender systems mainly focus on the product-to-product recommendations and user-to-product recommendations even during the time of events rather than modeling the typical recommendations for the target event (e.g., festivals, seasonal activities, or social activities) without addressing the multiple aspects of the shopping demands for the target event. Product recommendations for the multiple aspects of the target event are usually generated by human curators who manually identify the aspects and select a list of aspect-related products (i.e., product carousel) for each aspect as recommendations. However, building a recommender system with machine learning is non-trivial due to the lack of both the ground truth of event-related aspects and the aspect-related products. To fill this gap, we define the novel problem as the event-based product carousel recommendations in e-commerce and propose an effective recommender system based on the query-click bipartite graph. We apply the iterative clustering algorithm over the query-click bipartite graph and infer the event-related aspects by the clusters of queries. The aspect-related recommendations are powered by the click-through rate of products regarding each aspect. We show through experiments that this approach effectively mines product carousels for the target event.

cs.IR

Sunspot area catalogue revisited: Daily cross-calibrated areas since 1874

Long and consistent sunspot area records are important for understanding the long-term solar activity and variability. Multiple observatories around the globe have regularly recorded sunspot areas, but such individual records only cover restricted periods of time. Furthermore, there are also systematic differences between them, so that these records need to be cross-calibrated before they can be reliably used for further studies. We produce a cross-calibrated and homogeneous record of total daily sunspot areas, both projected and corrected, covering the period between 1874 and 2019. A catalogue of calibrated individual group areas is also generated for the same period. We have compared the data from nine archives: Royal Greenwich Observatory (RGO), Kislovodsk, Pulkovo, Debrecen, Kodaikanal, Solar Optical Observing Network (SOON), Rome, Catania, and Yunnan Observatories, covering the period between 1874 and 2019. Mutual comparisons of the individual records have been employed to produce homogeneous and inter-calibrated records of daily projected and corrected areas. As in earlier studies, the basis of the composite is formed by the data from RGO. After 1976, the only datasets used are those from Kislovodsk, Pulkovo and Debrecen observatories. This choice was made based on the temporal coverage and the quality of the data. In contrast to the SOON data used in previous area composites for the post-RGO period, the properties of the data from Kislovodsk and Pulkovo are very similar to those from the RGO series. They also directly overlap the RGO data in time, which makes their cross-calibration with RGO much more reliable. We have also computed and provide the daily Photometric Sunspot Index (PSI) widely used, e.g., in empirical reconstructions of solar irradiance.

astro-ph.SR