arXiv ScienceSearch

arXiv subjects

Diganta Mukherjee

Publications and source records attributed to Diganta Mukherjee.

At least 19 recordsLinked to original sources

Measuring Tail Dependence in Linear Processes: Theory and Empirics

The quantitative analysis of financial time series often reveals two distinct features that standard Gaussian frameworks fail to capture: heavy-tailed marginal distributions and the phenomenon of extreme co-movements.While extreme value theory characterizes marginal behavior, Copulas provide a functional bridge to describe the dependence structure independently of the marginals. We are proposing a different way of looking at the joint extremes on the basis of a dependence measure. The proposed idea incorporates both the non-identical and identical regularly varying distributions. Informed by the analysis of some high-frequency cryptocurrency datasets, the effect of persistence property have been thoroughly studied under these setups. A detailed simulation study confirms our intuition and findings.

math.ST

An Augmented Rating System for Test cricket: adapting the Glicko rating system

The International Cricket Council's (ICC) Test cricket ratings use match and series results alone, stating no allowance about the precision of a rating, or about home advantage and toss. We fold both into the Glicko rating system, which gives a probabilistic expected score, and recalibrate its scale for Test cricket. The two enter as covariates whose significance and weights are estimated from match data. Home advantage is worth approximately 13 rating points and the toss roughly 8. They add, with no significant interaction. Over the two completed World Test Championship cycles (2021-23 and 2023-25) the model predicts 77.6% of decisive matches correctly in the first, and matches whichever of standard Elo and unmodified Glicko rating system is more accurate, at a lower Brier score and log loss in both. Permuting each cycle's match order 1000 times leaves every final rating unchanged. The model shows robustness to ordering, not fairness towards the fixture list, which the unbalanced calendar prevents judging. Our ordering follows the ICC's (Spearman 0.979 and 0.983). We add not a different ranking but a calibrated win probability per match, a deviation on every rating, and home and toss adjustments of estimated size, none of which the ICC supplies.

stat.AP

MinDist is less than 7

The metric MinDist, introduced recently to quantify the distance of an arbitrary Rummy hand from a valid declaration, plays a central role in algorithmic hand evaluation and optimal play. Existing results show that the MinDist of any $13$-card Rummy hand from a single deck is bounded above by $9$. In this paper, we sharpen this bound and prove that the MinDist of any hand is at most $7$. We further show that this bound is tight by explicitly exhibiting a hand whose MinDist equals $7$ for a suitable choice of wildcard joker. The proof combines elementary combinatorial arguments with structural properties of card partitions across suits and resolves the gap between the previously known upper bound and the true extremal value.

math.CO

Quantitative Rule-Based Strategy modeling in Classic Indian Rummy: A Metric Optimization Approach

The 13-card variant of Classic Indian Rummy is a sequential game of incomplete information that requires probabilistic reasoning and combinatorial decision-making. This paper proposes a rule-based framework for strategic play, driven by a new hand-evaluation metric termed MinDist. The metric modifies the MinScore metric by quantifying the edit distance between a hand and the nearest valid configuration, thereby capturing structural proximity to completion. We design a computationally efficient algorithm derived from the MinScore algorithm, leveraging dynamic pruning and pattern caching to exactly calculate this metric during play. Opponent hand-modeling is also incorporated within a two-player zero-sum simulation framework, and the resulting strategies are evaluated using statistical hypothesis testing. Empirical results show significant improvement in win rates for MinDist-based agents over traditional heuristics, providing a formal and interpretable step toward algorithmic Rummy strategy design.

cs.AI

Analyzing Skill Element in Online Fantasy Cricket

Online fantasy cricket has emerged as large-scale competitive systems in which participants construct virtual teams and compete based on real-world player performances. This massive growth has been accompanied by important questions about whether outcomes are primarily driven by skill or chance. We develop a statistical framework to assess the role of skill in determining success on these platforms. We construct and analyze a range of deterministic and stochastic team selection strategies, based on recent form, historical statistics, statistical optimization, and multi-criteria decision making. Strategy performance is evaluated based on points, ranks, and payoff under two contest structures Mega and 4x or Nothing. An extensive comparison between different strategies is made to find an optimal set of strategies. To capture adaptive behavior, we further introduce a dynamic tournament model in which agent populations evolve through a softmax reweighting mechanism proportional to positive payoff realizations. We demonstrate our work by running extensive numerical experiments on the IPL 2024 dataset. The results provide quantitative evidence in favor of the skill element present in online fantasy cricket platforms.

cs.GT

Adapting Skill Ratings to Luck-Based Hidden-Information Games

Rating systems play a crucial role in evaluating player skill across competitive environments. The Elo rating system, originally designed for deterministic and information-complete games such as chess, has been widely adopted and modified in various domains. However, the traditional Elo rating system only considers game outcomes for rating calculation and assumes uniform initial states across players. This raises important methodological challenges in skill modelling for popular partially randomized incomplete-information games such as Rummy. In this paper, we examine the limitations of conventional Elo ratings when applied to luck-driven environments and propose a modified Elo framework specifically tailored for Rummy. Our approach incorporates score-based performance metrics and explicitly models the influence of initial hand quality to disentangle skill from luck. Through extensive simulations involving 270,000 games across six strategies of varying sophistication, we demonstrate that our proposed system achieves stable convergence, superior discriminative power, and enhanced predictive accuracy compared to traditional Elo formulations. The framework maintains computational simplicity while effectively capturing the interplay of skill, strategy, and randomness, with broad applicability to other stochastic competitive environments.

cs.GT

Empirical parameterization of the Elo Rating System

This study aims to provide a data-driven approach for empirically tuning and validating rating systems, focusing on the Elo system. Well-known rating frameworks, such as Elo, Glicko, TrueSkill systems, rely on parameters that are usually chosen based on probabilistic assumptions or conventions, and do not utilize game-specific data. To address this issue, we propose a methodology that learns optimal parameter values by maximizing the predictive accuracy of match outcomes. The proposed parameter-tuning framework is a generalizable method that can be extended to any rating system, even for multiplayer setups, through suitable modification of the parameter space. Implementation of the rating system on real and simulated gameplay data demonstrates the suitability of the data-driven rating system in modeling player performance.

stat.AP

Causal Analysis of Health, Education, and Economic Well-Being in India -- Evidence from the Young Lives Survey

This study investigates the dynamic and potentially causal relationships among childhood health, education, and long-term economic well-being in India using longitudinal data from the Young Lives Survey. While prior research often examines these domains in isolation, we adopt an integrated empirical framework combining panel data methods, instrumental variable regression, and causal graph analysis to disentangle their interdependencies. Our analysis spans five survey rounds covering two cohorts of children tracked from early childhood to young adulthood. Results indicate strong persistence in household economic status, highlighting limited intergenerational mobility. Education, proxied by Item Response Theory-based mathematics scores, consistently emerges as the most robust predictor of future economic well-being, particularly in the younger cohort. In contrast, self-reported childhood health shows limited direct impact on either education or later wealth, though it is influenced by household economic conditions. These findings underscore the foundational role of wealth and the growing importance of cognitive achievement in shaping life trajectories. The study supports policy approaches that prioritize early investments in learning outcomes alongside targeted economic support for disadvantaged households. By integrating statistical modeling with development policy insights, this research contributes to understanding how early-life conditions shape economic opportunity in low- and middle-income contexts.

stat.AP

Statistical inference for core-periphery structures

Core-periphery (CP) structure is an important meso-scale network property where nodes group into a small, densely interconnected {core} and a sparse {periphery} whose members primarily connect to the core rather than to each other. While this structure has been observed in numerous real-world networks, there has been minimal statistical formalization of it. In this work, we develop a statistical framework for CP structures by introducing a model-agnostic and generalizable population parameter which quantifies the strength of a CP structure at the level of the data-generating mechanism. We study this parameter under four canonical random graph models and establish theoretical guarantees for label recovery, including exact label recovery. Next, we construct intersection tests for validating the presence and strength of a CP structure under multiple null models, and prove theoretical guarantees for type I error and power. These tests provide a formal distinction between exogenous (or induced) and endogenous (or intrinsic) CP structure in heterogeneous networks, enabling a level of structural resolution that goes beyond merely detecting the presence of CP structure. The proposed methods show excellent performance on synthetic data, and our applications demonstrate that statistically significant CP structure is somewhat rare in real-world networks.

stat.ME

Skill vs. Chance Quantification for Popular Card & Board Games

This paper presents a data-driven statistical framework to quantify the role of skill in games, addressing the long-standing question of whether success in a game is predominantly driven by skill or chance. We analyze player level data from four popular games Chess, Rummy, Ludo, and Teen Patti, using empirical win statistics across varying levels of experience. By modeling win rate as a function of experience through a regression framework and employing empirical bootstrap resampling, we estimate the degree to which outcomes improve with repeated play. To summarize these dynamics, we propose a flexible skill score that emphasizes learning over initial performance, aligning with practical and regulatory interpretations of skill. Our results reveal a clear ranking, with Chess showing the highest skill component and Teen Patti the lowest, while Rummy and Ludo fall in between. The proposed framework is transparent, reproducible, and adaptable to other game formats and outcome metrics, offering potential applications in legal classification, game design, and player performance analysis.

cs.GT

Capturing Perception to Poverty using Conjoint Analysis & Partial Profile Choice Experiment

The objective of this study is applying a utility based analysis to a comparatively efficient design experiment which can capture people's perception towards the various components of a commodity. Here we studied the multi-dimensional poverty index and the relative importance of its components and their two-factor interaction effects. We also discussed how to model a choice based conjoint data for determining the utility of the components and their interactions. Empirical results from survey data shows the nature of coefficients, in terms of utility derived by the individuals, their statistical significance and validity in the present framework. There has been some discrepancies in the results between the bootstrap model and the original model, which can be understood by surveying more people, and ensuring comparative homogeneity in the data.

stat.AP

Skill Dominance Analysis of Two(Four) player, Three(Five) dice Variant of the Ludo Game

This paper examines two different variants of the Ludo game, involving multiple dice and a fixed number of total turns. Within each variant, multiple game lengths (total no. of turns) are considered. To compare the two variants, a set of intuitive, rule-based strategies is designed, representing different broad methods of strategic play. Game play is simulated between bots (automated software applications executing repetitive tasks over a network) following these strategies. The expected results are computed using certain game theoretic and probabilistic explanations, helping to understand the performance of the different strategies. The different strategies are further analyzed using win percentage in a large number of simulations, and Nash Equilibrium strategies are computed for both variants for a varying number of total turns. The Nash Equilibrium strategies across different game lengths are compared. A clear distinction between performances of strategies is observed, with more sophisticated strategies beating the naive one. A gradual shift in optimal strategy profiles is observed with changing game length, and certain sophisticated strategies even confound each other's performance while playing against each other.

cs.GT

Significance of Anatomical Constraints in Virtual Try-On

The system of Virtual Try-ON (VTON) allows a user to try a product virtually. In general, a VTON system takes a clothing source and a person's image to predict the try-on output of the person in the given clothing. Although existing methods perform well for simple poses, in case of bent or crossed arms posture or when there is a significant difference between the alignment of the source clothing and the pose of the target person, these methods fail by generating inaccurate clothing deformations. In the VTON methods that employ Thin Plate Spline (TPS) based clothing transformations, this mainly occurs for two reasons - (1)~the second-order smoothness constraint of TPS that restricts the bending of the object plane. (2)~Overlaps among different clothing parts (e.g., sleeves and torso) can not be modeled by a single TPS transformation, as it assumes the clothing as a single planar object; therefore, disregards the independence of movement of different clothing parts. To this end, we make two major contributions. Concerning the bending limitations of TPS, we propose a human AnaTomy-Aware Geometric (ATAG) transformation. Regarding the overlap issue, we propose a part-based warping approach that divides the clothing into independently warpable parts to warp them separately and later combine them. Extensive analysis shows the efficacy of this approach.

cs.CV

Analysis and Estimation of Consumer Expenditure assuming uniform prices across FSUs: A step towards cost effectiveness

Analysis and estimation of consumer expenditure and budget shares are important for understanding quantitatively the expenditure based behaviour of the people of a country or region. The costs attached with performing consumer expenditure and budget shares survey are significant. These surveys are quite time consuming and even a slight increase in sample size could lead to a significant increase in the cost of the survey under the current structure. This high cost for the survey also reduces its flexibility. Hence it is of paramount importance to be able to provide a statistically sound and relatively cheaper facilitation method of survey to be able to understand the consumption expenditure distribution of a country or a region without much loss of inferential power. In the context of the Indian National Sample Survey, in this paper we perform analysis and estimation of consumer expenditure assuming uniform prices across First Stage Units (FSUs) of a sampling design, and check its feasibility in estimating the consumer expenditure distribution of the country. We also compare it with the existing methodology and infer that there is no significant loss of information in the estimated consumer expenditure distribution and other inferences like the Lorenz curve and the Gini Index, when uniform distribution of prices within any FSU is assumed.

stat.AP

Significance of Skeleton-based Features in Virtual Try-On

The idea of \textit{Virtual Try-ON} (VTON) benefits e-retailing by giving an user the convenience of trying a clothing at the comfort of their home. In general, most of the existing VTON methods produce inconsistent results when a person posing with his arms folded i.e., bent or crossed, wants to try an outfit. The problem becomes severe in the case of long-sleeved outfits. As then, for crossed arm postures, overlap among different clothing parts might happen. The existing approaches, especially the warping-based methods employing \textit{Thin Plate Spline (TPS)} transform can not tackle such cases. To this end, we attempt a solution approach where the clothing from the source person is segmented into semantically meaningful parts and each part is warped independently to the shape of the person. To address the bending issue, we employ hand-crafted geometric features consistent with human body geometry for warping the source outfit. In addition, we propose two learning-based modules: a synthesizer network and a mask prediction network. All these together attempt to produce a photo-realistic, pose-robust VTON solution without requiring any paired training data. Comparison with some of the benchmark methods clearly establishes the effectiveness of the approach.

cs.CV

Exploring Financial Networks Using Quantile Regression and Granger Causality

In the post-crisis era, financial regulators and policymakers are increasingly interested in data-driven tools to measure systemic risk and to identify systemically important firms. Granger Causality (GC) based techniques to build networks among financial firms using time series of their stock returns have received significant attention in recent years. Existing GC network methods model conditional means, and do not distinguish between connectivity in lower and upper tails of the return distribution - an aspect crucial for systemic risk analysis. We propose statistical methods that measure connectivity in the financial sector using system-wide tail-based analysis and is able to distinguish between connectivity in lower and upper tails of the return distribution. This is achieved using bivariate and multivariate GC analysis based on regular and Lasso penalized quantile regressions, an approach we call quantile Granger causality (QGC). By considering centrality measures of these financial networks, we can assess the build-up of systemic risk and identify risk propagation channels. We provide an asymptotic theory of QGC estimators under a quantile vector autoregressive model, and show its benefit over regular GC analysis on simulated data. We apply our method to the monthly stock returns of large U.S. firms and demonstrate that lower tail based networks can detect systemically risky periods in historical data with higher accuracy than mean-based networks. In a similar analysis of large Indian banks, we find that upper and lower tail networks convey different information and have the potential to distinguish between periods of high connectivity that are governed by positive vs negative news in the market.

q-fin.ST

Demand Analysis with a Thin Price Sample

For about 125 items of food, the Consumer Expenditure Survey (CES) schedule of the Indian National Sample Survey asks the interviewer to obtain both quantity and value of household consumption during the reference period from the respondent. This would appear to put a great burden on the respondent. But it is likely that the price usually paid is almost the same within each first stage unit (fsu). The present work proposes a new sampling scheme to estimate demand elasticities of essential food items. While the conventional sampling method used in practice (e.g. in NSS consumer expenditure survey) involves seeking price information from many households sampled from a fsu, the proposed procedure involves only one household chosen randomly from every fsu for price data collection and thus requires much less interview burden. Using unit records for vegetable items in the NSS's 2011-12 CES, our results show that in spite of requiring much less data, the new scheme captures the household food consumption behavior as precisely as before.

stat.AP

Multi-asset Generalised Variance Swaps in Barndorff-Nielsen and Shephard model

This paper proposes swaps on two important new measures of generalized variance, namely the maximum eigenvalue and trace of the covariance matrix of the assets involved. We price these generalized variance swaps for Barndorff-Nielsen and Shephard model used in financial markets. We consider multiple assets in the portfolio for theoretical purpose and demonstrate our approach with numerical examples taking three stocks in the portfolio. The results obtained in this paper have important implications for the commodity sector where such swaps would be useful for hedging risk.

q-fin.MF