arXiv ScienceSearch

arXiv subjects

Patrícia Martinková

Publications and source records attributed to Patrícia Martinková.

11 recordsLinked to original sources

Identifiability of Partial-Mastery Cognitive Diagnostic Models

Partial-mastery (PM) cognitive diagnostic models (CDMs) extend traditional CDMs by replacing binary latent attribute mastery indicators with continuous mastery scores for multiple latent attributes. In PM-CDMs, each subject is characterized by a fixed continuous latent mastery vector, from which item-specific binary attribute profiles are independently generated. This formulation provides a bridge between classical CDMs and continuous latent variable models. Despite growing interest in PM-CDMs, their identifiability properties remain unexplored. In this work, we establish the first identifiability results for PM-CDMs. We derive sufficient conditions for identifiability that are direct analogues of established conditions for traditional CDMs. To develop the main argument, we use symbolic computation on a minimal example with five items and two latent attributes to show that the Jacobian of the model parameterization is generically nonzero. Combining tools from real analysis and algebraic statistics, we prove that this local property implies generic finite-to-one identifiability of the item parameters and the marginal distributions of the relevant latent attributes. We further show that if the $Q$-matrix contains such identifiable local structures for all attribute pairs, identifiability extends to the full PM-CDM. These findings provide a rigorous theoretical foundation for estimation and inference in partial-mastery cognitive diagnostic models.

math.ST

Asymptotically exact threshold for detecting anomalies in multivariate Gaussian data with application to time series

In this paper, we propose a new thresholding technique for detecting anomalies in multivariate normal random samples, under the assumption that anomalous observations are sparse and differ from the rest of the data in their mean. The mean vector of the non-anomalous data is assumed to be zero, while the covariance matrix is unknown. We derive conditions on the mean shift of the anomalous observations, as well as on the covariance matrix and its estimator, under which the proposed procedure achieves asymptotically exact detection, meaning that the expected number of misclassified observations converges to zero as the sample size increases. In addition, we establish conditions under which exact anomaly detection is impossible for any procedure. The performance of the proposed method is illustrated through an extensive simulation study and compared with other widely used anomaly detection methods. Real-data analyses involving wearable activity measurements and air pollution time series provide an assessment of its performance in real-world settings.

stat.ME

Refining Effect-Size Measures and Classification for Differential Item Functioning: Toward Unified Guidelines Across Methods

Differential Item Functioning (DIF) analysis is used to identify potentially biased items in multi-item measurements. In addition to testing the statistical significance, it is essential to evaluate the practical significance of DIF through effect-size measures. We review existing DIF effect-size measures and cut-off values used to classify the effect-size magnitudes for the Mantel-Haenszel test, SIBTEST, and model-based methods for binary items, and introduce a refinement of area-based effect-size measures. A simulation study is conducted to investigate the properties of these effect-size measures and existing classification guidelines, and to assess their comparative performance. The results indicate that some commonly used effect-size measures exhibit undesirable properties, including inconsistent classifications, systematic underestimation of the magnitude of the underlying DIF, and strong dependence on design factors. To address these issues, we introduce usage restrictions for some effect-size measures, revise cut-off values that unify results across different methods, and propose new cut-off values for area-based effect-size measures. The methods are demonstrated using two real data examples. Implementation is provided in the R software.

stat.ME

Sequential generalized kernel equating: Providing comparable scores across multiple test forms with nonequivalent groups and differently measured covariates

Test equating using covariates may be applied to provide comparable scores from multiple test forms when no anchor items are available. However, its performance may be compromised if some of the covariates themselves are measured using different test forms. In this work, we propose sequential generalized kernel equating to account for possible differences in the distribution of covariates used in the NEC design. We evaluate the proposed approach through a simulation study within the kernel equating framework. Results indicate that equating the covariate reduces bias in equated test scores, particularly when the covariate distributions differ and the correlation between the covariate and the test score is strong. A real data example from a national high school leaving examination further demonstrates the practical application.

stat.ME

Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning

Item difficulty must often be estimated before test administration, when no responses are yet available for calibration. While most response-free difficulty modelling approaches derive item-text features by hand for a separate statistical model, we fine-tune a transformer end-to-end on the wording, avoiding the theory-based feature design and the preprocessing that discards information. We address reading-comprehension multiple-choice items, whose difficulty depends on inferential demands spanning passage, question, and options, yet the simplest model sees one undifferentiated sequence and is trained on difficulty alone. We introduce and investigate two extensions to the joint-encoding baseline: a component-wise variant, which encodes the wording parts separately, and a multi-task variant, which adds an auxiliary task of question answering. We compare the methods across three training-set sizes sampled from a corpus of nearly 30,000 items whose labels approximate response-based Rasch difficulty. At the smallest training size, both extensions improve on the baseline, the multi-task variant across every metric, and component-wise encoding in rank ordering. Further research may ground the auxiliary supervision in observed responses and extend the approach to other item types.

cs.CL

A novel nonparametric framework for DIF detection using kernel-smoothed item response curves

This study introduces a novel nonparametric approach for detecting Differential Item Functioning (DIF) in binary items through direct comparison of Item Response Curves (IRCs). Building on prior work on nonparametric comparison of regression curves, we extend the methodology to accommodate binary response data, which is typical in psychometric applications. The proposed approach includes a new estimator of the asymptotic variance of the test statistic and derives optimal weight functions that maximise local power. Because the asymptotic distribution of the resulting test statistic is unknown, a wild bootstrap procedure is applied for inference. A Monte Carlo simulation study demonstrates that the nonparametric approach effectively controls Type I error and achieves power comparable to the traditional logistic regression method, outperforming it in cases with multiple intersections of the underlying IRCs. The impact of bandwidth and weight specification is explored. Application to a verbal aggression dataset further illustrates the method's ability to detect subtle DIF patterns missed by parametric models. Overall, the proposed nonparametric framework provides a flexible and powerful alternative for detecting DIF, particularly in complex scenarios where traditional model-based assumptions may not be applicable.

stat.ME

Enhancing Psychometric Analysis with Interactive SIA Modules

ShinyItemAnalysis (SIA) is an R package and shiny application for an interactive presentation of psychometric methods and analysis of multi-item measurements in psychology, education, and social sciences in general. In this article, we present a new feature introduced in the recent version of the package, called "SIA modules", which allows researchers and practitioners to offer new analytical methods for broader use via add-on extensions. SIA modules are designed to integrate with and build upon the SIA interactive application, enabling them to leverage the existing infrastructure for tasks such as data uploading and processing. They can access and further use a range of outputs from various analyses, including models and datasets. Because SIA modules come in R packages (or extend the existing ones), they may come bundled with their datasets, use object-oriented systems, or even compiled code. We illustrate the concepts using sample modules from the newly introduced SIAmodules package and other packages. After providing a general overview of building Shiny applications, we describe how to develop the SIA add-on modules with the support of the new SIAtools package. Finally, we discuss possibilities of future development and emphasize the importance of freely available, interactive psychometric software for dissemination of methodological innovations.

cs.HC

Bridging Item Response Theory and Factor Analysis: A Four-Parameter Mixture-Dichotomized Model with Bayesian Estimation

Item Response Theory (IRT) and Factor Analysis (FA) are two major frameworks for modeling multi-item measurements of latent traits. While the relationship between two-parameter IRT models and dichotomized FA models is well established, FA formulations for IRT models with additional parameters are less common. We focus on the four-parameter factor-analytic (4P FA) model that extends the traditional dichotomized single-factor FA model through a hierarchical mixture formulation accounting for guessing and inattention effects. We analytically establish the equivalence of the 4P FA and 4P IRT models, extending the FA--IRT correspondence beyond the two-parameter case. A Bayesian estimation procedure is developed for model estimation, to estimate the four item parameters, the respondents' latent scores, and the scores adjusted for guessing and inattention effects. The proposed algorithm is implemented in \texttt{R} and \texttt{Python}. A simulation study compares estimation under the FA and IRT formulations of the 4P model and evaluates the practical implications of the FA parametrization. Empirical examples based on an admission test and an anxiety inventory demonstrate the correspondence between the 4P FA and 4P IRT models and illustrate the application of the proposed methodology.

stat.ME

New iterative algorithms for estimation of item functioning

This paper explores innovations to parameter estimation in generalized linear and nonlinear models, which may be used in item response modeling to account for guessing/pretending or slipping/dissimulation and for the effect of covariates. We introduce a new implementation of the EM algorithm and propose a new algorithm based on the parametrized link function. The two novel iterative algorithms are compared to existing methods in a simulation study. Additionally, the study examines software implementation, including the specification of initial values for numerical algorithms and asymptotic properties with an estimation of standard errors. Overall, the newly proposed algorithm based on the parametrized link function outperforms other procedures, especially for small sample sizes. Moreover, the newly implemented EM algorithm provides additional information regarding respondents' inclination to guess or pretend and slip or dissimulate when answering the item. The study also discusses applications of the methods in the context of the detection of differential item functioning and addresses the measurement error. Methods are offered in the difNLR package and in the interactive application of the ShinyItemAnalysis package; demonstration is provided using real data from psychological and educational assessments.

stat.ME

Assessing quality of selection procedures: Lower bound of false positive rate as a function of inter-rater reliability

Inter-rater reliability (IRR) is one of the commonly used tools for assessing the quality of ratings from multiple raters. However, applicant selection procedures based on ratings from multiple raters usually result in a binary outcome; the applicant is either selected or not. This final outcome is not considered in IRR, which instead focuses on the ratings of the individual subjects or objects. We outline the connection between the ratings' measurement model (used for IRR) and a binary classification framework. We develop a simple way of approximating the probability of correctly selecting the best applicants which allows us to compute error probabilities of the selection procedure (i.e., false positive and false negative rate) or their lower bounds. We draw connections between the inter-rater reliability and the binary classification metrics, showing that binary classification metrics depend solely on the IRR coefficient and proportion of selected applicants. We assess the performance of the approximation in a simulation study and apply it in an example comparing the reliability of multiple grant peer review selection procedures. We also discuss possible other uses of the explored connections in other contexts, such as educational testing, psychological assessment, and health-related measurement and implement the computations in IRR2FPR R package.

stat.ME

Assessing inter-rater reliability with heterogeneous variance components models: Flexible approach accounting for contextual variables

Inter-rater reliability (IRR), which is a prerequisite of high-quality ratings and assessments, may be affected by contextual variables such as the rater's or ratee's gender, major, or experience. Identification of such heterogeneity sources in IRR is important for implementation of policies with the potential to decrease measurement error and to increase IRR by focusing on the most relevant subgroups. In this study, we propose a flexible approach for assessing IRR in cases of heterogeneity due to covariates by directly modeling differences in variance components. We use Bayes factors to select the best performing model, and we suggest using Bayesian model-averaging as an alternative approach for obtaining IRR and variance component estimates, allowing us to account for model uncertainty. We use inclusion Bayes factors considering the whole model space to provide evidence for or against differences in variance components due to covariates. The proposed method is compared with other Bayesian and frequentist approaches in a simulation study, and we demonstrate its superiority in some situations. Finally, we provide real data examples from grant proposal peer-review, demonstrating the usefulness of this method and its flexibility in the generalization of more complex designs.

stat.ME