arXiv ScienceSearch

arXiv · 2310.00417

Teaching at the Intersection of Social Justice, Ethics, and the ASA Ethical Guidelines for Statistical Practice

Abstract

Case studies are typically used to teach 'ethics', but when the content of a course is focused on formulae and proofs, a case analysis and the knowledge, skills, and abilities they require can be distracting. Moreover, case analyses are typically focused narrowly on research issues: obtaining consent, dealing with research team members, and/or research policy violations. Not all students in quantitative courses plan to become researchers, and ethical practice of mathematics, statistics, data science, and computing is an essential topic regardless of the learner's career plans. While it is incorrect to treat 'social justice' as a proxy for 'ethical practice', the topic of 'social justice' may be more interesting to both students and instructors. This paper offers concrete recommendations for integrating social justice content into quantitative courses in ways that limit the burden of new knowledge, skills, and abilities but also support reproducible and actionable assessments. Five tools can be utilized to integrate social justice into a course in a way that also meets calls to integrate 'ethics'; minimizes the burden on instructors to create and grade new materials and assignments; minimizes the burden on learners to develop the skill set to complete a case analysis; and maximizes the likelihood that the ethics content will be embedded in the learners' cognitive representation of the knowledge being taught in the quantitative course. These tools are: a. Curriculum Development Guidelines b. 7-task Statistics and Data Science Pipeline c. ASA Ethical Guidelines for Statistical Practice d. Stakeholder Analysis e. 6-step Ethical Reasoning paradigm This paper discusses how to use these tools in quantitative courses. The tools and frameworks offer structure, and facilitate ensuring that changes made to any course are evaluable and generate actionable assessments for learners.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rochelle E Tractenberg. 2023-09-30. Teaching at the Intersection of Social Justice, Ethics, and the ASA Ethical Guidelines for Statistical Practice. https://arxiv.org/abs/2310.00417

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Uncertainty-Aware Missing-Data Multimodal Latent for Fetal-Growth Analysis

Objective: Routine third-trimester examination yields fetal biometry, maternal, Doppler and fetal-cardiac measurements, acquired at clinical discretion and therefore often incomplete. We show that these four measurement blocks are close to mutually uninformative, and that this single property determines what a representation of them can impute, what it can audit, and what fetal size alone cannot indicate. Methods: A linear-Gaussian factor model (K = 8 by parallel analysis, VARIMAX-rotated) was fitted to 25 measurements in four blocks from 977 fetuses (169 SGA, 61 severe; 77 LGA). Posterior precision sums contributions from observed measurements only, so missing values are marginalized rather than imputed. Data quality was screened using the standardized residual between each measurement and its reconstruction. Results: Predicting any one block from the other three gives an out-of-fold R2 of 0.023. The representation is a continuous growth spectrum with no cluster structure (Hartigan dip p = 0.99, gap statistic k = 1, three-cluster silhouette 0.07) ordering fetuses by birthweight centile (Spearman rho = 0.55). Among 169 SGA fetuses the haemodynamic redistribution axis separated the 25 adverse outcomes (AUC 0.70, 0.585-0.808) where measured size did not (0.60, 0.451-0.738). With Doppler censored, the marginalized interval covered held-out measurements in 97% and 93% of cases against nominal 95% and 90%. The reconstruction residual flagged 38 of 977 records, 36 confirmed transcription errors in the registry. Conclusion: Marginalizing missing measurements yields a representation whose uncertainty reflects the available data, and whose reconstruction residual doubles as a data-quality screen. Because the blocks are nearly independent, confirmed flags are within-block errors, and a synthetic benchmark gives the coupling needed before cross-block detection becomes available.

stat.AP

Auditing the Global Carbon Budget: Exploring the 2024--2025 Vintage Shift

The Global Carbon Budget (GCB), the community reference dataset for the carbon cycle, is reissued annually. The 2025 release introduces several adjustments to the published series that we compare with prior releases starting in 2017. On a common 1959-2016 sample, the mean of the GCB budget imbalance jumps from within +/-0.17 GtC/yr of zero for every vintage 2017-2024 to 0.61 GtC/yr in 2025, the only vintage whose 95% confidence interval for the imbalance mean excludes zero. The size of the imbalance changes much less: its mean absolute value rises from 0.61 to 0.76 GtC/yr. It is the mean, the quantity the budget identity constrains, that moves. We document and explore this shift in two ways. First, we conduct a model-free analysis, where we attribute the shift to a new adjustment that places the published land sink 0.40 GtC/yr below its ensemble mean (the average of the underlying models), a smaller adjustment in the ocean sink in the opposite direction, and a change in the composition of the bookkeeping ensemble. Second, we consider a dynamic statistical GCB model augmented with climate covariates. Its parameters are estimated for every GCB vintage 2017-2025. The coefficients of atmospheric concentrations in the sink equations shift in opposite directions on the 2025 issue, mirroring the model-free findings. A constant in the budget equation, statistically unnecessary in every vintage from 2017-2024, is required in 2025 and is estimated at -0.59 (0.09) GtC/yr. There is a persistent drifting imbalance across the entire sample in the budget equation. Each of the three adjustments is documented in the 2025 release and rests on evidence about the component it corrects. Their joint effect is a budget that closes over the last ten years and carries a mean imbalance of 0.61 GtC/yr over the full record. We argue that this cost to the full sample outweighs the gain on the last ten years.

stat.AP

Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights

The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.

stat.AP