arXiv Science⌕ Search

arXiv subjects

Thomas B. Michelon

Publications and source records attributed to Thomas B. Michelon.

2 recordsLinked to original sources

Scientific Data Analysis for Class-Informatics in Computational Taxonomy

In this A.I. era, Computational Taxonomy is proposed to study complex systems by analyzing their databases under taxonomic hierarchies abiding the Principle of Science by providing "good explanations". Comparisons among branches or classes are carried out by Scientific Data Analysis (SDA) paradigm that explores all potential associative patterns, including interacting effects of all high orders, and evaluate finite sample precisions for all information pieces individually by effectively making use of all variables' categorical nature. Under each comparison, all confirmed information pieces are collected and displayed along row-axis of a heatmap with all involved study-subjects on the column-axis. Each comparison's heatmap individually characterizes participating classes and study-subjects and simultaneously provides a scientific basis for outlier detection upon all non-participants. All these heatmaps then collectively constitutes so-called Class-informatics that offers good explanations based on characteristic of all classes and study-subjects. Computational Taxonomy's Class-informatics indeed resolves multiple fundamental issues: Tukey's more than 60 years outlier detection problem, issue of self-correction annotation, and a crucial check on assumption of information-content equality between testing and training data sets in Machine Learning. A showcase of Computational Taxonomy is exclusively illustrated on Iris data.

stat.CO↗

Penguin data reanalyzed via Computational Taxonomy

We employ Computational Taxonomy (CT) to reanalyze the penguin data set penguins_lter by validating and addressing two biological issues: Sexual Size Dimorphism (SSD) and mate-selection criteria. Via Scientific Data Analysis (SDA) computing, CT constructs a Taxonomic Hierarchy by splitting Species first and then Sex, without involving Island, to achieve less complexity. This Taxonomic Hierarchy validates SSD as a branch comparison: (Species, Sex = Male)-vs-(Species, Sex = Female), upon which SDA explores all potential pieces of associative information from all covariate feature-sets, including interacting effects from order-2 to order-4, and then confirms them via their idiosyncratic reliability checks. The collective of confirmed information pieces are displayed on a heatmap platform to manifest underlying dynamics of SSD with explicit block-structured heterogeneity found within males and females. SSD dynamics is explained through mechanistic dependence pertaining to one chief factor consisting of up to 8 feature-sets: Body-Mass coupled by combinations of {Culmen-length,Culmen-depth, Flipper-length}, and two minor factors consisting of low-order combinations of {Culmen-length,Culmen-depth, Flipper-length}. Such Intra-Sex heterogeneity invalidates all Logistic regression modeling on SSD in the original paper. Further, we explore potential mate-selection criteria through the data-frame of Nest-ID within-species homogeneity.

stat.CO↗