arXiv · 2609.13201
Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies
Abstract
Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \emph{main} data points when the number of anomalies is small and is ``taken" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ghurumuruhan Ganesan. 2026-08-14. Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies. https://arxiv.org/abs/2609.13201
Cite the original work for its findings. Save a collection to share your selection of sources.