Directional Linear Separability of Neural Representations: Geometry and Transformations
Neural networks build representations through affine maps and nonlinear activations. Injective affine maps preserve linear separability, raising the problem of how they prepare data for nonlinear improvement and how much gain can be guaranteed before complete separation. We introduce the directional linear separability measure (D-LSM), which quantifies unavoidable competing-sample intrusion over affine halfspaces retaining every target sample, characterize its supporting geometry, and prove invariance under injective affine embeddings. For gated activations including ReLU, GELU, and SiLU, pre-activation projection bounds yield sufficient conditions for preserving all previous exclusions and recovering additional samples, with a gain bound determined by the certified recovery count. Under an aggregate-tube condition, an explicit affine construction realizes recovery with sufficient width, scaling conditions, and simultaneous multiclass guarantees through a shared layer. Exact controlled experiments compare certified and realized gains, assess certificate coverage, and exhibit bound attainment before complete separation and in affine-tube constructions. In learned Vision Transformer (ViT) representations, a feasible lower-bound estimator yields earlier post-GELU saturation certificates of exact separability, while boundary transport numerically supports affine invariance.