Machine Learning and Knowledge Discovery in Databases Research Track(2026)
Rheinische Friedrich-Wilhelms-Universität Bonn
被引用0|浏览0
摘要
We show that grokking, the sudden generalization of overfit neural networks, coincides with a geometric phase transition in representation space. Using Archetypal Analysis (AA), we decompose residual-stream activations of single-layer Transformers trained on modular arithmetic into convex combinations of extreme points. Before grokking, activations are diffuse. At the transition, they reorganize into a simplex whose P vertices correspond one-to-one with output classes, and this co-emergence serves as a mechanistic signature of grokking. The simplex structure is then causally validated by direct intervention: swapping a sample’s AA mixture weights to those of a target class redirects the model’s prediction with 98– 100% accuracy, while k-means ( ∼6% ) and unstructured baselines fail. The decomposition requires no label supervision: overcomplete AA with behavior-based merging recovers exactly P groups with 100% label alignment across all seeds. Archetype permutations further reveal cyclic group structure, implementing y ↦ y+d P at 80– 85% with k=P and 100% with the behavior-merged k = 2P→P representation, connecting the simplex geometry to the Fourier mode theory of grokking.