Singular learning theory and information geometry have studied the same parameter spaces in mostly separate vocabularies: the former computes Bayesian invariants in resolved coordinates, the latter works in original coordinates under a non-degeneracy assumption that overparameterised models routinely violate. We bridge them through one primitive, the dead direction: a unit vector along which the Fisher metric degenerates, equivalently a tangent to the analytic singular set with a definite KL order, set by how fast the KL divergence vanishes. The two readings name the same vector; our central move shows its KL order is recoverable as the decay rate of the directional Fisher curvature approaching the singularity, in original parameter coordinates and without a Hironaka resolution. A selection rule on smooth fibres translates this rate into Watanabe's single-direction contribution to the real log canonical threshold, and we extend the recovery to multi-component crossings, multiplicity m, the singular fluctuation ν (universal in the KL order for 1D directions), prior-RLCT shifts, and tempered posteriors. We then lift this rate to a deep network: a multi-layer K-FAC factorisation writes each Fisher block as a product of activation- and gradient-side rates with a duality between them, instantiated at modern-network primitives (residual streams, layer normalisation, attention). A quotient theorem carries the rate to the gauge quotient Θ/G under gradient flow on a G-invariant metric; SGD qualifies, standard Adam does not, and we construct a G-equivariant Adam-family preconditioner (DDCAdam) that does. The bridge yields a parameter-coordinate handle on singular geometry, closed-form per-architecture predictions, and a trajectory-rate readout of Watanabe's triple (λ, m, ν) from one checkpoint's forward and backward passes, without posterior sampling.
Pretrained transformers sit near singular minima of the loss, where the Fisher information metric degenerates along dead directions: directions in parameter space along which the directional Fisher vanishes. Locating such a direction normally needs a forward pass and an eigendecomposition of activations, or a sampling-based complexity estimate; none returns a direction computable from the network's parameters alone. We give one, for LayerNorm transformers. The inverse-scale direction γ^-1/γ^-1 of the LayerNorm affine is an exact algebraic kernel of the post-final-norm centred activation covariance, for any input distribution, and induces a corresponding dead direction in parameter space. It is read from the LN scale parameter alone, with no forward or backward pass and no eigensolve: the cheapest dead-direction read, specific to LayerNorm. We test it on 14 pretrained transformers (9 LayerNorm, 5 RMSNorm; 160M-35B; language and vision objectives). At random initialisation the predicted direction matches the measured bottom singular direction (one forward pass, direct SVD) to four decimal places on 9/9 LayerNorm models, and is correctly absent on 5/5 RMSNorm models, which lack the mean-subtraction projector that creates it. On the trained checkpoint the covariance eigenvalue along this direction deepens by ∼10^3× and further dead directions open; the random-init-to-trained gap is a one-forward-pass, per-checkpoint readout of singular structure along the predicted coordinate. Two consequences follow in closed form: the residual stream's smallest singular value is preserved block-to-block on 13/14 transformers measured on their own input distribution, the one exception (Gemma4-31B) a genuine dead direction the same read pinpoints; and the kernel direction's presence classifies a transformer's normalisation from the parameters alone.
Singular learning theory characterises the complexity of a deep network through the geometry of its loss singularities. The local learning coefficient (LLC), the standard estimator of Watanabe's real log canonical threshold (RLCT, λ), reads this geometry as an integrated Bayesian scalar through SGLD, which needs per-task calibration and 10^4-10^6 forward-backward passes per checkpoint. We introduce Dead-Direction Signatures (DDS), a family of cheap closed-form spectral readings of singular structure: each reads a network's activation matrix or per-sample-gradient Fisher-Gram at a chosen layer, replacing the SGLD posterior chain with spectral linear algebra. The readings rest on a dead-direction framework that predicts a structural correlation between activation- and Fisher-side spectra at any singular minimum, and a rank-multiplicative volume identity that single-eigenvalue monitors cannot produce: the active-volume log^+(G) slope counts the dead directions, tracking the rank-deficit r across r ∈{1,2,3,4} (slope ratios 2.0, 3.1, 4.0 at r=2,3,4 against the predicted 2,3,4), where the smallest eigenvalue is rank-blind. On reduced-rank regression with closed-form λ, calibrated LLC recovers λ at 99% mean and the DDS observables rank-track it at the framework-predicted sign; on a non-linear modular-addition transformer DDS separates d_model across eighteen orders of magnitude where calibrated LLC at the protocol budget is rank-flat. Complementary to LLC's integrated posterior reading, DDS gives a directional, layer-local handle on a network's dead directions, read in closed form from its activation and gradient spectra.
A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable. We build DDC, a Dead-Direction Conditioner that lifts a base optimizer into a G-equivariant one: it conditions the optimizer's state in the orbit decomposition of a G-invariant metric, so the trajectory stays a preconditioned gradient flow on the quotient = Θ/G. The construction carries four architectural gauges (cross-entropy shift, ReLU and SwiGLU rescaling, LayerNorm and RMSNorm scale, and a per-head O(d_ head) attention rotation matched to RoPE), proves exactly equivariant on an Adam base, and composes with a Muon base through a gauge-equivariant orthogonaliser. Respecting the symmetry changes both the minimum the optimizer reaches and what it leaves measurable there. On a language model trained past the point of fit, DDCAdam resists the over-training collapse AdamW falls into, holding a validation-train loss gap of 0.67 against 5.88, and reads the dead-direction rate in 32 of 65 layer-by-observable cells where AdamW reads it in 7. A vision transformer trained from scratch reaches lower validation loss (1.71 against 2.12) while compressing spare feed-forward capacity a matched AdamW leaves intact. On a Muon base, where the rotation gauge composes exactly, DDCMuon groks ten of eleven seeds at depth 24 that a plain Muon never reaches. Built into the optimizer, a network's gauge symmetry sharpens the minimum it finds and turns that minimum's geometry into something the trajectory can measure.
This paper introduces EvoGabor, a hardware-efficient approach for implementing Gabor filters (GFs) in resource-constrained medical image processing applications. While GFs effectively extract orientation- and frequency-selective features, they require high-precision multipliers, making implementation computationally demanding. To address this problem, We propose an evolutionary optimization framework that explores different precision levels for Gabor kernel coefficients, reducing hardware complexity while preserving image quality measured by Structural Similarity Index (SSIM). We quantized the original full-precision fixed-point 16 (FI-16) bit coefficients with graded precision to significantly reduce dynamic power consumption, silicon footprint, and design delay. Additionally, we employ clustering techniques to reduce the effective dimensionality of Gabor kernels, thereby reducing hardware redundancy and enabling kernel reuse. The proposed architecture, characterized on standard-cell synthesized in NANDgate 45 nm PDK, reported area and power savings of 79% and 6%, respectively, compared to the conventional baseline FI-16 precision-based GF architecture. This work successfully bridges the gap between algorithmic accuracy and hardware feasibility by combining optimization strategies with architectural simplification. The hardware design files are made freely available for further usage to the research and design community.