Tag
This paper introduces the first method for continuous gradient descent optimization in machine learning models with p-adic parameters, using the Berkovich affine line to enable effective learning on tasks like modular arithmetic.
This paper develops an ℓ0-type stability theory for subdominant (minmax) ultrametrics, proving that sparse edits propagate only through the minimum spanning tree and deriving Hamming–Lipschitz bounds on changed ultrametric entries. Experiments on deep-embedding graphs and clustering tasks demonstrate the utility of the resulting structural scores as vulnerability diagnostics.
This paper introduces Dynamic Ultrametric Attention, a framework where Transformers learn per-head block-sparse routing topologies during training, which are then offloaded to a custom Triton block-sparse kernel at inference time, achieving up to 28x speedup and 98.4% memory reduction over dense attention.