Are Flat Minima an Illusion?

preprint OA: closed
View at publisher

Abstract

Flat minima are widely believed to generalise better than sharp ones. I present a theoretical and empirical case that this belief is wrong. On the theoretical side, I prove that weakness, the measure-theoretic volume of compatible completions in the learner's embodied language, is reparameterisation-invariant, minimax-optimal, and tracked by the PAC-Bayes bound. Sharpness is none of these things. On the empirical side, I show that the large-batch generalisation advantage vanishes as a function of training set size. On MNIST, large-batch networks generalise better than small-batch networks when training data is scarce, and the advantage becomes negligible when data is abundant. A quantity that correlates, anticorrelates, or does not correlate depending on the data regime is not a cause. It is a confound. I confirm reparameterisation non-invariance of the Hessian trace on trained networks (up to 99x change with zero change in function or generalisation). I construct a formal vocabulary for frozen-partition ReLU networks and measure the pair proxy via linear feasibility on 100 networks. Weakness outperforms sharpness at model selection (ρ = +0.374, p = 0.00012 vs ρ = -0.226, p = 0.024).

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00