Training Integrable Parameterizations of Deep Neural Networks in the Infinite-Width Limit

Hajjar, Karl; Chizat, Lénaïc; Giraud, Christophe

doi:10.48550/arxiv.2110.15596

Cited by 1 publication

(1 citation statement)

References 12 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Large learning rate in the initial steps can impact the conditioning of loss surface [JSF + 20, CKL + 21] and potentially improve the generalization performance [LWM19, LBD + 20]. Under structural assumptions on the data, it has been proved that one gradient step with sufficiently large learning rate can drastically decrease the training loss [CLB21], extract task-relevant features [DM20,FCB22], or escape the trivial stationary point at initialization [HCG21]. While these works also highlight the benefit of one feature learning step 2 , to our knowledge this advantage has not been precisely characterized in the proportional regime (where the performance of RF models has been extensively studied).…”

Section: Related Workmentioning

confidence: 99%

High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

Ba¹,

Erdogdu²,

Suzuki³

et al. 2022

Preprint

View full text Add to dashboard Cite

We study the first gradient descent step on the first-layer parameters W in a two-layer neural network:, where W ∈ R d×N , a ∈ R N are randomly initialized, and the training objective is the empirical MSE loss: 1 n n i=1 (f (xi) − yi) 2 . In the proportional asymptotic limit where n, d, N → ∞ at the same rate, and an idealized student-teacher setting, we show that the first gradient update contains a rank-1 "spike", which results in an alignment between the first-layer weights and the linear component of the teacher model f * . To characterize the impact of this alignment, we compute the prediction risk of ridge regression on the conjugate kernel after one gradient step on W with learning rate η, when f * is a single-index model. We consider two scalings of the first step learning rate η. For small η, we establish a Gaussian equivalence property for the trained feature map, and prove that the learned kernel improves upon the initial random features model, but cannot defeat the best linear model on the input. Whereas for sufficiently large η, we prove that for certain f * , the same ridge estimator on trained features can go beyond this "linear regime" and outperform a wide range of random features and rotationally invariant kernels. Our results demonstrate that even one gradient step can lead to a considerable advantage over random features, and highlight the role of learning rate scaling in the initial phase of training.

show abstract

Section: Related Workmentioning

confidence: 99%