Spurious local minima in neural networks: a critical view
Chulhee Yun, Suvrit Sra, Ali Jadbabaie
Abstract
Chulhee Yun, Suvrit Sra, Ali Jadbabaie
Abstract
We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with slightest nonlinearity, the empirical risks have spurious local in most cases. Our results thus indicate that in general no spurious local minima is a property limited to deep linear networks, and insights obtained from linear networks may not be robust. Specifically, for ReLU(-like) networks we constructively prove that for almost all practical datasets there exist infinitely many local minima. We also present a counterexample for more general activations (sigmoid, tanh, arctan, ReLU, etc.), for which there exists a bad local minimum. Our results make the least restrictive assumptions relative to existing results on spurious local optima in neural networks. We complete our discussion by presenting a comprehensive characterization of global optimality for deep linear networks, which unifies other results on this topic.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
We investigate the loss surface of neural networks. We prove that even for one-hidden-layer networks with slightest nonlinearity, the empirical risks have spurious local in most cases. Our results thus indicate that in general no spurious local minima is a property limited to deep linear networks, and insights obtained from linear networks may not be robust. Specifically, for ReLU(-like) networks we constructively prove that for almost all practical datasets there exist infinitely many local minima. We also present a counterexample for more general activations (sigmoid, tanh, arctan, ReLU, etc.), for which there exists a bad local minimum. Our results make the least restrictive assumptions relative to existing results on spurious local optima in neural networks. We complete our discussion by presenting a comprehensive characterization of global optimality for deep linear networks, which unifies other results on this topic.
Key concepts: Maxima and minima, Spurious relationship, Artificial neural network, Counterexample, Sigmoid function, Computer science, Nonlinear system, Deep neural networks