arXiv · 1605.07154
Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations
Abstract
We investigate the parameter-space geometry of recurrent neural networks (RNNs), and develop an adaptation of path-SGD optimization method, attuned to this geometry, that can learn plain RNNs with ReLU activations. On several datasets that require capturing long-term dependency structure, we show that path-SGD can significantly improve trainability of ReLU RNNs compared to RNNs trained with SGD, even with various recently suggested initialization schemes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Behnam Neyshabur, Yuhuai Wu, Ruslan Salakhutdinov, Nathan Srebro. 2016-05-23. Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations. https://arxiv.org/abs/1605.07154
Cite the original work for its findings. Save a collection to share your selection of sources.