Understanding Gradient Descent Convergence in ImageNet Training
(i) It's possible for Stochastic Gradient Descent to converge faster than Batch Gradient Descent. This is because SGD uses a single data point at a time, which allows for faster updates and can escape local minima more effectively. However, SGD can be noisy and might oscillate around the optimal solution.
原文地址: https://www.cveoy.top/t/topic/pd7o 著作权归作者所有。请勿转载和采集!