What is Gradient Descent Algorithm?
Gradient Descent is an optimization algorithm used to find the minimum of a function. In machine learning and deep learning, it is the fundamental engine used to train models by minimizing their loss function (which measures the error between the model's predictions and actual reality). The Intuition: Lost on a Mountain Imagine you are standing near the top of a foggy mountain, and your goal is to find your way down to the lowest point of the valley (the minimum). Because of the heavy fog, you cannot see the path or the bottom of the valley. To get down, you have to use your feet to sense the slope of the ground immediately around you: You feel which direction slopes downward the steepest. You take a step in that downward direction. You repeat this process step-by-step until the ground flattens out, indicating you have reached the bottom. In this analogy, the mountain represents the loss function, the steepness of the slope is the gradient, and your step size is the learning rate. A 3D Loss Surface with local and global minima. Kaynak: Towards Data Science How It Works Mathematically To train a machine learning model, we want to adjust its parameters (such as weights w and biases b) to make the error as close to zero as possible. The update rule for a parameter w is defined by: w new =w old −α⋅ ∂w ∂L Let's break down the components of this formula: w (The Parameter): The weight we want to adjust to improve our model. ∂w ∂L (The Gradient): The derivative of the Loss function L with respect to w. It tells us the slope of the function at our current position. If the slope is positive, subtracting it moves us backward (to the left). If the slope is negative, subtracting it moves us forward (to the right). α (The Learning Rate): A small positive number (usually between 0.1 and 0.0001) that defines the size of the steps we take. The Crucial Role of Learning Rate (α) Choosing the right step size (learning rate) is one of the most critical decisions when training a model: Too Small: The steps are tiny. The algorithm will take a massive amount of time to find the minimum, consuming high computational power. Too Large: You take giant leaps. The algorithm might overshoot the lowest point, bounce back and forth, and fail to settle (diverge). Three Main Variants of Gradient Descent Depending on how much data we use to calculate the gradient at each step, there are three types: Variant How it works Pros Cons Batch Gradient Descent Calculates the gradient using the entire dataset for every step. Stable path to the minimum; easy to converge. Incredibly slow and memory-intensive for large datasets. Stochastic Gradient Descent (SGD) Calculates the gradient using just one random sample at a time. Extremely fast; can escape local minima due to its noisy path. Highly erratic path; never truly settles at the exact minimum. Mini-batch Gradient Descent Splits the data into small groups (batches) (e.g., 32, 64, or 128 samples) to update parameters. Best of both worlds: faster than Batch, more stable than SGD. Requires tuning the batch size. The Ultimate Challenge: Local Minima In complex machine learning models (like deep neural networks), the loss landscape is not a simple smooth bowl; it looks like a rugged mountain range with many dips and valleys. As shown in the 3D surface image above, there can be multiple valleys: Local Minima: A dip that looks like the bottom, but is actually just a crater on the side of the mountain. The gradient becomes zero here, which can trick the basic algorithm into stopping prematurely. Global Minimum: The absolute lowest point on the entire landscape (the ultimate goal). Modern optimizers like Adam, RMSProp, and Momentum build on top of standard Gradient Descent by adding "momentum" (like a rolling ball gathering speed) to help the algorithm roll right through small craters and find the true global minimum.