Gradient Descent: How It Works, Types & Learning Rate

Jump to

Key Summary

Gradient descent is one of the fundamental optimization algorithms used in machine learning and deep learning. It helps models learn by adjusting their parameters to reduce the difference between predicted and actual results.

When a machine learning model is trained, its parameters need to be continuously updated so that the model produces better predictions. Gradient descent provides a systematic way to determine how those parameters should change.

The algorithm is used in linear regression, logistic regression, neural networks, deep learning models, and many other machine learning applications.

Understanding gradient descent is therefore essential for anyone learning machine learning or artificial intelligence because it explains one of the core processes through which models learn from data.

What is gradient descent?

Gradient descent is an optimization algorithm that minimizes a mathematical function by repeatedly adjusting its parameters in the direction that reduces the function’s value.

In machine learning, the function being minimized is usually a loss function. Suppose a model predicts house prices. If its predictions are significantly different from the actual prices, the model has a high loss. During training, gradient descent calculates how changes to the model parameters would affect that loss and updates the parameters accordingly.

A simplified implementation for a single parameter can be written in Python:

x = 5
y = 20

weight = 0.0
learning_rate = 0.01

for _ in range(100):
    prediction = weight * x
    error = prediction – y

    gradient = 2 * x * error

    weight -= learning_rate * gradient

print(weight)

The important operation is:
weight -= learning_rate * gradient

The gradient indicates how the loss changes with respect to the parameter. The learning rate determines how large the update should be.

1. Batch gradient descent

Batch gradient descent calculates the gradient using the entire training dataset before updating the parameters. For a dataset containing thousands of examples, the model calculates the loss across all examples and then performs one parameter update. A basic implementation can look like this:

import numpy as np

X = np.array([1, 2, 3, 4, 5])
y = np.array([2, 4, 6, 8, 10])

weight = 0.0
learning_rate = 0.01

for epoch in range(1000):
    predictions = weight * X

    errors = predictions – y

    gradient = (2 / len(X)) * np.sum(
        X * errors
    )

    weight -= learning_rate * gradient

print(weight)

The entire dataset is used to calculate the gradient during every iteration.

2. Stochastic gradient descent

Stochastic gradient descent updates the model after processing one training example.

import numpy as np

X = np.array([1, 2, 3, 4, 5])
y = np.array([2, 4, 6, 8, 10])

weight = 0.0
learning_rate = 0.01

for epoch in range(100):
    for x, target in zip(X, y):
        prediction = weight * x
        error = prediction – target

        gradient = 2 * x * error

        weight -= learning_rate * gradient

print(weight)

Because the update happens after each example, stochastic gradient descent can make frequent parameter updates.

3. Mini-batch gradient descent

Mini-batch gradient descent combines characteristics of batch and stochastic approaches. Instead of using the entire dataset or one example, it processes a small group of examples.

For example:

import numpy as np

X = np.array([1, 2, 3, 4, 5, 6, 7, 8])
y = np.array([2, 4, 6, 8, 10, 12, 14, 16])

weight = 0.0
learning_rate = 0.01
batch_size = 2

for epoch in range(100):
    for start in range(0, len(X), batch_size):

        X_batch = X[start:start + batch_size]
        y_batch = y[start:start + batch_size]

        predictions = weight * X_batch
        errors = predictions – y_batch

        gradient = (
            2 / len(X_batch)
        ) * np.sum(
            X_batch * errors
        )

        weight -= learning_rate * gradient

print(weight)

Mini-batches are particularly useful for deep learning because they work efficiently with modern GPUs and allow frequent parameter updates without requiring an update after every individual example.

What are the types of gradient descent?

The main types of gradient descent differ primarily in how much training data is used to calculate each parameter update.

1. Batch gradient descent

Batch gradient descent uses the complete dataset for every update. Its major advantage is that each update is based on the overall dataset rather than a single example. However, processing the entire dataset for every update can become computationally expensive when datasets are large.

2. Stochastic gradient descent

Stochastic gradient descent uses one training example for each parameter update. This makes updates much more frequent and can allow the model to start learning quickly. However, the updates can be noisy because each update is based on only one example. A simple implementation using a neural network framework is:

import torch
import torch.nn as nn

model = nn.Linear(1, 1)

optimizer = torch.optim.SGD(
    model.parameters(),
    lr=0.01
)

loss_function = nn.MSELoss()

for x, y in zip(
    torch.tensor([[1.0], [2.0], [3.0]]),
    torch.tensor([[2.0], [4.0], [6.0]])
):
    optimizer.zero_grad()

    prediction = model(x)

    loss = loss_function(
        prediction,
        y
    )

    loss.backward()
    optimizer.step()

Here, the model parameters are updated after each individual example.

3. Mini-batch gradient descent

Mini-batch gradient descent processes a selected number of examples at a time.

PyTorch’s DataLoader makes this approach straightforward:

from torch.utils.data import DataLoader, TensorDataset

X = torch.tensor([
    [1.0],
    [2.0],
    [3.0],
    [4.0],
    [5.0]
])

y = torch.tensor([
    [2.0],
    [4.0],
    [6.0],
    [8.0],
    [10.0]
])

dataset = TensorDataset(X, y)

loader = DataLoader(
    dataset,
    batch_size=2,
    shuffle=True
)

Training can then be performed batch by batch:

model = nn.Linear(1, 1)

optimizer = torch.optim.SGD(
    model.parameters(),
    lr=0.01
)

loss_function = nn.MSELoss()

for epoch in range(100):

    for X_batch, y_batch in loader:

        optimizer.zero_grad()

        prediction = model(X_batch)

        loss = loss_function(
            prediction,
            y_batch
        )

        loss.backward()
        optimizer.step()

This approach is widely used in neural network training.

Why is Gradient Descent so Important in Machine Learning?

Machine learning models typically contain parameters that need to be learned from data. A neural network, for example, can contain thousands, millions, or billions of parameters. Manually determining the correct value for each parameter is impossible for a model of this scale. Gradient descent provides a systematic optimization process. Consider a simple linear model:

prediction = weight * x + bias

The model needs to learn both weight and bias.

A loss function measures how different the prediction is from the target:

loss = (prediction – target) ** 2

Gradient descent calculates how changes to weight and bias affect the loss. In a neural network, the same basic idea is applied to many parameters simultaneously using backpropagation.

For example:

loss.backward()

calculates gradients for the parameters involved in the computation.

The optimizer then uses those gradients:

optimizer.step()

to update the model.

This combination of forward propagation, loss calculation, backpropagation, and parameter updates forms a central part of neural network training.

How to Implement Gradient Descent

Gradient descent can be implemented from scratch before using machine learning libraries. Consider a simple linear regression problem:

import numpy as np

X = np.array([1, 2, 3, 4, 5], dtype=float)
y = np.array([3, 5, 7, 9, 11], dtype=float)

weight = 0.0
bias = 0.0

learning_rate = 0.01
epochs = 1000

The model prediction is:

predictions = weight * X + bias

The loss can be calculated using mean squared error:

errors = predictions – y

loss = np.mean(
    errors ** 2
)

The gradients for the weight and bias can then be calculated:

dw = np.mean(
    2 * X * errors
)

db = np.mean(
    2 * errors
)

The parameters are updated:

weight -= learning_rate * dw
bias -= learning_rate * db

Putting everything together:
import numpy as np

X = np.array(
    [1, 2, 3, 4, 5],
    dtype=float
)

y = np.array(
    [3, 5, 7, 9, 11],
    dtype=float
)

weight = 0.0
bias = 0.0

learning_rate = 0.01
epochs = 1000

for epoch in range(epochs):

    predictions = (
        weight * X + bias
    )

    errors = predictions – y

    loss = np.mean(
        errors ** 2
    )

    dw = np.mean(
        2 * X * errors
    )

    db = np.mean(
        2 * errors
    )

    weight -= learning_rate * dw
    bias -= learning_rate * db

    if epoch % 100 == 0:
        print(
            f”Epoch: {epoch}, “
            f”Loss: {loss:.4f}”
        )

print(“Weight:”, weight)
print(“Bias:”, bias)

This example demonstrates the fundamental process of gradient descent without relying on an optimization library.

Choosing the learning rate

The learning rate determines how much the parameters change during each update.

For example:

learning_rate = 0.001

makes relatively small updates.

A larger value:

learning_rate = 0.1

results in much larger updates.

If the learning rate is too small, training can take a very long time. If it is too large, the optimization process may overshoot useful parameter values and fail to converge. You can experiment with different learning rates:

learning_rates = [
    0.0001,
    0.001,
    0.01,
    0.1
]

for lr in learning_rates:
    print(
        f”Testing learning rate: {lr}”
    )

In real projects, the appropriate learning rate depends on the model, dataset, optimizer, batch size, and training objective.

What are the Challenges with gradient descent?

Gradient descent is powerful, but optimization is not always straightforward.

1. Local minima and saddle points

A loss function can contain multiple regions where the gradient becomes very small. A local minimum is a point where the loss is lower than nearby points but may not represent the lowest possible value across the entire function. Saddle points are different. They can have gradients that are close to zero while the point is neither a local maximum nor a local minimum. Modern neural networks often operate in extremely high-dimensional parameter spaces, making optimization considerably more complex than a simple two-dimensional curve.

2. Vanishing and Exploding Gradients

During backpropagation, gradients are propagated through multiple layers. If gradients become extremely small, earlier layers may receive very little information about how their parameters should change. This is known as the vanishing gradient problem. If gradients become extremely large, parameter updates can become unstable. This is known as exploding gradients. Gradient clipping can help control extremely large gradients:

loss.backward()

torch.nn.utils.clip_grad_norm_(
    model.parameters(),
    max_norm=1.0
)

optimizer.step()

This limits the magnitude of gradients before the optimizer updates the parameters.

3. Choosing an unsuitable learning rate

Learning rate selection is another important challenge. A very small learning rate can make training unnecessarily slow. A very large learning rate can make optimization unstable.

Learning-rate scheduling can help:

optimizer = torch.optim.SGD(
    model.parameters(),
    lr=0.01
)

scheduler = torch.optim.lr_scheduler.StepLR(
    optimizer,
    step_size=10,
    gamma=0.5
)

After each training epoch, the scheduler can update the learning rate:

for epoch in range(50):

    for X_batch, y_batch in loader:
        optimizer.zero_grad()

        prediction = model(X_batch)

        loss = loss_function(
            prediction,
            y_batch
        )

        loss.backward()
        optimizer.step()

    scheduler.step(

4. Optimization can be computationally expensive

Large datasets and neural networks can require significant computing resources. This is why modern machine learning systems often combine gradient descent with optimizers such as Adam, RMSprop, or momentum-based SGD.

For example:

optimizer = torch.optim.Adam(
    model.parameters(),
    lr=0.001
)

Adam adapts parameter updates using information from previous gradients and is commonly used when training neural networks.

Which Fields Use Gradient Descent?

Gradient descent is used across many areas of machine learning and artificial intelligence.

1. Natural Language Processing

NLP models use gradient-based optimization to learn language representations and patterns.

Applications include:

  • Text classification
  • Machine translation
  • Sentiment analysis
  • Question answering
  • Text generation
  • Large language models

2. Computer Vision

Neural networks used for image classification, object detection, and image generation are commonly trained using gradient-based optimization.

For example, a convolutional neural network can be trained using PyTorch:

import torch.nn as nn

model = nn.Sequential(
    nn.Conv2d(
        3,
        32,
        kernel_size=3
    ),
    nn.ReLU(),
    nn.Flatten(),
    nn.Linear(
        32 * 222 * 222,
        10
    )
)

The model parameters can then be optimized using gradient descent:

optimizer = torch.optim.Adam(
    model.parameters(),
    lr=0.001
)

3. Speech Recognition

Speech recognition models learn relationships between audio features and language representations through gradient-based training.

4. Recommendation Systems

Recommendation models can use gradient-based optimization to learn user and item representations from interaction data.

5. Reinforcement Learning

Many reinforcement learning algorithms use gradient-based optimization to train policies or value functions.

For example, policy-gradient methods directly optimize a parameterized policy using gradient information.

6. Generative AI

Modern generative models rely heavily on gradient-based optimization during training.

Large language models, diffusion models, and other neural generative systems are trained by repeatedly calculating losses and updating their parameters.

Conclusion

Gradient descent is a fundamental optimization technique that enables machine learning models to learn from data. By calculating gradients and adjusting parameters in the direction that reduces the loss, it provides a practical method for optimizing models with many trainable parameters.

The three primary approaches are batch gradient descent, stochastic gradient descent, and mini-batch gradient descent. Each approach uses a different amount of training data for parameter updates, with mini-batch training being particularly common in modern deep learning.

Understanding gradient descent also provides a foundation for learning more advanced optimization techniques such as momentum, RMSprop, Adam, and learning-rate scheduling. These methods build upon the same fundamental principle while addressing some of the challenges encountered during model training.

For anyone learning machine learning or artificial intelligence, gradient descent is an essential concept because it connects mathematical optimization with the practical process of training predictive models.

Frequently Asked Questions (FAQs)

1. What is gradient descent in machine learning?

Gradient descent is an optimization algorithm used to minimize a machine learning model’s loss function. It repeatedly adjusts the model’s parameters based on the gradients of the loss with respect to those parameters.

2. How does gradient descent work?

Gradient descent calculates the gradient of a loss function with respect to the model’s parameters. The parameters are then updated in the direction that reduces the loss. This process is repeated during training until the model reaches an appropriate solution or the training process is stopped.

3. What are the different types of gradient descent?

The three main types are batch gradient descent, stochastic gradient descent, and mini-batch gradient descent. Batch gradient descent uses the entire dataset for each update, stochastic gradient descent uses one example, and mini-batch gradient descent uses a small group of examples.

4. What is the learning rate in gradient descent?

The learning rate determines the size of the parameter updates made during optimization. A small learning rate results in smaller updates, while a larger learning rate produces larger updates. Selecting an appropriate learning rate is important for stable and efficient model training.

5. What is the difference between batch gradient descent and stochastic gradient descent?

Batch gradient descent calculates a parameter update using the complete training dataset, while stochastic gradient descent calculates an update using one training example at a time. Batch gradient descent generally produces more stable updates, while stochastic gradient descent makes more frequent and potentially noisier updates.

Leave a Comment

Your email address will not be published. Required fields are marked *

You may also like

Vector Embeddings

Vector Embeddings: What They Are, How They Work & Uses

Key Summary A vector embedding is a numerical representation of data that captures meaningful relationships between pieces of information. Instead of representing a word, sentence, image, or other data as

Attention Mechanism

Attention Mechanism: How It Works, Types & Applications

Learn how the attention mechanism works, including queries, keys, values, attention weights, and major types such as additive, dot product, and scaled dot product attention. Explore its role in Transformers, NLP, AI, and modern deep learning.

Transformer architecture

Transformer Architecture: Components, Working & Applications

Learn how Transformer architecture works, from self-attention and multi-head attention to embeddings, positional encoding, and encoder-decoder workflows. Explore its applications, limitations, modern variants, benchmarks, and role in today’s AI systems.

Categories
Interested in working with AI, Artificial Intelligence ?

These roles are hiring now.

Loading jobs...
Scroll to Top