Weights and Biases in Neural Networks

|
7 min read
|
15 views
Weights and Biases in Neural Networks

The two types of numbers learned by the neural network are the weights and biases. All other components of the neural network, including layers, activation function, and optimizers, serve the purpose of modifying these two parameters to make predictions accurate.

Weights and Biases in Neural Networks

This tutorial covers the purpose of weights and biases, their changes during training, and the actual form of weights and biases in code.

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

What Are Weights in a Neural Network?

Weight is a learned numerical value which determines how much an input value contributes to the output of the neuron. Each neuron has a weight corresponding to each input it receives. Weight close to zero indicates very little impact on the outcome, whereas, large positive or negative weight, indicates a considerable influence.

Each weight is first multiplied by the corresponding input value and then added to other values. When a neuron has three inputs (x₁, x₂, x₃), it also has three different weights (w₁, w₂, w₃).

The optimization algorithm changes those values during the training procedure. Gradient descent is the most used algorithm for changing weights, and its application to multi-layer neural networks training was brought into popularity by the backpropagation algorithm developed by Rumelhart, Hinton, and Williams (1986).

What Is a Bias in a Neural Network?

The bias is an additive constant that has been acquired by the neuron along with the output from its weighted sum, and it is independent of any input data point.

Bias makes it possible for the neuron to have an output regardless of whether all of the input values are zero. The lack of bias term makes the neuron’s output zero despite any change in the weightings regardless of how the weights are trained. Positive bias increases the output while negative bias decreases the output.

Biases are initialized as 0 while the weights are initialized randomly, usually small values (Goodfellow, Bengio & Courville, Deep Learning, MIT Press, 2016). Just like the training of the weights, the bias is also updated repeatedly through the gradient descent algorithm.

Can a neural network work without a bias term? Yes, but flexibility is lost. If there is no bias in the neuron, then the neuron will produce an output value of zero only if the input value to it is zero.

Weights vs. Biases: What’s the Difference?

Weights and biases both change during training, but they control different things. The table below shows how they compare.

AttributeWeightBias
What it doesScales one input’s contributionShifts the neuron’s total output
Depends on input?Yes — multiplied by input valueNo — added regardless of input
Typical starting valueSmall random number (e.g., 0.01–0.1 range)0
Effect if set to zeroThat input is ignored entirelyNeuron output depends only on weighted inputs
Geometric effectChanges the slope of the output functionMoves the output function up or down
Number per neuronOne per input connectionOne per neuron

The Weighted Sum Formula, Step by Step

The pre-activation raw output of a neuron is the sum of all inputs multiplied by their weights, plus the bias. Mathematically, this can be represented as:

z=(w1x1+w2x2+w3x3+...+wnxn)+bz = (w₁x₁ + w₂x₂ + w₃x₃ + … + wₙxₙ) + b

This can be further simplified to z = Σ(wᵢxᵢ) + b for n inputs.

Here is a complete worked example with real numbers.

A neuron receives three inputs: x₁ = 2, x₂ = 0.5, and x₃ = -1. Its trained weights are w₁ = 0.8, w₂ = -0.3, and w₃ = 1.2. Its bias is b = 0.1.

The calculation proceeds in four steps:

  1. Multiply each input by its weight: 2 × 0.8 = 1.6; 0.5 × -0.3 = -0.15; -1 × 1.2 = -1.2.
  2. Add the three products together: 1.6 + (-0.15) + (-1.2) = 0.25.
  3. Add the bias: 0.25 + 0.1 = 0.35.
  4. Pass 0.35 through an activation function. Using ReLU, which outputs the input directly if it’s positive and 0 otherwise, the result stays 0.35.

The neuron’s final output is 0.35.If we change the value of any one weight or the bias, then the output will also change accordingly. This property makes them learnable.

This example illustrates the importance of the bias term. If all the inputs in the previous example were 0, then the output of the neuron would be 0 regardless of the values of the weights. However, in this case, the bias is 0.1, and hence the neuron produces 0.1 as output.

Why Random Initialization Matters

Weights are initialized with small random values rather than zeros or identical values because identical values lead to learning the same information by neurons. If all the neurons have identical weights in a layer, then all the neurons receive identical gradients during training and stay identical to each other. This phenomenon is called the symmetry problem.

There are two solutions to this problem. Xavier or Glorot initialization scales the initial values of weights according to the number of inputs and outputs in a layer, as proposed by Glorot and Bengio (2010) to stabilize the variance of signal between layers. On the other hand, He initialization, which was proposed by He et al. (2015), uses a different scaling depending on whether there is ReLU activation in a network or not. 

This problem does not exist for biases because each bias impacts one neuron and not inputs shared with others. This is why bias is usually set to 0 in most frameworks.

How Weights and Biases Update During Training

The weights and biases are updated through two processes that are related to each other: forward propagation and backpropagation. Forward propagation gives the output of the network; backpropagation computes the contribution of each weight and bias in the error produced.

Weights and Biases Update During Training

Forward propagation comes first. The data is fed into the network, and then every neuron calculates the weighted sum, adds the bias term, and passes it through an activation function. This is done layer by layer until an output is achieved.

This output is then compared to the actual result through a loss function. The loss function used can either be the mean square error or cross-entropy loss function, depending on whether it is numerical or classification-based prediction respectively.

Then comes backpropagation, where the algorithm determines the gradients of the loss function with respect to all the weights and biases of the neural network starting from the output layer and moving to the input layer. After that, gradient descent updates all the weights and biases according to the corresponding gradients multiplied by the learning rate.

This process of moving forward through the network, calculating the loss function, backpropagating, and updating is repeated over and over again for many iterations.

Common Mistakes and Misconceptions About Weights and Biases

Several assumptions about weights and biases lead to real training problems. The list below states each one directly.

  1. “Bias does not matter much” is false. The removal of bias terms limits the types of functions that can be represented by the network and might stop the model from fitting data that does not cross the origin.
  2. “More weights always improve a model” is false. Increasing weights increases the ability of the model to memorize the training data and hence increases the probability of overfitting, which is the situation where a model works well for the training data but poorly for new data.
  3. “Weights must be manually initialized for a particular task” is false for most architectures. Manually initializing the weights removes the statistical advantage gained by using the Xavier or He initialization methods and makes training slower and more unstable.
  4. “A bias value close to zero implies that bias is unnecessary” is false. A near-zero bias is a valid trained bias value.
Professional Certificate

Agentic AI Course

Go beyond prompting. Learn to design, build and deploy autonomous AI agents with LangChain, CrewAI, AutoGen, LangGraph and RAG — from single-agent workflows to production multi-agent systems.
4.9 (7,352 ratings)  •  Beginner to Advanced level
Class Starts on 3 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time:6 month(s) + Lifetime Access

Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving

Weights and Biases in Real Code

In PyTorch, a linear layer stores its weights and biases as two separate tensors, accessible directly after creating the layer.

import torch.nn as nn

layer = nn.Linear(in_features=3, out_features=1)

print(layer.weight)   # shape [1, 3] — one weight per input

print(layer.bias)     # shape [1] — one bias for this neuron

Running this code prints two tensors filled with small random numbers, which reflect PyTorch’s default initialization scheme. These are the exact numbers that gradient descent updates during training. Inspecting layer.weight and layer.bias before and after a training loop shows the values changing directly, which makes the abstract formula from the previous section observable in a real model.

Frequently Asked Questions

Q1. Can a neural network function without bias terms?

Ans. Yes, but then you lose the capability to model functions that do not equal 0 when all the inputs are 0.

Q2. What happens if all weights in a layer start at the same value?

Ans. The gradient updates would be identical for every neuron in that layer, keeping them identical throughout training and making it impossible to learn different features.

Q3. How many weights and biases does a typical neural network layer have?

Ans. A fully connected layer with m inputs and n output neurons will have m×n weights and n biases; 100 inputs and 10 neurons will result in 1,000 weights and 10 biases.

Q4. Are weights and biases the same as parameters?

Ans. Yes. In usual terminology, parameters mean the whole set of weights and biases that a model learns during the training process.

Q5. What is a typical starting value for a bias term?

Ans. In most frameworks, the bias is initialized as 0, while the weights are initialized as small random values with Xavier or He initialization.

Gyansetu offers top professional training certification courses designed to enhance your skills and advance your career, providing industry-relevant knowledge and practical expertise.