Lesson 44 of 48
Forward pass from scratch
A network is a stack of matrix multiplies with a nonlinearity between them.
Written out, a network is smaller than its reputation. The nonlinearity is the entire reason depth buys anything — without it, the whole stack collapses into a single matrix.
Do this
Write a two-layer forward pass in NumPy with no framework: weights, biases, one hidden nonlinearity, softmax at the output. Feed it a batch and check the output rows sum to one.
The question that unlocks the next lesson
Remove the nonlinearity between two linear layers. What do you have?
- AA deeper network with the same capacity
- BA single linear layer — composing linear maps gives a linear map
- CA network that cannot be trained
- DA convolutional layer