Consider a recurrent network with a single layer GRU module…
Consider a recurrent network with a single layer GRU module with an input size of 5 and a hidden size of 7, processing sequences of 20 time steps. Fill in the blanks in the following table, breaking down the number of weights and biases per layer. GRU Layer Weights Biases #1