Requirements
- Write a valid-mode 1-D convolution routine,
convolve_valid(input, kernel, bias), which produces results only where the kernel fits entirely within the input.
- The routine receives an input array, a kernel array, and one scalar bias value.
- At every legal starting position, calculate the input window's dot product with the kernel and then add the bias.
- Be prepared to discuss CPU or hardware optimisation approaches, along with a multithreaded implementation, for these cases:
- an input of roughly 1 million elements with a kernel of length 3;
- an input of roughly 1 million elements with a kernel also near 1 million elements long;
- no more than 100 available threads.
Examples
input = [2, 1, 4, 3]
kernel = [1.5, -1]
bias = 1
output[0] = (2 * 1.5) + (1 * -1) + 1 = 3
output[1] = (1 * 1.5) + (4 * -1) + 1 = -1.5
output[2] = (4 * 1.5) + (3 * -1) + 1 = 4
Each result comes from one full input window of the same length as the kernel, followed by addition of the bias.
Notes
- Provided that the kernel length does not exceed the input length, the standard result size is
len(input) - len(kernel) + 1.