Algorithm · Amazon · Medium
Given a collection of $$n$$ scalar observations $$(x_i, y_i)$$, where $$x_i$$ is a real-valued input feature and $$y_i$$ is a real-valued target, fit the linear model $$ \hat{y}_i = a x_i + b $$ so that the mean squared error computed over the full dataset is as small as possible. Complete the following parts: Write the exact scalar loss $$L(a,b)$$ used for training. Derive the two partial derivatives $$\frac{\partial L}{\partial a}$$ and $$\frac{\partial L}{\partial b}$$…
Checking your access…