Source: Annotated lecture PDF. Page references use PDF page numbers.
Overview
- Gradients describe local change; directional derivatives explain steepest ascent and descent.
- Convexity and Taylor approximations provide the language for optimisation.
- Matrices represent linear transformations relative to a chosen basis.
Gradients and directional derivatives
For a differentiable scalar function ,
Each gradient component measures change along one coordinate axis. For a unit vector , the directional derivative is the rate of change per unit distance in that direction. Without normalisation, the vector’s length also scales the rate.
Short derivation: why the dot product?
Define . The chain rule gives
The annotated chain-rule calculation is on p. 9. The useful idea is to turn movement through several coordinates into a function of one step parameter, .
For and ,
Thus the unit steepest ascent/descent directions are . The comparison assumes equal-length directions.
Geometry: gradient perpendicular to a level curve
Key takeaway: along a level curve, stays constant, so its derivative along a tangent is zero. Consequently, a nonzero gradient is perpendicular to that tangent. The diagram and handwritten explanation preserve this geometry.
Worked example: direction versus distance
From p. 10, let . Then
This is the rate with respect to the step parameter. Per unit distance along the same direction it is . The greatest possible unit-direction increase is , attained along .
Convexity
A set is convex if every segment between two of its points stays in :
A function on a convex domain is convex if
Its graph lies below the chord joining two graph points. See the diagrams on p. 13. This becomes useful in Lecture 2: convexity makes local minima global.
Taylor approximations and the Hessian
For a small displacement from , first-order and second-order approximations are
The Hessian has entries . The gradient supplies slope; the Hessian supplies curvature. means continuous first derivatives; means continuous second derivatives. These are local approximations, not exact descriptions of an arbitrary function far from .
Source: Taylor approximations, p. 14 and Hessian, p. 15.
Bases, transformations and similarity
A basis lets us express each vector uniquely as . Its coordinate column depends on the basis and its ordering.
A transformation is linear when . A matrix describes it once bases are chosen; matrix–vector multiplication applies it.
If contains the new basis vectors in old coordinates, then
The similarity formula represents the same transformation in different coordinates: convert to old coordinates, apply the transformation, then convert back. must be invertible. See p. 22.
Geometry: moving a point versus changing coordinates
Key takeaway: can describe moving a point with fixed axes. The same numerical coordinate change can instead arise by keeping the point fixed and shrinking the coordinate units to and . Keep the geometric object separate from its coordinates; p. 26 develops the basis interpretation.
Important results
- Directional derivative: ; use unit vectors when comparing slopes per unit distance.
- Negative gradient gives Euclidean steepest descent when the gradient is nonzero.
- Taylor’s linear term predicts change; its quadratic term captures curvature.
- A matrix representation depends on the basis, while the underlying transformation need not change.