Source: Annotated lecture PDF. Page references use PDF page numbers.

Overview

  • Gradients describe local change; directional derivatives explain steepest ascent and descent.
  • Convexity and Taylor approximations provide the language for optimisation.
  • Matrices represent linear transformations relative to a chosen basis.

Gradients and directional derivatives

For a differentiable scalar function ,

Each gradient component measures change along one coordinate axis. For a unit vector , the directional derivative is the rate of change per unit distance in that direction. Without normalisation, the vector’s length also scales the rate.

Short derivation: why the dot product?

Define . The chain rule gives

The annotated chain-rule calculation is on p. 9. The useful idea is to turn movement through several coordinates into a function of one step parameter, .

For and ,

Thus the unit steepest ascent/descent directions are . The comparison assumes equal-length directions.

Geometry: gradient perpendicular to a level curve

Key takeaway: along a level curve, stays constant, so its derivative along a tangent is zero. Consequently, a nonzero gradient is perpendicular to that tangent. The diagram and handwritten explanation preserve this geometry.

Worked example: direction versus distance

From p. 10, let . Then

This is the rate with respect to the step parameter. Per unit distance along the same direction it is . The greatest possible unit-direction increase is , attained along .

Convexity

A set is convex if every segment between two of its points stays in :

A function on a convex domain is convex if

Its graph lies below the chord joining two graph points. See the diagrams on p. 13. This becomes useful in Lecture 2: convexity makes local minima global.

Taylor approximations and the Hessian

For a small displacement from , first-order and second-order approximations are

The Hessian has entries . The gradient supplies slope; the Hessian supplies curvature. means continuous first derivatives; means continuous second derivatives. These are local approximations, not exact descriptions of an arbitrary function far from .

Source: Taylor approximations, p. 14 and Hessian, p. 15.

Bases, transformations and similarity

A basis lets us express each vector uniquely as . Its coordinate column depends on the basis and its ordering.

A transformation is linear when . A matrix describes it once bases are chosen; matrix–vector multiplication applies it.

If contains the new basis vectors in old coordinates, then

The similarity formula represents the same transformation in different coordinates: convert to old coordinates, apply the transformation, then convert back. must be invertible. See p. 22.

Geometry: moving a point versus changing coordinates

Key takeaway: can describe moving a point with fixed axes. The same numerical coordinate change can instead arise by keeping the point fixed and shrinking the coordinate units to and . Keep the geometric object separate from its coordinates; p. 26 develops the basis interpretation.

Important results

  • Directional derivative: ; use unit vectors when comparing slopes per unit distance.
  • Negative gradient gives Euclidean steepest descent when the gradient is nonzero.
  • Taylor’s linear term predicts change; its quadratic term captures curvature.
  • A matrix representation depends on the basis, while the underlying transformation need not change.