LeoMath

Foundations of linear algebra

Matrices

Coordinate representation of linear maps; why matrix multiplication is defined the way it is.

about 8 min

Start from a problem

Problem

A linear map is determined by the images of a basis. But to evaluate a map, store it, and compose two maps, we need a notation that turns these operations into mechanical arithmetic.

Matrices are that notation. The "row times column" rule for multiplication looks odd; this section shows it is the only natural choice.

Observe

In the previous experiment you dragged T(e1)T(\mathbf e_1) and T(e2)T(\mathbf e_2), and the matrix AA below changed with them. Watch how: the coordinates of T(e1)T(\mathbf e_1) are the first column, those of T(e2)T(\mathbf e_2) the second. A matrix is nothing but the two image vectors written side by side.

Interactive experimentLinear transformation
A=(1001)A=\begin{pmatrix}\textcolor{#c2452d}{1}&\textcolor{#1f7a4d}{0}\\\textcolor{#c2452d}{0}&\textcolor{#1f7a4d}{1}\end{pmatrix}
Determinant (signed area) 1
Eigenvalues 1, 1
Conjecture

Since a linear map is determined by the images of a basis, writing those images as columns of a table records the map completely. "Matrix times vector" should be "evaluate the map", and "matrix times matrix" should be "compose the maps". If so, the odd-looking definition of matrix multiplication stops being odd.

Definitions

Definition 4.1Matrix of a linear map

Let T:Rn→RmT:\mathbb R^n\to\mathbb R^m be linear and e1,…,en\mathbf e_1,\dots,\mathbf e_n the standard basis. The matrix of TT is the m×nm\times n array AA whose jj-th column is the coordinate vector of T(ej)T(\mathbf e_j): A=[ T(e1)  T(e2) ⋯ T(en) ].A=\bigl[\,T(\mathbf e_1)\ \ T(\mathbf e_2)\ \cdots\ T(\mathbf e_n)\,\bigr].

Definition 4.2Matrix times vector

For A=[aij]A=[a_{ij}] and x=(x1,…,xn)x=(x_1,\dots,x_n), AxAx is the linear combination of the columns of AA: Ax=x1A⋅1+x2A⋅2+⋯+xnA⋅n,i.e. (Ax)i=∑j=1naijxj.Ax=x_1A_{\cdot1}+x_2A_{\cdot2}+\cdots+x_nA_{\cdot n},\qquad\text{i.e.}\ (Ax)_i=\sum_{j=1}^na_{ij}x_j.

Proposition 4.1

If AA is the matrix of TT, then T(x)=AxT(x)=Ax for all xx.

Proof

x=∑xjejx=\sum x_j\mathbf e_j, so by linearity T(x)=∑xjT(ej)=∑xjA⋅j=AxT(x)=\sum x_jT(\mathbf e_j)=\sum x_jA_{\cdot j}=Ax.

Why matrix multiplication is defined this way

Definition 4.3Matrix product

For AA of size m×nm\times n and BB of size n×pn\times p, ABAB is the m×pm\times p matrix whose jj-th column is A(B⋅j)A(B_{\cdot j}), i.e. (AB)ij=∑k=1naikbkj.(AB)_{ij}=\sum_{k=1}^na_{ik}b_{kj}.

Theorem 4.2Matrix product = composition

If S:Rp→RnS:\mathbb R^p\to\mathbb R^n has matrix BB and T:Rn→RmT:\mathbb R^n\to\mathbb R^m has matrix AA, then T∘ST\circ S has matrix ABAB.

Proof

T∘ST\circ S is linear (check both properties). Its jj-th column is (T∘S)(ej)=T(S(ej))=T(B⋅j)=A(B⋅j)(T\circ S)(\mathbf e_j)=T(S(\mathbf e_j))=T(B_{\cdot j})=A(B_{\cdot j}), the jj-th column of ABAB.

That is the whole origin of "row times column". Immediate consequences:

  • Associativity (AB)C=A(BC)(AB)C=A(BC), because composition of maps is associative. No sums need expanding.
  • Non-commutativity AB≠BAAB\ne BA: rotating then stretching differs from stretching then rotating. Try it in the experiment.
  • The identity matrix II is the identity map, AI=IA=AAI=IA=A.

The determinant

In the experiment the unit square becomes a parallelogram whose area depends on the matrix.

Definition 4.42×2 determinant

det⁡(abcd)=ad−bc.\det\begin{pmatrix}a&b\\c&d\end{pmatrix}=ad-bc.

Theorem 4.3The determinant is signed area

The signed area of the parallelogram spanned by T(e1)=(a,c)T(\mathbf e_1)=(a,c) and T(e2)=(b,d)T(\mathbf e_2)=(b,d) is ad−bcad-bc. It is negative exactly when TT reverses orientation (the turn from e1\mathbf e_1 to e2\mathbf e_2 changes from counter-clockwise to clockwise).

Proof

Let u=(a,c)u=(a,c), v=(b,d)v=(b,d), and u⊥=(−c,a)u^\perp=(-c,a) the counter-clockwise quarter-turn of uu. The area is ∣u∣|u| times the signed projection of vv onto u⊥u^\perp: ∣u∣⋅v⋅u⊥∣u⊥∣=v⋅u⊥=−bc+ad|u|\cdot\dfrac{v\cdot u^\perp}{|u^\perp|}=v\cdot u^\perp=-bc+ad. The projection is positive when vv lies on the counter-clockwise side of uu.

Theorem 4.4Multiplicativity

det⁡(AB)=det⁡A⋅det⁡B\det(AB)=\det A\cdot\det B.

Proof

Geometrically: BB turns the unit square into a parallelogram of area det⁡B\det B, and AA multiplies the area of any region by det⁡A\det A (it sends every small grid cell to the same parallelogram, and any region's area is a limit of sums of cells). So ABAB multiplies area by det⁡Adet⁡B\det A\det B; orientation signs multiply likewise.

Proposition 4.5Invertibility

AA is invertible iff det⁡A≠0\det A\ne0, and then A−1=1ad−bc(d−b−ca)A^{-1}=\dfrac1{ad-bc}\begin{pmatrix}d&-b\\-c&a\end{pmatrix}.

det⁡A=0\det A=0 is the "Collapse" preset: the plane is flattened onto a line (or a point), information is lost, nothing can be undone.

Common mistake

AB≠BAAB\ne BA. Matrix multiplication is composition: ABAB means apply BB first, then AA (acting on a vector to the right, BB touches it first). Writing the factors in the wrong order is the most common mistake.

Application

ApplicationRotation matrices

Counter-clockwise rotation by θ\theta: e1↦(cos⁡θ,sin⁡θ)\mathbf e_1\mapsto(\cos\theta,\sin\theta), e2↦(−sin⁡θ,cos⁡θ)\mathbf e_2\mapsto(-\sin\theta,\cos\theta); as columns, Rθ=(cos⁡θ−sin⁡θsin⁡θcos⁡θ),det⁡Rθ=1.R_\theta=\begin{pmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{pmatrix},\qquad\det R_\theta=1. Expanding RαRβ=Rα+βR_\alpha R_\beta=R_{\alpha+\beta} yields the angle-addition formulas: matrix multiplication turns them into the triviality "rotate by β\beta, then by α\alpha".

Remark

One sentence suffices: the columns of a matrix are where the basis vectors go. Every definition and theorem about matrices can be re-derived from that sentence and linearity. The next section, Eigenvalues, asks: is there a vector whose direction the map leaves unchanged?

Exercises

01
Compute det⁡(2134)\det\begin{pmatrix}2&1\\3&4\end{pmatrix}.
02
The matrix (in the standard basis) of the counter-clockwise rotation by 90∘90^\circ is