LeoMath

Foundations of linear algebra

Vectors

Axioms of a vector space; bases and coordinates.

about 7 min

Start from a problem

Problem

Arrows in the plane, polynomials, matrices, continuous functions: their addition and scaling obey exactly the same rules. Proving the "same" theorem separately for each kind of object is wasted work.

We need a definition that depends only on the rules of the operations, not on what the objects "are". Then one proof holds everywhere.

Observe

Arrows in the plane can be added (tip to tail) and stretched (scalar multiplication). So can polynomials. So can 2×22\times2 matrices, continuous functions, sequences…

These objects look completely different, yet the rules obeyed by their addition and scaling are identical: addition commutes and associates, scaling distributes, there is a zero, there are negatives.

Conjecture

If a theorem uses only these rules, then whenever it holds for arrows it holds for polynomials, matrices and functions too. So it pays to isolate those rules, name every object satisfying them, and study only the name.

Definition

Definition 1.1Vector space

Let VV be a set with an addition u+vu+v and a real scalar multiplication cvcv such that for all u,v,w∈Vu,v,w\in V and a,b∈Ra,b\in\mathbb R:

  1. u+v=v+uu+v=v+u; 2. (u+v)+w=u+(v+w)(u+v)+w=u+(v+w);
  2. there is 0∈V0\in V with v+0=vv+0=v; 4. each vv has −v-v with v+(−v)=0v+(-v)=0;
  3. a(u+v)=au+ava(u+v)=au+av; 6. (a+b)v=av+bv(a+b)v=av+bv;
  4. (ab)v=a(bv)(ab)v=a(bv); 8. 1v=v1v=v.

Then VV is a (real) vector space and its elements are vectors.

R2\mathbb R^2, Rn\mathbb R^n, polynomials of degree at most nn, all m×nm\times n matrices, continuous functions on [0,1][0,1]: all vector spaces. The exercises contain a non-example: polynomials of degree exactly 2 are not, since x2+(−x2+x)x^2+(-x^2+x) has degree 1 and leaves the set.

Definition 1.2Linear combination, span, independence

A linear combination of v1,…,vkv_1,\dots,v_k is c1v1+⋯+ckvkc_1v_1+\cdots+c_kv_k. The set of all of them is span⁡(v1,…,vk)\operatorname{span}(v_1,\dots,v_k). If c1v1+⋯+ckvk=0c_1v_1+\cdots+c_kv_k=0 only for c1=⋯=ck=0c_1=\cdots=c_k=0, the vectors are linearly independent.

Independence means no vector is redundant: none can be built from the others.

Definition 1.3Basis and dimension

If v1,…,vnv_1,\dots,v_n are independent and span⁡(v1,…,vn)=V\operatorname{span}(v_1,\dots,v_n)=V, they form a basis of VV. The number of vectors in a basis is the dimension dim⁡V\dim V.

For dimension to make sense we must prove that any two bases have the same size.

Theorems and proofs

Theorem 1.1Uniqueness of coordinates

If v1,…,vnv_1,\dots,v_n is a basis of VV, every v∈Vv\in V can be written uniquely as v=c1v1+⋯+cnvnv=c_1v_1+\cdots+c_nv_n. The tuple (c1,…,cn)(c_1,\dots,c_n) is the coordinate vector of vv in this basis.

Proof

Existence is spanning. For uniqueness, if v=∑civi=∑diviv=\sum c_iv_i=\sum d_iv_i then ∑(ci−di)vi=0\sum(c_i-d_i)v_i=0, so ci=dic_i=d_i by independence.

This theorem is the entire justification for the word "coordinates". Once a basis is fixed, every vector in any vector space corresponds to one list of numbers, a point of Rn\mathbb R^n.

Theorem 1.2Dimension is well defined

Any two finite bases of a vector space have the same number of elements.

Proof

It suffices to show: if u1,…,umu_1,\dots,u_m are independent and v1,…,vnv_1,\dots,v_n span VV, then m≤nm\le n (swap the roles to get equality).

Exchange argument. u1∈span⁡(v1,…,vn)u_1\in\operatorname{span}(v_1,\dots,v_n); write u1=∑aiviu_1=\sum a_iv_i. Since u1≠0u_1\ne0 some ai≠0a_i\ne0, say a1a_1. Then v1v_1 is a combination of u1,v2,…,vnu_1,v_2,\dots,v_n, so these still span VV. Now expand u2u_2 in this new spanning set; some vjv_j (j≥2j\ge2) must have nonzero coefficient, otherwise u2∈span⁡(u1)u_2\in\operatorname{span}(u_1), contradicting independence. Replace that vjv_j by u2u_2. Continue. Each step trades one vv for one uu, and the vv's cannot run out before the uu's do (or some uku_k would lie in the span of earlier uu's). Hence m≤nm\le n.

Common mistake

Linear independence is not "no two are collinear". (1,0,1),(0,1,1),(1,1,2)(1,0,1),(0,1,1),(1,1,2) are pairwise non-collinear yet dependent, since the third is the sum of the first two. The definition is about the whole family having only the trivial zero combination; it must be checked all at once.

Application

ApplicationWhy pictures in ℝ² speak for every 2-dimensional space

Any two-dimensional vector space, once a basis is chosen, corresponds one-to-one with R2\mathbb R^2, with addition and scaling preserved. So every picture drawn in the plane in Linear maps is true for every two-dimensional space, whether its vectors are arrows, linear polynomials, or solutions of a second-order linear differential equation.

Remark

A vector is not "a quantity with magnitude and direction"; that is one concrete picture valid in R2,R3\mathbb R^2,\mathbb R^3. A vector is an element of a vector space, and a vector space is defined by eight axioms. Defining things by the rules their operations obey, rather than by what they "are", is the basic move of all of algebra.

Exercises

01
What is the dimension of the subspace of R3\mathbb R^3 spanned by (1,0,1),(0,1,1),(1,1,2)(1,0,1),(0,1,1),(1,1,2)?
02
Under the usual operations, which of these is not a vector space?