Skip to content

Chapter II: Front Matter (2)

Text size

and then, if h be so small that the terms after the second may be
neglected, [f]([gamma]) + h[f]'([gamma]) = 0, that is, h =
{-[f]([gamma])/[f]'([gamma])}, or the new approximate value is x =
[gamma] - {[f]([gamma])/[f]'([gamma])}; and so on, as often as we
please. It will be observed that so far nothing has been assumed as to
the separation of the roots, or even as to the existence of a real
root; [gamma] has been taken as the approximate value of a root, but
no precise meaning has been attached to this expression. The question
arises, What are the conditions to be satisfied by [gamma] in order
that the process may by successive repetitions actually lead to a
certain real root of the equation; or that, [gamma] being an
approximate value of a certain real root, the new value [gamma] -
{[f]([gamma])/[f]'([gamma])} may be a more approximate value.

Referring to fig. 1, it is easy to see that if OC represent the
assumed value [gamma], then, drawing the ordinate CP to meet the curve
in P, and the tangent PC' to meet the axis in C', we shall have OC' as
the new approximate value of the root. But observe that there is here
a real root OX, and that the curve beyond X is convex to the axis;
under these conditions the point C' is nearer to X than was C; and,
starting with C' instead of C, and proceeding in like manner to draw a
new ordinate and tangent, and so on as often as we please, we
approximate continually, and that with great rapidity, to the true
value OX. But if C had been taken on the other side of X, where the
curve is concave to the axis, the new point C' might or might not be
nearer to X than was the point C; and in this case the method, if it
succeeds at all, does so by accident only, i.e. it may happen that C'
or some subsequent point comes to be a point C, such that CO is a
_proper_ approximate value of the root, and then the subsequent
approximations proceed in the same manner as if this value had been
assumed in the first instance, all the preceding work being wasted. It
thus appears that for the proper application of the method we require
_more_ than the mere separation of the roots. In order to be able to
approximate to a certain root [alpha], =OX, we require to know that,
between OX and some value ON, the curve is always convex to the axis
(analytically, between the two values, [f](x) and [f]"(x) must have
always the same sign). When this is so, the point C may be taken
anywhere on the proper side of X, and within the portion XN of the
axis; and the process is then the one already explained. The
approximation is in general a very rapid one. If we know for the
required root OX the two limits OM, ON such that from M to X the curve
is always _concave_ to the axis, while from X to N it is always convex
to the axis,--then, taking D anywhere in the portion MX and (as
before) C in the portion XN, drawing the ordinates DQ, CP, and joining
the points P, Q by a line which meets the axis in D', also
constructing the point C' by means of the tangent at P as before, we
have for the required root the new limits OD', OC'; and proceeding in
like manner with the points D', C', and so on as often as we please,
we obtain at each step two limits approximating more and more nearly
to the required root OX. The process as to the point D', translated
into analysis, is the ordinate process of interpolation. Suppose OD =
[beta], OC = [alpha], we have approximately [f]([beta] + h) =
[f]([beta]) + h{[f]([alpha]) - [f]([beta])} / ([alpha] - [beta]),
whence if the root is [beta] + h then h = - ([alpha] -
[beta])[f]([beta]) / {[f]([alpha]) - [f]([beta])}.

Returning for a moment to Horner's method, it may be remarked that the
correction h, to an approximate value [alpha], is therein found as a
quotient the same or such as the quotient [f]([alpha]) / [f]'([alpha])
which presents itself in Newton's method. The difference is that with
Horner the integer part of this quotient is taken as the presumptive
value of h, and the figure is verified at each step. With Newton the
quotient itself, developed to the proper number of decimal places, is
taken as the value of h; if too many decimals are taken, there would
be a waste of work; but the error would correct itself at the next
step. Of course the calculation should be conducted without any such
waste of work.

_Imaginary Theory_.

7. It will be recollected that the expression _number_ and the correlative epithet _numerical_ were at the outset used in a wide sense, as extending to imaginaries. This extension arises out of the theory of equations by a process analogous to that by which number, in its original most restricted sense of positive integer number, was extended to have the meaning of a real positive or negative magnitude susceptible of continuous variation.

If for a moment number is understood in its most restricted sense as meaning positive integer number, the solution of a simple equation leads to an extension; ax - b = 0 gives x = (b/a), a positive fraction, and we can in this manner represent, not accurately, but as nearly as we please, any positive magnitude whatever; so an equation ax + b = 0 gives x = -(b/a), which (approximately as before) represents any negative magnitude. We thus arrive at the extended signification of number as a continuously varying positive or negative magnitude. Such numbers may be added or subtracted, multiplied or divided one by another, and the result is always a number. Now from a quadric equation we derive, in like manner, the notion of a complex or imaginary number such as is spoken of above. The equation x^2 + 1 = 0 is not (in the foregoing sense, number = real number) satisfied by any numerical value whatever of x; but we assume that there is a number which we call i, satisfying the equation i^2 + 1 = 0, and then taking a and b any real numbers, we form an expression such as a + bi, and use the expression number in this extended sense: any two such numbers may be added or subtracted, multiplied or divided one by the other, and the result is always a number. And if we consider first a quadric equation x^2 + px + q = 0 where p and q are real numbers, and next the like equation, where p and q are any numbers whatever, it can be shown that there exists for x a numerical value which satisfies the equation; or, in other words, it can be shown that the equation has a numerical root. The like theorem, in fact, holds good for an equation of any order whatever; but suppose for a moment that this was not the case; say that there was a cubic equation x^3 + px^2 + qx + r = 0, with numerical coefficients, not satisfied by any numerical value of x, we should have to establish a new imaginary j satisfying some such equation, and should then have to consider numbers of the form a + bj, or perhaps a + bj + cj^2 (a, b, c numbers [alpha] + [beta]i of the kind heretofore considered),--first we should be thrown back on the quadric equation x^2 + px + q = 0, p and q being now numbers of the last-mentioned extended form--_non constat_ that every such equation has a numerical root--and if not, we might be led to _other_ imaginaries k, l, &c., and so on _ad infinitum_ in inextricable confusion.

But in fact a numerical equation of any order whatever has always a numerical root, and thus numbers (in the foregoing sense, number = quantity of the form [alpha] + [beta]i) form (_what real numbers do not_) a universe complete in itself, such that starting in it we are never led out of it. There may very well be, and perhaps are, numbers in a more general sense of the term (quaternions are not a case in point, as the ordinary laws of combination are not adhered to), but in order to have to do with such numbers (if any) we must start with them.

8. The capital theorem as regards numerical equations thus is, every numerical equation has a numerical root; or for shortness (the meaning being as before), every equation has a root. Of course the theorem is the reverse of self-evident, and it requires proof; but provisionally assuming it as true, we derive from it the general theory of numerical equations. As the term root was introduced in the course of an explanation, it will be convenient to give here the formal definition.

A number a such that substituted for x it makes the function x1^n -
p1x^(n - 1) ... [+-]p_n to be = 0, or say such that it satisfies the
equation [f](x) = 0, is said to be a root of the equation; that is, a
being a root, we have

a^n - p1a^(n - 1) ... [+-]p_n = 0, or say [f](a) = 0;

and it is then easily shown that x - a is a factor of the function
[f](x), viz. that we have [f](x) = (x - a)[f]1(x), where [f]1(x) is a
function x^(n - 1) - q1x^(n - 2) ... [+-]q_(n - 1) of the order n - 1,
with numerical coefficients q1, q2 ... q_(n - 1).

In general a is not a root of the equation [f]1(x) = 0, but it may be
so--i.e. [f]1(x) may contain the factor x - a; when this is so, [f](x)
will contain the factor (x - a)^2; writing then [f](x) = (x -
a)^2[f]2(x), and assuming that a is not a root of the equation [f]2(x)
= 0, x = a is then said to be a double root of the equation [f](x) =
0; and similarly [f](x) may contain the factor (x - a)^3 and no higher
power, and x = a is then a triple root; and so on.

Supposing in general that [f](x) = (x - a)^[alpha] F(x) ([alpha] being
a positive integer which may be = 1, (x - a)^[alpha] the highest power
of x - a which divides [f](x), and F(x) being of course of the order n
- [alpha]), then the equation F(x) = 0 will have a root b which will
be different from a; x - b will be a factor, in general a simple one,
but it may be a multiple one, of F(x), and [f](x) will in this case be
= (x - a)^[alpha] (x - b)^[beta] [Phi](x) ([beta] a positive integer
which may be = 1, (x-b)^[beta] the highest power of x - b in F(x) or
[f](x), and [Phi](x) being of course of the order n - [alpha] -
[beta]). The original equation [f](x) = 0 is in this case said to have
[alpha] roots each = a, [beta] roots each = b; and so on for any other
factors (x - c)^[gamma], &c.

We have thus the _theorem_--A numerical equation of the order n has in
every case n roots, viz. there exist n numbers, a, b, ... (in general
all distinct, but which may arrange themselves in any sets of equal
values), such that [f](x) = (x - a)(x - b)(x - c) ... identically.

If the equation has equal roots, these can in general be determined,
and the case is at any rate a special one which may be in the first
instance excluded from consideration. It is, therefore, in general
assumed that the equation [f](x) = 0 has all its roots unequal.

If the coefficients p1, p2, ... are all or any one or more of them
imaginary, then the equation [f](x) = 0, separating the real and
imaginary parts thereof, may be written F(x) + i[Phi](x) = 0, where
F(x), [Phi](x) are each of them a function with real coefficients; and
it thus appears that the equation [f](x) = 0, with imaginary
coefficients, has not in general any real root; supposing it to have a
real root a, this must be at once a root of each of the equations F(x)
= 0 and [Phi](x) = 0.

But an equation with real coefficients may have as well imaginary as
real roots, and we have further the _theorem_ that for any such
equation the imaginary roots enter in pairs, viz. [alpha] + [beta]i
being a root, then [alpha] - [beta]i will be also a root. It follows
that if the order be odd, there is always an odd number of real roots,
and therefore at least one real root.

9. In the case of an equation with real coefficients, the question of the existence of real roots, and of their separation, has been already considered. In the general case of an equation with imaginary (it may be real) coefficients, the like question arises as to the situation of the (real or imaginary) roots; thus, if for facility of conception we regard the constituents [alpha], [beta] of a root [alpha] + [beta]i as the co-ordinates of a point _in plano_, and accordingly represent the root by such point, then drawing in the plane any closed curve or "contour," the question is how many roots lie within such contour.

This is solved theoretically by means of a theorem of A.L. Cauchy
(1837), viz. writing in the original equation x + iy in place of x,
the function [f](x + iy) becomes = P + iQ, where P and Q are each of
them a rational and integral function (with real coefficients) of (x,
y). Imagining the point (x, y) to travel along the contour, and
considering the number of changes of sign from - to + and from + to -
of the fraction corresponding to passages of the fraction through zero
(that is, to values for which P becomes = 0, disregarding those for
which Q becomes = 0), the difference of these numbers gives the number
of roots within the contour.

It is important to remark that the demonstration does not presuppose
the existence of any root; the contour may be the infinity of the
plane (such infinity regarded as a contour, or closed curve), and in
this case it can be shown (and that very easily) that the difference
of the numbers of changes of sign is = n; that is, there are within
the infinite contour, or (what is the same thing) there are in all n
roots; thus Cauchy's theorem contains really the proof of the
fundamental theorem that a numerical equation of the nth order (not
only has a numerical root, but) has precisely n roots. It would appear
that this proof of the fundamental theorem in its most complete form
is in principle identical with the last proof of K.F. Gauss (1849) of
the theorem, in the form--A numerical equation of the nth order has
always a root.[3]

But in the case of a finite contour, the actual determination of the
difference which gives the number of real roots can be effected only
in the case of a rectangular contour, by applying to each of its sides
separately a method such as that of Sturm's theorem; and thus the
actual determination ultimately depends on a method such as that of
Sturm's theorem.

Very little has been done in regard to the calculation of the
imaginary roots of an equation by approximation; and the question is
not here considered.

10. A class of numerical equations which needs to be considered is that of the binomial equations x^n - a = 0 (a = [alpha] + [beta]i, a complex number).

The foregoing conclusions apply, viz. there are always n roots, which,
it may be shown, are all unequal. And these can be found numerically
by the extraction of the square root, and of an nth root, of _real_
numbers, and by the aid of a table of natural sines and cosines.[4]
For writing

/ [alpha] [beta] \
[alpha] + [beta]i = [root]([alpha]^2 + [beta]^2) ( ---------------------------- + ----------------------------i ),
\[root]([alpha]^2 + [beta]^2) [root]([alpha]^2 + [beta]^2) /

there is always a real angle [lambda] (positive and less than 2[pi]),
such that its cosine and sine are = [alpha] / [root]([alpha]^2 +
[beta]^2) and [beta] / [root]([alpha]^2 + [beta]^2) respectively; that
is, writing for shortness [root]([alpha]^2 + [beta]^2) = [rho], we have
[alpha] + [beta]i = [rho](cos[lambda] + i sin[lambda]), or the
equation is x^n = [rho](cos[lambda] + i sin [lambda]); hence observing
that (cos [lambda]/n + i sin [lambda]/n )^n = cos[lambda] + i
sin[lambda], a value of x is = [root n][rho] (cos [lambda]/n + i sin
[lambda]/n). The formula really gives all the roots, for instead of
[lambda] we may write [lambda] + 2s[pi], s a positive or negative
integer, and then we have

/ [lambda] + 2s[pi] [lambda] + 2s[pi] \
x = [root n][rho] ( cos ----------------- + i sin ----------------- ),
\ n n /

which has the n values obtained by giving to s the values 0, 1, 2 ...
n - 1 in succession; the roots are, it is clear, represented by points
lying at equal intervals on a circle. But it is more convenient to
proceed somewhat differently; taking one of the roots to be [theta],
so that [theta]^n = a, then assuming x = [theta]y, the equation
becomes y^n - 1 = 0, which equation, like the original equation, has
precisely n roots (one of them being of course = 1). And the original
equation x^n - a = 0 is thus reduced to the more simple equation x^n -
1 = 0; and although the theory of this equation is included in the
preceding one, yet it is proper to state it separately.

The equation x^n - 1 = 0 has its several roots expressed in the form
1, [omega], [omega]^2, ... [omega]^(n - 1), where [omega] may be taken
= cos 2[pi]/n + i sin 2[pi]/n; in fact, [omega] having this value, any
integer power [omega]^k is = cos 2[pi]k/n + i sin 2[pi]k/n, and we
thence have ([omega]^k)^n = cos 2[pi]k + i sin 2[pi]k, = 1, that is,
[omega]^k is a root of the equation. The theory will be resumed
further on.

By what precedes, we are led to the notion (a numerical) of the
radical a^(1/n) regarded as an n-valued function; any one of these
being denoted by [root n]a, then the series of values is [root n]a,
[omega][root n]a, ... [omega]^(n - 1)[root n]a; or we may, if we
please, use [root n]a instead of a^(1/n) as a symbol to denote the
n-valued function.

As the coefficients of an algebraical equation may be numerical, all
which follows in regard to algebraical equations is (with, it may be,
some few modifications) applicable to numerical equations; and hence,
concluding for the present this subject, it will be convenient to pass
on to algebraical equations.

_Algebraical Equations._

11. The equation is

x^n - p1x^(n-1) + ... [+-]p_n = 0,

and we here _assume_ the existence of roots, viz. we assume that there are n quantities a, b, c ... (in general all of them different, but which in particular cases may become equal in sets in any manner), such that

x^n - p1x^(n - 1) + ... [+-] p_n = 0;

or looking at the question in a different point of view, and starting with the roots a, b, c ... as given, we express the product of the n factors x - a, x - b, ... in the foregoing form, and thus arrive at an equation of the order n having the n roots a, b, c.... In either case we have

p1 = [Sigma]a, p2 = [Sigma]ab, ... p_n = abc ...;

i.e. regarding the coefficients p1, p2 ... p_n as given, then we assume the existence of roots a, b, c, ... such that p1 = [Sigma]a, &c.; or, regarding the roots as given, then we write p1, p2, &c., to denote the functions [Sigma]a, [Sigma]ab, &c.

As already explained, the epithet algebraical is not used in
opposition to numerical; an algebraical equation is merely an equation
wherein the coefficients are not restricted to denote, or are not
explicitly considered as denoting, numbers. That the abstraction is
legitimate, appears by the simplest example; in saying that the
equation x^2 - px + q = 0 has a root x = 1/2{p + [root](p^2 - 4q)}, we
mean that writing this value for x the equation becomes an identity,
[1/2{p + [root](p^2 - 4q)}]^2 - p[1/2{p + [root](p^2 - 4q)}] + q = 0;
and the verification of this identity in nowise depends upon p and q
meaning numbers. But if it be asked what there is beyond numerical
equations included in the term algebraical equation, or, again, what
is the full extent of the meaning attributed to the term--the latter
question at any rate it would be very difficult to answer; as to the
former one, it may be said that the coefficients may, for instance, be
symbols of operation. As regards such equations, there is certainly no
proof that every equation has a root, or that an equation of the nth
order has n roots; nor is it in any wise clear what the precise
signification of the statement is. But it is found that the assumption
of the existence of the n roots can be made without contradictory
results; conclusions derived from it, if they involve the roots, rest
on the same ground as the original assumption; but the conclusion may
be independent of the roots altogether, and in this case it is
undoubtedly valid; the reasoning, although actually conducted by aid
of the assumption (and, it may be, most easily and elegantly in this
manner), is really independent of the assumption. In illustration, we
observe that it is allowable to express a function of p and q as
follows,--that is, by means of a rational symmetrical function of a
and b, this can, as a fact, be expressed as a rational function of a +
b and ab; and if we prescribe that a + b and ab shall then be changed
into p and q respectively, we have the required function of p, q. That
is, we have F([alpha], [beta]) as a representation of [f](p, q),
obtained as if we had p = a + b, q = ab, but without in any wise
assuming the existence of the a, b of these equations.

12. Starting from the equation

x^n - p1x^(n - 1) + ... = x - a.x - b. &c.

or the equivalent equations p1 = [Sigma]a, &c., we find

a^n - p1a^(n - 1) + ... = 0,
b^n - p1b^(n - 1) + ... = 0;
. . .
. . .
. . .

(it is as satisfying these equations that a, b ... are said to be the roots of x^n - p1x^(n - 1) + ... = 0); and conversely from the last-mentioned equations, assuming that a, b ... are all different, we deduce

p1 = [Sigma]a, p2 = [Sigma]ab, &c.

and

x^n - p1x^(n - 1) + ... = x - a.x - b. &c.

Observe that if, for instance, a = b, then the equations a^n - p1a^(n - 1) + ... = 0, b^n - p1b^(n - 1) + ... = 0 would reduce themselves to a single relation, which would not of itself express that a was a double root,--that is, that (x - a)^2 was a factor of x^n - p1x^(n - 1) +, &c; but by considering b as the limit of a + h, h indefinitely small, we obtain a second equation

na^(n - 1) - (n - 1)p1a^(n - 2) + ... = 0,

which, with the first, expresses that a is a double root; and then the whole system of equations leads as before to the equations p1 = [Sigma]a, &c. But the existence of a double root implies a certain relation between the coefficients; the general case is when the roots are all unequal.

We have then the _theorem_ that every rational symmetrical function of the roots is a rational function of the coefficients. This is an easy consequence from the less general theorem, every rational and integral symmetrical function of the roots is a rational and integral function of the coefficients.

In particular, the sums of the powers [Sigma]a^2, [Sigma]a^3, &c., are rational and integral functions of the coefficients.

The process originally employed for the expression of other functions
[Sigma]a^[alpha] b^[beta], &c., in terms of the coefficients is to
make them depend upon the sums of powers: for instance,
[Sigma]a^[alpha] b^[beta] = [Sigma]a^[alpha] [Sigma]a^[beta] -
[Sigma]a^([alpha] + [beta]); but this is very objectionable; the true
theory consists in showing that we have systems of equations

p1 = [Sigma]a,

p2 = [Sigma]ab,
p1^2 = [Sigma]a^2 + 2[Sigma]ab,

p3 = [Sigma]abc,
p1p2 = [Sigma]a^2 b + 3[Sigma]abc,
p1^3 = [Sigma]a^3 + 3[Sigma]a^2 b + 6[Sigma]abc,

where in each system there are precisely as many equations as there
are root-functions on the right-hand side--e.g. 3 equations and 3
functions [Sigma]abc, [Sigma]a^2 b, [Sigma]a^3. Hence in each system
the root-functions can be determined linearly in terms of the powers
and products of the coefficients:

[Sigma]ab = p2,
[Sigma]a^2 = p1^2 - 2p2,

[Sigma]abc = p3,
[Sigma]a^2 b = p1p2 - 3p3,
[Sigma]a^3 = p1^3 - 3p1p2 + 3p3,

and so on. The other process, if applied consistently, would derive
the originally assumed value [Sigma]ab = p2, from the two equations
[Sigma]a = p, [Sigma]a^2 = p1^2 - 2p2; i.e. we have 2[Sigma]ab =
[Sigma]a.[Sigma]a - [Sigma]a^2,= p1^2 - (p1^2 - 2p2), = 2p2.

13. It is convenient to mention here the theorem that, x being determined as above by an equation of the order n, any rational and integral function whatever of x, or more generally any rational function which does not become infinite in virtue of the equation itself, can be expressed as a rational and integral function of x, of the order n - 1, the coefficients being rational functions of the coefficients of the equation. Thus the equation gives x^n a function of the form in question; multiplying each side by x, and on the right-hand side writing for x^n its foregoing value, we have x^(n + 1), a function of the form in question; and the like for any higher power of x, and therefore also for any rational and integral function of x. The proof in the case of a rational non-integral function is somewhat more complicated. The final result is of the form [phi](x)/[psi](x) = I(x), or say [phi](x) -[psi](x)I(x) = 0, where [phi], [psi], I are rational and integral functions; in other words, this equation, being true if only [f](x) = 0, can only be so by reason that the left-hand side contains [f](x) as a factor, or we must have identically [phi](x) - [psi](x)I(x) = M(x)[f](x). And it is, moreover, clear that the equation [phi](x)/[psi](x) = I(x), being satisfied if only [f](x) = 0, must be satisfied by each root of the equation.

From the theorem that a rational symmetrical function of the roots is
expressible in terms of the coefficients, it at once follows that it
is possible to determine an equation (of an assignable order) having
for its roots the several values of any given (unsymmetrical) function
of the roots of the given equation. For example, in the case of a
quartic equation, roots (a, b, c, d), it is possible to find an
equation having the roots ab, ac, ad, bc, bd, cd (being therefore a
sextic equation): viz. in the product

(y - ab)(y - ac)(y - ad)(y - bc)(y - bd)(y - cd)

the coefficients of the several powers of y will be symmetrical
functions of a, b, c, d and therefore rational and integral functions
of the coefficients of the quartic equation; hence, supposing the
product so expressed, and equating it to zero, we have the required
sextic equation. In the same manner can be found the sextic equation
having the roots (a - b)^2, (a - c)^2, (a - d)^2, (b - c)^2, (b -
d)^2, (c - d)^2, which is the equation of differences previously
referred to; and similarly we obtain the equation of differences for a
given equation of any order. Again, the equation sought for may be
that having for its n roots the given rational functions [phi](a),
[phi](b), ... of the several roots of the given equation. Any such
rational function can (as was shown) be expressed as a rational and
integral function of the order n - 1; and, retaining x in place of any
one of the roots, the problem is to find y from the equations x^n - p1
x^(n - 1) ... = 0, and y = M0x^(n - 1) + M1x^(n - 2) + ..., or, what
is the same thing, from these two equations to eliminate x. This is in
fact E.W. Tschirnhausen's transformation (1683).

14. In connexion with what precedes, the question arises as to the number of values (obtained by permutations of the roots) of given unsymmetrical functions of the roots, or say of a given set of letters: for instance, with roots or letters (a, b, c, d) as before, how many values are there of the function ab + cd, or better, how many functions are there of this form? The answer is 3, viz. ab + cd, ac + bd, ad + bc; or again we may ask whether, in the case of a given number of letters, there exist functions with a given number of values, 3-valued, 4-valued functions, &c.

It is at once seen that for any given number of letters there exist
2-valued functions; the product of the differences of the letters is
such a function; however the letters are interchanged, it alters only
its sign; or say the two values are [Delta] and -[Delta]. And if P, Q
are symmetrical functions of the letters, then the general form of
such a function is P + Q[Delta]; this has only the two values P +
Q[Delta], P - Q[Delta].

In the case of 4 letters there exist (as appears above) 3-valued
functions: but in the case of 5 letters there does not exist any
3-valued or 4-valued function; and the only 5-valued functions are
those which are symmetrical in regard to four of the letters, and can
thus be expressed in terms of one letter and of symmetrical functions
of all the letters. These last theorems present themselves in the
demonstration of the non-existence of a solution of a quintic equation
by radicals.

The theory is an extensive and important one, depending on the notions of _substitutions_ and of _groups_ (q.v.).

15. Returning to equations, we have the very important theorem that, given the value of any unsymmetrical function of the roots, e.g. in the case of a quartic equation, the function ab + cd, it is in general possible to determine rationally the value of any similar function, such as (a + b)^3 + (c + d)^3.

The _a priori_ ground of this theorem may be illustrated by means of a
numerical equation. Suppose that the roots of a quartic equation are
1, 2, 3, 4, then if it is given that ab + cd = 14, this in effect
determines a, b to be 1, 2 and c, d to be 3, 4 (viz. a = 1, b = 2 or a
= 2, b = 1, and c = 3, d = 4 or c = 3, d = 4) or else a, b to be 3, 4
and c, d to be 1, 2; and it therefore in effect determines (a + b)^3 +
(c + d)^3 to be = 370, and not any other value; that is, (a + b)^3 + (c
+ d)^3, as having a single value, must be determinable rationally. And
we can in the same way account for cases of failure as regards
particular equations; thus, the roots being 1, 2, 3, 4 as before, a^2 b
= 2 determines a to be = 1 and b to be = 2, but if the roots had been
1, 2, 4, 16 then a^2 b = 16 does not uniquely determine a, b but only
makes them to be 1, 16 or 2, 4 respectively.

As to the _a posteriori_ proof, assume, for instance,

t1 = ab + cd, y1 = (a + b)^3 + (c + d)^3,
t2 = ac + bd, y2 = (a + c)^3 + (b + d)^3,
t3 = ad + bc, y3 = (a + d)^3 + (b + c)^3:

then y1 + y2 + y3, t1y1 + t2y2 + t3y3, t1^2 y1 + t2^2y2 + t3^2y3 will
be respectively symmetrical functions of the roots of the quartic, and
therefore rational and integral functions of the coefficients; that
is, they will be known.

Suppose for a moment that t1, t2, t3 are all known; then the equations
being linear in y1, y2, y3 these can be expressed rationally in terms
of the coefficients and of t1, t2, t3; that is, y1, y2, y3 will be
known. But observe further that y1 is obtained as a function of t1,
t2, t3 symmetrical as regards t2, t3; it can therefore be expressed as
a rational function of t1 and of t2 + t3, t2t3, and thence as a
rational function of t1 and of t1 + t2 + t3, t1t2 + t1t3 + t2t3,
t1t2t3; but these last are symmetrical functions of the roots, and as
such they are expressible rationally in terms of the coefficients;
that is, y1 will be expressed as a rational function of t1 and of the
coefficients; or t1 (alone, not t2 or t3) being known, y1 will be
rationally determined.

16. We now consider the question of the algebraical solution of equations, or, more accurately, that of the _solution of equations by radicals_.

In the case of a quadric equation x^2 - px + q = 0, we can by the
assistance of the sign [root]( ) or ( )^1/2 find an expression for x
as a 2-valued function of the coefficients p, q such that substituting
this value in the equation, the equation is thereby identically
satisfied; it has been found that this expression is

x = 1/2{p [+-] [root](p^2 - 4q)},

and the equation is on this account said to be algebraically solvable,
or more accurately solvable by radicals. Or we may by writing x =
-1/2 p + z reduce the equation to z^2 = 1/4(p^2 - 4q), viz. to an
equation of the form x^2 = a; and in virtue of its being thus
reducible we say that the original equation is solvable by radicals.
And the question for an equation of any higher order, say of the order
n, is, can we by means of radicals (that is, by aid of the sign [root
m]( ) or ( )^(1/m), using as many as we please of such signs and with
any values of m) find an n-valued function (or any function) of the
coefficients which substituted for x in the equation shall satisfy it
identically?

It will be observed that the coefficients p, q ... are not explicitly
considered as numbers, but even if they do denote numbers, the
question whether a numerical equation admits of solution by radicals
is wholly unconnected with the before-mentioned theorem of the
existence of the n roots of such an equation. It does not even follow
that in the case of a numerical equation solvable by radicals the
algebraical solution gives the numerical solution, but this requires
explanation. Consider first a numerical quadric equation with
imaginary coefficients. In the formula x = 1/2{p [+-] [root](p^2 -
4q)}, substituting for p, q their given numerical values, we obtain
for x an expression of the form x = [alpha] + [beta]i [+-]
[root]([gamma] + [delta]i), where [alpha], [beta], [gamma], [delta]
are real numbers. This expression substituted for x in the quadric
equation would satisfy it identically, and it is thus an algebraical
solution; but there is no obvious _a priori_ reason why
[root]([gamma]+[delta]i) should have a value = c + di, where c and d
are real numbers calculable by the extraction of a root or roots of
real numbers; however the case is (what there was no _a priori_ right
to expect) that [root]([gamma] + [delta]i) has such a value calculable
by means of the radical expressions [root]{[root]([gamma]^2 +
[delta]^2) [+-] [gamma]} : and hence the algebraical solution of a
numerical quadric equation does in every case give the numerical
solution. The case of a numerical cubic equation will be considered
presently.

17. A cubic equation can be solved by radicals.

Taking for greater simplicity the cubic in the reduced form x^3 + qx -
r = 0, and assuming x = a + b, this will be a solution if only 3ab = q
and a^3 + b^3 = r, equations which give (a^3 - b^3)^2 = r^2 -
(4/27)q^3, a quadric equation solvable by radicals, and giving a^3 -
b^3 = [root](r^2 - (4/27)q^3), a 2-valued function of the
coefficients: combining this with a^3 + b^3 = r, we have a^3 = 1/2{r +
[root](r^2 - (4/27)q^3)}, a 2-valued function: we then have a by means
of a cube root, viz.

a = [root 3][1/2{r + [root](r^2 - (4/27)q^3)}],

a 6-valued function of the coefficients; but then, writing q = b/3a,
we have, as may be shown, a + b a 3-valued function of the
coefficients; and x = a + b is the required solution by radicals. It
would have been wrong to complete the solution by writing

b = [root 3][1/2{r - [root](r^2 - (4/27)q^3)}],

for then a + b would have been given as a 9-valued function having
only 3 of its values roots, and the other 6 values being irrelevant.
Observe that in this last process we make no use of the equation 3ab
= q, in its original form, but use only the derived equation 27a^3 b^3
= q^3, implied in, but not implying, the original form.

An interesting variation of the solution is to write x = ab(a + b),
giving a^3 b^3(a^3 + b^3) = r and 3a^3 b^3 = q, or say a^3 + b^3 =
3r/q, a^3 b^3 = (1/3)q; and consequently

3/2 4 3/2 4
a^3 = --- {r + [root](r^2 - --q^3)}, b^3 = --- {r - [root](r^2 - --q^3)},
q 27 q 27

i.e. here a^3, b^3 are each of them a 2-valued function, but as the
only effect of altering the sign of the quadric radical is to
interchange a^3, b^3, they may be regarded as each of them 1-valued; a
and b are each of them 3-valued (for observe that here only a^3 b^3,
not ab, is given); and ab(a + b) thus is in appearance a 9-valued
function; but it can easily be shown that it is (as it ought to be)
only 3-valued.

In the case of a numerical cubic, even when the coefficients are real,
substituting their values in the expression

x = [root 3][1/2{r + [root](r^2 - (4/27)q^3)}] + (1/3)q /
[root 3][1/2{r + [root](r^2 - (4/27)q^3)}],

this may depend on an expression of the form [root 3]([gamma] +
[delta]i) where [gamma] and [delta] are real numbers (it will do so if
r^2 - (4/27)q^3 is a negative number), and then we _cannot_ by the
extraction of any root or roots of real positive numbers reduce [root
3]([gamma] + [delta]i) to the form c + di, c and d real numbers; hence
here the algebraical solution does not give the numerical solution,
and we have here the so-called "irreducible case" of a cubic equation.
By what precedes there is nothing in this that might not have been
expected; the algebraical solution makes the solution depend on the
extraction of the cube root of a number, and there was no reason for
expecting this to be a real number. It is well known that the case in
question is that wherein the three roots of the numerical cubic
equation are all real; if the roots are two imaginary, one real, then
contrariwise the quantity under the cube root is real; and the
algebraical solution gives the numerical one.

The irreducible case is solvable by a trigonometrical formula, but
this is not a solution by radicals: it consists in effect in reducing
the given numerical cubic (not to a cubic of the form z^3 = a, solvable
by the extraction of a cube root, but) to a cubic of the form 4x^3 - 3x
= a, corresponding to the equation 4 cos^3 [theta] - 3 cos[theta] = cos
3[theta] which serves to determine cos[theta] when cos 3[theta] is
known. The theory is applicable to an algebraical cubic equation; say
that such an equation, if it can be reduced to the form 4x^3 - 3x = a,
is solvable by "trisection"--then the general cubic equation is
solvable by trisection.

18. A quartic equation is solvable by radicals, and it is to be remarked that the existence of such a solution depends on the existence of 3-valued functions such as ab + cd of the four roots (a, b, c, d): by what precedes ab + cd is the root of a cubic equation, which equation is solvable by radicals: hence ab + cd can be found by radicals; and since abcd is a given function, ab and cd can then be found by radicals. But by what precedes, if ab be known then any similar function, say a + b, is obtainable rationally; and then from the values of a + b and ab we may by radicals obtain the value of a or b, that is, an expression for the root of the given quartic equation: the expression ultimately obtained is 4-valued, corresponding to the different values of the several radicals which enter therein, and we have thus the expression by radicals of each of the four roots of the quartic equation. But when the quartic is numerical the same thing happens as in the cubic, and the algebraical solution does not in every case give the numerical one.

It will be understood from the foregoing explanation as to the quartic
how in the next following case, that of the quintic, the question of
the solvability by radicals depends on the existence or non-existence
of k-valued functions of the five roots (a, b, c, d, e); the
fundamental theorem is the one already stated, a rational function of
five letters, if it has less than 5, cannot have more than 2 values,
that is, there are no 3-valued or 4-valued functions of 5 letters: and
by reasoning depending in part upon this theorem, N.H. Abel (1824)
showed that a general quintic equation is not solvable by radicals;
and _a fortiori_ the general equation of any order higher than 5 is
not solvable by radicals.

19. The general theory of the solvability of an equation by radicals
depends fundamentally on A.T. Vandermonde's remark (1770) that,
supposing an equation is solvable by radicals, and that we have
therefore an algebraical expression of x in terms of the coefficients,
then substituting for the coefficients their values in terms of the
roots, the resulting expression must reduce itself to any one at
pleasure of the roots a, b, c ...; thus in the case of the quadric
equation, in the expression x = 1/2{p + [root](p^2 - 4q)},
substituting for p and q their values, and observing that (a + b)^2 -
4ab = (a - b)^2, this becomes x = 1/2{a + b + [root](a - b)^2}, the
value being a or b according as the radical is taken to be +(a - b) or
-(a - b).

So in the cubic equation x^3 - px^2 + qx - r = 0, if the roots are a,
b, c, and if [omega] is used to denote an imaginary cube root of
unity, [omega]^2 + [omega] + 1 = 0, then writing for shortness p = a +
b + c, L = a + [omega]b + [omega]^2 c, M = a + [omega]^2 b + [omega]c,
it is at once seen that LM, L^3 + M^3, and therefore also (L^3 -
M^3)^2 are symmetrical functions of the roots, and consequently
rational functions of the coefficients: hence

1/2{L^3 + M^3 + [root](L^3 - M^3)^2}

is a rational function of the coefficients, which when these are
replaced by their values as functions of the roots becomes, according
to the sign given to the quadric radical, = L^3 or M^3; taking it =
L^3, the cube root of the expression has the three values L, [omega]L,
[omega]^2 L; and LM divided by the same cube root has therefore the
values M, [omega]^2M, [omega]M; whence finally the expression

(1/3)[p + [root 3]{1/2(L^3 + M^3 + [root](L^3 - M^3)^2)} + LM /
[root 3]{1/2L^3 + M^3 + [root](L^3 - M^3)^2}]

has the three values

(1/3)(p + L + M), (1/3)(p + [omega]L + [omega]^2 M),
(1/3)(p + [omega]^2 L + [omega]M);

that is, these are = a, b, c respectively. If the value M^3 had been
taken instead of L^3, then the expression would have had the same
three values a, b, c. Comparing the solution given for the cubic x^3 +
qx - r = 0, it will readily be seen that the two solutions are
identical, and that the function r^2 - (4/27)q^3 under the radical
sign must (by aid of the relation p = 0 which subsists in this case)
reduce itself to (L^3 - M^3)^2; it is only by each radical being equal
to a rational function of the roots that the final expression _can_
become equal to the roots a, b, c respectively.

20. The formulae for the cubic were obtained by J.L. Lagrange (1770-1771) from a different point of view. Upon examining and comparing the principal known methods for the solution of algebraical equations, he found that they all ultimately depended upon finding a "resolvent" equation of which the root is a + [omega]b + [omega]^2 c + [omega]^3 d + ..., [omega] being an imaginary root of unity, of the same order as the equation; e.g. for the cubic the root is a + [omega]b + [omega]^2 c, [omega] an imaginary cube root of unity. Evidently the method gives for L^3 a quadric equation, which is the "resolvent" equation in this particular case.

For a quartic the formulae present themselves in a somewhat different form, by reason that 4 is not a prime number. Attempting to apply it to a quintic, we seek for the equation of which the root is (a + [omega]b + [omega]^2 c + [omega]^3 d + [omega]^4 e), [omega] an imaginary fifth root of unity, or rather the fifth power thereof (a + [omega]b + [omega]^2 c + [omega]^3d + [omega]^4 e)^5; this is a 24-valued function, but if we consider the four values corresponding to the roots of unity [omega], [omega]^2, [omega]^3, [omega]^4, viz. the values

(a + [omega]b + [omega]^2 c + [omega]^3 d + [omega]^4 e)^5,
(a + [omega]^2 b + [omega]^4 c + [omega]d + [omega]^3e)^5,
(a + [omega]^3 b + [omega]c + [omega]^4 d + [omega]^2e)^5,
(a + [omega]^4 b + [omega]^3 c + [omega]^2 d + [omega]e)^5,

any symmetrical function of these, for instance their sum, is a 6-valued function of the roots, and may therefore be determined by means of a sextic equation, the coefficients whereof are rational functions of the coefficients of the original quintic equation; the conclusion being that the solution of an equation of the fifth order is made to depend upon that of an equation of the sixth order. This is, of course, useless for the solution of the quintic equation, which, as already mentioned, does not admit of solution by radicals; but the equation of the sixth order, Lagrange's resolvent sextic, is very important, and is intimately connected with all the later investigations in the theory.

21. It is to be remarked, in regard to the question of solvability by radicals, that not only the coefficients are taken to be arbitrary, but it is assumed that they are represented each by a single letter, or say rather that they are not so expressed in terms of other arbitrary quantities as to make a solution possible. If the coefficients are not all arbitrary, for instance, if some of them are zero, a sextic equation might be of the form x^6 + bx^4 + cx^2 + d = 0, and so be solvable as a cubic; or if the coefficients of the sextic are given functions of the six arbitrary quantities a, b, c, d, e, f, such that the sextic is really of the form (x^2 + ax + b)(x^4 + cx^3 + dx^2 + ex + f) = 0, then it breaks up into the equations x^2 + ax + b = 0, x^4 + cx^3 + dx^2 + ex + f = 0, and is consequently solvable by radicals; so also if the form is (x -a)(x - b)(x - c)(x - d)(x - e)(x - f) = 0, then the equation is solvable by radicals,--in this extreme case rationally. Such cases of solvability are self-evident; but they are enough to show that the general theorem of the non-solvability by radicals of an equation of the fifth or any higher order does not in any wise exclude for such orders the existence of particular equations solvable by radicals, and there are, in fact, extensive classes of equations which are thus solvable; the binomial equations x^n - 1 = 0 present an instance.

22. It has already been shown how the several roots of the equation
x^n - 1 = 0 can be expressed in the form cos 2s[pi]/n + i sin
2s[pi]/n, but the question is now that of the algebraical solution (or
solution by radicals) of this equation. There is always a root = 1; if
[omega] be any other root, then obviously [omega], [omega]^2, ...
[omega]^(n - 1) are all of them roots; x^n - 1 contains the factor x -
1, and it thus appears that [omega], [omega]^2, ... [omega]^(n - 1) are
the n - 1 roots of the equation

x^(n - 1) + x^(n - 2) + ... x + 1 = 0;

we have, of course, [omega]^(n - 1) + [omega]^(n - 2) + ... + [omega]
+ 1 = 0.

It is proper to distinguish the cases n prime and n composite; and in
the latter case there is a distinction according as the prime factors
of n are simple or multiple. By way of illustration, suppose
successively n = 15 and n = 9; in the former case, if [alpha] be an
imaginary root of x^3 - 1 = 0 (or root of x^2 + x + 1 = 0), and [beta]
an imaginary root of x^5 - 1 = 0 (or root of x^4 + x^3 + x^2 + x + 1 =
0), then [omega] may be taken = [alpha][beta]; the successive powers
thereof, [alpha][beta], [alpha]^2 [beta]^2, [beta]^3, [alpha][beta]^4,
[alpha]^2, [beta], [alpha][beta]^2, [alpha]^2[beta]^3, [beta]^4,
[alpha], [alpha]^2 [beta], [beta]^2, [alpha][beta]^3, [alpha]^2
[beta]^4, are the roots of x^14 + x^13 + ... + x + 1 = 0; the solution
thus depends on the solution of the equations x^3 - 1 = 0 and x^5 - 1
= 0. In the latter case, if [alpha] be an imaginary root of x^3 - 1 =
0 (or root of x^2 + x + 1 = 0), then the equation x^9 - 1 = 0 gives
x^3 = 1, [alpha], or [alpha]^2; x^3 = 1 gives x = 1, [alpha], or
[alpha]^2; and the solution thus depends on the solution of the
equations x^3 - 1 = 0, x^3 - [alpha] = 0, x^3 - [alpha]^2 = 0. The
first equation has the roots 1, [alpha], [alpha]^2; if [beta] be a
root of either of the others, say if [beta]^3 = [alpha], then assuming
[omega] = [beta], the successive powers are [beta], [beta]^2, [alpha],
[alpha][beta], [alpha][beta]^2, [alpha]^2, [alpha]^2[beta], [alpha]^2
[beta]^2, which are the roots of the equation x^8 + x^7 + ... + x + 1
= 0.

It thus appears that the only case which need be considered is that of
n a prime number, and writing (as is more usual) r in place of
[omega], we have r, r^2, r^3, ... r^(n - 1) as the (n - 1) roots of
the reduced equation

x^(n - 1) + x^(n - 2) + ... + x + 1 = 0;

then not only r^n - 1 = 0, but also r^(n - 1) + r^(n - 2) + ... + r +
1 = 0.

Comments

Log in to leave a comment.

Encyclopaedia Britannica, 11th Edition, "Equation" to "Ethics"Chapter II: Front Matter (2)

0%35 min left in chapter