This content was uploaded by our users and we assume good faith they have the permission to share this book. If you own the copyright to this book and it is wrongfully on our website, we offer a simple DMCA procedure to remove your content from our site. Start by pressing the button below!
)=
jpf
*(¥>,z).
iGSol(vj)
Here Sol(<^) denotes the set of solutions of
0 k = l,...,q
1070
where fi, gj, h^ are polynomials in x\,..., xn with real coefficients. There is a machine M over R which decides, on input
all * € H. Then ||£|| - |)o|| and Av |Lx| - Av |
| - | L | Av <*. x>, then \f(E{(xc))\<$, Proof. If p,S 1 then |/(z«)|<(p/2/2)cp, Thus 0. PROOF. E ^ r - 1 < d E o ^ ' - ' < dZd0r<. The following is just an application of the inequality on arithmetic and geo metric means. LEMMA 4. 1 possesses will confuse the picture. and constants from R). A sentence (as opposed to a "formula") has no free variables, i.e., all variables must be quantified. "First order" means that quantification is over elements (e.g., "there exists an element x in R such that for all elements y in /?"), not over subsets. 0&-&Qimj(x) > 0 where x = (x 0 ,..., x,,.!) and c e Rm is a vector of all the coefficients occurring in the g's. So, there is an r e R" such that (c,r) is true in R, then f(r,z) = 0 has no solution in R. Note if R has strong Property D, then without loss of generality we may assume c e Qm. Then, given the hypotheses of our theorem, the following lemma will yield a contradiction. Lemma 2. Suppose R is dense in R, its real closure, or that c e Q m . Suppose \p(c, r) is true in R for some r e R". Then there is an r' e R" such that \j/(c, r') is true in R and f(r', z) = 0 is solvable in R. PROOF. Let
*mH
*
mmH
where e - o/)|p|t- Hence (tee D w A y )
This yields the proposition. We wfll simplify the proof of Theorem D slightly by working with the subspaces Jt* - { / « ^ V ( 0 ) - 0} and JT0J - { / € JTV(0) - / ' ( 0 ) - 0}. Let / : •*? -» * be the integral, / ( / ) - / / . and denote by )n)ejT0l duaLThus
the
•/(/)-<J a \/>. Similarly, let./™ e JT02 denote the dual of / : JT02 - R. Next, for i - 1,2, and 0 < t < 1, define £,<'>: JT0* -»J? by £,">(/) - /(r). Let £/" e jf0' be the corresponding duals of this evaluation map. LEMMA 1.
(i)
£/»
(Hi) £ , » « ( , ) - ! - , + £,>.
<*> £«■<•>-{ir - - ; ; i The proof is easy calculus: For (iX it amounts to checking
For(ii),
For (Hi),
1209 116
STIVE SMALE
For(iv),
<£w./)^-j['('-*)r(')*-/(')-*,<2>(/). Towards the proof of Theorem D, use the proposition to see that
£ K/ - *,)/l - (f p / - *»U. - (f)'V' - *FkV • Now
**(/)- *I/(rt)-*Z^*(/). 1-1 <-i and it wfll be shown that
l/o) _ _ y *o)| . JL By Lemma 100, tct{j - D* < * < A
£x4P<«)-«-(y-i>. Thus, using Lemma 1(1),
|^-*£jg>|-{££ w [i-i-*(»-; +i)]2*} . But An - 1, so this is !»,'"*>
which proves the first of four parts of Theorem D. Next I wiQ carry out the same process for the "Trapezoidal Rule". Since *"*> - - **(/(D +/(0» + *»(/) « d / ( 0 ) - 0,
which proves the second part of Theorem D. A couple of more lemmas help pave the way for proving the third part of Theorem D. The first is well known. LEMMA 2.
fo-p'.<'-i>y-i>.
1210 EFFICIENCY OF ALGORITHMS OF ANALYSIS LEMMA
l.Forse
117
[(j - l)A, jh)
f^,?(>).Mi*ii".Miiil./(..y+1). 3. Write out the sum using Lemma l(iv) to obtain it as (jh - s) + (j + 1)A - * + • • • + nA - s.
PKOOF OF LEMMA
The rest follows. Let .4 stand for the quantity in Lemma 3. Then
using Lemma 1. This quantity may be written as (using Lemma 3)
M?id)*' ,, ~ 2j °" 1) * +0 '" 1)V+( * ,(y " 1) "* ) l1 • One can integrate it, using t - * - (j - 1)A; then a little calculation using Lemma 2 yields 2^l
1 +
2
+
10J *
which is the formula for eJ(A) in Theorem D. Next,
*»(/)-§ /(I)+ 2 'I- i/ ( A ) + 2 tl / ( 2 ( / - J ) * ) 1
"'+ 2 £* S 1>(,) + 4 £4S?-i>»(,)}
&S>M - 3
LEMMA 4.© On (2* - 2)A < s < (2* -
I
"2J^(')
1)A,
- *(»(2« - 1) - ( 2 * - 1)(* - 1)) -*(2» - Ik + 1)
dnrf <-i«
«*
4SW
f ^5?-i>*(') - *(«' - * 2 ) -'(- - * ) • i
*
1211 118
STEVE SMALE
The proof is a straightforward calculation as before, keeping in mind 1 + 3 + 5+ •••+(2/-l)-/2. Finally, a straightforward calculation using the above expression for Sk, and Lemma 4, yields the formula
*™-(sr&3^15"
finishing the proof of Theorem D.
2. Questions of precision. I wiD not attempt to give the proof of Theorem C, but instead refer the reader to Kosttan and Ocneanu. There are some aspects that I would like to comment on, Since the singular values of a matrix play an important role, it is worthwhile to say what they are and how they are used. If A is any matrix (say n x n f o r simplicity), the singular values of A are the nonnegative square roots of the eigenvalues of AAT. This makes sense because the composition of a matrix with its transpose is positive semidefinite and the eigenvalues are positive or zero. Singular values are discussed at length in Forsyth-Moier and Wilkinson I. They measure the distortion of a linear map. Let A be an n x n real matrix and 0 < » t < » } < • • • < nm denote its singular values. It is easily shown that ||^|| - pm and M" l || * 1/Mi- Thus the condition number KA of A equals Ma/fi- The number LA whose average is estimated in Theorem C therefore satisfies LA - log^a/Mi)- Since LA is in fact a function of the singular values of A, the average over A is equal to an average over the space of singular values. More precisely, the following is true. PROPOSITION.
^(")-c^^^)nw-»»?)«p(-5 E»»?) *.•■•*.. 0 < p t < • • • < Ma> where em is the rtciprocal of the same integral with the log factor deleted. The proof of the proposition uses the previous discussion together with a well-known result on the probability density for the Gaussian measure on matrices in terms of singular values. See Kosttan tot reference and details. This proposition is the basis for the estimates given in Theorem C. The above analysis of the system Ax - b involves a worst-case hypothesis for the error in 6. Kostlan studies the loss in precision for this system by averaging over b and the error in b. This yields zero for a given matrix A. So in this form, on the average, the amount of precision lost is the same as that gained. Thus the variance becomes crucial to estimate. Kostlan bounds this variance by a polynomial in n and conjectures that this quantity satisfies
VarU) < »7«. The above comments on precision are independent of algorithms. But the limitations of precision di
1212 EFFICIENCY OF ALGORITHMS OF ANALYSIS
119
One can also consider the problems of precision for nonlinear problems. If / : R" -» R" is a continuously differentiable map then the equation /(x) - y may be looked at from two points of view: Given x find y (evaluation), or given y find x. For the first, the natural definition of condition number is \\Df(x)\\, where Df(x) is the matrix of partial derivatives and, as usual, the norm is the operator norm. For the second problem, of solving/(at) - y for x, this is replaced by \\Df(x ) ' l \ \ . See Wilkinson II lot discussion of these things. Quantities reflecting the loss of precision for these two problems are thus log||£>/(jc)|| and logWDfix)'^ respectively. For an interesting analysis of the average loss of precision in the evaluation of rational functions, see Bhan-Shub. If I is a zero of a complex polynomial/, the loss of precision for the problem of finding it is
For the case of a double root this is infinite (for an almost double root, arbitrarily large). One has the paradoxical situation of polynomials with double roots, but no algorithm (with the best machine even to be built) can affirm it. REFERENCES L Adler, R_ M. Karp and R Shamir, A simplex oariant soloing an m x d linear ptegrmm at 0(raia(m 2 , D2)) expected number of pivot tups. Report UCB CSD 83/158, Computer Science Division, University of California, Berkeley (December 1983). L AdJer and W. Megiddo, A simplex algorithm whose average number of steps is bounded between two quadratic functions of the smaller dimension. Report, Dec 1983. A. Abo, J. Hopcraft and J. Unman, The design and analysis of computer algorithms, AdducetWesley, Readini, Mass., 1979. K- Atkinson, An introduction to numerical analysts, Wiley, N. Y„ 1978. B Barna, Ober die Diuergenzpunkle des Sewtonschen Verfahrtns tur Bestimnamg oon Wurzebi Atgebraischer Gleichungen. U, PubL Math. Debrecen 4 0956), 384-397. P. Blanchard, Complex analytic dynamics on the Riemann sphere, BuIL Amer. Math. Sec (N.S.) 11 0984). 85-141. L Blum and M Shob, Eooiuating rational functions: Infinite precision is finite cost and tractable on average, SIAM J. Comput. (to appear). K_ R Borgwardl, The overage number of pivot steps required by the simplex-method is pnfymumal, Z. Oper. Res. 26 0982), 157-177. J. Carry, On the periodic behaviour of Newton's Method for a family of real cubic polynomials, preprint, 1983. J. Curry, L. Garnett and D. Sullivan, Om the iteration of a rational function: Computer experiments with Newton's method, Comm. Math. Pbys. 91 (1983), 267-277. C. Dantzig. Linear programming and extensions, Princeton Univ. Press, Princeton, N J., 1963. B. Dejon and P. Henrici, Constructive aspects of the fundamental theorem of algebra, Wiley, N. Y„ 1969. J. Drmmri, A numerical analyst's Jordan canonical form. Thesis, Univ. of California, Berkeley, 1983. R Dorfman, P. Samurison and R Solow, Linear programming and economic analysis, McGrawHiB, N. Y , 1958. A. Douady, Sysiema aynamiqua holomorphes, Sennnaire Bourbaki, No. 399,1982. C. Eaves and R Scarf, The solution of systems ofpiecewue hnear equations. Math. Oper. Res. 1 0976),l-27.
1213 130
STEVE SMALE
D. Bworthy, G M M mmtwu m Mama* mom mi mmufoidi, Global Anaryait and in Application*, VoL H, uieraatioail Atomic Energy Agency, Vienna. 1974, pp. 151-166. P. Fatou. Sm m centringfimrrirmmtlln.Bui Soc Math. France 47,41 (1919-1920), 161-211; 33-94; 201-314. G. Fonyth and C. Mokr, Campmu tmtim tf t w r m^gthmc tyumm, Prance-HalL Ea■>c»oodClifh.N.J..lM7. H Goidsane, <4 history af manorieol mmlyaU, from dt* Xtok dumafr dm 19» cemary. SpringerVerlag, Bcriia aad New York, 1977. J. Guckcahcimcr. Emaemwrpkimm tfdm lUammm mkart. Global Aaar/eis (Chen aad Soak. eds.). Amcr. Math. Sot. Providenceft.L, 1970, pp. 95-123. W. Haymaa, HvttwaUntfmeaam, Cambridge Uaiv. Proa, ramtirirlgi, 1951. P. Hcarici. Appli* omd amptaatiamd eempia amdym, Wary, N. Y, 1977. E Hnk,/
«a*M).u-2a J. Kcacgar I. Om dm compiatty a/a ;iK>nii tuaar algorithm/er apprmrimmimg rota ofommpkx J. B^aaaar n.Oa»Wcawe/«j/iu.uiiiiata|af roc«t^ai^u^»o/|iia^W, Maa^rYngriiinmng (to appear). J. Bcnegar Dt, Om aW tfficitmcy tf Nnrtm't method m ajipwiTiwerfaj off Mror of a tyttam tf ematkx ptdyntimiah, preprint, Colorado Suat Uaiv, 19*4. D. Saari aad i. Ureako, Ntwtm't mtthod, ebxk mam, mtd cfc—He mottm, Amcr. Math. Monthly 91 (I'M). 3-17. G. Sauadert, lurmtim tf rational fmaicm af mm cciiata aanenlr mtd batimi tf attrmetm* fixtd poima, Then*. Uaiv. of California, Btrkcky, 1964. R. Shaaar. 7VqgfcM»c*c/aWji»aWfata«c«ae«\-i<jaTa^ 1964.
M. Saab, Tat gtomttry mm tootioty tf tymmmcm tysttmtt mtdtlftwwmtfar awmartcmnwMevv, notes prepared for iecturs pvea at DJJ.4 Pcaaat Uaivcnity, Beqiat, Caiaa, Aap-Sept 1983. M. Shub aad S. Soaalt I, Computational compkxay: On oW gntnturj afpolymmtiah and m dtaery tfeott, fan /, Ann. ScL took Norm. Sup. (4) (to appear). M Shub aad S. Smak P. Comptttaiiimal itmitit.\ilr- On dtt gutnttry ofpofymmuali ande dmery tfettr. 'art II, SIAM i. Coaputiat (to appear). S. Smak I, DtfftrtmoaUe dynmmtwl tyttmm, lac Mathcautici of Time, Spriaaer-Vertae, New York aad BcrKa, 1910. S. Smak H, A cmvtrgaM praeam af prier odjmtmtnt mtd ghoul Ntwtm mtthtm, J. Math. Econoai 3 0970.107-120. S. SmakDI, Tmtfmdmmaml aumvm tfargtbra mtd tompltxity aWerj'. Bull Amcr. Math. Sec (N.S.)4a*>lXl-36. S. Smak IV, On aW aotragr IMIMW afttapt m die nmpkx mmhad af Mmmrprqgrmwnmg. Math. PwtrimmiH " 0*t3). 241-261
1214 EFFICIENCY OF ALGORITHMS OF ANALYSIS
121 ,
S. Souk V, The problem of me tpeeioftht Simplex method. Mathematical i*" i"»""i"| The Suie of Ibc An (Bam, 1982) (Bachem ct at, ed*.), Serinjer-Veriaa, Bcriia aad New Yocfc, 1983. D. Suffivan, Quaticonformal humeumoiphamt ani iynomta III: Topotogical canjugacy c h a d of an*J)fae*iomorphumi,AaiLOtMAth.(*>appar). M. J. Todd, Polynomial expected behavior tf a pwodng algorithm for tutor complementary ami hnaar programming probhnu, Tcdmical Report No. 595, School of Operabotu Research aad Industrial Enajacerint, Cornell Univtmty, Ithaca, New York, 19(3. J. Traub aad H. Wccaiakowaki, Information mi computation, m VoL 23, Advaacei ia Coav patcn, vol 23 (M. C. Yowl*. ed.), Academic Pre*, N. Y„ 1914. A. Venhik aad P. Sporythcv, An animate of the avtragt member of tups m the timpkx mtthoi ami problem in arymptoac integral gtomttry, Soviet Mala. Dokl 3B (1983). (RiMStaa) J. Von Neumann. Collected works, voL V (A. Taob. ad.), MacMfflaa, N. Y , 1963. J. Wakiaton 1, The algebraic etgeiwahtt problem, Oxford Uaiv. Prett, Oxford, 196S. J. WOkmaao n . Homing arras m algebraic proceam, Prentice-Hall. Eafleweod Cliff*, N. J-, 1973. DEPAKTassn- OF MATHEMATICS, U i o v n s n or CAiirotNiA,
tamn,
CAUPOCMA 94720
^
1215
RAM I C O M W I ¥•1 15. M* I. Nkrav) IMS
f , t«M Socwty ft» MuMru! art A p p M MaihtaHa
COMPUTATJOSAL COMPLEXITY: ON THE CEOMETRY OF POLYNOMIALS AND A THEORY OF COST: II* M IHUB+ AND S. SMALEt AkatracL This paper dealt with traditional algorithms, Newton'• method tod ■ higher order generalizatioa due to Evict. The*e iterations schemes ami their modifications have had a great success in solving •onlinear systems of equations We give tome anderstartding of this phenomenon by giving estimate* of efficiency The problem we focus on u thai of Boding a zero of a complex polynomial. Key wards. Newton, Euler, approximate zero, polynomial, average
1. This paper deals with traditional algorithms, Newton's method and a higher order generalization due to Euler. These iteration schemes and their modifications have had a great success in solving nonlinear systems of equations. We give some understanding of this phenomenon by giving estimates of efficiency. The problem we focus on is that of finding a zero of a complex polynomial. Following the work of Newton and Euler we define a rational map (or iteration scheme) £:C-*C (C the complex numbers) which depends on three parameters: (a) /, a polynomial, ./"(*) = £,_ 0 iv', fl**0. Often times we take / to be in the space Pj(\) where
PA\)~\f /(z)« I «,*-,*-i.talsil; (b) a positive integer k (which amounts to the number of derivatives used); and (c) A number h, 0 < f c S l . Then define E = £A,A./> E:C-»C by E(z)=7i(/T'((l-*)/(*))). 1
Here /f is the branch of the inverse o f / which takes /(z) into z, given as an analytic function in a neighborhood of/(z) (provided/"(zJ^O). Tk is the truncation of the power series expansion in h about h = 0 at degree it. It is easy to check that £,.,,/ is Newton's method. One can see a full discussion in Shub-Smale (1982) (hereafter referred to as [S-SI]). Consider first the problem: Given (/ e), / e P^(\), e > 0 , produce a z e C with [/(z)| < t. For this we particularize the Newton-Euler iteration scheme by choosing k and h to depend only on / and e, in a certain way. Let Jc = [max (log (log e|, log d)] where \x] is the least integer gjeater than or equal to x We will define in § 2, universal constants H and X, approximately jh and 512 respectively. Then we will take
• Received by the editors November 17, 1983, and in revised form September IS, 1984 This work was supported in part by grants from the Nation*.' Science Foundation t Mathematics Department, Queens College and the Graduate School, City University of New York, New York, New York 11367. : Mathematics Department, University of California, Berkeley, California 94720. 145
1216
146
M. SHU* AND.S. SMALE
Thus with these specializations the Newton-Euler iteration scheme £:C-»C depends only on («,/) and we write £, ■ £ With * > 0 de6nc (N-E), Let f€ PAD u d «i-JC(rf+|log e|). (1) Choose ZgCC, | z j » 3 at random and act for i - 1 , 2 , 3 , • • • (an Iteration) z,m £.(z,_i) terminating if ever |/(z,)| < e
ALGORITHM
(2) i n - n , g o t o ( l ) ( a c v d e ) . THEOREM A. For each/, t, (N-E), terminates with probability one and produces a t satisfying |/(z)| < e. 7he average number of cycles is less than or equal to 6. Hence the average number of iterations is less than 6K(d + \\og e\). Here average and probability refer to the choice of the sequence of z« in (1) of (N-E).. Remark With certainty h only takes about twice as long. See 12 for an elucidation of this remark. In practice one can obviously do better by trying and testing h ■ 1, J, • • •, H. We have not analyzed this. Also see i 2 for the total number of arithmetic operations required. Next consider sharpening the goal |/(x)| < e. Machine or discrete processes will not generally succeed infindingexact zeros of polynomials. For our theory we use the notion of approximate zero z of a polynomial / Smale (1981), [S-SI]. This complex number z is one close to an actual zero, where closeness is defined without any arbitrary choices. The justification of close is given in both theoretical and practical terms. More precisely define
P/« min [f(9)\.
There is a universal constant c (about A) and z is an approximate zero of / if l/(z)l
|/(£'(z))|<6""V The extremely rapid convergence gives some good justification for "approximate zero". In Algorithm (N-E), there was a random element, the choice of z0. Now probability enters into our analysis in a second way. We average over/e P4{\) with respect to a uniform distribution; that is we normalize Lebesgue measure on ft(l)cC*«RM. We use these probabilities since speedy algorithms are not usually infallible. Define for each / e P4(\)
''-(27r|D'' where D, is the discriminant o f / (see Lang (1965)). With K as above let
««r*<
*-
1217 ON THE OEOMETRY OP POLYNOMIAL! AND A THEORY OP COST
147
(N-E). U t / c ^ ( l ) , satisfy « / >0. (I) S c t m - 1 ; (2m) Choose «,€C,|»o!-3 at random and set x.-£"(««)• lfl/(».)l<«/terminate and print: "x, is an approximate S*TO;*° (3) Otherwise let m - m +1 and go to (2.). ALGORITHM
THEOREM B. Algorithm (N-E) terminates (and hence produces an approximate xerv) with probability 1 and the average number of iterations is less than Kxd log d where X| Is a universal constant We make the probability considerations a bit more precise. Let Sit be the circle in C defined by |*|-Jt and endow h with the uniform probability measure (Lebesgue measure normalized to 1). Set R ■ 3 and denote by ft the product of S* with itself a countable number of times. Thus a point s, of ft is a sequence f «(! u fj, • • •) with |5,|» 3. Endow ft with the product measure as well as P4(\) x n . L e t T : i » , ( l ) x n - » Z + be defined by: T(J,l)ii1hc first m such that E"(Im)<
Thus the total number of iterations of Algorithm (N-E) for a given / is of the form S(/, I) - nT(f. i), n « K(d + (log tf\). Theorem B asserts than when e,> 0, S(/, I) is defined for almost all 5eft. Moreover S ( / ) » J f < n S ( / f ) is defined and finite for almost all / and I
S(/)SX,rflogi
By Fubini's theorem, we could equally well assert that [
St/, *)£*><> log 4
Remark 1. We are assuming exact arithmetic is the theory here. In general, because of the robust properties of Algorithms (N-E), and (N-E), this is reasonable. However the calculation of t} in Algorithm (N-E) is not so robust. In that respect. Theorem A is more satisfying than Theorem B. Remark 2. Our work emphasizes the theoretical side, and the understanding of classic algorithms, rather than the design of new practical algorithms. Yet the results do have some implications for the latter. For example they suggest calculating deriva tives up to order flog d] and/or [log [log e|] could give speedier routines, especially for one complex polynomial. We have not tested our algorithms on the machine. Remark 3. The number of arithmetic operations in contrast to the number of iterations is approximately quadratic in d This is proved in S 2. Remark 4. Questions of variance arising in these theorems can be handled. See 12. At this point we review some of the motivation from Smale (1981), Hirsch and Smale (1979) and [S-SI]. Let / b e a polynomial,/: C-» C, let zeC and w«/(z). The ray from w to 0 is the segment in the target space from w to 0. Let X. denote this ray and /, the branch of the inverse of / uking /(z) back to r If/,"' is defined on all of /t» then { « / f ' ( 0 ) is one of the zeros o f / (Fig. 1). Since a polynomial maps a
1218 1**
M. SHUB AND •. SMALC
neighborhood of infinity to a neighborhood of infinity and has onlyfinitelymany critical poinu it is a fairly simple caJculus exercise to see that except for afinitenumber of rays (at most d -1) /,"' is defined on all of R, Thus we attempt to follow the curves /7 (RJ from an initial starting point x to a zero ( of t One way to do this is to parameterize the R, as (1 - k)f(z) for OS k 31. Then try to follow the ray by analytic continuation in k,f;\(l - k)f(z)). Finally, truncate the power series at degree kink to make the computationfinite,Wr'CO - A)/(r)). This is EKK/{z). If P(k) u positive for small positive * then (1 -P(k)) is also further down the ray and we may try Tkf?\l-P(k))f(z)). The inverse images of the rays are also solution curves of the differential equations f - -/U)//(*) and i--igrad|/(r)| 2 (see Smale (1981)). Thus we may attempt to solve these equations numerically with step size n to attempt to follow the ray. These examples are given in greater detail in [S-SI]. We analyze a class of fast algorithms broader than the Euler iterations, but which still agree with the inverse of the ray to high order. First we extend the E^z)Tk(f, ((l-k)f(z))) by replacing A on theright-handside by P(k)«£*.,cf\' where c, is real for all i and et > 0. These generalized Euler iterations are GE,w/(:)-r,(/;l((l-f(»))/(:)). Thus £*.*., is given by P(k) - k. The GE nM>/ are polynomials in * of degree k. We allow modifications of these polynomial iterations by addition of a well bounded remainder term of order * +1 in k. We denote this largest class of iterations we consider by GEMt (GEM - Generalized Euler with Modification). These iterations are fast An important ingredient in the analysis is the function t4 xC-»R*. {f,z)-*BA, which was introduced in [S-SI]. We recall the definition of this function Given/ z andOSaSs-/2, let »>^.-{*€C|0
2/U)
FIG. 2. /,-' it defimd em thii wtdgt.
In 6 3 we prove Theorem C. C. Suppose tkat z' - lK/(z) is a GEM4 iteration. Then there a a constant k depending only on I suck that. If BAH>0 and |/(z,)| > L > 0 then there is an k given explicitly suck tkat THEOREM
l/U.)|
1219 ON THE OEOMETRY OF POLYNOMIALS AND A THEORY OF COST
149
Theorem C can be used to thow that any GEM* iteration can be adapted to produce fast algorithms as in Theorems A and B of this paper. Finally, in 13 we ahow that the GEM* iterations are precisely the efficiency ft incremental algorithms defined in [S-SI] which satisfy an additional "amallness" condition. 2. The main goal of this section is to prove Theorems A and B of 11. We require some of the main result* of [S-SI]. First Proposition 1 is a special case of [S-SI, Thm. 2] at least after a short translation of constant*. The constants JC, K' are universal, not very large and well estimated via [S-SI]. Let ft «1,2, • • •, arid e > 0 . PROPOSITION 1. There exist JC, X ' > 0 to that tf 0
|/(£'(Zo))|<( for tome 0 S « n . Here £ ' is E • • • • • E, i times and E is the Eu!erk iteration of 11. Next by specializing k to ft ■ max (flog d], [log [log e|"|) we obtain COROLLARY. There exist universal constants H, K so that for n « K(d + |log e|), fePM) end |zo|-3 with ©/-JoS w/12 |/(E'(zo))|<e for some
Oii
Finally we require the following proposition which is obtained from [S-SI, Proposition 3, 14]. For/€ Pt(\), let V,«{2||z|«3 and BA,> v/12}. Then using the uniform probability measure on S « { i j | z | - 3 } we have Proposition 2. PROPOSITION 2. The measure of\t^\ for any f. We recall a bit of probability theory. The set 5 has a probability measure. Impose the product measure on f) the (ordered) counuble product of S with itself. Suppose V c s has measure c For f e l l , f - ( f „ f,,• • •). Let m(£) be the m such that i,tV for i<m but f„e V. PROPOSITION 3.
f »"(?)«-. Jfcfl
C
Proof. Let V,»{f|m(f)»i}, i - 1 , 2 , • • • . Let e, be the measure of V, Then v,« u(l - »)'"' and moreover |
m(f)-X
fe,«£
re(l-r)'-'-i.
Now we can prove Theorem A of § 1. In fact it is an immediate consequence of the corollary of Proposition 1, and Propositions 2 and 3. Q.E.D. We will need a few more facts to prove Theorem B. We begin with another elementary result of probability theory. DEFINITION. Let (X, M) be a probability space with no atoms. Let S:X-» JL, be a real valued nonnegative measurable funaion and let / : (0,1) ■+ R be decreasing and Riemann integrable. We say that S(x)Sf(n) with probability l-p if fi{x\s(x)£ /(v))fcl-vforallO
1) E(S)=\
S(x) M (dx)sj
fMdp.
2) Var(S)« f ( S ( X ) - E ( S ) ) V < * ) S ] V ( M ) * M - ( £ ( $ ) ) ' •
1220 150
M. fHUB AND S. IMAU
Proof. We only prove 1). Let y,e(0,1) for -oo< Kao be a decreasing sequence with 0 and 1 as limit points. Let •••=> M* => M^ ,=>••• be constructed so that M(M,,) ■ 1 - / , and S(x)Sf(yt) for *eM„ Then J S ( * ) M ( * 0 * I / W M ( M » "**.,) « Zfiji)iyi-i~yt) which converges to the Riemann integral off. Recall that p, ■ mn9j-l9)m0 \f{9)\. From Smale (1981) we have Proposition 5. PROPOSITION 5.
Volt/c^Dl^aKsfa1. Here Vol means nonnalized volume so that Vol (>\(1))«1 and Vol is a probability measure on P*(l). PROPOSITION 6.
I,
llogp/lailogrf+l.
ifr*A I)
Proof. Let p-da*
so from Proposition 5, VOI{/CP - (1)|P /
Vol{/eP - (l), P/
a*(/./7rf) L (2d-i)<.'lj(d y-i
.l)\aJ'-\ V J I
2) VfePM) then (2
**(/./Y<0 S(2rf-l)!V^^,)|^-. Let.T(*)-0and|/W|-p, L e t / . - / - / ( # ) . Thin aJLU) « * - / ( • ) . « . ( / , ) - « , ( / ) for i> 1 and R(A,fJd) - 0.
1221 ON THE OEOMETRV OF POLYNOMIAL! AND A THEORY OF COST
1SI
Applying the mean value theorem to the line tegment between / e and / grvet
I *(/./7<0l-l*(/./7rf) - *(/*/w/')l , rf I i(2rf-i)!i /( 7 )(W+*yV-,l/-*l i-i \ J I
i(2j-i)!*i^( 4 ': , )(p / +iy-V/
—«5(S(';V)U
«(2
rf'(2
For the proof of c), note that . , < ! . By definition Pof«/l- log
(2d) 4-
Therefore by using Proposition 7a, lloge/ISXjdllogd-l-llogp/ll and J|loge,|SKj(log
PogP/l)-
Now the first term on the left is estimated by Proposition 6. The second is estimated using the fact that pf < d +1. This yields Proposition 7. With these preparations we now prove Theorem B from Theorem A. Note first that by Proposition 7b, Algorithm (N-E) does terminate with an approximate zero if it terminates. Now apply Theorem A with Algorithm (N-E), where t it chosen to be th and then integrate over P*(l). This yields
f
5(/*)S6JH + 6Jt I |log«y|.
Now apply Proposition 7c. Q.E.D. Remark 1. There are various ways to compute the Euler iterations z'-£ M (r 0 )= Tk(/r'((l-*)/(*))). a) The Taylor series expansions of / at z, /, may be calculated by the algorithm of Shaw and Traub with (2d -1) multiplications, d-l divisions and rf2+ d/2 additions, after writing /(w)-(w-Zo)Q(») + *- See Knuth (1981, p.470) or Borodin-Munro (1975, p. 33). Moreover/, may be calculated in 0(d log d) arithmetic operations and perhaps in O(d) operations, tec Borodin-Munro (1975, p. 106). b) T»(/,"') may then be computed in k*/6 multiplications by Lagrangian power aeries reversion (Knuth (1981, p. 508)). The algorithm of Brent-Kung (1976) (tee alto Knuth (1981, p. 510)) computes Tj;' in 0(k log k) ,/3 operations and the value 7fc/7'((l-*)/(')) u> 0(k log k) operations once the Taylor series of / a t x is known.
1222 152
M. SHUI AND S. SMALE
c) Putting together a) and b) to get a good asymptotic estimate gives that Ti/T'O *!/"(*)) u computable in 0( + k log*) operations. d) Thus in Theorem A the average of the total number of operations is 0((
n(/;'((i-p(A))/u))). DEFINITION 1. z'« /»,/(z) is a GEM* iteration iff there is a polynomial P(h) and constants c>0, 6>0 such that: IK,(z) «GE„ kA/ (z) + FR^iKA i)
where
_
f(z)
m
and | / W * . / *)l* Ck"*' max (1,1/Af) for 0< k < 8 min (1, A,). The number V - M / z ) u the radius of convergence of/,"'((1-A)/(z)) as a power series ia k around zero. Examples of GE iterations without modification are described in [S-SI] these include incremental Newton's method,fcthorder incremental Euier, andfcthorder Taylor's method for the solution of the differential equation dx/di - -f(z)/f(z).
1223 ON THE OEOMETRY OF POLYNOMIALS AND A THEORY OF COST
133
It is sometimes more convenient to express the CE iterations and the modifications in terms of the polynomials au introduced in [S-SI].
*(«0-**.(»)-I>,w' where *.«0.
»,-!. c
j
^
^
,
We recall tome basic facts about a. Given the iteration i'*IK/(t)
. write
We frequently write /*(«,/, z) or just R{h), R(z) or R. By Taylor's formula (tee [S-Sl.ll]) or
Thus R(Kf){z)€ o-~'(l -f(z')/f(z)). a is a polynomial of the same degreerfa s / thus there are points 2' which give the tame value for/(z')//(r). If / K /
j
H' m
where a" is the branch of the inverse of a taking 0 to 0. The radius of convergence of
Q**jK/M-*+FTk{a-lmh))). Consequently we may restate Definition 1. DEFINITION 1'. z'^ IKf(z) is a GEMk iteration iff there is a polynomial P(ft) and constants c > 0 , 8>0 such that IKf(z) = z + F(r»(a-(P(fc))+ **♦,(«,/ z)) where I**,,(«,/ z)|S eft**' max (1,1/ftf) for 0
GE^W-z+FlV'*'. j-o Po« c„ P1 = eJ-vxe7u Pi = c, - 2a,c, c2 - (o-j - 2a\)c\. One can write down Pt explicitly inductively, in terms of P„ i<j, crj4., and e^,. Proo/ The coefficients of a"' arc computed from the a, and P/c, 0-) is defined by
I/A^'-r.^^^)').
1224
154
M. SHUt AND ». SMALE
The GEM* iterations may be similarly written with the addition of a remainder term. Examples of Generalized Euler with Modification are the simple Runge-Kutta approximation to the solution of dx/dt - -f(z)/f(z) (see [S-SI]) and an incremental Laguerre method. For the latter do the following. Instead of letting R - T^r~l(h) as for Euler we let (T& • R) - A and solve for R. We can do this for Ic - 2. That is we solve a2R3+R * A for X yielding R - (-1 W l + 4a,A)/2ir,. For A ■ 1 this is Laguerre's method of Henrici (1977, p. S3) with y - 2 . We now turn to an analysis of the GEM iterations with the goal of showing that GEM's are cheap, Theorem C. Given P(A) there is a 6 > 0 such that P is injective on the disc of radius 8, D{8). Let K be the Lipschitz constant of P on D(8), K « tup,tDi$, |P'(z)|- Let y*(c,) be the first positive root of ( l - y ) , - 4 e 1 y ( l + « ( * + l)y J ) where 1S B< 1.07 and £ is a constant which makes the Bieberbach conjecture true (see [S-SI]). Finally, let y*(P)«min(«.— ,y»(c,)V LEMMA 1. Suppose that 0< A, 3 min (1, A,(/ z)) and that A - ah^ for some complex number a with \a\ - y and 0
a) b) 0 d)
k- , P(A)|Sc,yA # /(l-y) 1 ; k- , P(A)-^ t< ^- , P(A)|Sc 1 fc,B(k+l)y*♦ , /(l-r)^ |Tfc<
Here D(h) is the disc of radius A around 0. Proof. Since A,S(1/IC)A, P(D(AJ)cD(A,)and since A, < «A, tr"1. Pis defined and injective on D(A,). (o , ~'»P)o«c, so (l/c,)(
|rt(a"'(P(A))|S-f^+£^±^ 0-y)
(i-y)
and theright-handside is < A,/4 since y < y t (c,). This proves c). Since |Tfca",(P(A)|< A,/4, [S-SI, Lemma 5, 12] proves d). LEMMA 2. Let Ikjiz) -z + F(z)RiKf)(z) be a GEM iteration then there is a 8 such that
UW>I<^ for 0
Proof. By the Kocbc theorem r-'(!Xft,))3!>(j)
1225 ON THE GEOMETRY OF POLYNOMIALS AND A THEORY OF COST
1SS
where D(r) denotes the open disc ofradiusr around 0 in C. By Lemma 2 there is • f such that
IW)I<*^ for 0 < * < « min (1, «,(/.*))• Thus *tK/Ai)~9~\h') for some *'cD(fc,) 9• * (fc>/) (z) - V and o-\o • ft(Mrfi)) - o'\W) - *(Kr>(z). , > PROPOSITION 2. //z'-/ < l > / (z)-»+F(r J ^" (i (*))+Jl»*,(fc)) is a GEMk ttrra(ion, then there are constant K> 0, e> 0 «w/i ffcat <•)
^«1-P(k)+S^,(k)
"*ew |S»,,(fc)|< *«**' max (1,1/fcJ) /or 0 < * < t min (1, »,). Proof. ^ii.
1 - a • R - 1 - v{T*q~-'P(h) + **♦,(*)).
There are K, 5 > 0 such that for OS h < 5 min (1, h,) S^W-aiv-^PihV-riTtV-ipW+RtM) and
k-(/»(fc))|<2cA |a"P(k) -(T fc a-'(P(A)) + K*.,(k))|< W
1
max ( l , £ )
by Lemma 1, the definition of X t4 ,(h) and the triangle inequality. Now by the generalized Loewner theorem (see Smale (1981) say) W " " ' <4/fc,. Now apply [S-SI, Lemma 3, $2] with a«=4/fc„ fc-2c,fc, c^k/Mc,)*** 1 max (1,!/«{). Make sure that 5 is small enough to guarantee (1 + c)ab < 1. Iterations satisfying (•) were called efficiency k in [S-SI]. One of the fundamental estimates on speed is proven for iterations of efficiency k [SS-I, Thm. 4]. Since a polynomial / of degree d is generally d to 1 a point z' is only determined by f(z') up to this d to 1 ambiguity. Efficiency k iterations are determined by the ratios/(z')//(z) and thus are determined up to this same ambiguity. We impose a continuity condition on iterations. DEFINITION. The iteration IKf(z)~z +FR{h,fz) is called small iff there is a £>0such that a~\tr(R)) = R for all 0
^J-l-(P(k) + St
1226 136
M. SHUi AND f. SMALE
where |Sk-,(A)|< KA**» max (1,1/A J) for 0< A < « min (1,1/Af) then
for 0 ( A ) + S^1(A))-r»
k"(P(*))-r«a-(P(*))|g C '***[**j a ) ^' so for 0< A < y»(P)/2 min (1, A,) \a-\P\h))-Tki,-l(P(k))\S4e^m4clkk*imMx^l,-pj. Now estimate \
fl|
ForO
*P{h){
and
l
by Lemma la and there are 6U K such that |S^,(A)| < Kkk" max ( l , ^ ) - K ^ r for 0
(l-(4/y*(^)A/A 1 )(l-(l + #:y4(/»)(A/A,)t)((4/y4(P))A/A1))• Now it is easy to produce a 6 such that for 0
1227 ON THE GEOMETKY OF POLYNOMIAL* AND A THEORY OF COST LEMMA
157
4. Suppou !/(»,)!«Lf(x)| and that s,-/r , ((1-tM(i)) for 0<\k\<
/*DO/ Since |/(x,)| < |/(z)| itradicesto see that z, 6/;'( W/>f) tee [SS-1,13] for then/,",1 is defined on a wedge of angle at least the angle of W/t minus |arg (/(z,)//(x)|. But the open disc of radius \f(z)\ tin Blt centered at /(z) is conuined in WA, as Fig. 3 ahows. Thus (l-*)/(z) u in this disc for 0*|fc|<«n •,., and ttef7\WA,).
Fie. 3 LEMMA 5. Let /^/(z) be a GEMk iteration. Then there is a constant a, 1fca >0 depending only on I such that: If 0< h < a sin 6/,,, then
e**e,.-|arg^2]. Proof. By Lemma 4 we need only show that there is an a such that if 0 < h < osine,, then *'=/;'((!-A)/(z)) for some h with 0<|A[<sine,.r N o w / ; ' ( ( ! h)f(z)) = z + Fa-\h). Thus h suffices to show for z'- IK/(z)-z+ FR that R(h,f, z) =
forO
""( , -Sj)"«ind.J. 1 )» < 5^Thus for 0 < h < n sin 0 M
^{k^S**
tin 8j.r
1228
15S
M. SHUB AND f. SMALE
We may assume M < min (1/21,1/2) for 0 < A < M tia 6 ^ |P(A)| < | sin 9/., |5k«i(A)| < (J)**' tin 6/., and we are done. D We now can prove that OEM's are cheap. This is the analogy of [S-SI, Theorem 4] but for GEMs A/,, can be replaced by &£%. THEOREM C. Suppose that z'■ Ikj(z) is a GEM* iteration. Then there is a constant K depending only on I such that: If Bj^,> 0 and |/(xb)| > L> 0 then there isank given explicitly such that |/(z.)|
Proof. The proof is the same as [S-SI, Thm. 4]. Take a to be the min of the a considered there and the a of Lemma 5 above, and replace \/M by 6 />l0 . Remark. It is easy to adapt the algorithms of Theorem A and B to GEM* iterations and prove analogous theorems. We state a particular theorem which is the analogue of the "Main Theorem" of [SS-I] and which has the same proof starting from Theorem C. THEOREM E. Ifz'- IKf(z) is a GEM* iteration, there are positive constants K,, K2 depending only on P(h) with the following true: Given
1229 ON THE OEOMETRY OP POLYNOMIALS AND A THEORY OP COST
159
known I, to zero* of / One obtains good algorithms by following these curves simultaneously. We suspect that the speed in this cue can be understood by the sethods in these papers. t4) We have used the space
J»*(l)-{(«»,*".*-,.*)lkl*1.4-l) of coefficients of/(z) - J f - o V*. While this parameterization has a simple immediacy, the protective space C - *7(C - 0) is more natural. Here (a* ••-,«*) c C**• is equivalent to (Aflo. •*•.*«<). csch nonzero complex A. It would seem reasonable for the results to go over. Also we have used one particular probability measure, the uniform one. It would be useful to make a generalization to a wide class of probability measures, say given axiomatically as in Smale (1982b). Finally the problem suggest* itself to replace polynomials by other classes of comlex analytic functions. For example much of the analysis applies to rational functions which are also invertible up to the "first" critical value. (5) We recall a problem, yet open, from Smale (1981) which has received some attention. For any polynomial/, d e g / > 1, and complex number z,f(z) + 0, it seems likely that there exists a critical point 0(/(*)~O) such that
z-9 It is proved there that
1/(2) ~/(g) £ K | / ( z ) | with X - 4 .
min • I
z-l
ro-o One can express the conjecture in a slightly sharper form by making K a function of 4, K-K*md-l/d. This conjecture is the best possible as can be seen by choosing /(z) ~z* -dz, and z « 0. The conjecture is false for entire functions such as /(z) ■ e7. Dick Palais first pointed out to Smale that the estimate (X = 1) was true for / with real zeros. Nan Boultbee confirmed Smale's early calculations for degree / 5 4. And recently David Tischler (1982) has proved the conjecture when one root o f / is zero and the others have the same absolute value. Linda Keen and Tischler have produced some supportive numerical evidence, but as mentioned above, the general conjecture remains open. (6) We remark that although detailed techniques are different, there are basic similarities between the main theorems here and in Smale (1982a). Each gives a good estimate for the average number of iterations of a well-known algorithm or variation thereof. Moreover the underlying geometry of the algorithm in each case is following the inverse image of a segment in the target space. The work of Kuhn-Zeke-Senlin (1982) and Renegar (1982) on the speed of piecewise linear algorithms to find zeros of polynomials relates to both our paper here and Smale (1982a). (7) The algorithms in [S-SI] and in this paper start with z<>eC satisfying Izd » 1 , e g |z©|«3. This is necessary for our analysis since the rough behaviour o f / € P ^ ( l ) on points z0 with IzJ large enough is independent of / On the other hand the large starting value contributes eventually to the d in the estimates of the main theorems. This suggests that if one surted with |ze| 5 1 , e.g. z« » 0 at least that factor of d would be eliminated, sharpening the theorem drastically.
1230
160
M. SHUB AND f. SMALE
The problem here U to obtain information on the behavior of 8 ^ for |x| a 1 or even z «0. Consider the integral Q* of ©j, over the space P*(l) xSj with the usual uniform probability measure. Is there an e > 0 independent of d such that f}4 > t >0? A related question is: Do there exists universal constants «,,tj>0 such that \ft,4W (measure of {z€ S'|©/.> «i})"'<«j'- An affirmative answer would imply that the d in d log d of Theorems A and B could be eliminated. In the case of Theorem A this is very direct; Theorem B actually requires a slightly different algorithm. The idea is to switch to A * 1 at some point We develop such an algorithm a bit. Let p/«min(}c,}cp'/3) and./«./»- flog»+,(80. If |/(z,)|
cp4f m_cpj_ 2**2WK*»' (Id)*4
a
. '
J
If Pz 11 then |/(£1(Zo)l < c/ild)*4 < ts by a similar calculation. We are now ready to describe the Algorithm (N - E)': 0) m-1 1) k«max(riog
I
">(/) < C, log d for Cx a constant
This shows that Theorem B applies to Algorithm (N-E)'. The extra computations involved in increasing k do not seriously effect the total number of arithmetic operations either. REFERENCES A. BORODIN AND I. MuNao, (1973), Tk* Ctmifviatiomal Cemjttxitj mfAlfibnk and Numtrical PrvbUmx, American Eltcvicr, Ne* York.
1231 ON THE OEOMETRY OF POLYNOMIALS AND A THEORY OF COST
161
R. B U N T AND KUNO (I97»). 0 ( ( * log nf") algorithms for companion ami reversion of power term, la Analytic Computational Complexity,J. f. Traut, ad., Academic Proas. Nrw York, pp. 217-22$. (1*71), Foil algorithms for manipulating formal tenet, i AMOC. Coaput Mack., 2S, pp. $81-595. • . D U O N AND P. HENRICI (1969), CmiminlM Aspects of Ike Fundamental Theorem of Algebra, John Wiley. New York. t. HENRICI 1977, Applied and Computational Complex Analysis, John Wiky, New York. M. HIRSCH AND S SMALE (1979). On algorithms for totting f(x)~ 9, C o m . Put Appl. Math., 32. pp. 2SI-J12. D. KNUTH (I9tl), TV Art of Computer Ftogrtunmmg, VoL 2, Semimimerieol Algorithms, 2ad ad., AddisonWesley, Reading. MA. H. KUHK. W. ZEKE AND X. SENLIN (IM2), On the tmt of computing roots of polynomials, preprint. S. U N O (1965), Algebra, AddisonWaeley. Reading MA. J. RENEOAR (1912), On the complexity aft pitcrwtst hnear algorithm for approximattng roots of complex polynomials. A. SCHONHAOE (19t2), The fundamental theorem of algebra in terms ofcomputationalcomplexity-preliminary report. Math. last. der Uait. Tibingen M. SHUR AND S. SMALE (19S2), Computational complexity, on me gtometry of polynomials and a theory of eoti (Fan I), Ann Sctenl. Ec Nonn. Sup. 4 aerie. It (1915). S. SMALE (1981), The fundamental theorem of algebra and complexity theory, tall Amer Math Soc., (New Series). 4, pp 1-36 (1982a), On the average speed of the simplex method of linear programming, preprint. (1982b), The problem of the average speed of the simplex method, to appear in the Proceedings of the International Symposium of Mathematical Programming, Boon, 1982, Springer, Heidelberg. D. TISCHLER (1982I, preprint.
J. TRAUR AND H. WOZNIAKOWSKI (1980), A General Theory of Optimal Algorithms, Academic Press, New York. (1982), Complexity of linear programming, Oper. Res Lett., 1. pp. 59-62.
1232 JOURNAL OF COMPLEXITY 2 , 2 - 1 1 (1986)
On the Existence of Generally Convergent Algorithms MICHAEL SHUB* AND STEVE SMALE* IBM, Thomas J. Watson Research Center, Yorktown Heights, New York 10598, and Department of Mathematics, University of California, Berkeley, California 94720 Received September 6, 1985
1
To motivate our result, consider Newton's method N for solving the equa tion/(z) = 0, where/is a complex polynomial,/(z) = Sf_0 Oiz'. We write N: &d x S -+ S, where % is the space of polynomials of degree -&d and S is the Riemann sphere C U <*. Then N(f, z) = Nf(z) = z - /(z)//'(z) is ra tional over C in/and z; that is, N can be formed from the complex rational operations (+, - , x , -=-) from the coefficients of/and z. If z is sufficiently close to a zero £ of / , then the iterates zk = Nj (z) converge to ( as k tends to °°. However, as is well known there is an open set U in 3j, x C (if d > 2) such that this convergence will not happen for (/, z) in U. See, e.g., Smale (1985). In this paper it was conjectured that no such algorithm could be generally convergent. Curt McMullen settled the question by proving the following result. THEOREM (McMullen). Let d > 3 and T: &d x 5 -*■ S be any map ratio nal over C in f and z. Then there is no open set U C % X S of full measure with this property: / / ( / , z) 6 U, then Tkf{z) = zk converges to a root of fas k~* °°.
Here a "set of full measure" means one whose complement has Lebesque measure zero. McMullen's result can be paraphrased as saying there is no generally convergent purely iterative algorithm, rational over C, for finding roots of polynomials of degree ^ 4 . Here "purely iterative" means that the algorithm can be expressed as a discrete dynamical system on 5 parameterized by the polynomial. Equivalently, the algorithm is one point stationary. The goal of this paper is to show that if one adds the operation of complex *We would like to acknowledge partial support from the National Science Foundation and express our appreciation to IMPA, Rio de Janeiro, for its hospitality. 2 0885-064X/86 $3.00 Copyright C 1986 by Academic Press, Inc. AHrightsof reproduction in any form reserved.
1233 GENERALLY CONVERGENT ALGORITHMS
3
conjugation, then there do exist generally convergent purely iterative algo rithms for finding zeros of polynomials. This gives a complement of McMullen's theorem. Moreover our theorem works for n variables while McMullen's result, which depends on a recent one-variable theorem of Mane, Sad, and Sullivan (1983), remains unproved for two or more variables. THEOREM. For any d, there is a map T: &d x S —> S formed from the complex rational operations and complex conjugation from the coefficients of f E % and z G S with the following property: there is an open set of full measure U C % X S such that if (/, z) E U, then the iterates zk = T) (z) converge to a zero off.
Theorem 2 (complemented by Theorem 1) in Section 2 below is a slightly sharper version of this result and moreover contains the n-variable case. In Section 3 we give another example of a generally convergent purely iterative algorithm which is presumably more efficient. This second example is a modification of Newton's method so that it has quadratic convergence near a zero of a polynomial of multiplicity one. However, this algorithm uses square roots of positive numbers as well as complex conjugation. Also, we have not been able to extend the general convergence proof to more than one variable, thus leaving open a problem which seems to us important and challenging. Of course there is a long history of results related to our work, a few of which are mentioned in Smale (1985). Also, there are works of Kim (1985), Hirsch and Smale (1979), Murota (1982), Wasilkowski (1983), and Wisniewski (1984). In Kim (1985) an algorithm similar to the one-variable case of Section 3 is proposed and studied with respect to general convergence.
2
Let 9d be the linear space of all polynomial maps C" —*■ C" of degree -^d, d > 1 (more abstractly one could say: let E, F be complex Hilbert spaces of dimension n and % the space of all maps E —» F whose (d + l)st derivative is identically zero). Let Ud be the subset of 9^ of those / : C" -» C" which satisfy these three conditions: (a) The dth homogeneous parts of the coordinate functions / , i = 1, . . . , n, off have no common zeros except the origin. This implies that/is proper (see Hirsch and Smale, 1979, for example). (b) If/(z) = 0, then the derivative Df{z): C" -> C" is nonsingular (our calculus notation follows Lang, 1983). (c) The map g: C" -* R defined by g(z) = \\f(z) ||2 is a Morse func tion. Here we use the Hermitian inner product and norm on C". A Morse
1234 SHUB AND SMALE
4
function is one with nondegenerate critical points. (Milnor, 1963, is a good reference for Morse theory.) Note that (c) implies grad g has finitely many zeros. THEOREM 1. Ud is an open set of&d containing the complement of a real algebraic subvariety; thus Ud is an open set offull measure. Proof. First work over the complex numbers. Let A C 9i be the set of/such that the dth homogeneous parts of the/ have a nontrivial common zero. Let B C Sj, be the set of/such that the/ and Det Df(z) have a common zero. By elimination theory of algebraic geometry (see Van der Waerden, 1950, p. 15), A and B are each contained in algebraic subvarieties of 9»d of complex codimension 1. See Renegar (1984) (also Smale, 1981) for this (in particular Renegar's Proposition 5.1). Thus it remains to deal with (c). For this, the same procedure works, using the real numbers instead of the complex numbers. The equations (polynomial, real) this time are given by (i) Dg(z): U2n-+U,Dg(z) (ii) DetD2#(z) = 0.
= 0 (real derivative), Q.E.D
For/ E 9j/, define an endomorphism 7}: C —» C by 7}(z) = T(z) = z-hz
grad g{z),
g(z) = ||/(z) ||2,
where £- = Z l 1
+
^i— H1
+
llgradg(z)||2)' 2.
Here || ||o denotes the sum of the squares of the corresponding components, which is greater than or equal to the operator norm squared || ||2. The follow ing argument shows this. Let V, W be inner product spaces. Express L (V, W) as matrices with respect to an orthonormal basis of V and W. Let || ||E denote the Euclidean norm and || |op the operator norm. If A: V—* L(V, W) is linear, and L(V, W) has the operator norm, then the multilinear norm of A is the operator norm of A, HI, ||A (lollop HAtaXt^ll p IIA |op = sup ,, ,, = sup ■\ ||t>i| B ,f, \\V\\\ «2 2
Now since || ||E ^ || Hop on L(V, W) the operator norm of A: V—* LE(V, W) is ^ operator norm of A and ||A||E ^ ll^llop- Now induction finishes the argu ment.
1235 GENERALLY CONVERGENT ALGORITHMS
5
THEOREM 2. For each f E Ud, there is an open set Vf C C of full measure such that for z E Vf, Tk(z) = zk converges to a zero offas k —* o°.
The proof of Theorem 2 uses two propositions. PROPOSITION
1. Let z E C with grad g(z) * 0 and let z' = T{z). Then
g(z')
Expand g by a Taylor series about z, and evaluate it at z' = T(z),
g(z') = g(z) - Ajgrad g(z) I2 + 2 ( - ^ ) ' ^ T ^ grad *(*)*'• 1=2
'■
Then Proposition 1 is a consequence of this lemma: LEMMA
1. //grad g{z) * 0, f/wTt /i r |gradg(z)| 2 >
2-l'^^(grad^(z))' i=2
Proof of Lemma 1. Dividing by the left-hand side, it is sufficient to show Aj^A|-2!l^iWJ!||gnidg(z)||,-2
i=2
(Here we use the operator norm on D'g(z); cf. Lang, 1983.) Since hz < 1, this amounts to d
/!j2
II D'e(z) II
^lM||gradg(z)||.-2<1
1=2
'•
The last follows from the definition of h2, the fact that 1 + JC2 > x for any x > 0, and the fact that \\A\\ < ||A||0. PROPOSITION 2. Ler/ E Ud, 6 E C satisfy f (6) * 0, and grad g(0) = 0. 77ien f/ie ser WJ(0) of all z such that Tk(z) —* 0ask^> °° /ws measure zero.
For the proof we use some lemmas. LEMMA
Proof
2.
8 is not a local minimum of g.
This is a consequence of the maximum principle.
LEMMA
3. DT{8) has an eigenvalue greater than 1.
Proof
DT(8) = I - heD2g(8), so Lemma 3 follows from Lemma 2.
From the center manifold theory (see, e.g., Hirsch, Pugh, and Shub, 1977), it follows that there are arbitrarily small neighborhoods U of 8 such
1236
SHUB ANDSMALE
6
that W'(6) fl U is contained in a closed set of measure 0, in fact a differentiable disc of codimension one Wf(6, U) = Wcf(U) = WC(U) with the property that T-l(We{U)) H U C WC(U). Next note that T is real algebraic and nondegenerate in that its image contains an open set by checking near the roots as below; in particular the Jacobian determinant det(D7) vanishes on a real subvariety of codimension one. Proposition 2 now follows. W'(0) C Wc($) = Uo T-k(Wc(U)), which has measure zero since it is the countable union of measure zero sets. Now for the proof of Theorem 2. Let/ G Udandg = ||/|p. If z0 is a critical point of g let W'(z0) be the set of z E C" such that Tk(z) -+ z0 as k -*■ «. Then define Wf =
U
W'(z)
and ty = C" - Wf.
gr»d«tt=0 f(z)*0
Let z E Vf. We claim that z* = Tk(z) converges to a zero of/as it -* ». By property (a) of/(since/ E t/d) the set of z4 is bounded; therefore by Propo sition 1, zk must converge to a zero of grad g. Since z £ Wf, this zero of grad g is also a zero off. By property (c) of/any zero ZQ of/is a sink of —grad g; that is, all the eigenvalues of — D(grad g)(z0) are negative. By the definition of hz all the eigenvalues of —h!(jD grad g(zo) are negative but greater than —1. Thus all the eigenvalues of DT(z0) = I — hI0D grad g(z0) are greater than zero but less than one. Thus z0 is a sink for T. This shows that W'(z0) is open, and Vf is open. Moreover, there is a disc D0 around z0 mapped into its interior by a contraction for any g in Uj close enough t o / It follows by continuity that if z E W(z0) and/ E Uj then (/, z) is in the interior of U = {(/, z) | / E (/rf and z &V{}. Thus (/ is open and of full measure in % x C\ Q.E.D 3 Let/be a polynomial of one variable,/(z) = So a,z', z E C U °° = 5. Define
_ Kz
2*'(|z|)2|/(z)lll/IU'
where d
and
U/H^ = max|a,-|.
Let p(z) = min(l, it,) and define Tf:S^S
by Tf(z) = z - p(z) x
1237 GENERALLY CONVERGENT ALGORITHMS
7
(/(z)//'(z)). Here note that the max and min of positive numbers may be expressed in terms of square roots, e.g., . ,. la — b\ \a + b\ max(fl, b) = ' 2 ' + ' 2 ',
.n rrx . ,, V ( A - b)2 = \a - b\.
Next let Gj be the space of polynomials of one variable, degree ^d, with zeros and critical points all distinct. THEOREM 3. Letf E Gd. Then there is a closed set Wf of measure zero such that if z & Wf, then Tkf[z) converges to a zero of f as k tends to «>. Moreover 7} is Newton's method in a neighborhood of each zero off.
Proof. The last statement follows from the definitions. We now prove the rest. Define for each z E C , / a polynomial with/'(z) # 0, a(z,f)
= sup *&2
PROPOSITION
/'(z) ±0.Leth
fk)(z)
i/(* - i)
\f'(z)l
*! /'(*)
3. Let f be a polynomial and z E C, with f(z) # 0, satisfy 0 < h < l/2a,
a =
a(z,f).
Then for z' = z-hj^y
|/(z')|<|/(z)|.
Proof. Expand/by Taylor's series about z so
f(z>)
= (i - H)f(z) +
ih^irij
and
< 1 - /» + /i 2(/ia)'-' i-2
-(r^)
< 1 - A + M M ( - — h r ) < 1.
Q-E.D
1238 8
SHUB AND SMALE
This generalizes easily to polynomial maps from one Banach space to another. PROPOSITION
4.
(a) Let f be a polynomial, z G 5, and
Then
,
n
- !/(*)! ll/IU»'(l* I)2
a(Z f)
' -
|/W*(|x|)
•
(b) Let r > 0. Then 4>{rW(r) >(r\i < 1. 24>'{r) This is proved in Smale (1986), where in fact (a) is proved for Banach spaces. PROPOSITION 5.
Let f G Gd,
f'(6)
= 0,
= {z \ Tk(z) -> 0,
W'(0)
as
it —* °°}. Then W'(0) is a closed set of measure 0. Proof. By the argument of Proposition 2 of Section 2, it is sufficient to prove Proposition 5 locally, in a neighborhood of 0. To that end we calculate the derivative DT{0): R2 -*■ R2, where R2 is just C. For z in a neighborhood of 0, p{z) = k2 and we may write T(z) as T{Z)
Z
<MH)/(z)
2*'(|z|) 2 ||/|| m |/(z)| / ' ( z )
so
DW,W
-■ -
mrnmm\{Dru))-M
for v G R2. Now Dif^i^v) = FW)v. Thus the linear map DT(0) has the form DT{0)(v) = v - /3tJ, where
ft = AoJFWtd
0\)/2
4. The linear map R2 —* R2 given by v —* /3tJ has trace 0 and determinant —1/3|2 ^ 0. Thus its eigenvalues are ±|/3|. LEMMA
The proof is simple and direct. From the lemma it follows that the eigenvalues of DT(6) are 1 ± | jS |. But ln<
l/"(0)l<M|fl|)^l/"(0)l
""
2 | | / | U ' ( M ) - ||/u
2
i
1239 GENERALLY CONVERGENT ALGORITHMS
9
by Proposition 4(b), which can be seen to be less than 1. Therefore 6 is a saddle point for 7), proving Proposition 5. For the proof of Theorem 3, note that kz < 1 /2a (/, z) by using Proposition 4(a). Thus Proposition 3 applies to 7). Now the same argument used in the end of the proof of Theorem 2, using Proposition 5, yields Theorem 3. One needs to remark that the operations involved in the definition of 7}(z), besides the complex rational operations only require complex conjugation and the square root of a positive real number. The preceding arguments for Theorem 3 extend to the n-variable case except for the local argument of Proposition 5. We give a short discussion. The Newton vector field N(z) = -Df{z)~lf(z) is not generally globally defined on C because Df(z) may not be invertible. We desingularize N as follows: Given the n x n complex matrix A, let A be the n x n matrix whose (/, y')th entry is ( - \)'+J det A,,, where An is the (j, j)th cofactor of A. The standard proof of Cramer's rule for inverting a matrix gives AA = AA = (det A)I.
(a)
Now define ft(z) = -Df(z)f{z). ft(z) is globally defined, and
, * M) = N(z).
dett Df(z)
Note that #(z) is zero in the following cases. (i) Df(z) has corank 2 or more; then Df(z) is identically zero. (ii) Df(z) has corank 1 and/(z) £ Image Df(z) = kernel Df(z). (iii) /(*) = 0. DEFINITION.
Let
A )
2||/(Z)||||/|U^'(||Z|HD/-'(Z)|P'
where | | / | U = max,(||D'/(o)||/i!) and let p(z) = Pf(z) = min(l, K,{z)). For a polynomial/such that Df(z) is invertible at the roots off, p(z) extends continuously to be identically one in a neighborhood of the roots of/and p(z)||D/"'(z)|| extends to be zero on the variety 2 of z such that Det Df(z) = 0. Now let T(z) = 7,(2) = z -
p(z)Df(z)-,f(z).
For/with nondegenerate roots, T is Newton's method near the roots of/and the identity on X. Near X,
1240 SHUB AND SMALE
10
T(z) = z - Kf(z)Df-\z)f(z)
= z +
h(z)DctDf(z)ft(z),
where »(r).
* «
^
.
2ll/(r)IHI/IU*'(H)WMf Kf{z) ^ \/2a(z,f), so Proposition 3 still applies. Add the additional hypoth esis that/is proper. Then the crucial question for the global behavior of T is the nature of the set of points which tend to X under the iteration of/. PROBLEM 1. For all / i n the complement of an algebraic subvariety of % of codimension ^ 1, is it true that W(2.) = {z \ Tk(z) —*■ Xas k —* +»} is in a closet set of measure zero? If we assume that 8 G X, that Df has corank one at 8 and is transversal to the corank one matrices there, and moreover that/(0) g Image Df{8) and Ker Df{6) is not tangent to X then T'(8)v = v + h(8)L(8)vft(8), where L(0)isalinearmapL(0): C"-> C,h(8) # O,and#(0) # Oandthus7'(0) has an eigenvalue larger man one. Now the theory of partially hyperbolic fixed points (Hirsch, Pugh, and Shub, 1977) shows that locally near 8, W'(1) has measure zero; i.e., there is a neighborhood U of 8 such that {z EU\ f{z) E U for all k > 0 and/*(z) -► X} has measure 0. This takes care of most of the points in X, but generically there are points not satisfying these hypotheses even for two variables. PROBLEM 2. If 8 E C" and Df(8) is singular, under generic conditions on /, is the set of z such that Tk(z) -* 8 as jfc —» » contained in a closed set of measure 0? REFERENCES HIRSCH, M., PUGH, C , AND SHUB, M. (1977), "Invariant Manifolds," Lectures Notes in Mathematics, Vol. 583, Springer-Verlag, New York. HIRSCH, M., AND SMALE, S. (1979), On algorithms for solving/M = 0, Comm. Pure Appl. Math. 32, 281-312. KIM, M. -H. (198S), Ph.D. thesis, C.U.N.Y. Graduate School, to appear. LANG, S. (1983), "Real Analysis," Addison-Wesley, Reading, Mass. MANE, R., SAD, P., AND SULLIVAN, D. (1983), On the dynamics of rational maps, Ann. Sci. Ecole Norm. Sup. 16, 193-217. MCMULLEN, C. (1985), "Families of Rational Maps and Iterative Root-Finding Algorithms," Ph.D. thesis, Harvard University, May. MILNOR, J. (1963), "Morse Theory," Annals of Mathematics Studies, Vol. 51, Princeton Univ. Press, Princeton, N.J. MUROTA, K. (1982), Global convergence of a modified Newton iteration for algebraic equa tions, SIAMJ. Numer. Anal. 19, 793-799. RENEGAR, J. (1984), On the efficiency of Newton's method in approximating all zeros of a
1241 GENERALLY CONVERGENT ALGORITHMS
11
system of complex polynomials, Department of Mathematics, Colorado State University, Sept., preprint. SMALE, S. (1981), The fundamental theorem of algebra and complexity theory, Bull. Amer. Math. Soc. (N.S.) 4, 1-36. SMALE, S (1985), On the efficiency of algorithms of analysis, Bull. Amer. Math. Soc. SMALE, S. (1986), On the convergence and efficiency of Newton's method, in preparation. VAN DER WAERDEN, B. (1950), "Modern Algebra," Vol. II, Ungar, New York. WASILKOWSKI, G. (1983), Any iteration for polynomial equations using linear information has infinite complexity, J. Theoret. Comput. Sci. 11, 195-208. WISNIEWSKI, H. (1984), Rate of approach to minima and sinks—The Morse-Smale case, Trans. Amer. Math. Soc. 284, 567-581.
1242 NEWTON'S METHOD ESTIMATES F R O M DATA AT O N E P O I N T STEVE
Department University
of
SMALE
Mathematics
of California,
Berkeley, California
Berkeley 94720
In honor of Gail Young
Newton's method and its modifications have long played a central role in finding solutions of non-linear equations and systems. T h e work of Kantorovieh has been seminal in extending and codifying Newton's method. Kantorovieh's approach, which dominates the literature in this area, has these features: (a) weak differentiability hypotheses are made on the system, e.g., the m a p is C2 on some domain in a Danach space; (b) derivative bounds are supposed to exist over the whole of this domain. In contrast, here strong hypotheses on differentiability are made; analyticity is assumed. On the other hand, we deduce consequences from d a t a at a single point. This point of view has valuable features for computation and its theory. Theorems similar to ours could probably be deduced with the Kantorovieh theory as a starting point; however, we have found it useful to s t a r t afresh. T h e results in this paper are important for our construction of global algorithms based on Newton's method, and for estimation of the efficiency of those algorithms. T h e idea is simply to apply the theorems here to a finite sequence of equations of the form f{z)
- t,f[zo)
■= 0, 0 < tt < 1, to solve f(z)
-- 0. T h a t development will l>e
exposed ill a subsequent article. T h e classical work can be found in Kantorovieh and Akilov (1001). as well as Ostrowski (1073).
More recent work in this direction includes Cragg and Tapia
(107-1), Kail (I07-1), and Trauh and Wozniakow.ski (1070). Some or Hie spirit here is reflected ill Smale (1081, 108.1). Important for me have been conversations with Mike Slinh and our two joint, papers (108.1, ION(i). Trying to understand I In- important
1243 1X6
work of Renegar [8] was a strong motivation for this paper. Ideas of Kim (1985) are closely related also. S e c t i o n 1. To indroduce the ideas, we give some of the results first for the simple case of a single polynomial. Consider a complex polynomial / ( * ) = X ) 0 a t 2 ' . To solve / ( f ) — 0, following Newton's method, let «o £ C and inductively z* = * * - i - / ( * * - 1 ) / / ' ( « * - 1 ) Precision doubling motivates the following definition. D e f i n i t i o n , ZQ C C is an a p p r o x i m a t e s c r o of f \\zk-zk-l\\<(\)2k"-1\\*i-«>l
fc
provided = l,2,... .
Note the ultrafast convergence due to the exponent. Toward giving a test for an approximate zero, we define a function o(z, / ) which is central to our account. Definition.
«(*,/)
/(*)
sup k>l
Here, /'*'(«) is the kth derivative
/<"(*) «/'(*)
l/k-l
of f it z.
Theorem (Special Cases). ( A ) There is a constant cto between 1/8 and 1/7 such that if a(z,f) an approximate
(D) Q(z n < 11/11
< oro, then z is
zero of f.
-!/WL M i l 2
(B) a U , / ) < | | / | | m a » | / f ( z ) | 2 M]z]) Here, ||/||m.-»x = s u p | o , | and
R e m a r k . The theorem gives two tests for ensuring that Newton iteration converges fast starting immediately and continuing indefinitely. Combining (B) with (A) gives a test, while cruder than (A) alone, involving only the first derivative. These tests ran be used to terminate zero finding algorithms, as well as for constructing global algorithms. Kim (1985) has independently given a proof of Theorem (A) with r»o
1/54.
Since her proof uses the theory of Schlicht functions, it does not extend to several variables. The theorem ran be used to study the likelihood of z being an approximate zero lor / as follows, bet /',( be the spare of polynomials / of degree less than or e«pial lo whose roellirients satisfy \n,\ •'• I, i
0
tl. Then /*( i
use .is a probability measure on /'j normalized behesgue measure. I,et ;>(.?,<') be I lie probability that z is an approximate zero of / drawn randomly from /'/.
1244 187
Corollary. That z — 0 ia an approximate sero for / has a probability bounded below by a positive constant e independent ofd. That ia, p(0,d) > e > 0. Proof. Put * = 0 in (B) of the theorem to see that |oo|/|ai| 2 < «o implies 0 is an approximate zero of f(z) = X3oa'*'- But |ao|/|ai| J is independent of d. Remark. In fact one computes c = a*/3 by integration. The corresponding bound for real polynomials is ao/3. After I found the above result I raised the question for general z. Rcnegar then showed using the above theorem that p[*,d) > e(\z\) > 0 if \z\ < 1 and showed that no such bound existed, independent of d, with \z\ > 1. Section 2. A standing hypothesis in this section is that f : £ -* 7 is an analytic map from one Banach space to another, both £ and 7 are real or both are complex. Main examples are the finite dimensional cases £ = C n , 7 — C , where n could be one as in Section 1. The map / could be given by a system of polynomials. Our calculus notation follows [5] or [1], where the basic theorems of calculus on Banach spaces which we use can also be found. Taylor series expansions of an analytic map are the most important of these. The derivative o f / : £ - » / a t * e £ i s a linear map Df[z) : £ -» / . If Df(z) is invertible, Newton's method provides a new vector z' from z by *' = * - Df[z)-*f(z)
= N,[z) = N(z).
Let 0 denote the norm of this Newton step z' — z, i.e., 0{zJ)=0{z)
=
\\Df{z)-lf(z)\\.
In case Df(z) is not invertible, let 0(z) = oo. For a point «o 6 £, define inductively the sequence zn = *„-i Df(zn-i)~'/(*«-1) (if possible). Say that *o is an approximate *cro of / if zn is defined for all n and satisfies:
II*--*—ill <(J)"" , -'|l*i-«oll.
»»•»•
Clearly this implies that z„ is a Cauchy sequence with a limit, say, f f /(<;) r 0 ran !><• soon as follows. Sinro ;„ , \
f.
That
zn = /•'/(in) '/(-«)•
ll/(*»)ll = ||D/(* n )(*» + , - zH)\\ < \\Df{zn)\\\\zn„
- zn\\.
Tako the limit as H —• oo, so
ll/(<)ll'-||«/(c)l|Mri.||*.,,
a.||
0.
Nolo that for an approximate zero, Newton's tuotliod is supi-n onvrrnrttl sl;«rlIIIR with tin- lirsl iteration. There is no l.iri'.r ronslanl on the rii'lil II.HHI .i«l
1245 IXX
Proposition 1. If zo is an approximate zero and zn —» ( as n —» oo, then
ik»-fii<(ir"'iki-*bii^ wAere
For the proof, sum both sides in the definition of approximate zero. ||rw~x,||<
£ n
| | x B - x n . , | | < ||x, - x o | | f ] ($)*""' '
Let N -» oo and factor (1/2)*
n -1
rI I
from the right-hand side to obtain Proposition 1.
Toward giving criteria for * to be an approximate sero define
7(«, / ) = »«P
«
*'«"
or, if Z?/(z)~' or the sup does not exist, let *i(z,f) = oo. Here £>*/(*) is the /bth derivative of / at z as afc-linearmap. See [l] or [5]. Abo, Df(z)~lDkf(z) denotes the composition. The norm is the norm as a multilinear map as defined in those references. We sometimes shorten i(z, f) to *i(z) or 1 in situations which make the abbreviation clear. Define a(z, f) = 0{z, f)i{z, f) where 0 is defined earlier in the section. Theorem A. There is a natura//y defined number a0 approximate/y equa/ to 0.130707 such that if a(z, f) < a0, then z is an approximate zero of f. The issue of sharpness is discussed later. Suppose now / : £ -» 7 is a map which is expressed as f(z) = X^_ 0 «***, *ll z € £ ,0 < d < oo. Here £ and 7 are Banach spaces and o* is a bounded symmetric it-linear map from £ x • • • x £ (k times) to 7. Thus a*** is a homogeneous polynomial of degree k, so to speak. For £ - C , this is the case in the usual sense, and in one variable a* is the fcth coefficient (real or complex) of / . Then if d is finite, / is a polynomial. Define 11/11 = 8«P ll«*ll *>o
(general case)
where ||
..n.l 4>(r)
J
1246 ]**
Theorem B .
^XIWW-'IIII/-^,,,, Here, if Df(z) is not invertible interpret \\Df[z)~x || = oo as usual. Thus combining Theorems A and B, we have a first derivative criterion (at z) for z to be an approximate zero. Corollary. If \\f\\\\Df(z)-1\\%^0(z,f)
ll*--fll<(J),"",H«»>-fll
for n = 1,2,3,... _,
*n = *n-l - D / ( « „ - l ) / ( « n - l ) While the first definition of approximate zero deals with information at hand, and computable quantities (in principle, eventually), an approximate zero of the second kind can often be studied statistically or theoretically more handily. Theorem C. Suppose that / : £ -» 7 is analytic, f C. £, /(f) - 0 and z f £ satisfies
Then z is an approximate zero of the second kind. This jesult gives more evidence for the importance of the invariant i(f). Section 3. Here we prove some lemmas and propositions from which the main results will follow c;isily. Suppose £ and 7 arc Manach spaces, both real or both complex. Lemma 1. Let A,B : £ -> 7 be bounded linear maps with A invertible such that ||/I ' I) - l\\ < c < 1. Then B is invertible and \\n~l A\\ < 1/(1 r). Proof. Compare [I, 8.3.2.1]. Lot v ■ I - A* 11. Since ||i>|| - 1, ^ . T •'' r x i s l s with norm < 1/(1 c). Note (/ - u ) ^ 0 % " ( / - v) ■- / t / n " . Uy t;ikinK limits, A ' li I v is seen to lie invertible with inverse ^3!i° "'• • t! '" rr ' '1 ' " ^ , 7 •'' '• // ■ nn l>r will It'll its tilt' i ouipo'ition of invert idle in.ip'.. I
1247 IWI
Lemma 2. Suppose f : £ -» 7 is analytic, z\z
e £ such that \\z' - z\\*t(z) <
1 - \Z5/2. Then fa; Df(z')
is invertible. l
(b) \\Df{z')- Df{z)\\
<
2-*(||*»-,«,(«)) 3
(C)
Tl* * ™ 2-*>(\\z>-zUz))
(l-|l*'-*h(*))
Here 4>'(r) = 1/(1 - r ) s could be replaced by
Proof. Take a Taylor series expansion of Df about z (i.e., the map £ — ► L(£,7)) as follows (see [1], [5]).
» * i - £ ^ <••-.>' From this,
and
||D/(z)-«D/(z')-/|| < £ > + !)
{* + !)!
*=i
+1
)H*)"Z'-*II
<0'(-,(z)||z'-z||)-i. Observe, that since -T(z)||z' - z|| < 1 - v/2/2 a |i the series converge. Moreover, note that
0(0 Moth quantities on the right are seen to be equal to the £th derivative. V2
I W I \ *»(r)M(| r ) J
\\ IHMT
l/>(r)
2r2
\r I I.
I V'(r)
1248 IVI
Now we prove (c) of Lemma 2. Let 1k = lk{z) =
Dfi'Y
i P(*}/(*)
t/t-i
and
fc!
i — sup ik-
Then by Taylor's theorem (see [5])
P /(.r' P /c)E
p/M D>
" ;'/" ) " , -' ) '
«=o
'
t
< l | P /(.r- D /MiiE( ^ii
'
p/M
'',i /( ,f-^
lk{t,)
j
x(k+l)/(k-l)
- 2 - ^ C v l l * ' - *ll) V* —»ll** —
The supremum is achieved at Jk = 2, yielding the statement of Lemma 2c. Lemma 3. (a; Leta = a(z, f) < 1 and x> = z - Df(z)~xf(z),
0 = /?(«, / ) . Then
1
\\Df{z)- f(z')\\<0(T^). (b) LetzyeC
with f{z) = 0, and ||*' - z\\-y{z) < 1. Then
mvw W)i < n i ^ H ^ We first prove part (a). The Taylor series yields
*=o
The first two terms drop out. Since 0(z) = \\z' - z\\,
l|o/W-I/{*')ll < /»5>fl* _I < 0vQn k 2
proving (a). For part (b), wc start as above and now the first term is zero since f(z) Wc have
i D/(*)-'/(*') ii < ii(*' - *)n(i + J > J - v - *n*~') »r2
ll»' * l l » l l * ' I
Tliis finishes the proof of l,emmn 3.
-y||*'
0.
1249
Proposition 1. Let a = a(«) = <*(«,/), 0 = 0[z), 0' = ^(z') where f : £ -* T is an analytic map from the Banach space £ to 7 as usual. (a) if a < 1 - v ^ / 2 , then
'**(T==)(i=iiw) (bj if/(z) = 0 and -r||*' - <|| < 1 - \Z2/2 then
Proof of Proposition 1. Write 0(z') = \\Df(z')-ll{*')\\
= ||D/(« , )- 1 D/(«)I?/W- , /(«')ll
<||D/(«')-I0/(*)||||D/W-,/(«')ll
The last step uses Lemmas 2b and 3a. Similarly for the second part of the proposition. If f(z) — 0, use Lemmas 2b and 3b as follows. 0(z')<\\Df{z')-iD/(z)\\\\Df{z)-lf(z')\\ ~ 2 - * ' ( | | * ' - z h ( z ) ) "*' " *" (1 - \\z> -
zWz))
This proves Proposition 1. Proposition 2. Recalling tp(r) = 2r2 - 4 r + l and using the notation of Proposition 1, (a) if a < 1 - v/2/2,
(b) iff(c) = 0 and \\z - f||-r(f) < 1 - >/2/2, a,^
For the proof of (a), note that a'
7(f)ll«-fll 0(l(f)ll*-fll) 2 ft'l'-
Use Lomma 2r and Proposition la to
nlilain
Tli<> proof of Proposition 2l> is similar IIS'IIIR lamina 2r and Proposition Hi. Tin following is |>rov<>
1250
Proposition 3. Suppose that A > 0, a, > 0, » = 0 , 1 , . . . satisfy: an i < Aa,, a//1. Then a* < (i4o 0 ) 2 ' - , ao, «H fc. We end this section with a short discussion of sharpness. Lemma 2b can be seen to be sharp by taking z — 0, f(z)=2z-
+ l-a,
0
Then z' = a, 7(0, / ) = 1 and \z' - z-y(0, f)\ = o at * = 0. The same example may be used to see that Lemma 2c is sharp. One only needs to make the easy computation that
*»w= (rb) 3 (2^y) Again, the same example shows that Lemma 3a is sharp, just observing that o(0,/) = a , / ( a ) = o V ( l - « ) . One can see that Lemma 3b is sharp with the example f(z) =
* = 0,
0 < z' < 1.
Proposition la is sharp with the example of Lemma 2. The same applies to Proposition 2a. Section 4. In this final section we finish the proofs of our main results. Toward the proof of Theorem A, consider our polynomial t/»(r) = 2r s - 4r + 1, and the function (a/xl>(a)) of Proposition 2a. In the range of concern to us, 0 < r < 1 - \Jij2, 0(r) is a parabola decreasing from 1 to 0 as r goes from 0 to 1 - v^/2. Therefore r/V»(r)2 increases from 0 to 00 as r goes from 0 to 1 - v/2/2. Let a0 be the unique r such that r/ip(r)3 = 1/2. Thus ao is a zero of the real quartir polynomial V'(r)7 - 2r. Using Newton's method one calculates approximately «o = 0.130707. With this discussion, Theorem A is a consequence of the following proposition where a --- 1/2. Proposition 1. Let f : £ -* 7 be analytic, z - z 0 f £ and «(*)/t/'(«(c)) o < 1. Let zk - z*_, - D(zk_x)"'/(**1), *=-• 1,2 then (a) zk it defined for all k. Lot Qfc = (t(zk), 0* - il>(a(zk)), k --- 1,2 (b) ak
**
=
1,2,...,
1||||.«n*.
Proof. Note that (a) follows from (l>) and that (l>) is a consequence of I'roposil ions 2;i nml .'{ of Section .'I. Il remains to chock (c). The case it I is trivial so assume k I.
1251 I'M
Wc may write using Proposition la and our relation between <j>' and 0 , II II ^ II II a n - » U ~ ° n - j ) ll*n - *n-l|| < ||*n-l - *n-*\\ ;
Vn-2
Now use part (b) and induction on this inequality to obtain n - «n-ill < a ' " " - ' ! ! . ! " * o | | < » 2 " " - , a ( * o ) ( L - ^ ) K a * "
1
«(*>)
- ^ - * , ^
But ~—V>n-2
<-ro<\VO
R e m a r k . Theorem A and most of the lemmas and propositions of Section 3 can bo slightly sharpened in case that / is a polynomial map £ --> 7 of Danarh spares of decree d < oo.
Replace (p(r) by
conclusions. For example, going through proofs this way yields the following generalization of Proposition 2a of Section 3. P r o p o s i t i o n 2 . If f : £ —• / has degree d, then (using notations ,
2 &>-2(<»)/
of Section 3) if
If d — oo, this reverts to Proposition 2a, Section 3. E x a m p l e . The following shows that a0 must be less than or equal 3 - 2\/2 in Theorem A. Let / 0 : C -» C be /„(z) = 2z - z / ( l - z) - o, o > 0. Then a(0, /„) = a and /„(f) = 0 where t = ((1 + a) ± v/(l + o ) 2 - 8 a ) / 4 . If a = a > 3 - 2v/2, these roots are not real, so that Newton's method for solving / 0 ( f ) = Oi starting at zo = 0 will never converge. Toward the proof of Theorem C, wo have: P r o p o s i t i o n 3 . Let f : £
> T,c,zC
£ satisfy /(<)
0 ami iU)\\z
c\\ < I
y/'i/'l.
Then r\ \\nf{z)
')
l « , l I,
/
A\\<
^C)H*'
Cl1
'
W-(*-c)ll^^(f)|„...f||)
Note that this proposition gives an estimate on how well the Newton vector '-*/(')
' f(') approximates <
z, the exact vector from z to <.
I'Or I lie proof wc consider the two Taylor series:
/(-> L ' T " o'
1252 W
Now apply the second to (z — f) and subtract it from the first to obtain:
Then
W W ' / W -{'-()
= -Df{z)-*DfU) J > - I) P 'fr>~ lo yri('-f>*
Take norms and apply Lemma 2b to obtain:
II0/M-7M - (« - Oil < ( j r ^ ) ) ( D * " »)«*-*)!l« - ell where to = i{s)\\z - f||. The right hand is estimated by ||z - $\\tv/ip(w) proving the proposition. Corollary. Suppose f, (, z are as in the proposition. Let 1=
-rtolk-fll g . tf(-r(?)ll*-?ll)
or equivaiently
Then
lf«»-Cll<-**"- , ll*-ffH where * = «o, *» = *»-« - 0 / ( * » - i ) _ , / ( * i » - i ) This follows from Proposition 3 using Proposition 3 of the previous section. Now Theorem C follows by choosing A = 1/2 in the corollary. Tratib and Woxniakowski (1979) may be reinterpreted as lying in the direction of this corollary. For the sharpness of the corollary, consider f(z) - z/(l z) with c 0. Then lU)
' » n d zH =■■ « * _ , .
Acknowledgment. This work was supported in part by NSF and DOK contracts. H.eforoncm.
11) J. l)icudonn
1253
[3] L. Kantorovich and G. Akilov, Functional Analysis in Normed Spaces. MacMillan, New York, 1064. |4] M. Kim, Ph.D. Thesis, City University of New York, lo appear, 1985. [5] S. Lang, Real Analysis, Addison-Wesley, Reading, Massachusetts, 1983. |6| A. Ostrowski, Solutions of Equations in Euclidean and llannch Spares, Academic Press, New York, 1973. (7) L. Rail, "A note on the convergence of Newton's method," SIAM J. Numrr. Anal., 11(1074), 34-36. [8] J. Renegar, "On the efficiency of Newton's method in approximating all zeroes of a system of complex polynomials," Mathematics of Operations Research, to appear. [9] M. Shub and S. Smale, "Computational complexity: on the geometry of poly nomials and a theory of cost, part I," Ann. Set. Ecole Norm. Sup. 4 serir t, 18(1085), 107-142. (10] M. Shub and S. Smale, "Computational complexity: on the geometry of polyno mials and a theory of cost, part II," SIAM J. Computing, 15(1986), 145 161. [11] S. Smale, "The fundamental theorem of algebra and complexity theory," Bull. Amer. Math. Soc. (N.S.), 4(1081), 1-36. (12] S. Smale, "On the efficiency of algorithms of analysis," Bull. Amer. Math. Soc. (N.S.), 18(1085), 87-121. |13] J. Traub and II. Wozniakowski, "Convergence and complexity of Newton itera tion for operator equations," J. Assoc. Comp. Mach., 20(1979), 250-258. The author haa been at Berkeley since 1064. Ilia work ia in the analysis ttf algorithms. Ilr *f»rit<J* bis spare time ia the pursuit of collecting fine minerai specimens, and tailing in San Franciaco Hay and the adjoining ocean.
1254 JOURNAL OF COMPLEXITY 3, 81-89 (1987)
On the Topology of Algorithms, I STEVE SMALE
University of California, Berkeley, California 94720
1
This paper deals with the structure of algorithms forfindingapproxima tions of the zeros of a complex polynomial, especially lower bound esti mates. Consider the problem: Poly(<0: Data, a complex polynomial of degree d, leading coefficient 1 and e > 0. Find all the roots of/within e. So if £i,. . . , k are the roots of/, perhaps multiple, the problem is to find z\, . . . , Zd such that \zi - (,| < e, each i. Eventually we will specify e(d) and require e < e(d). For the purposes of this paper, an algorithm will be a rooted tree: root at the top (!) for the input, leaves at the bottom for the output. Internal nodes will be of two types: Computation nodes, r, which transmit a program of real numbers, mod ified by a rational operation +, - , x, + ; Branching nodes. A, which go right or left according to whether an inequality is true or false (precision will be given in Section 2). We call such an algorithm a computation tree. A computation tree for the problem Poly() has input the coefficients of a polynomial / (in terms of real and imaginary parts). The output must consist of (z, zj) (again given in terms of real and imaginary parts), each zt being within e of £,, the & being the roots of/. The computation nodes do not contribute to the topology of the compu tation tree, so we define the topological complexity of the tree, as the number of branching nodes. The topological complexity of problem Poly(d) is the minimum of the topological complexity of all computation trees for that problem. 81 0B85-O64X/87 $3.00 Copyright O IM7 by Academic P K I S , Inc. t l rifhu ofreproductionin lay form reserved.
1255
82
STEVE SMALE
Our main result is: MAIN THEOREM. For all e < e(d), the topological complexity of the problem Poly(d) is greater than (log2rf)M.
The proof goes by topology, especially algebraic topology. Eventually Fuchs' results on the cohomology ring of the braid group play a decisive role. Some of the ideas of the proof seem quite universal, but unsolved problems in algebraic topology prevent extension of the result to several variables. Steele and Yao (1982) used algebraic topology to study decision trees for very different problems. Subsequently, Ben-Or (1983) extended this work. The braid group enters into McMullen's work (1985; 1986a, b) on algorithms for zero finding. His negative results and those of the present paper are different in character. Two conversations with Emery Thomas were very helpful to me in understanding the work of Fuchs. 2 We now formally state what we mean by an algorithm. The notion of a computation tree of Section 1 is made precise (some of the computation nodes of Section 1 are collapsed, but the number of branching nodes is the same). The following foundational account is a little more systematic than necessary here, but it will be useful later. The definition of a Flowchart Program in Manna (1974, p. 163) is modi fied in this way. No loops are allowed (for the present paper), the vari ables are real numbers, and "predicates" of Manna are defined in terms of rational functions. Thus the input domain, denoted here by £, the program domain 9, and the output domain 0, are each real cartesian spaces of some dimension. The set of usable inputs (satisfying an input predicate in the terminology of Manna, 1974) is supposed to be a real semialgebraic set Y in $. There fore Y has the form Y = {y G $|*,(y) = 0, tj(y) < 0, uk(y) s 0} for some finite set of rational functions, {$/, tj, uk}. Moreover we always suppose that rational functions have integer coefficients in this paper. The set of acceptable outputs is defined by another semialgebraic set X C Y x 0. Define/: X-* Y as the restriction of the projection Y x C-* Y; we require that/be surjective. Nodes of the computation tree are of four types: root (or start), compu tation (or assignment), branching (or test), and leaf (or halt). Each has an associated rational map.
1256 ON THE TOPOLOGY OF ALGORITHMS, I
The root is defined by a rational map/: f-*9 rational function).
(each coordinate of/is a or
y-fix)
83
y «-/M
—r~ A computation node is described by a rational map g: S x ? - » > . or
y' - *(«. y)
y «- gix, y)
1 To a branch node is associated a rational function * : > x f - . R ,
1
r
Ju
*(JT,
y)< 0
n
(could be kix, y) s 0)
Finally a W i s defined by a rational map /: * x 9 -* 0.
_L /(Jr. y)
or
z «- /(*, y)
Each or E £ defines a path starting down the tree. We require that if x £ Y, then division by zero is not encountered along the path. This con dition on the computation tree ensures that each such path leads to a leaf. A final requirement is that for x E Y, the endpoint z of this path satisfy (x, z) G X C i x 0. A further reference on algorithms with an extensive up-to-date bibliog raphy is Purdom and Brown (1985). The number of paths equals the number of leaves equals the number of branches plus one.
Let/: X-* Kbe a continuous map. Define the covering number of/as the least k with this property; there is an open covering %,. . . , % of K and continuous maps gt:%-* X with/(#<(?)) = y, each i and y € %. Note that if/ is not surjective, the covering number is infinite. Next let 9d be the space of complex polynomials of degree d with leading coefficient /. Thus a point of 9
1.
Let O be complex (/-dimensional space and «-: Cd -* 94 the map which assigns to ({1,. . . , &) the polynomial with roots ( , £*• Thus w has as coordinates the symmetric functions a, (cf. Lang, 1984) in the Q.
1257 STEVE SMALE
84 Let A = {{ = (£
W e C% = &, some i * j]
and w(A) = I C ? j . One may describe X as the algebraic variety of polynomials / whose discriminant (Lang, 1984) is zero. Note that ir"'(X) = A. 9 a is the input space of problem Poly() and C the output space. Moreover, the set of usable inputs (Y of Section 2) is BK and the set of acceptable outputs is X = {(/,(z,
Zrf)) E BK x C\ \z,-ii\<
e,f(z) = f l (z - Id), i-i
where BK = { / e 9d\ \at\ s K, i = 0 d - 1} and K = *() is chosen large enough that if/has all roots in the unit disk, then/E BK. A. The covering number of the restriction IT: Cd - A -*■ 9i - X is less than or equal to the topological complexity of problem Poty(d), for all s < t{d), e(d) described in the proof. THEOREM
Proof. Let a computation tree for Poly(rf) be given with leaves num bered i = l , . . . , k. Denote by V, the subset o(BK (inputs) which arrives at leaf i. Then BK = U*.|V, and V, n V, = B if i * j . (The V, are real semialgebraic subsets of 9j.) These input-output maps, denoted by fa: Vi-*C, are continuous real rational maps with integer coefficients in the variables (Re at, Im a,). The values satisfy fa(f) = (zi, . • . , Zw). \zi - &| < e, where the (j are the roots off. For our purposes, we only need the fa to be continuous. The V, may be described by V, = {a = (<*>, • • • .flrf-i)e **! gj(a) < 0, > = 1
/; Ma) * 0, £ = 1, . . . , m),
where the gj and /i* are continuous (even rational) functions. Thus Vt is a closed subset of an open set V} in BK. By the Tietze Extension Theorem (see Munkres, 1975), fa can be extended to an open set % of V, and this map still denoted by fa; fa: % -» C will satisfy: fa(f) = (zi Zj), \zt - &| < e, & the roots of/. These sets % are open in BK and cover BK since the V, do. If Y is a subspace of a space X, it is called a deformation retract of X provided there is a homotopy h,: X-* X, 0 s t s 1, satisfying: /io is the identity, ht{X) C r, and /«,(y) = y for y E Y.
1258
ON THE TOPOLOGY OF ALGORITHMS, 1
85
The following well-known lemma is implicit in Spanier (1966, pp. 290291). LEMMA 1. Let Ybea closed subspace of a compact space X such that the pair (X, Y) can be triangulated. That is, there is a homeomorphism h: (X, Y) -* (K, L), where L is a subcomplex of a simplicial complex K. Then there is a neighborhood NofY such that X - N is a deformation retract ofX - Y.
Let S = {z £ C^l Hz|| = 1} using the Hermitian inner product on C . LEMMA
2. The pair (ir(S), 2 fl ir(S)) can be triangulated.
For the proof, see Lojasciewicz (1964). LEMMA
3. n(S) - X D ir(S) is a deformation retract of'9\/ - 2.
Proof. First define h,: Cd - A -♦ Cd - A by h,(x) = (1 - t)x + tr/|x|. The homotopy is invariant under the group S(d) of covering transforma tions, hence induces the required homotopy of 9d - Z. As a consequence of Lemmas 1, 2, and 3, we have: LEMMA 4. There is a neighborhood Nofl,r\ ir(S) - N is a deformation retract of9d - 2.
n(S) in ir(S) such that
Let h,\ 9d - 2 -+ 9* - 2 be the retraction. Thus h0 is the identity, hi(9d - 2) C ir(S) - N, and ht(y) ■ y for all y e ir(S) - N. Choose TJ = rj(d) with this property if / £ ir{S) - N', then the roots o f / a r e separated by at least TJ. Next let P, = % D (it{S) - N), and suppose e < TJ(,if) = (zi, . . . , zA each zi has a closest root fc of / defined unambiguously. Let M / ) = (d, . . . , &). Then tji,: Pt-+ O is continu ous and in^if) = /. We have found a covering {Pt} of rr{S) - N showing that the covering number of ir: S - ir~l(N) -* ir(S) - N is at least d. The final step in the proof of Theorem A is to use the deformation retraction to define the appropriate covering {Qi} of 9d - 2. Let Qi = Af'(P/) and extend 4>i to Q( using the covering homotopy property. This finishes the proof of Theorem A. Remark. It is clear from the proof that Theorem A holds in considera bly greater generality. 4
The cup length of a ring & is defined as the maximum number k such that y, U . . . U yk * 0, y< E 91, where "U" denotes the product. For a continuous map f: X -* Y, let K{f) be the kernel (an ideal) of
1259 86
STEVE SMALE
/*: A W ) - / / • ( * ) , i.e.. K(/) = {yGtf*(K)|/*(y) = 0}. Here H*(X) is the singular cohomology ring of X, and / * is the induced map. PROPOSITION
1.' The covering number off is greater than the cup
length ofK{f). The cup length depends on the coefficients in cohomology, but Proposi tion 1 is true for any coefficients. Later the coefficient ring will be the integers mod 2. This proposition is related to category theory of Lusternik and Schnirelman; see Schwartz (1967) or Spanier (1966, p. 279). Proof of Proposition 1. We proceed by supposing the proposition is false. In that case there exist y , y* G K with y\ U • • • U yk ± 0 and there is an open covering Vlti = 1,. . . , k of Y, with associated continu ous maps O-J: Vi -* X having the property /(
ir*: H?(9d - X, ZJ - Hf(C - A, Z2) is trivial for i > 0 (and an isomorphism for i = 0, of course). For this and the next proposition, we use the work of Fuchs (1970), but also the works of Arnold (1968), Birman (1974), Brieskom (1973), Cohen in Cohen et al. (1976), and Fadell and Newwirth (1962) are also quite 1 Note added in proof. Moe Hirsch pointed out to me that by taking X as the path space of Y, Proposition 1 contains the Lustemik-Schnirelman result.
1260 ON THE TOPOLOGY OF ALGORITHMS, I
87
pertinent. One definition of the braid group is the fundamental group of 9* - 2 and these papers all deal with the topology of the braid group. The cohomology of the braid group is the same thing as the cohomology of 94 -2. Proof of Proposition 2. Consider certain spaces as follows. Let 0(d) be the orthogonal group and ZJcw> the corresponding classifying space (see Husemoller, 1966). Let S(d) be the symmetric group on d elements and let Bsu) be the Eilenberg-MacLane space K(S(d), 1) let u4 -» A M ) be the universal covering (see Spanier, 1966). According to Fadell and Newwirth (1962) 9 - 2 is an EilenbergMacLane space, tf(n,(* - X), 1). The map IT: C* - A -> 9 - 2 is a regular covering (see Spanier, 1966) with group S(d) since the map «■ is given by the symmetric functions. Thus there is a natural map from cover ing space theory n,(» - 2) -> S(d). This map can also be given by interpreting geometrically rii(
S(d)->0(d) by the symmetric group permuting the coordinates. Ring homomorphisms in cohomology over Z2 are induced by the group homomorphisms, so we have H*(9i - 2) «- H*{Bsut) *- H*(Bou))- Ac cording to Fuchs (1970) the map H*(8ow, zd ~* H*(94 - 2 , Zj) is surjective. Thus LEMMA.
The map H*(Bs(4i), Zfi -» H*(9i - 2, Z$ is sur jective.
Since the composition n , ( C - A) -» Il,(9rf - 2) -» n,(BMllt) =« S(d) is zero, by covering space theory (Spanier, 1966), there is a map h with the cummutative diagram: C - A -*■
i
i
9d-X-*Bsw. Since H*^, ZH - JSflflU Zj), If ftBjw, ZJ -» Hf(C - A) is trivial for / > 0 (either way around the diagram). Proposition 2 follows, using the lemma. Let HU9< - 2 , Zj) be the ring 2 f t J W * - 2 , Zj).
1261 STEVE SMALE
88 PROPOSITION
3. The cup length of Ht(9j - 2, ZJ is greater than
(log2)M. Proof. According to Fuchs (1970), the generators of Ht(94 - 1, Z$ are am%k, k = 0, 1, 2, . . . , m = 1, 2, 3, . . . , degree am%k = 2*(2"-'), relations a2mJl = 0 and otherwise, amiJtt am,jk, = 0 just when 2<m+"+*f+*i+"+*f > ti%
So we want to find a sequence of distinct pairs, (mi, k\),.... with t as large as possible and
(m,, k,)
E mt + T kt « lofcrf.
(*)
Consider now the set of all distinct pairs (m,, A:,) such that m, + A, £ A/. An easy counting shows that there are / = M(M + l)/2 of these pairs. A second easy counting shows that (*) will be satisfied provided 2f0' 2 = M(M + IK2M + l)/6 s logid. It is not difficult to check that / = (log2)2/3 satisfies these conditions. Actually there is a universal e > 0 with t = (1 + eY}ogid)2n satisfactory. This proves Proposition 3. The proof of the Main Theorem now follows: Topological complexity Poly() a Covering Number (w: Cd - A -» 9d - Z) > Cup length ker w* = Cup length Hi(94 - 2, Zj) > (log^)"
Theorem A Proposition I Proposition 2 Proposition 3.
REFERENCES ARNOLD, V. I. (1968), On braids of algebraic functions and cohomologies of swallowtails, Uspekhi Mat. Nauk 23, 247-248. BEN-OH, M. (I983), Lower bounds for algebraic computation trees, in "Proceedings, 15th ACH STOC," pp. 80-86. BIKMAN, i. (1974), "Braids, Links, and Mapping Class Groups," Annals of Math. Studies, Princeton Univ. Press, Princeton, NJ. BUESKOKN, E. (I973), Sur les groupes de tresses (d'apres V. I. Arnold), in "Seininaire Bourbaki," Lecture Notes in Mathematics, Vol. 1971/72, No. 317, Springer-Verlag, New York. COHEN, F., LADA, T., AND MAY, P. (I976), The bomology of iterated loop spaces, in "Lecture Notes in Mathematics, Vol. 533," Springer-Verlag, New York. FADELL, E., AND NEWWIRTH, L. (1962), Configuration spaces. Math. Scand. M, III-II8. FUCHS, D. (I970), Cohomologies of the braid groups mod 2, Functional Anal. Appi. 4,14315I.
1262 ON THE TOPOLOGY OF ALGORITHMS, I
89
HUSEMOLLER, D. (1966), "Fibre Bundles," 2nd ed., Graduate Texts in Mathematics, No. 20. Springer-Verlag, New York. LANG, S. (1984), "Algebra," 2nd ed., Addison-Wesley, Reading, MA. LoJASCiEWicz, S. (1964), Triangulation of semi-analytic sets, Ann. Scuola Norm. Sup. Pisa CI. Sci. (3) 18, 449-474. MANNA, Z. (1974), "Mathematical Theory of Computation," McGraw-Hill, New York. MCMULLEN, C. (1985), Families of rational maps and iterative root-finding algorithms, preprint. MCMULLEN, C. (1986a), Automorphisms of rational maps. I. Nielson realization and dy namics on the ideal boundary, MSRI, Berkeley. MCMULLEN, C. (1986b), Automorphisms of rational maps. II. Braiding of the attractor and the failure of iterative algorithms, MSRI, Berkeley. MUNKRES, J. (1975), "Topology," Prentice-Hall, Englewood Cliffs, N.J. PURDOM, P., AND BROWN, C. (1985), "The Analysis of Algorithms," Holt, Rinehart & Winston, New York. SCHWARTZ, J. (1967), "Non-linear Functional Analysis," Gordon A Breach, New York. SPANIER, E. (1966), "Algebraic Topology," McGraw-Hill. New York. STEELE, M., AND YAO, A. (1982), Lower bounds for algebraic decision trees, J. Algorithms 3, 1-8.
1263 Proceedings of the International Congrna of Mathemati. iant Berkeley, California, USA. 1086
Algorithms for Solving Equations STEVE SMALE The main goal of this work is an attempt to understand the efficiency of algorithms for solving systems of equations. For example, consider a map / from complex Cartesian space C to C n given by polynomials. How fast can one find good approximations to one or all of the zeros of /? However, for clarity of exposition and development we will usually swing from one extreme to another about this example, proving theorems on one hand for a single complex polynomial and on the other hand for an analytic map /: F -» F from one Banach space to another (both real or both complex). This subject should undoubtedly be considered as numerical analysis. How ever, my point of view has been especially influenced by complexity theory of computer science. Thus the emphasis is on the algorithms themselves and a global study of their speed. Hence less emphasis is given to the results obtained by algorithms and the asymptotic criteria of efficiency that are often found in numerical analysis literature. This global study, or complexity theory, gives a more systematic, more ab stract, more theoretical flavor to our approach, and thus we are less concerned about producing immediately faster methods for problem solving. A basic under standing of the tried and effective is sought. The global study of the algorithm forces the introduction of topology and geometry into the subject. If any algorithm has proved itself for the problem of nonlinear systems, it is Newton's method and its many modifications. The Greeks essentially used it for finding square roots; and for that purpose it is still a top method today. On the other hand, for nonlinear functional equations framed in a Banach space setting, Newton'8 method finds a central place. In the account here, the one theme is Newton's method. We will prove the orems about it for one variable (dealing with the fundamental theorem of alge bra) and for maps of Banach spaces; we will approximate Dantzig's [8] simplex method for the linear programming problem by Newton's method. Supported in part by National Science Foundation grants and by the Mathematical Sciences Research Institute. © 1987 International Confreia of Mathematicians 1988
172
1264 ALGORITHMS FOR SOLVING EQUATIONS
173
The contributions and point of view of Kantorovich dominate the modern treatment of Newton's method. For background see Berger [2], Henrici [19], Kantorovich-Akilov [23], and Ostrowski [38]. This treatment has great beauty derived from its simplicity, its generality, and its weakness of its hypotheses. The assumptions in a Banach space context are first and second derivative bounds on a domain of a map. What we do is to start afresh with analytic maps and derive estimates at one point. In the execution of an iterative algorithm, one finds oneself at some point in the domain space and is able to compute derivatives at that point. From that informtion, it is important to make a judgment for the next step, or perhaps even terminate the algorithm. These considerations motivated our paper, [52], referred to subsequently as [point estimates] or just [P.E.]. Since the results of [P.E.] are used frequently below, we review a bit of that theory. One aspect of a theoretical study of problems where only approximate so lutions are possible, is the omnipresence of a small parameter e > 0, which measures the distance to a solution. However, for a well-conditioned solution of /(f) = 0, / a complex polynomial, one can dispense with such arbitrariness using the notion of an approximate zero. One variable, Newton's method for solving f(c) = 0, with starting point zo € C, is the sequence z* = Zk-i - f(zk-i)/f'(zk-i) defined inductively. The point ZQ will be called an approximate zero if
|**-**-i|<(i)a*",-1|*i-*|.
fc
= 1,2,3
Thus if k = 5, for example, the coefficient is already (j) 1 5 . In machine com putations one sees this phenomenon with doubling of precision at each step. However, it would be satisfying to have a test that one could make with infor mation at zo to insure this estimate for all k. That motivates the introduction of our invariant a(z, f) and Theorems A and B as follows. Let 0(z, / ) be the length of the "Newton vector"
/*(*./) = Here fk(z) is the ifcth derivative. Then the invariant product 0{zj)i(z,f) or 0(z)i(z).
Q(Z,/)
is defined as the
THEOREM A (SPECIAL CASE). There is a constant c*o equal to approxi mately .130707 such that ifa(z,f) < ao, then z is an approximate zero of f. The general case of Theorem A has the same statement, but / and the defi nition of approximate zero are generalized as follows. Now / : E — ► F is an analytic map from one Banach space to another (e.g., E = Cn=F). Newton's method starting from ZQ € E is given by *k = Zk-l~
£>/(Zfc_l) - 1 /(2fc-l)
1265 174
STEVE SMALE
>ofof / / ifif \\z zk-i\ << (kf-'-'W*! and ZQ is an approximate zero \\zkk -- Zk-i\\ (s) 2 ~ - 1 ||*i -*o\\, -- «o||, kk = 1,2,.... Let
fi(*J) = \\Dnz)-lf{z)
and if(z, / ) = sup
r
vn
k
«.ri^/W
">
i/(t-i)
Id
k
where D f{z) is the fcth derivative of / as a Jk-linear functional (see Dieudonn6 [10] or Lang [30] for the calculus). In case Df(z)~1 is not defined, we make a, 0, and i infinite. Now Theorem A makes sense for analytic /: E —► F and is true; this is proved in [RE.]. M. Kim [27] independently has a proof of Theorem A with one variable and Qo = J-J. I wrote in [P.E.] that theorems similar to ours could probably be deduced with the Kantorovich theory as a starting point. Subsequently, H. Royden [42] showed that that indeed was the case with a slight improvement of oo in Theorem A. Moreover, recently J. Curry [7] has extended the onedimensional version of Theorem A to higher order generalizations of Newton's method. One might object to the feature of Theorem A that requires knowledge of all derivatives for the 7 factor of a. In fact only the first derivative is necessary, together with some crude information on / . To make this precise, define the function
~(, n < ML 4M! 1l,/
' - | / ' W I M\*\)'
For the general case, let /: E —► F be an analytic map of Banach spaces which can be expressed as f(z) = ]C*lo a * zfc w n e r e t n e a * a r e symmetricfc-linearmaps and OkZk is a* evaluated at the Jt-tuple (z,..., z) of elements of E. Let ||ajt|| be the usual norm as for example in Lang [30]. Then define ||/|| = supfc ||a*||. THEOREM B. With } : E -» F, z e E as above
7(«./)< 11/11 IIWW-MI^BTheorem B is announced in [P.E.] and proved in §1 below. It is used in our proofs of the theorems below on the Tractability of the Random Algorithm and Average Area of Approximate Zeros. By providing a termination test, Theorem A allows us to introduce a simple Random Algorithm for polynomial zero finding. This goes as follows: Given a complex polynomial / , (1) Choose z eC, \z\ < 3, at random. (2) Is a(z, f) < ao? If yes, terminate (or apply Newton's method, say 5 times, then terminate). If no, go to (1). It is natural to ask, for a given / : How many steps will the random algorithm take on the average? For this to make sense let (1 be the set of all possible
1266 ALGORITHMS FOR SOLVING EQUATIONS
175
sequences Z = (21,23,23,. ..), 1**1 < 3. Give the disk of radius 3 normalized Lebesgue measure and endow 0 with the infinite product measure as a proba bility measure, as in Shub-Smale [44]. Define the function a: 0 —► Z+ by o(Z), which is the first a such that 0t(zo, f) < ato- Then the average number of steps needed to terminate the random algorithm is defined as at = average o{Z). zen THEOREM (THE TRACTABILITY OF THE RANDOM ALGORITHM).
The set
of polynomials f in Pi(l), such that average number of steps is larger than 0, has measure less than cd6/a for an appropriate constant c. Of course one hopes for a better bound eventually. This theorem is proved in §2. Here Pd(l) is the set of all complex polynomials f(z) = Yloaizt w ' t n ad = 1 and |a,| < 1. We have imposed uniform probability measure on Pd(l) (i.e., normalized Lebesgue measure) using the fact that Pd(l) is a bounded set in C d . The set of / € P<j(l) with a/ = 00 has measure 0. These / are the ones where zero finding is an ill-posed problem in some reasonable sense. For the next result, let 20 6 C, / a complex polynomial, and consider
\*n-<\<(W"-1\>0-<\
(*)
where /(c) = 0 and z„ = zn-x - f(zn-i)/f'{zn-i)Thus (*) is another form of a superconvergent property independent of a small parameter. Let Qf be the set of |2oI < 1 satisfying (*), and let Aj be its area. THEOREM (AVERAGE AREA OF APPROXIMATE ZEROS).
average Af > c > 0 where c is independent of d. This result is equally valid if / is averaged over the set {f\f(z) = Z)f=o ''i**' |oi| < 1, i = 0 , 1 , . . . , d). The proof is given in §3. Points ZQ satisfying (*) are called approximate zeros of the second kind. Gen erally if/: E-+ F is an analytic map, zo € E, z„ = z„-i-Df(zn-i)~1(F(zn-i)),
ll*.-fll<(i) a "- 1 ll*-?ll,
/(f) = 0,
then zo will be called an approximate zero of the second kind. For the proof of the previous theorem we use Theorem C which was proved in (P.E.). THEOREM C. Suppose that /: E -♦ F is analytic, c € E, /(f) = 0, and z 6 E satisfies
„
3 - V7
n
'-
1
—wjj-
Then z is an approximate zero of the second kind.
1267 176
STEVE SMALE
Next, consider an analytic map / : E —» P from one Banach space to an other. We will study an algorithm based on Newton's method for approximating solutions of /(f) = 0. Complexity aspects of this algorithm will be proved. We suppose / given as above and z§ € E. It is important for our purposes to understand the lifted path a = a(zo, f) C E, defined as follows. If the derivative Df(zo): E -» F is an isomorphism, let f~^ be the local inverse of / defined from a neighborhood of f(zo) to E and taking f(zo) to ZQ. For nonnegative t 0 sufficiently close to one, f~* maps {t/(zo)| to < t < 1} into E. It may happen that f~0y extends to a map of the ray {tf(zo)\ 0 < t < 1} -» E, remaining an inverse to / and such that the derivative along this ray is always an isomorphism. In this case we say that a = a(zo,f) is defined and is the set /«!{«/(*>)|0<<
. There exist
(small) positive constants, c real, I an integer with this property. Let / : E —► F be analytic, ZQ € E, with M(zo,f) < oo. Suppose n is an integer, n>c||/(z0)||M(zo,/),
A = l/n.
Let wt = (1 - iA)/(zo), t = 0, . . . , n . Then, inductively, z, = #}_„, (*t-i) is well defined and z„ is an approximate zero of f. Here Nf-Wl(zt-i) is Newton's method for solving /(?) — to, = 0, applied to z,_i, and Nlj_Wi(zi-i) is the same applied iteratively / times. The constants c and / are about 4. We give the proof in §4. The algorithm in the above theorem is a version of the Global Newton in Smale [47], Hirsch-Smale [20], Keller [26], Garcia-Gould [15], Abadie-Guerrero (1). Related results can be found in Kung [28], and the book of Dejon-Henrici [29] (especially the article by Dejon-Nickel in the last volume). Surely the theorem has roots in the one-dimensional case of Shub-Smale [43, 44] and Smale [50, 51]. The ideas of this proof extend to deal with other situations, for example, homotopy problems as in Keller [25] and Chow-Mallet-Paret-Yorke [4]. More over, if a problem of zero finding is ill-posed, the corresponding residual problem |/(z)| < e could be well-posed and the complexity theory of the above theorem applies (again similar to Smale [50, 51] and Shub-Smale [43, 44]). Both Newton's method and piecewise linear algorithms can be thought of as path following algorithms. This is a theme in Eaves-Scarf [11], Smale [47], and Hirsch-Smale [20]. One can interpret Dantzig's [8] self-dual variant of the
1268 ALGORITHMS FOR SOLVING EQUATIONS
in
FIGURE I
simplex method as a path following method (Smale [48]) in this way. Thus a relation between the simplex method of linear programming and Newton's method, is no surprise. In fact, the relationship turns out to be very simple and natural in the framework of the linear complementary problem or LCP. The LCP starts with data (M, q) where M is an N x N matrix and q € RN. Define the functions z* = (z ± |z|)/2 for a real variable z (more formally one could write x + (x) = (z + |z|)/2 etc.). Let **f: RN -* RN be the map ♦«(*) = z + + Mx~ where x± = (zf,..., x%). The LCP is: given (M, q) solve for z; *M(x)=q(LCP) For background, see Cottle-Dantzig [6], Smale [48, 49], and the cited refer ences. We are especially concerned here with how the LCP can be considered as containing the linear programming problem, or LPP, as a special case. The LPP with data (A, b, c), A an m x n matrix, b € R m , e € Rn is: minimize c • x subject to z > 0, Ax > b.
(LPP)
Here c • z = (c, z) is the inner product. Then the LPP data (4, b, c) generate LCP data (M, q) by N = n + m, q = (c, —6), and M = [^ "£ ]. It will be assumed that M has this special form. The LPP has a solution if and only if the corresponding LCP does, and in that case the solutions correspond in a natural way. We propose here an analytic approximation to $M, which at the same time desingularizes the LPP and makes analytic methods—Newton's method, in par ticular—available for its solutions. In this way, Dantxig's [8] self-dual algorithm, a version of his simplex method, can be approximated by Newton's method. We start with an approximation of the functions of one variable, x±. Let
1269 178
STEVE SMALE
As a -* 0, pf approaches z* uniformly, and as z -» ±00 the approximation only improves. Next let * a : RN -► R " be the map *«(z) = *+(*) + A / * ; (z) where *Z(z) =
(
From the above observations, ♦„ tends to 4>w uniformly on all of R N as a -» 0; and for each a, 4 a (z) tends to * M ( Z ) as ||z|| —► 00. Our next goal is to give some analysis of ♦„ for a > 0. Define l i M c RN as the cone of all positive linear combinations of the column vectors of -M and the coordinate vectors t\,..., e#. A diffeomorphi8m is a differentiate map with a differentiable inverse. It is analytic if, in a neighborhood of each point, it can be defined by a convergent power series. THEOREM (ON THE REGULARIZATION OF THE N
the map Qa: R * a : RN ^UM.
-♦ R
N
LPP). For each a > 0,
A<»* image UM and is an analytic diffeomorphism
This theorem is proved in §5. REMARK 1. Actually the proof does not need the special form of M. The proof and the theorem are valid for any square matrix satisfying {Mx, x) > 0. The same is true for the proposition in §5. REMARK 2. As a consequence of the theorem, given M for any q € UM, and any a > 0, there is a unique solution of the problem $M(X) = q. This solution depends analytically on the data (A,b,c) of the LPP and a, as long as q € UM and a > 0. Say that the LPP with Data(i4,6,c) is "extremely ill-posed" if q € OUM (3UM = 12M — UM)- Here is some justification for our terminology. The LPP has a solution if and only if q € UM (see Cottle [5]). Thus if q € dU, a small perturbation of the data means that the problem has no solution. The above theorem implies that all LPP which are not extremely ill-posed can be desingularized at once, by taking any a > 0. Newton's method for $„ -q follows the curves * ~ l (qq) where qq is the segment from an initial value q to q. Dantzig's method follows the curve ^~1(qq) for a = 0. Thus as a —» 0, Newton's method, with an appropriate step size, tends to the Dantzig Algorithm. It is convenient to introduce a map * a : RN -* RN by ♦ a (x) = ♦a(z) -*o(0). Then as a goes to zero, $„ approaches the LPP. Moreover, as is shown in §5, Theorem A can be applied to solve 4„(z) = q for all a such that a > c\\q\\. Then the original LPP problem can be obtained by reducing a to zero. Related in one way or another to the previous result is the work of Mangasarian [31], Wierzbicki [53], Karmarkar [24], Gill-Murray-Saunders-Tomlin-Wright [17], Blum [3], Renegar [41], and Megiddo-Shub [36].
1270 ALGORITHMS FOR SOLVING EQUATIONS
179
More generally, there are many other papers related to the subject of this paper. In particular, in the area of information based complexity, one can see Wofniakowski [55] for a recent survey. Smale [51] has many other references. Papers not listed in [51], but which have interesting related results, include Pan [39], Wong [54], Gao [13, 14], and Wright [56]. I would like to express my thanks to a number of mathematicians for their help in this paper. In particular, a conversation with Dick Cottle on the LCP was useful. Lenore Blum's MSRI talk on a condition number for linear programming via the LCP helped put the LCP back in my mind. Her comments and those of Jim Curry and Feng Gao have been generally useful to me. Especially important through all of this has been the work of, and conversa tions with, Jim Renegar and Mike Shub. That contribution from Mike Shub, to me, has persisted over many years indeed. 1. Suppose a is defined as above and
■£(;)
where p, d, j are nonnegative integers with d > p (but recall the binomial coeffi cient (p) = 0 if i < p). LEMMA
l.
Q {d)
>
_ (d+i)d(d-i)-(d-P
~
(pTTjl
+ i)
•
PROOF OF LEMMA 1. Note that Qp(d) is a polynomial in d of degree p + 1. Then it is sufficient to check the formula of the lemma at p+2 values of d. Note thenQp(d) = 0for d = 0 , 1 , . . . , p - l , Qp(p) = 1, andQ p (p+l) = p + 2 . Q.E.D. LEMMA 2. For0<m
1271 180
STEVE SMALE
In these terms the estimates of the theorem is Vo~ V * < i>k, k = 1,2, For fc = 1, this is trivial. We proceed to prove the theorem by induction on Jfc. Thus it is sufficient to show rpo^Pk < i>iipk-i, k = 2 , 3 , . . . , x > 0. Write out this last estimate and multiply both sides by xk to obtain
The product on the left can be written in the form £m=o c m* m while that on the right has the form J^L=oamXm- It is sufficient to show that c m <
-s(0-«
(m),
°n> =
JL(r»-3)(ki_l)
If d < m < 2d, then
j=m—d \
/
]=m—d
\
/
Define Pfc_,(fl) = j:?=o>(fc-i)From Lemma 3 and the above evaluations, we have Qk(m) = mQk-i(m) — Pfc_i(m) for a// m, since this formula is independent of d. To finish the proof of the theorem it is sufficient to show that for d < m < 2d, Cm - Cm > 0 or
mQfc_!(d) - P fc -i(d) - mQk^(m
- d - 1)
+ P fc _i(m - d - 1) - Qk(d) + Qfc(m - d - 1) > 0. Use the previous equation to simplify, obtaining as sufficient (m — d)Qk-\ (d) — ( d + l ) Q f c _ 1 ( m - d - l ) > 0. Let mi = m - d - l ; this becomes (mi+l)Qk-i(d) > (d+ l)<3fc-i(mi) for 1 < mi < 2 d - 1. Apply Lemma 2; this proves Theorem 1. We now prove Theorem B of the Introduction, or THEOREM 2. - Y ( o , / ) < | | Z ? / ( « ) - 1 | | | | / | | ^ ( | N | ) a / ^ ( l k l | ) . PROOF. First note Dkf(z) - J2,=o »(* - 1) •••(»'- it + l)a,zi~k where aizi~k is the fc-linear form defined by putting z into (i — k) places of the i linear form a,. From this picture it follows that
Dkf(z) < |(/|| ^(iWJ) it!
Ik!
1272 ALGORITHMS FOR SOLVING EQUATIONS
181
Since \\Df(z)-l\\\\Df(z)\\ > \\Df(z)~lDf(z)|| = 1, using k = 1 we obtain 1 ll^/C*)"" !! ll/ll!Pi(IMI) > 1- From Theorem 1, it follows for k > 1 that ^
( f )
Dkf(z)
< *><(')*
A:!
< ll/ll
^(11*11)* *>-(IMI)*- x '
Now we use the above estimate to prove Theorem 2 as follows Tf(*,/)<sup \\Df{z)-
Jt!
*>i||
||
< T (IWW-'II|^M||) !/(*-») # ..
*
UJ ■ l l l ^ l l l
1
< sup I fc>l V
< ^(IWDsupdiP/w 1 !! ||/||»/d(||*||)),/(*-1). 1Pd
fc>l
The sup is over the (k - 1) roots of a number which is independent of k and not less than 1. Thus the sup is achieved at k — 2. This yields Theorem 2. Remark. There is a kind of UL2 version" of Theorem 1, which I have not been able to resolve. Problem. Is a(r, ipd) < 1 for all r > 0 where Vd(r) = (E?=o r 2 ') 1 / 2 ? 2. The goal of this section is to prove the tractability of the Random Algo rithm by giving a probability estimate of the function a/ (the theorem is stated in the Introduction). Let Bf be the area of the set of z e D$ such that a(z, f) < ao- The following is similar to Proposition 3, §2 of Shub-Smale [44] (and very simple). PROPOSITION 1. af = 9ir/Bj. (Note area of £>3 = 9ir.) We next prove a proposition which shows that our criterion for approximate zeros of the second kind (Theorem C) also works for approximate zeros (of the first kind). PROPOSITION 2. There is a constant c > 0 with the following property. Let /: E —► F be an analytic map from one Banach space to another and f on element o/E with /(?) = 0. Ifz 6 E satisfies \\z - ?|| < c/i{s, / ) then a(z,}) < a 0 and z is an approximate zero. PROOF. Suppose c < \(\ - v^/2) and z satisfies the hypothesis of the proposition. Then by Lemma 2c, §3 of [P.E.], 7(2) < 017(f) for a suitable constant c\. Next use Proposition 3, §4 of [P.E.], to obtain that 0(z) < C2||z-f|| for a suitable constant c3. Choose
-(K'-fl-a)-
c < mm
So 0(2, / ) = #2)7(2) < cic2\\z - f 117(f) < a 0 .
Q.E.D.
1273 182
STEVE SMALE
For each polynomial / , define Pf=
nun \f(0)\. /'(«)=o Let Tf = min(l,p/). From Smale [50, p. 29], LEMMA 1. m e a s { / e Pd(l)|r/ < p)
Bf>c
\ \Jm\)) \m\)) • /(f)=0
PROOF. It follows from Proposition 2 that
Bf>c
? W7? /(f)=0
for some constant c > 0 (compare Lemma 1 of §3). Now apply Theorem B. Since / € Pd{l), Il/H = 1. This yields Lemma 2. LEMMA 3.
/(O=o
/(c)=o
LEMMA 5. There is a positive constant c such that
II
f'Ak\)l/d
< cdb'\
allf€Pd(l).
Id)=o PROOF. We use a theorem of Specht (see Marden [32, p. 129]) which easily implies the following assertion. If / 6 Pd(l) has roots Ci,... ,fa, then for any subset ii,...,ifc of { l , . . . , d } , |ft,fi, •••?,»! < %/d+l.
1274 ALGORITHMS FOR SOLVING EQUATIONS
183
Then
i_i nt ri(kD=nE>ifti t=ij=i /(c)=o
=
*'ikir,"1«'aifti
£ (M
<«)
l
d
s
1
s( w+v(a^ii)'. Then take the dth root to obtain
I I (^(kl)) 1/d < c«*'2.
Q.E.D.
/(«)=o
LEMMA 6. Let f € Pd(l)- and Df be the discriminant off (Lang [29]). Then \Df\V«>dTf. For this see Shub-Smale [44, Proposition 7 of §2] for the argument. From Lemmas 2 to 6 we prove Proposition 3. By Lemmas 2 and 3
Then by Lemmas 4 and 5 we obtain
t
Recalling Dj = flf (/'(?)) ( Lan K f29))- Lemma 6 applies to yield Bf > proving Proposition 3.
crf/d4,
Now the theorem is proved as follows. Define Gj: Z+ -* [0,1] by Gd(a) = meas{/ € Pd{\)\af
> a).
The theorem may be restated: Gd{o) < cdb/a. But using Propositions 1 and 3 Gd(a) = meas{/ G Pd(l)\9ir/Bf
> a), 2
Gd{a) < meas{/ € Pd{l)\cd*/T
}
> a}
or Gd{a) < meas{/ e Pd(l)\rf Now use Lemma 1, so Gd(o~) < cdP/o.
<
Q.E.D.
(cd4/*)1'*}.
1275 184
STEVE SMALE
3. The goal of this section is to prove the theorem on the Average Area of Approximate Zeros. The quantity average^-gp^^A/ by definition is equal to (-)
/
\*/
^|o„_i|
••• /
Afdao-dad.
J\a0\
Our first task will be to make the estimate
Then, by (I), for the proof of the theorem, it is sufficient to show that
/
■"/
J\a„.
i\
/
|/'(f)|V>W-
(II)
«/|a,|
Of course K and L are positive constants, independent of d. Some estimates of these constants can be given from the proof. First we deal with (I). Let ^ = max(l,7(c,/)) where 7(c,/) is as in the Introduction. LEMMA
l.
*>(^)''E V/(O=o UI
n/3
U "/(?)• kl<»/2
From Theorem C, it follows that z e fl/(?) if ||z - c\\ < (3 - v / 7)/27 3 . Therefore the area of 0/(f) is greater than or equal to JT((3 — \/7)/27<) 2 . Putting these facts together yields Lemma 1. Now we prove the following change of variable formula. LEMMA 2. Let /o = £?=i a«2*> ""'^ la«l - *> so ^ai /o(0) = 0> / = /o + ao, ao varying in C with \a0\ < 1. Then
I 1 0l
an<
* consider
2. Note that |/ 0 (c)| < 1 if |c| < ±. Let V =
f0{D1/2)
E ^ 2 ^ o = / e/o -. (Dl) 7r 2 i/o(c)i 2 ^ - /(cj-o
M<1
»
kl
and let g,: V —» £>!/2, t = 1,... ,d, be branches of JQ1 restricted to £>i/2- If Hi = J/»(V), then the Zi, are disjoint measurable subsets of D i / 2 whose union is D1/2, and /o in V.
1276 ALGORITHMS FOR SOLVING EQUATIONS
185
Note that c = ft(-ao) if and only if /o(f) = -ao, or equivalently, f(c) = 0 when f = fo + ao- Then
/
7r 2 i/o(f)i 2 *=E/ 4
-£/
->r2i/o(?)i2*
t
(l{f>i{~ao)))2 dao
<=iJa
-i -L
«o€V
£
(change of variable)
T(f)-a*>o
/(f)=„
l«ol<» Ki /( C )=0 LEMMA 3. Tftere is a constant > 0 aaefc tfort 7~3 > /Ci|/'(f)|2 »/|f| < 5. kl
PROOF. Suppose 7t = 7(f,/). Use Theorem B so 7f
" MM)
Imi
!/'(?)!
Therefore
^(l/2) a ITOI2 for/€P<jand|f|<$. The case of max(l,7(f,/)) = 1 is easily checked directly. This yields Lemma 3 with K\ easily estimated explicitly. Then Lemmas 1, 2, and 3 yield (I); just note / ' = / 0 . Toward the proof of (II), we have LEMMA 4. For any complex polynomial g,
I,
\t[
PROOF. Use polar coordinates. The odd powers of 2 give a zero contribution and the even powers a positive contribution. The lemma follows easily. Now apply Lemma 4 to obtain (II). Here ?(?) = /'(f) 2 and g(0) = (01)2. This gives thefirststep into the integration of (II). Next one integrates over |ax | < 1 using polar coordinates. The remaining integrations are trivial. This proves (II) and hence the theorem. Another perhaps even simpler proof of the analogous theorem, using approx imate zeros (of the first kind), goes by Theorem A, Lemmas 1 and 3 above, and LEMMA 5. There exist universal positive constants 6 and e such that if /(*) = E ? - o 0 ^ ' M ^ 1> l«ol/l«ila < e, and \z\ < 6, then a(zj) < «o and so z is an approximate zero (of the first kind).
1277 186
STEVE SMALE
4. Here we give the proof of the theorem on the speed of a Global Newton Method. LEMMA 1. There is a positive constant K with this property. Given any (zo,f) with Q(ZQJ) < a0, let zx = W}(zo), / = 1,2,..., and c = lim,-,*,*,. PROOF OF LEMMA 1. Let a, = a(z,)> Vi = ^(a/), / = 1,2,..., where tl>(r) = 2 r 2 - 4 r + l. Then aj decreases as / increases and a; < ( | ) 2 ~1a(z0) (see [P.E.], §4, Proposition 1 and §3, Proposition 2). Therefore 1/^j decreases to 1 from above (and very fast) with / (see the beginning of §4 in [P.E.]). Next define *0=(l-a(zo)M"(*o))
"*
Kl=
l = h2
{i^Wr
By Lemma 2c, §3 of [P.E.], -yfo) < #fi_i-y(«»-i) *> l(*i) < 7(«b)Il}«J>ff>Then it is sufficient to show that Ylj^o ^i increases to a finite constant K as / —oo. This follows from
io
i-i
i-i
K
i-i
i-i
« n i = n 8 > = E - «& - E io«(i - a ') 0
io K
0
io
0
0
and our estimate on Q| above. Q.E.D. Let Ki < \\ be as in Proposition 1 of §2 of [P.E.], and define L$ = KK^ao where K is as in Lemma 1. LEMMA 2. As in Lemma 1, let a(zo,f) < 2
OQ, Z\
=
NJ(ZQ),
and c = limzj.
,
T/.en||^-c||7(c)<(i) '- L0. PROOF OF LEMMA 2. By Proposition 1, §3 of [P.E.],
ll*-fll
II* - slhr(?)< K*{\)2"1 \\*i - *o\M*o)K. Lemma 2 follows from the definition of LQ and the fact that \\z\ - z0\h{zo) = a(zo) < a0Next we will prove the theorem. It is sufficient to prove that a(zi, f-Wi+i) < aro for t = 0 , 1 , . . . , n — 1, and for this we proceed inductively. First observe a(z„ f - wl+1) < iMWDHztr'ifizi) +
- Wi)\\ 1
l(zt)\\Df(zi)- (wt-wl+l)\\.
It is sufficient to show that (a) and (b) are true for a suitable fixed /, where (a) 7 ( * ) l | 0 / ( * ) - 1 ( / ( * ) - VH)\\ < a o ( i ) 2 ' - 1 , (b) ^zMDflzi)-1^ - wi+l)\\ < oo(l - (I) 2 '" 1 ).
1278 187
ALGORITHMS FOR SOLVING EQUATIONS
For i = 0, (a) is true since f(zo) = «*>. For i > 0, use [P.E.], Proposition lb of §4, with o = \ to obtain a(zi, f - u>i) < (£) 2 - 1 a ( z , _ i , / - vii) which, by the induction hypothesis, is less than (j) a ~lato- That proves (a). For the proof of (b), define ft = limfc_oo Nf-w(zi-i) so that /(ft) = W{, and let T = H*, - ft||7(ft). By Lemma 2, r < (*) 2 " £,0, and it can be seen that for / larger than some small /o (= 4 for example), r < 1 - \/2/2. By Lemma 2c, §3 of [P.E.),
Also using Lemma 2b, §2 of [P.E.], | | £ > / ( * ) - > , - wi+l)\\ < WDMr'DfMW I{>(T)
\\Df{qi)-l{wi
- tOi+1)||
\\wi\\
Therefore, for (b) it is sufficient to show
«'J5«-»(™)<-('-(if) or Q(?.,/) ( 1 - 0 A l l / / -
\\\^n
or ,(1-0 MAH/tzoJIIi^T^ < a 0 V-(r) or, for some c> 0,1, 1-r
V>(r)2J T = ( 1 / 3 ) 3 ' -
(-©"I
(-0 (-en
< a0 ,
I,O
The last equation is clear and defines a choice of constants c and / of the theorem. Q.E.D. 5. We prove here the theorem on the regularization of the LPP. Then we add some discussion related to Newton method algorithms for this problem. Toward the proof of the theorem we have LEMMA 1. The inverse of the derivative D$a(x): RN -» KN exists and has a bound on its norm independent of M, ||0*aa(zr1||<4max(^±^) REMARK. This already shows the continuity of the solution of our perturbed LPP in (A, b, c). PROOF OF LEMMA 1. First note that M is antisymmetric so that MTx = —Mx, MT the transpose of M. This is clear from the structure of M as
1279 188
STEVE SMALE
M
= [A ~o 1 • T h u s f o r any v G R N , (Mv, v) = 0 where (•, •) is the usual inner product on R n . This fact is crucial for all the subsequent analysis (although the milder con dition (Mv, v) > 0 is sufficient). Fix the following notation for the rest of this section. A = diagmatrix
({x*+*a)i/a)^
so 2D*a(x) = I + A + M(I - A). The above observation yields (27?* 0 (z)(u), (7 - A)u) = ((/ + A)u, (/ - A)«). Applying the Schwartz inequality to the left side gives us |{(/ + A)«, (7 - A)u)| < 2||£>* a (*)(«)|| ||(/ - A)(u)||. The left side of the last estimate equals £ u ? ( a 2 / ( * ? + | | ( 7 - A H | < 2 | | u | | . Thus
fl2
)); moreover
m in
. ( ? T ? ) I H I - ^l^-WMH-
The last estimate is sufficient to yield Lemma 1. LEMMA 2. For real numbers a, x,y, (x-y)2-((x2+a2)1/2-(2+«2)1/2)2 = 2((xy + a 2 ) 2 + a 2 (x - y) 2 ) 1 ' 2 - 2(xy + a 2 ). The proof just goes by expanding out and comparing terms. Note that the right-hand side is always positive unless x = y. Thus the same is true for the left-hand side. LEMMA 3. The map $„: R v -» R N is one to one if a > 0. PROOF. Let A, = ((x2 + a 2 ) ' / 2 , . . . , (x% + a 2 ) 1 / 2 ), so 2*„(A) = x + \ x + Af(i-Aa). We have 2(* 0 (*) - «.(»), (x - A,) - (y - A,)) = <x + Az - (y + A„), (x - A«) - (y - Av))
= Dx.-y.) a -((x?+aY /a -(K?+« a ) 1/ ») a 1=1
Now apply Lemma 2 to finish the proof of Lemma 3. LEMMA 4. The image of$a:
RN -* RN is contained in U\t-
PROOF. Since $ 0 (x) = *a (x) + A/*~(x), *o(x) is a positive linear combi nation of the coordinate vectors e\,..., e;v and the columns of -M. Even more explicitly, note that
= -a2/4,
•= !
TV.
1280 ALGORITHMS FOR SOLVING EQUATIONS
189
If Mk is the Jfcth column of AS,
Now consider the usual LCP which, as always, we suppose is generated by the LPP. A result of Cottle [52] (the Theorem on p. 663) can be interpreted as LEMMA 5. Given M, $M{X) = q has a solution x if and only ifq belongs to the closure UM of UM . LEMMA 6. Given M and q € UM, there is an a" such that for each 0 < o < a* there is a solution x ofQa(x) = q. The proof of Lemma 6 goes by using a topological degree argument and Lemma 5. Let V = Vs{q) be the closed ball of radius 6 about q in RN, with 6 chosen so that V«(g) c UM- Then from standard linear programming theory, the restriction * M / : * M 1 ( V ) ~~* V 18 w e " defined, proper, and of degree 1. There fore, since $ a uniformly approximates * M , we obtain the conclusion of Lemma 6. To finish the proof of the theorem, it is sufficient to show that the map *„: RN -» UM has image all of UM- This argument proceeds as follows. Fix M,q and consider the solution i(o) of $a(x(a)) = q, given by x(o) = *~*(q) for small a, by Lemma 6. Thus x(o) is defined for a < a*(q) where a* (q) may be supposed maximal. We will prove that a* (q) = oo. For each q (fixed A/), x(a) satisfies an ordinary differential equation by dif ferentiating Qa(x(a)) = q with respect to a. This equation is dx
n * / \-i
d
*<»
Here
»£-<'-«> 0 ^ ) 1 and 2£>*„ = (/ + A) + M(J - A) as above. Let u = dx/da. Then u satisfies
Take the inner product with the quantity in braces to obtain
or ||u|| 2 = ||A« -I- a/(x? + a2)1'2!*!!2. This implies
1281 190
STEVE SMALE
or yet
E
(u, - z,/a) 2
Thus for each t, | | m - * . / o | < > / N ( l + (* 4 /a) a ) 1 / a and so u, satisfies an estimate of the form |u,| < K\ + Jf 2 |a: t /a|. This bound implies our a*(q) = oo, the surjectivity of * a : RN -* UM and hence the theorem. Next we show how the theorems on approximate zeros can be used to initiate algorithms for solving our approximation of the LPP. Note that 1»a(0)=a(J^)«?0,
90 = {1
I)-
Define an associated map $<,(*) = *o(z) - *o(0). Then if 40(a:(a)) = q, as a —► 0, x(a) tends to a solution of $M{X) = q and hence the original LPP. Moreover, x(o) satisfies the ordinary differential equation dx
i
x-i#*a/ ^
To find an initial value for this equation (to let a decrease to 0) it is sufficient to solve ia(x) = q for some value of a. This motivates a PROPOSITION. If a > 8||g||/ao, then 0 e R N is an approximate zero of Thus for o = 8||g||/ao, one can solve $a(x) = q with great precision in a few (e.g., 5) Newton iterations starting at 0. Here recall ao > f • We use Theorem A and show that Q(0, * a - q) < QQ. LEMMA 7.
Let ±
* ± (x3 + a 2 ) 1 / 2
i/(*-i)
/2|»(*)(0)|\I/t'-1,;,4_ PROOF. Our proof uses Schlicht function theory (in particular, Loewner's work on the Bieberbach Conjecture)! An elementary estimate of the higher derivatives can be done, but is not trivial. The proof assumes the plus sign, but with
1282 ALGORITHMS FOR SOLVING EQUATIONS
191
Thus a(x,
KN be an analytic map of the form G(x) =
\\DkG(x)\\ =
(EW*) fc/a
where Xi = g+ (it). PROOF.
For
v€RN,
DkG(x)(v") = (Ai»f,
,\NvkN),
\\DkG(x)(vk)\\ = (£(A.t,*) 2 )
1/2
and the maximum over \\v\\ = 1 is at vk = KXi for some K > 0. So 1/2
IH*7<«)(«*)|| = K ( £ A?) ' ,
AT = 1/ ( £ A3'*)
.
Q.E.D.
Now we can easily make the estimate on o(0, Qa — q). We have 4o(0) = 0, £>*a(0) = (I + M)/2, ||D*a(0) _1 |l < 2, so that 0(0, * a -q) < 2\\q\\. Moreover, 7 ( 0 , * o " 9) < SUP
,P fc »,(0) M « l (0)-
1/( _1)
*
ifc!
fc>l
(I + ^fc)(0)
fc!
M)-l{I-M)
D*-»A fc!
i/(fc-i)
i/(*-J)
a
We have used Lemmas 7, 8, and 9. Thus a = 0i < 8||g||/a < a0, proving the proposition. 6. In Smale [51], one dozen problems in the algorithms of analysis were ex plicitly listed. Since that paper was written, solutions or progress for half of them have resulted. In this section we report on these results. Problem 1 deals with extending work of Smale [50], Shub-Smale [43, 44], and on zero finding of polynomials. This work proves probability bounds for algo rithms related to Newton's method, and eventually, that the number of iterations is proportional to the degree of the polynomial. This work used the theory of Schlicht functions (related to the Bieberbach Conjecture) and did not extend to finding zeros of polynomial systems / : C -♦ C n where n > 1. Problem 1
1283 192
STEVE SMALE
asked for such results. Renegar [40] has now found polynomial bounds, statisti cally, for each n. This important work thus gives some answer to our problem. However, deep questions, including those on the conceptual level, remain. The bounds of Renegar, while polynomial in the degrees, are crude. For one variable, for example, the number of iterations is on the order of d26 (in contrast to d of Shub-Smale [44]). Renegar's results are one of the inspirations for this paper, and I hope that the theorems here will make some contribution to the problem of complexity for polynomial systems. Problem 6 deals with the question of understanding systematically the set of polynomials of one variable, where Newton's method can cycle on open sets. Cayley showed that for quadratic polynomials, this could not happen generically. Janet Head [18] of Cornell has given a good analysis of the situation of cubics and has shown that cycling in that case is a rather rare occurrence. Her work does not include / of degree higher than 3. Problem 7 raises the problem as to how often (simple) Newton's method con verges for polynomials of one variable. Let A/ be the normalized area of the set of points of £>j which converges to a zero of / under Newton's method. Let Ad =
min
/e^(i)
At.
Joel Friedman [12] at the University of California at Berkeley proved my con jecture that Ai > 0, all d, and gave some estimates on Ad as a function of d, dealing well with Problem 7A. For Problem 7B, see Problem 9 below. Problem 8. In it, a graph Tf is assigned to each polynomial / which reflects the topology of / as a map, especially related to Newton's Method. The problem itself asks what graphs can occur. Work dealing with this had been already initiated by Jongen-Jonker-Twilt [22]. See their two articles for references and an account of their work. Moreover, Shub-Tischler-Williams [46] have given a good study, and one could say that the problem is essentially solved. Problem 9, on the average area of the set of approximate zeros being bounded below, is answered completely by the theorem proved in §3 of the present paper (incidentally, answering Problem 7B as well). Problem 10 conjectured that there exists no purely iterative generally conver gent algorithms for polynomial zero solving over C. Curt McMullen [33] solved the problem completely. More precisely he showed this THEOREM. Let d > 3 and T: PdX S -* S be any map rational over C in f and z. Then there is no open setUcPdxS of full measure with this property. If (/, z) € U, then Tf(z) = z^ converges to a root of f as k —» oo. Here Pd is the space of all polynomial of degree < d and S is the Riemann sphere. That T is rational over C means that T can be formed from the complex rational operations (+, - , x,-r) in the coefficients of / and z. For d = 2, New ton's method or T/(z) = z-f{z)/f'{z) is such an algorithm as proved by Cayley.
1284 ALGORITHMS FOR SOLVING EQUATIONS
183
For d = 3, McMullen [33] found a new algorithm which is purely iterative and generally convergent. In McMullen [34, 36] he deepens his investigation. In Shub-Smale [45], Shub and I showed that if one adds the operation complex conjugation, then one can construct purely iterative generally convergent algorithms. Moreover, our result holds for polynomial maps of several variables, / : C -+ C . REFERENCES 1. J. Abadie and G. Guerrero, Mtthodc du GRG, mithode de Newton globale et ap plication a la programmation matMmoHqut, RAIRO Rech. Oper. 18 (1984), 319-351. 2. M. Berger, Non-linearity and functional analyst*, Academic Preaa, New York, 1974. 3. L. Blum, Toward* an asymptotic analysis of Karmakar's algorithm, Inform. Proceaa. Lett. 38 (1988), 189-194. 4. S. Chow, J. Mallet-Paret, and J. Yorke, Finding zeros of map*: homotopy methods that are constructive with probability one, Math. Comp. S3 (1978), 887-899. 5. R. Cottle, Note on a fundamental theorem in quadratic programming, J. SIAM 13 (1984), 683-665. 6. R. Cottle and G. Dantaig, Complementary pivot theory of math, programming in math, of the decision sciences, G. Dantaig and A. Veisott Jr., eda., Amer. Math. Soc., Prov idence, R.I., 1988, pp. 115-138. 7. J. Curry, On tero finding methods of higher order from data at one point, MSRI Preprint, 1986. 8. G. Dantaig, Linear programming and extensions, Princeton Univ. Preaa, Princeton, N.J., 1963. 9. B. Dejon and P. Henrici, Constructive aspects of the fundamental theorem of alge bra, Wiley, New York, 1989. 10. J. Dieudonnt, Foundations of modern analysis, Academic Press, New York, 1960. 11. C. Bavea and H. Scarf, The solution of systems of piecewise linear equations, Math. Oper. Res. 1 (1976), 1-27. 12. J. Friedman, Rough draft of a theorem on Newton's method, Univ. of California Preaa, Berkeley, Calif., 1985. 13. F. Gao, Nonasymptotic error of numerical integration—an average analysis, Univ. of California Preaa, Berkeley, Calif., 1986 (to appear). 14. , Probabilistic analysis of automatic integration (to appear). 15. C. Garcia and F. J. Gould, Relations between several path following algorithms and local and global Newton methods, SUM Rev. 33 (1980), 263-274. 16. G. Ghellinck, and M. P. Vial, A polynomial Newton method for linear programming, CORE, Louvain, Belgium, 1966. 17. P. Gill, W. Murray, M. Saundera, J. Tomlin, and M. Wright, On projected Newton barrier methods for linear programming and an equivalence to Karmarkar's projective method, Tech. Report SOL 85-11, Dept. of Op. Re*., Stanford Univ., Stanford, Calif., 1985. 18. J. Head (to appear). 19. P. Henrici, Applied and computational complex analysis, Wiley, New York, 1977. 20. M. Hirach and S. Smale, On algorithms for solving f(x) = 0, Comm. Pure Appl. Math. S3 (1979), 281-312. 21. H. Th. Jongen, P. Jonker, and F. Twilt, The continuous desingularized Newton's method for meromorphic functions, Memorandum No. 501, Dept. of Appl. Math., Twente Univ. of Technology, 1985. 22. , Non-linear optimisation theory in R n from a global point of view. VIII, Memorandum No. 559, Dept. of Appl. Math., Twente Univ. of Technology, 1986. 23. L. Kantorovkh and G. Alrilov, Functional analysis in normed spaces, Macmillan, New York, 1984.
1285 194
STEVE SMALE
34. N. Karmarkar, A new polynomial-time algorithm for linear programming, Combinatorica 4 (1984), 373-395. 26. H. Keller, Numerical solution of bifurcation and nonlinear eigenvalue problems, Application of Bifurcation Theory, Academic Press, New York, 1977. 36. , Global homotopie and Newton methods. Recent Advances in Numerical Anal ysis, Academic Press, New York, 1978. 37. M. Kim, Ph.D. Thesis, Graduate School, CUNY, 1985 (to appear). 38. H. Kung, The complexity of obtaining starting points for solving operator equa tions by Newton's method, Analytic Computational Complexity (J. Traut, ed.), Academic Press, New York, 1976. 29. S. Lang, Algebra, Addison-Wesley, Reading, Mass., 1963. 30. , Real analysis, Addison-Wesley, Reading, Mass., 1983. 31. O. Mangasarian, Equivalence of the complementarity problem to a system of non linear equations, SIAM J. Appl. Math. 31 (1976), 89-92. 32. M. Marden, Geometry of polynomials, Math. Surveys, no. 3, Amer. Math. Soc., Providence, R.I., 1966. 33. C. McMullen, Families of rational maps and iterative root-finding algorithms (to appear). 34. , Automorphisms of rational maps I: Nielsen realization and dynamics on the ideal boundary, MSRI, Berkeley, Calif., 1986. 35. , Automorphisms of rational maps II : braiding of the attractor and the failure of iterative algorithms, MSRI, Berkeley, Calif., 1986. 36. N. Megiddo and M. Shub, The boundary behavior of interior methods for linear programming (to appear). 37. L. Nazareth, Homotopy techniques in linear programming, Algorithmica (to appear). 38. A. Ostrowski, Solutions of equations in Euclidean and Banach spaces, Academic Press, New York, 1973. 39. V. Pan, Algebraic complexity of computing polynomial zeros, Tech. Report 85-27, SUNY, Albany, N.Y., 1985. 40. J. Renegar, On the efficiency of Newton's method in approximating all zeros of a system of complex polynomials, Math. Oper. Res. (to appear). 41. , A polynomial-time algorithm based on Newton's method for linear program ming, MSRI, Berkeley, Calif., 1986. 43. H. Royden (to appear). 43. M. Shub and S. Smale, Computational complexity: on the geometry of polynomials and a theory of cost. I, Ann. Sci. Ecole Norm. Sup. (4) 18 (1985), 107-142. 44. , Computational complexity: on the geometry of polynomials and a theory of cost. II, SIAM J. Comput. 15 (1986), 145-161. 45. , On the existence of generally convergent algorithms, J. Complexity 2 (1986), 2-11. 46. M. Shub, D. Tischler, and R. Williams, The Newton's graph of a complex polynomial (submitted to SIAM J. Math. Anal). 47. S. Smale, A convergent process of price adjustment and global Newton methods, J. Math. Econom. S (1976), 107-120. 48. , On the average number of steps in the simplex method of linear program ming, Math. Programming 27 (1983), 241-262. 49. , The problem of the average speed of the simplex method, Mathematical Pro gramming: The State of the Art (Bonn, 1982), Springer-Verlag, Berlin and New York, 1983, pp. 530-539. 50. , The fundamental theorem of algebra and complexity theory, Bull. Amer. Math. Soc. (N.S.) 4 (1981), 1-36. 51. , On the efficiency of algorithms of analysis, Bull. Amer. Math. Soc. (N.S.) IS (1985), 87-121. 52. , Newton's method estimates from data at one point, Proceedings of a Confer ence in Honor of Gail Young, Laramie, Springer, New York, 1986 (to appear). (Referred to as [P.E.] in text.)
1286 ALGORITHMS FOR SOLVING EQUATIONS
195
53. A. Wierzbicki, Note on the equivalence of Kuhn-Tucker complementarity condi tions to an equation, 3. Optim. Theory Appl. S7 (1982), 401-405. 54. S. Wong, Newton's method and symbolic dynamics, Proc. Amer. Math. Soc. 91 (1984), 245-353. 55. H. Woiniakowski, A survey of information-based complexity, 3. Complexity 1 (1985), 11-44. 56. Paul Wright, Statistical complexity of the power method for Markov chains, Univ. Calif., Berkeley, 1986. (to appear). UNIVERSITY OF CALIFORNIA, BERKELEY, CALIFORNIA 94720. USA
1287
The Newtonian Contribution to Our Understanding of the Computer S T E P H E N SMALE
Isaac Newton made a continuous mathematical model of a discrete universe using differential equations to explain how things in that universe move. Mathematicians can look at the computer today in the same way that Newton looked at the world in the seventeenth century to gain some understanding of "the machine." One can try to understand die computer in a way similar to the way a physicist understands the world, a different way from the engineers' ap proach. In order to understand the laws of computation one can design a mathematical model using continuous mathematics and calculus to lead to a deeper understanding of the process of computation. The model involves going from the discrete to the continuous. An example which helps explain die difference between the concepts of discrete and con tinuous is that of clock faces. An old watch widi a traditional face is a continuous object. The hands move continuously around the face and the whole spectrum of "possible" time. One gets a continuous picture rather than a discrete pic ture of the time. A modern watch using digits shows only the present minute. As time passes it will show you another and another, but each minute is separate and discrete. The world before Isaac Newton was perceived as fundamentally discrete in nature. Thomas Kuhn's book The Copemican Revolution discusses very well the background of science when Newton began to work. The understanding of the universe was based on the theory of atomism which was first promulgated by the Greeks, Democritus and Lucretius. Atomists described the matter in the universe as being composed of a finite - discrete - number of particles. Each particle is indivisible so they are called the elementary particles of matter. It was this picture of the nature of matter and the universe which was generally accepted when Newton began his work. Newton believed diat the universe was made up of a finite number of building blocks. He used the word corpuscular and he believed in a corpuscular universe. On the other hand, Newton's vision of the universe was based on the geometry of the Greeks with its notions of curves and lines, which are con tinuous objects. There is nothingfiniteat all about a smooth, continuous curve.
90
QUEEN'S QUARTERLY 9 5 / 1 (SPRINC 1 9 8 8 )
1288
The resolution of the problem of the discrete universe and his continuous mathematical models was a crucial step for Newton. According to Thomas Kuhn diis seeming contradiction delayed die publication of Newton's Principia. Real numbers are crucial to Newton's vision of a continuous universe. They are the numbers representable by infinite decimal expansion. Real numbers give the picture of a continuous set of numbers between zero and one. This concept is basic to an understanding of how Newton reinterpreted the world. The success of real numbers in his continuous picture of the universe contrasted greatly with his picture of the world as being finite. Newton saw the Earth as made up of a lot of particles; each has gravity, is affected by the others, and has different forces pulling upon it in a very complex situation. He resolved this dilemma by calculating the forces exerted on the accumulation of particles of the Earth in this finite picture. He doubled the number of particles he con sidered in his calculation, filling out the picture more and more densely and making one calculation after another. He did it again and again, each time making a new calculation. Since mathematically the number of particles goes to infinity, one can think of die Earth analogously as a single particle located at me centre of an infinite universe. Kuhn, in die Copemican Revolution, says: "In 1685 [Newton] proved Uiat, whatever the distance to the external corpuscle, all the earth corpuscles could be treated as though they were located at the earth's centre. That surprising discovery, which at last rooted gravity in the individual corpuscles, was die prelude and perhaps the prerequisite to die publication of Principia'' (258). And Kuhn adds: "At last it could be shown that both Kepler's Law and die motion of a projectile could be explained as the result of an in nate attraction between the fundamental corpuscles of which the world machine was constructed" (258). Once Newton had devised that mathematical framework he could use the tools of calculus to understand the motions of the planets. Calculus is the mauiematics of ordinary differential equations, on die space of states of physics. The "states* in physics describes the position of a particle widi its velocity, or the positions and velocities of a number of particles. The basic laws of physics define differential equations on die space of states which govern how a state moves in time. For example, one might want to calculate die path of a planet widi a specific velocity. Newton's Law, which was diat force equals mass times die action (f = ma), explains it in mechanical terms. But a differential equa tion explains it in madiematical terms, such as dx/dt equals some function of x, so x is a state and x will move in time. This equation says diat motion is unique when one knows where die particular state is at time zero. In some sense this idea is die principle one diat inspired die Principia and the way Newton cal culated, for example, die ellipses of die planets. He explained die ellipses of
UNDERSTANDING THE COMPUTER
91
1289
Kepler in terms of differential equations using the laws of gravity. Kuhn says: "The construction of Newton's corpuscular world machine com pletes the conceptual revolution that Copernicus had initiated a century and a half earlier" (361). But Kuhn emphasizes the fact that Newton was assuming a corpuscular, atomistic, discrete world. And there were things that Newton did not deal with .sufficiently in the Prittcipia, things which have to do widi par tial differential equations. These are die laws by which fluids and solids move, the laws of motion of terrestrial objects in general. Newton attempted to deal with these matters in his book, but the attempt was successful only in oudine, and it was left to such madiematicians as Daniel Bernoulli, Leonhard Euler and J.L. Lagrange to deal with diese matters direcdy. They carried out die pro gramme that was implicit in Prittcipia by formally discovering the partial differential equations diat show how fluids move. Eider's equations offluidmo tion are an example. Alan Turing was a very interesting and tragic figure, who about 1936 for malized our notion of a computer in a madiematical model which provides die theoretical foundation of digital computer science. His great contribution to computer dieory is known as the Turing machine. A luring machine is not a stationary object widi moving parts and levers and buttons. It is a "machine" on paper, a theoretical model. (An example of a machine which is a paper idealization is die flow chart, which depicts die se quence of work in a machine, die programme of a machine used for solving problems and scientific computations.) A simplified, idealized way to under stand how a digital computer works might be to picture it as an input tape widi information encoded on it consisting of either a zero or a one. As it moves through die computer a counter moves back and forth according to a system of instructions. What is on the tape when die process is completed is called die output of die machine. One sees in diis picture of die luring machine die prop erty of discreteness. There is a long string of zeros and ones, and die machine can encode a lot of information because of die many possible combinations, but it is a very discrete process in die sense diat die input of die machine con sists of die limited set of numbers diat I have been talking about. Aho, Hopcroft, and Ullman discuss die random-access machine in a way which is much closer to die way we dunk of die computer today. But even in random-access computers one dunks of die inputs as discrete numbers and diey carry out die same input/output functions as luring described. I am quite critical of diis idea of die computer as afiniteinstrument, lur ing modelled his theory on logic which is a very specific and narrow part of madiemadcs. This has kept it away from die mainstream of madiematics and hindered its development. John von Neumann, die madiematician who built
92
QUEEN'S QUARTERLY
1290
some of thefirstcomputing machines, writes frequently about the foundations of computer science and the general and logical theory of automata. In his opin ion, "We are very far from possessing a theory of automata which deserves that name." He said the reason for this lack was that logic has very little contact with the continuous concept of a real or complex number, that is, with mathematical analysis. He went on to say: "A detailed, highly mathematical and more specifically analytical, theory of automata and of information is needed." Computer scientists still rest their theories on the discrete idea expressed by Turing machines. Why I object to that, and why von Neumann expected more has to do with scientific computing, which is the most complex use of the computing machine. Like differential equations, scientific computation ad dresses especially complex problems, such as predicting weather and providing aerodynamic models. Numerical analysis, the theoretical side of scientific com putation, provides some basis for solving differential equations using algorithms. Some scientific people - engineers in scientific computation or theoreticians of numerical analysis - have very little respect for the Turing machine. The Turing theory is seen as old-fashioned and limited and remote from the kind of explanation they need. There is also a move away from the use of calculus in the discrete mathematics of computer science. What this has caused is an unhappy lack of unity between two disciplines, computer science and numerical analysis, two subjects which should be intimately related. What I would like to see is a programme for changing the foundations of computer science theory in the same way that Newton changed physics in the Principia, to idealize models for the computer which will relate to the algorithms of numerical analysis, and to explain mathematically the use of the idealization for scientific computation. What I want to suggest is a different model for computer theory based on the idea of using real numbers. It will be an idealization just as the Turing machine model is an idealization, and it provides a different model to explain die same phenomena. In physics one has the Newtonian picture, a mechanical picture of the universe, but one has also a relativistic picture, and a quantum picture. Relativity is useless to explain most of engineering, which is based on Newton's mechanics. Different models explain different aspects of the same physical reality. One should not look for a single model of the computer, but rather look for models which will help us understand scientific computation on the one hand and classical computer science on the other. Why is there this dichotomy between the two? What makes the computer discrete? Digital machines make calculations using data represented by the digits between zero and one. Machines have a great deal of precision but it is still afiniteprecision: they will enter numbers with accuracy, io"8, for exam-
UNDERSTANDING THE COMPUTER
93
1291
pie. Translated into this picture it means that they will put in, between zero and one, io8 numbers, about a hundred million. The computer will allow us to enter without doing any special programming a hundred million numbers in this interval. What this means is that while mere is a discrete, finite, pic ture in terms of the input, the numbers that are inputted into the computer approximate the real numbers very well. To extrapolate the input from a finite number to a continuous picture of the inputs, to a real number input, is not difficult. It is an idealization, but it is a plausible idealization. We know that the universe is a discrete universe, and die laws of differential equa tions - Newton's mechanics - are still adequate to explain how matter moves within this discrete universe. In die same way I would suggest that we can understand at least one major facet of the computer better if we smooth over the discrete set of numbers to make a continuous set, like Newton did when he expanded a discrete particle dieory to describe a continuous universe. One can describe a machine in very abstract mathematical language but I will do itfirstin a simple way. I will start by describing a "tame" machine by means of the real numbers. Those of you who have studied computer science will say "that's nothing new," but let us see how a general foundation for com puting functions over the real numbers develops. In the domain of discrete mathematics, computer scientists have system atized the notion of a machine into recipes for solving discrete problems. What one hopes is that that could develop into a way of systematizing recipes for mainstream mathematics. An example of this is the fundamental theorem of algebra. You can imagine a function being given, say, z into some polynomial p (z) by such a machine (theory) where we do not allow any division. We can imagine z going into/) (z) by addition, subtraction, and multiplication. That describes the polynomial. Is there some point which goes into zero under this particular machine? There is and it is given by the fundamental theorem of algebra. One has to allow things that are called complex numbers. How can one find the solution to the equation p (z) = o, where p is the polynomial az2 + bz + c = o? This would be a polynomial of degree 2. Then one can study polynomials of degree d which look like
94
QUEEN'S QUARTERLY
1292
tional algorithms won't work because one needs things like calculus and con tinuity. lb solve this kind of problem, die topological complexity of die method of solving die problem must grow at a certain rate. My colleagues, Lenore Blum and Mike Shub, and I have extrapolated from an ordinary computer die idea of a real number universal machine which will contain all ordinary machines as a special case. Such a universal machine idealizes the actual computer which can solve all kinds of problems. Develop ment in this dieoretical direction is veryfirmlybased on die idea of going from die discrete to the continuous as seen in Newton's Principia. Newton's passage in Principia from die corpuscular physical world to his fun damental differential equations of motion set die stage for centuries of developments in science. Three hundred years after Principia, die computer revolution is transforming society. We can wonder if die passage from die discrete picture of die computer to a continuous model will deepen our understanding of die laws of computation. WORKS CITED
Aho, A., J. Hopcraft andj. Ullman. The Design and Analysis of Computer Algorithms. Reading, MA: Addison-Wesley, 1979. Hodges, Andrew. Alan Turing: The Enigma. London: Burnett, 1983. Kuhn, Thomas S. The Copemican Revolution. Cambridge: Harvard UP, 1957. Von Neumann, John. Collected Works. 6 vols. New York: Pergamon Press, 1961-63.
UNDERSTANDING THE COMPUTER
95
1293 BULLETIN (New Scnn) OF THE AMERICAN MATHEMATICAL SOCIETY Volume 21. Number I. July 1989
ON A THEORY OF COMPUTATION AND COMPLEXITY OVER THE REAL NUMBERS: //^-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES' LENORE BLUM2, MIKE SHUB AND STEVE SMALE ABSTRACT. We present a model for computation over the reals or an arbitrary (ordered) ring R. In this general setting, we obtain universal machines, partial recursive functions, as well as A'P-complete problems. While our theory reflects the classical over Z (e.g., the computable func tions are the recursive functions) it also reflects the special mathematical character of the underlyingringR (e.g., complements of Julia sets provide natural examples of R. E. undecidable sets over the reals) and provides a natural setting for studying foundations! issues concerning algorithms in numerical analysis.
Introduction. We develop here some ideas for machines and computa tion over the real numbers R. One motivation for this comes from scientific computation. In this use of the computer, a reasonable idealization has the cost of multiplication independent of the size of the number. This contrasts with the usual theoretical computer science picture which takes into account the number of bits of the numbers. Another motivation is to bring the theory of computation into the do main of analysis, geometry and topology. The mathematics of these sub jects can then be put to use in the systematic analysis of algorithms. On the other hand, there is an extensively developed subject of the theory of discrete computation, which we don't wish to lose in our theory. Toward this end we define machines, partial recursive functions, and other objects of study over a ring R. Then in the case where R is the ring of integers Z, we have the same objects (or perhaps equivalent objects) as the classical ones. Computable functions over Z are thus ordinary computable functions. R.E. sets over Z are ordinary R.E. sets. But when the ring is specialized to the real numbers, we have computable functions which are reasonable for the study of algorithms of numerical analysis. R.E. sets over R are no longer countable and include, for example, basins of attraction of complex analytic dynamical systems. Received by the editors April 21, 1988. 1980 Mathematics Subject Classification (198S Revision). Primary 03D1S, 68Q15; Sec ondary 65V05. 1 Partially supported by NSF grants. Some of this work was done while Blum and Smale were visiting Shub at the IBM, T. J. Watson Research Center. 2 Partially supported by the Letts-Villard Chair at Mills College and the International Computer Science Institute, Berkeley. © 1 9 S 9 American Mathematical Society 027M>9?9/«9 11.00 ♦ S.2S per page 1
1294 2
LENORE BLUM. MIKE SHUB AND STEVE SMALE
There is another virtue of developing a theory of machines over a ring. It forces a more algebraic approach, closer to classical mathematics, than the approach from logic. Here is an abbreviated description of some of the results of this paper, in this context of machines over a ring R. (I) Most Julia sets are not R.E. over the reals, so their complements, the basins of attraction, provide natural ex amples of R.E. undecidable sets over R (§§1 and 10). The Julia set example provides an interesting link between the theory of computation and dynamical systems. A perhaps deeper link is the com puting endomorphism (§3) which is an important conceptual and technical tool used in our development. (II) An analogue of Cook's Af/'-completeness theorem is proved over the real numbers. The NP-complete prob lem over R is the 4-Feasibility problem, i.e. the prob lem of deciding whether or not a real degree 4 polynomial / : R" — R has a zero (§6). This result, in addition to focusing attention on the 4-Feasibility prob lem over R has some interesting consequences which point to the subtle differences between the theories of NP over R and over Z. For example, by straightforward counting arguments, any W/'-problem over Z is seen to be solvable in 2poiy("> time. (See eg. Garey-Johnson.) An analogous result over the reals is far from obvious since there are a continuum of possible guesses over R. It is not even clear a priori that iVP-problems over R are decidable. However, since the 4-Feasibility problem over R is decidable (by Tarski-Seidenberg) and since the current best upper bound for decidability of the existential theory of the reals is 40*"' (see Renegar, also Canny and Grigortv-Vorobjov), we also get exponential upper bounds for TV/'-problems over R but for much deeper reasons than the case over Z. For another interesting difference between the two theories, note that by Hilbert's Tenth Problem, the 4-Feasibility problem restated over Z is not even decidable over Z and so nor in NP over Z. PROBLEM. What is the relation between the problems P = NP over R, and P « NP over Z? (HI) Computable functions over R are characterized in trinsically by a class we call partial recursive junctions over R. For R — Z, these are the usual partial recursive func tions (§7). (IV) There exists a universal machine over R. This ma chine, inspired by the Universal Turing Machine, does the computation of any machine over R. The universal ma chine over R turns out to be independent of R. Moreover, by avoiding G6del coding via prime numbers, the algebraic structure of the universal machine remains intact (§8).
1295 Wf-COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
3
(V) Inspired by the work of Davis, Robinson and Put nam, and Matijasevic on Hilbert's tenth problem, we give a "diophantine-like" description of R.E. sets, for a certain class of machines (§9). There are a large number of contributions of mathematicians and com puter scientists which predate and relate to this work. A very brief survey, with some comparisons, follows. Of course, the work of Turing, Godel, Church and others in the thirties forms the core of the existing framework for our work. Although much of the classical theory of computation deals with computing over the natural numbers, certain approaches have considered other underlying domains. Gose to the classical approach, Rabin developed a theory of computable algebra andfieldsin which the underlying domains can be effectively coded by natural numbers and are thus, necessarily countable. On the other hand, the theories of computation over abstract structures, are perhaps more general than ours. See e.g., Friedman (or as discussed by Shepherdson in Harrington, et al.), Tiuryn, and Moschovakis. These general approaches both exploit and explore the logical properties of pro cedures. But, when applied to specific structures such as the reals, they do not yield the concrete mathematical results (e.g. about Julia sets or the 4-Feasibility problem) that quite naturally follow from the more mathe matical model developed in this paper. Recursive analysis provides yet another approach. See, e.g. FriedmanKo, Pour-El-Richards, Hoover and Kreitz-Weihrauch. Some tools here are recursive functionals, computable real numbers and oracle Turing ma chines where, roughly, one imagines a real number fed to the machine bit by bit. To contrast, we view a real number not as its decimal (or binary) expansion, but rather a mathematical entity as is generally the practice in numerical analysis. Thus, for example, we suppose Newton's algorithm for finding the zeroes of a polynomial / to be performed on an arbitrary real, not just a computable real; the fundamental components of the algorithm in our model, as in practice, are the rational operations Nf(x) * x - f(x)/f(x), not the bit operations. The development of algebraic complexity theory, in particular the work of Ostrowski, Pan. Winograd, Strassen and Schonhage (see von zur Gathen for a recent survey) gave rise to the "real number moder approach to com putation. Decision and computation tree models as in Rabin, Steele-Yao, Ben-Or, and the tame machines in Smale, are such real number models of computation but considerably less powerful or general purpose than ours (e.g., they have bounded halting time and none are universal; also they don't allow for uniform algorithms as do our infinite dimensional machines). More closely related are the register machines of Shepherdson-Sturgis and the RAM's or random access machines. (See Aho-Hopcroft-Ullman or Machtey-Young for discrete versions.) While a definition of a real RAM is given in Preparata-Shamos, the formal development of a theory is not
1296 4
LENORE BLUM. MIKE SHUB AND STEVE SMALE
pursued. Indeed, in their book, only a subclass of machines, equivalent to the class of decision trees, is utilized. Perhaps closest to our approach is the work of Herman-hard on computability over arbitrary fields. Also close in spirit is a theory of real Turing machines outlined by Abramson. Nimrod Megiddo has also considered an example of an AfP-complete over R in our model. Some other related papers are Borodin, Valiant, Endler, Lovdsz, and Eaves-Rothblum and Traub-Wozniakowski. Books having significant con tact with this paper include Davis, Eilenberg, Manin, Manna, Minsky and Rogers. Especially in §§5 and 6 below, the influence of complexity work of Cook and Karp (see Garey-Johnson) among many others, is evident. In our §8, Robinson, Matijasevic, Davis, and Putnam, and Denef have been influen tial. We would like to acknowledge helpful discussions with Martin Davis and Steve Simpson. Sections 1. Examples of machines over R 2. Machines over a ring R 3. The computing endomorphism and the register equations 4. Time T halting sets, equations, polynomials and computations 5. Complexity theory of machines over R 6. ^^-completeness and the analogue to Cook's theorem over R 7. Computable functions, normal forms and partial recursive functions over R 8. Existence of a universal machine over a ring 9. Characterizing R.E. sets as output sets and pseudo-Diophantine sets 10. Most Julia sets are undecidable 11. Some final remarks and problems References 1. Examples of machines over R. Even before defining our notion of a machine, we give some examples. The first examples are related to the theory of complex dynamical systems. We present them in some detail. EXAMPLE 1. Consider a complex polynomial map g: C — C. This map g will be considered as an endomorphism, mapping C into itself. Thus it makes sense to iterate it. That is g(g(z)) = g2{z) is defined as well as the fcth iterate gk(z), for each z eC. LEMMA. There is a real constant, C = Cg such that if\z\ > C, then \gk{z)\ — ► oo as k —> oo.
This is true because the highest order term of a polynomial dominates the others for \z\ sufficiently large. Moreover, if g0(z) = zd, |s£(z)| =
l*l'4.
Now we may define a "machine" M from g, by the following flow chart (see Figure 1). This M is a machine over R, not C, since it uses the real comparison \z\ < Cg; in the context of this machine, we view C as R2.
1297 ^^-COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
5
Compute
5
Input ziC
x—g{z)
\
g(z) and replace z by g{z) Branch
1*1 2 Ct Output 2
I \
I
\z\ < Ct
J
Z FIGURE 1
One can sec that the "halting set" QM of the machine M is precisely the set of points which eventually tend to oo under iterates of g. The halting set is analogous to the R.E. sets of recursive function theory, and eventually we will define a class of machines which contains not only this g machine, but machines equivalent to Turing machines as well. Thus we call Q.M an R.E. set over R. Note that it is certainly not a usual R.E. set since it is not countable. It is natural to ask, is Q-M "decidable" or, inspired by the classical tradition, is the complement of QM the halting set of some other machine over R?3 Of course at this point, not even having a definition of a machine over R, the question can't be answered. But later we will show PROPOSITION
algebraic sets.
1. Any R.E. set over R is the countable union of basic semi-
Here a basic semialgebraic set is a subset of Cartesian space R" defined by a set of polynomial inequalities of the form /»,(*)< 0, hj(x)<0,
/=1,...,/, j = l+l,...,m.
A general reference for semialgebraic sets is Becker. Using this proposition, we answer our question in the next example for a class of halting sets ClM. } If the complement of QM is the halting set of some other machine M\ we can construct a machine to decide for each z € C "Is r € fl«r' schematically as follows (see Figure 2).
Input
fTl
"Parallel proceu" Output
V hftsj
iftfbalti FlOUU 2
Jrl No
if W halts
1298 6
LENORE BLUM, MIKE SHU* AND STEVE SMALE
FIGURE 3
2. Specialize g to be a polynomial g(z) = z2 + c, where \c\ > 4. In that case the complement of QM is a Cantor set, a certain Julia set of complex dynamical system theory. Let / = C - £IMTo see that J is a Cantor set, first note that the exterior of the circle of radius \c\ - 1 is mapped into itself, in fact any point in it moves away from zero and tends to oo under iteration. As 0 is the only critical point of z2 + c and it maps to c which is in the exterior of the circle of radius \c\ - 1 we see that the inverse image of the interior of the circle is two discs interior to the circle (see Figure 3). Now the inverse image of these discs is four discs, 2 each in the interior of the two, etc., The intersection of the inverse images is precisely the set of points which don't tend to oo. At the very first stage we have the two inverse mappings of the disc into its interior, by the Schwarz lemma each of these is a strict contraction. Therefore any infinite nesting of inverse images contains exactly one point and the intersection of the inverse images is a Cantor set. Now since a basic semialgebraic set has a finite number of connected components, it follows from the previous proposition that J is not an R.E. set and hence EXAMPLE
2. ClM for the case g(z) = z2 + c, \c\ > 4, is an R.E. set over R which is not decidable over R. PROPOSITION
EXAMPLE 3. We now suppose, more generally, that g = p/q: C —» C is a rational endomorphism of degree at least 2 of the Riemann sphere C = Cu{oo}. Thus, (C, g) is a discrete complex analytic dynamical system. (See e.g. Blanchard.) Of primary interest is the long term behavior of points in C under the "action" of g. And so a key object of study is the orbit of a point zQ under g: z 0 ,Z| = g(zQ),...,zk = gk{z0),.... The simplest case is that of a fixed point of g, i.e. a point ZQ € C such that g(z0) = ZQ. We say a fixed point is attracting if the modulus of the derivative of g at z0 is less than I, i.e. \g'(z0)\ < 1. This implies there is a neighborhood U of z0 that is contracted into itself under g, i.e., g(U) c U; and so if the orbit of a point eventually enters U, it will asymptotically approach ZQ. A fixed point is repelling if |f'(z 0 )| > 1, so nearby points are
1299 ^-COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
7
pushed away by g. (In the nonhyperbolic case, i.e. when \g'(zo)\ = 1, the behavior of nearby points is not as clear cut.) Now more generally, a point z$ € C is periodic (of period n) if g"(zo) = 20 for some n € Z + . It is attracting (respectively repelling) if in addition, l(£")'(zo)l < I (respectively \(g")'(ZQ)\ > 1), where (gn)'(z0) is the deriva tive of the nth iterate of g at ZQ. By the Chain Rule, these properties remain the same for all points in the orbit of z0: zo. *i = ?(zo). -.., r« = g"(zo). (The corresponding periodic properties of the point at oo are usu ally determined by the properties of 0 after the change of coordinates 2-1/2.)
If ZQ is attracting of period n, the derivative condition implies there is a neighborhood U of ZQ such that g*(U) C U. So orbits of points that eventually enter U under the action of g will asymptotically approach the orbit of ZQ. Such points are said to be in the basin {of attraction) of z0; the basin (of attraction) of g is the union of all such basins. PROPOSITION
3. The basin of attraction of g is an R.E. set over R.
To show this we construct a machine M whose halting set is the basin of g. Since there are only a finite number of attracting periodic points for rational maps (see Blanchard) there is a real polynomial h (of 2 real variables) such that h(z) < 0 if and only if z belongs to a finite union of discs around the attracting periodic points which is contracted into itself by g. Thus, a point is in the basin of g if and only if for some z in its orbit, h(z) < 0. Now let the machine M be described by (Figure 4). Clearly, ClM, the halting set of M, is the basin of attraction of g. Again, it is natural to ask if the basin of attraction of g is decidable, or equivalently, if its complement is R.E. Input ziC
Compute g(z) and replace * by it
h(z)2
*(*) < 0
Output i FIGURE 4
0
1300 S
LENORE BLUM. MIKE SHUB AND STEVE SMALB
Quite generally, the complement of the basin of attraction is the Julia set of g. This is true at least for the set of hyperbolic rational maps. (These maps are open and nonempty in the set of rational maps of each degree, and conjectured to be open and dense. See Sullivan.) The Julia set of g, Jt, is the closure of the set of repelling periodic points of g. In Example 2 above, the Julia set of g is the complement of ft w and, as shown, not decidable. To contrast, it is an easy, but instructive exercise to show that the Julia set of the map g(z) = z2 is the unit circle, and hence decidable over R. In §10 below, we shall investigate systematically conditions under which a Julia set can be R.E. over R. Indeed we show that "most" Julia sets are not R.E., and hence they and their complements are not decidable. EXAMPLE 4. A more elementary example (due to Feng-Gao) of an R.E. set over R which is not decidable is the complement of the Cantor Middle third set in the unit interval. The demonstration is via the following "machine" (and the fact the Cantor Middle third set is an uncountable totally disconnected set) (see Figure 5). EXAMPLE S. Another type of example is the machine that computes the greatest integer in x, [x\, for x > 0 in R. (See Figure 6.) x
Input x€R
,
lA Branch
OSxSl
j S
/
yT
/
—
VSubtract 2 from x
Compute 3x and
j
/
OSxSl
x<0orx>l
| x*-3x j
I
/
1
replace x by 3x
Branch
2s/s3
x*-x - 2 FIGURE 5
\
Kx<2
x
Output x
^_______
1301 N/"-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
Input x?R as the second coordinate of a point in R2 with fim coordinate 0
(0^) 1 ■
Output*, the first coordinate
H* I 1
f—v.
if. m. Branch
\ \
I^)«-(a + U - l 1 / ^-r / \^ ^S
Replace^) by (A + 1.x-1).
FIGURE 6
Note that for x > 0 in R, the "cost" of computing |xj by the above machine is | x + 1J comparisons and [x\ (pairs of) basic arithmetic compu tations. Using binary search these costs could be reduced to 0(log(x + 1)) but essentially no more in our model (See Proposition 3 in §4). It is of interest to note that in models of computation where [xj as well as the basic arithmetic operations can be computed in constant time, seemingly hard problems such as factoring integers {Shamir) and testing satisfiabil ity of prepositional formulas (this follows from Schdnhage) can be solved efficiently, i.e., in polynomial time. EXAMPLE 6. Now let S c Z+, the positive integers. We construct a machine Ms over R that "decides" S. That is, for each input n € Z + , Ms outputs 1 (yes) if n € S and 0 (no) if n g 5. Ms has a built in constant j € R defined by its binary expansion f 1 if n € S , I 0 otherwise.
s = J , J 2 • -S*.... where s„ « {
Ms with its built in constant s, plays a role analogous to an "oracle" for a Turing machine that answers queries "Is n € ST" at a cost of n logn. (Using methods related to those used in Propositions 3 and 4 and the Remark at the end of §4, one can give an order n lower bound on Input n€Z+CR
[* 1 {(xltxt) «- (|2"«|. 212"-^Hl
Compute (ria subroutines)
■ m I Branch] xj - xj as 1 [ 1|
/
N.
xj - x 2 « 1 | 0I FIGURE 7
Output
9
1302 10
LENORE BLUM, MIKE SHUB AND STEVE SMALE
the cost of accessing the nth bit of binary representations of real numbers s e (0,1) by machines over R.) (See Figure 7.) EXAMPLE 7. As a final example we describe the Travelling Salesman Problem (TSP) over an ordered ring R. Here we are given a nonnegative symmetric n x n matrix A over R with entries Au denoting the distance between "cities" i and j . Thus A is a mileage chart for the n cities { 1 , . . . , n}. The "problem" is: Given instance A, find a tour of minimum distance. A tour is a cycle t of {1,...,«} and the distance function to be minimized is
r>(A,t) =
j:iiAm.
The Travelling Salesman "decision problem" over R is: Given (k,A), where k € R* and A is a mileage chart, decide if there is a tour of distance < k. Over Z, this is the famous W-P-complete problem of classical complexity theory. Of interest to us is the TSP over the reals, R. In §4 we show that the "topological complexity" of the TSP over R is at least {n - l)!/2. The topological complexity measures the branching necessary and sufficient to solve all instances of the n-city TSP over R by machines over R. In §5 we develop a notion of NP over an ordered ring R and show that Travelling Salesman decision problem is NP over the reals. 2. Machines over a ring R. Let R be a ring, commutative with unit, and we suppose that R is ordered. The main examples are R = Z, the integers, and R = R, the real numbers. In every case Z is a natural subring of R. Let R" denote the direct sum of R with itself n times. We often want to allow that n = oo, in which case R" is called the countable direct sum of R with itself. A point x = (x\,...,x„,...) in R°° satisfies xk = 0 for k sufficiently large. If n, m are both finite a map / : R" —• Rm is a polynomial map if the coordinate maps / are polynomials in the n variables for / = 1,..., m. If R is a field, then / is rational if the /) are rational functions in the n variables. The degree of / is the maximum of the degree of the
f,
In case n is oo we impose a further condition on / in order that it be called polynomial or rational. This condition is that there is a k such that fi{x) = Xj if / > k and dfj(x)/dXj m 0 if i < k and j > k. This means that at most k variables and coordinates are active in that computation. The least such k will be called the dimension of f. The degree of / is as above. In case R is a field we will write f:R" — Rm,f / may not be defined everywhere.
rational, even though
A machine M over R consists of an input space 7, output space O and state space S, together with a connected directed graph whose nodes labelled 1 A^ are of certain types and with associated functions. We proceedjnore precisely and at first in the finite dimensional case. Here /, O, and S are each R', Rm and R" respectively, with l,m,n < oo.
1303 ^■COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
11
The directed graph of the machine M has 4 types of nodes as follows: (1) Exactly one input node, node 1, characterized as having no incoming edge, and one outgoing edge. Associated to this input node is a linear injective map 1:1 — S (which just takes the input and puts it into the machine), and /?(1) the next node. (2) Output nodes characterized by having no outgoing edges. To each such node, n, is associated a linear map0„: S — O. (3) Computation nodes; each such node has a single out going edge, so that a next node 0(n) is denned. To n is associated a polynomial map g„: S —» 5 (an "endomorphism"). If R is a field then g„ could be taken rational. (4) A branch node n has two outgoing edges, giving us next nodes p~{n) and fl*{n). To n is associated a polynomial h„:S — R with fi~{n) associated to the condition h„(x)<0,p*(n)xoh„{x)>0. If M is a finite dimensional machine over R, we may define the inputoutput map
n
gH
A(x)«0 / \ to
AU) = 0 j
/
using a programming device to write
A(x)*0
/ \
FIGURE 8
A(x) = 0
1304 12
U N O t t BLUM. MIKE SHUB AND STEVE SMALE
Let a machine M over R be given and let y € 7. The computation on y by M goes in a natural way. First I(y) € 3 and we are at node 1 in the computation process. If £(1) » n is a computation node, the computation g„ is performed on the state x « l[y) to produce gm(x) which replaces x. Ifn is a branch node, and h„(x) < 0, the next node is fi~(n). Otherwise the next node is 0+(n). In both cases, the next state is still x. Thus the computation proceeds until an output node n is reached (if ever) and 0„{x) for some x € 3 is computed. In this case the computation is said to halt and produces fu(y) m $•(•*)• Denote QM C 7 as the set of y where the computation halts, ft* is the halting set of M. Thus M defines the input-output map fu '• &M — 3 . Compare the example of §1, which essentially are finite dimensional machines over R. It is important to allow infinite dimensional machines in the construc tion of universal machines and to analyze uniform algorithms (algorithms which solve problems with inputs of arbitrarily large size). We now discuss the modifications needed in the finite dimensional ma chines to make them infinite. In as infinite dimensional machine we take 7 ■ R1, D « Rm now allowing /, m to be < oo, together with a (finite) directed graph. The 4 nodes as before are the same from a graph theo retical view, but the definition of the associated maps must be extended to cover the oo-dimensional case. We have already defined a polynomial map R°° — It 40 . But this will not affect x, for large / in the expression x » (X|,X2,...) € R°°. This consideration forces us to introduce another type of node into the machine which we call *ffih node. A fifth node is de signed to access the x, with large i. For this we take 3 - Z ^ + Z * * / ? 0 0 with a typical element [i,j,x\,xi,...), i,j positive integers. A fifth node (one outgoing edge and unique next node 0(n) as in a computation node) op erates on this element by transforming it to the element i,i,j,X\,xi,...,x, replacing x, in the jib place in R°°, . . . ) € $ . No other changes are made in the computation of a fifth node. If n denotes this node, then that map will be written as g„: ~S -* IS. In the oo-dimensional case the input map / : 7 - » J « Z + + Z + + R°° can conveniently be taken as follows: non-zero-coordinates are given by J(*)2*+i - xk for k < I leaving starting values / » \,j » 1 in the Z* + Z* part of 3. This leaves "working space" in 3. For technical reasons, we also assume coordinate 7(x) 4 denotes the length of x. Here, the length of a nonzero element x € Rf is denned as the largest k such that xk # 0; the length of the 0 vector is 1. The gm: "5 -► "5 of a computation node is required to be of the form
gm(ij,x) m (i'{i,j),f(i,j),x'(ij,x))
with i'(ij) « /+ 1 or 1. (Similarly,
j'(i>j) « ; + 1 or 1) and x'(i,j,x) satisfying the previous defined condition for polynomial (or rational) maps on oo-dimensional spaces. The branch node polynomial h: Z* + Z* + R°° -» R, we suppose satisfies dh/dx, * 0, / > some k. (This condition defines a polynomial junction.) We let kM denote the maximum dimension of the maps and functions associated with the computation and branch nodes of hi. Similarly, du is the maximum degree of these maps.
1305 Nr-COMPLETENESS RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
11
The final modification for the oo-dimensional case is on the output map On: 5 — 5 . Suppose it to be of form On{i,j,x)k *x 2 *_i, k = 1,2 The input-output map 9» for the infinite dimensional machine M is defined just as in the finite dimensional case. For /, m < oo we say that a (partial) map f:/?'-» Rm is computable over R iff there is a machine M over R such that ft* » fi,, the domain of f, and p *(*) * p (x) for all x € £ V In such a case we say M computes p. Machines M and M' will be called equivalent if they compute the same map. A set Y c R" will be called R.E. over R if y « Q* for some machine A/ over A. It will be said to be deciaable if it and its complement are both R.E. over R. It is an easy exercise to show that Z is a decidable subset of R over R for any Archimedean ring R. Sometimes it is convenient to restrict the inputs of a machine from 7 to a subset Y c7. Thus Y would be considered the space of admissible inputs (for some problem, for example). Various notions in our theory may be relativized to such a space of admissible inputs. For example, if Y C 7 let CIM.Y = AM n Y. Then we would say CIM,Y is an R.E. set over R relative to Y. It is also convenient often to view R* as being contained in R' for k < I < oo via the natural injection j : Rk — R' where j{x) « ( x , 0 , . . . ) , / - k zeros after x. Thus, e.g. we may think of ( 1 , 0 , 0 , . . . ) € R°° as the element 1 € Rl C R°°. In the infinite dimensional case, it is often useful to write R°° = R°° + • + R°° (m times) in "S. When we talk about a machine over R in the sequel, this could mean either the finite dimensional or the infinite dimensional case. For x € J we will sometimes use the notation x]*, it ■= 1,2,... to denote the klh coordinate of x in the finite dimensional case, and the k + 2 coordinate of x if 5 is infinite dimensional. In the latter case, x]_i,x]o will denote the first 2 coordinates of x respectively. 3. The comparing fdomorpliiiw aad the register eovatkMS. In this and the following two sections we develop the machinery and concepts needed to prove our Main Theorem on ^/'-completeness in §6. First of all, we define a machine over R to be in normal form if it satisfies: (1) At each branch node, n, hH(x) « x)\, for x € 5. This is a minor condition making things a bit more convenient (2) There is one output node, and hence one output map O: 5 -► 7). (3) There is a given labeling of the nodes {1,2,..., AT} such that (a) 1 labels the input node, (b) N labels the output node. (4) In thefinitedimensional case, / is the natural injection and O is the natural projection. PROPOSITION 1. Given any machine M over R, there is an equivalent one in normal form. PROOF. One achieves property (1) by adding a computation node before and after each branch node using a little straightforward programming
1306 14
LENOKE BLUM. MIKE SHUB AND STEVE SMALE
exercise. To obtain (2), one just collapses ail of the output nodes to a single one. (And in thefinitedimensional case, perhaps expand the state space and number of nodes a bit to obtain Q, « O.) The existence of the labeling (3) is dear. For (4) we add compuution nodes immediately after the input and before the output nodes. For a machine M over R in normal form, there is a natural map H = HM from Tf x "S to itself, called the computing endomorphism of the machine. Here V = {1,..., N) is the set of nodes and J is the state space. We call U x 5 the full state space of M. The computing endomorphism H: TJx'S -»TfxS has the form H{n,x) — (0(n,x(x)),gH(x)) where fi describes the next node and g„ the new state. Here x '■ 5 — R is defined by
f X(x)-{
1 if*],>0, 0 ifx],»0,
[ - 1 i f x j i <0. We define 0:W x {0,±l} — 77as follows, depending on M\ N fi(n) Bin, a) ■ ' P*{n) . 0~(n)
if n « N, if n < N and n is a nonbranching node, if n is branching and a * 0 or 1, if n is branching and a = - 1 .
To define g„ as a function of n and x, let g„(x) « x for n = 1, N or a branch node. At a computation or fifth node we suppose g„(x) is the computation given by that node. Thus the computation of a machine M over R with input y € 7 is rep resented by a sequence zo,Z|,Z2,...,z*,... for z* e ^ x 5 , z0 = (l,I(y)) and H(zk.t) ■ z*, k « 1, In dynamical systems terminology, this se quence of z* « (n*, Jf/t) is the orbit of the computing endomorphism with starting point z<>. Note that if M is infinite dimensional and xk « (i,j,...), then (,y < k. We remark that if gH is a rational map, then f«(jc) may not be defined for some x € 5. Thus // is a "partial" map. However, by starting with Zo » (l./Cv)) we are assured (by our convention on machines) that H can be iterated without concern. A necessary and sufficient condition that a sequence zo, Z\,... be a com putation by the machine M is that (U)
z*-tf(z*_,),
*«1,2,...
and secondarily, z<> ■ («o.*b) has the form (lb)
no*l.
xo*f(y).
for some ^€7.
Let us call equations (la),(lb) the register equations of the machines M. Equation (la) may also be written (la')
4(n*_i,z(x t _,))«»i t
^.,(-Xt-i)«Xt
for*-1,2
1307 A^-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
Input
1i
yil
2 = (n,x) is replaced by H(2)
Output FIGURE 9
In the sequel, symbols such as nk or xk will sometimes represent specific elements (in ~N or $ say), and sometimes variables. The intended usage should be dear from context One can represent the computation process of a machine over R by the following flow chart. (See Figure 9.) (Compare this flow chart with the "Julia set machines" of §§1 and 10.) Thus for given y, one keeps computing H until node N is reached in which case the machine outputs O of that x in 5. For reasons to become apparent in the next sections, we move to putting the register equations of a machine M over R into a more algebraic form. To do this wefirstput the computing endomorphism into a more algebraic form, at the same time extending it to an endomorphism of R'xR'". Here R' is the ring generated by R and Q and, in the infinite dimensional case, R'" is to be interpreted as R1 + R1 + JT» (similarly for R"). LEMMA 1. 0 can be extended to a polynomial map (also denoted) 0: R' x R'-ffofdegreeN+l. PROOF.
For each n € U, let
a.(y)
n
an(y)
1 ify«n, 0 if not
iy-J) (n-j)'
So, for y € 7f,
1308 16
UNORE BLUM, MIKE SHUB AND SIZVESMALE
Let B « {branch nodes of M). Then fi(y,e)-
£
aniy)fi{n) + ( ^ J t i l
+ (a
+ i ) ( i _ a ) ) £ « , ( ) , ) * ♦ (n)
(Note that 0{n,o) is independent of a for all nonbranching nodes n.) It can be easily seen that the degree of 0 is N+1 and that 0 can be written as a sum of 6N monomials in the variables y and a. Now let g(n,x) « g*(x). As above, we can extend the definition of gn(x), or equivalently g(n,x), to all n in it. A formula for this is g(y,x)m 52
Here pi{x), qi(x) are polynomials if n is a rational computation node; otherwise, p{(x) « flj(x), the /th coordinate of f, and qj,(x) s 1. Then the above is modified simply to
Since E»«yv <*»00 * 1 for y € ^, the previous formula is recovered for y € 77, whenever qi(x) * 1, all n € # . It is dear that if M is finite dimensional, then g is a polynomial (or rational) endomorphism of R* x A". This is not necessarily the case if M is infinite dimensional due to the existence offifthnodes. Nevertheless, if V i s infinite dimensional, it is rA* case that for each k € 2*, there is a polynomial (or rational) endomorphism g(k) of K x R"1 of dimension X* that is identical to g on ^ x J (4) . Here Kk « max(J!rM, Jt+2) and ?(«) » {(/,./'...) € "S\i,j < k). Furthermore, these endomorphisms are uniform in k. This will follow from the above together with LEMMA 2. There is a universal map F: Z* x R** — R" such that for each k € Z"\ /"(*) (« F(k, •)) is a polynomial endomorphism ofK* ofdimension k + 2. degree 2k - 1 and F(k) is identical to the fifth node computation on PROOF.
Let a: Z* x 2* x Z+ x R* x R> - tf be defined by
«*.«...)- n ($3) n (fff
1309 ff^COMFlflZNES&RBClfltSIVE FUNCTIONS AND UNIVERSAL MACHINES
where* = {!,...,*}. So, for*€Z + and ij,v,w i(/c,/,;,«,«;)» |
1*
eE,
I if(v,w)>(ij), 0 otherwise.
Note, for k, i,j fixed, a is a polynomial in v, w of degree 2(/c - I) which can be written as the sum of k2 monomials. For each /, j € Z+, let
"Mi
otherwise.
Then the coordinate functions F{k)(x) = F'(k,x)
=
«(*,i.y.xl-i.xloH^t/M, + (I - rf'O'))*7)
£ («J)€*x*
where x1 is the /th coordinate of x, define an appropriate map: F[k) is identical to fifth node computation on J ( t ) since if x « (x_i,xo, XI,JC2,...) and x_i,xo < it, then the only nonzero coefficient a occurs when i = x_ ( and y * xo, and the only nonzero coefficient d' occurs when j + 2 « /. So fjg 2 (x) ■ x,, while for / # > + 2, F/k)(x) « x'. We also see that F(k) is polynomial of dimension it + 2 of degree 2k - 1, and can be written as a sum of it4 monomials (per coordinate function). Now, letting
g{k)(y,x)= £
a
*(y)i*(x) +
'52a*(y)F{k)(x)>
where T * {fifth nodes of M}t we get polynomial (or rational) endomorphisms of IF x R" identical to g on 77 x J(jk), uniformly in it. The dimension of g{k > is Kk + 1 , the degree is max(*, 2Jt - 1 ) + (N - 1 ) and the number of monomials needed to describe each coordinate function of g(k) is it4 plus a constant depending only on M. If there are no fifth nodes, each of these bounds can be replaced by constants depending only on M. 2. There is a universal map 77: Z* x If x R"1 -»If x IV such that for each it € Z + , 77 (t) (» 77(it, •, •)) is a composition ofpolynomials (or rational) maps and the characteristic function x. and 77,*) is identical to the computing endomorphism on 77 x 5 ( t ) . (If n < oo,5 (t) is to be interpreted as ~5.) PROPOSITION
Thus, we can rewrite equations (la') as equations involving composi tions of polynomials over R and the characteristic function x* d»")
c(*(«*-i.;t(**-.))-»*)«0, c*(*<*)(«*-i.x*-i)-**)-0,
1310 IS
LENORE BLUM. MIKE SHUB AND STEVE SMALE
for k » 1,2,..., where c, ck are the (smallest) positive integers sufficient to cancel denominators. (If n < oo, gik) is just g.) In case M has ra tional computation nodes, we replace the second set of equations by the coordinate polynomial equations:
ck\
J2 ^(nt.OpiUjk-iJ + ^ a i . ^ - , ) / ; ^ ^ - , ) \n47S-T
»€?
»€*
J
Note that the coordinate equations of the second set for / > Kk can be written as * { _ , - x( = 0. Now suppose R is a field with the property that any positive element is a square. Thus R could be the real numbers or any real closed field. With R satisfying the property we can eliminate the characteristic function x in the above and thus put the register equations into an even finer algebraic form. To do this we introduce new variables ultU2,...,uk,... fork = 1,2,... and replace (la") by the following polynomial equations over R. (la'")
xk.l)l(xk.l]iu2k_l
+ 1 ) ( * * - , W _ , - 1) = 0
fi{nk.uxk.x)xu\.x)
- nk = 0
£<*)(«*-I»■**-1) - xk « 0
for fr = 1,2, Again, if M has rational computation nodes, the last set of equations can be modified appropriately. PROPOSITION 3. Suppose the field R has the property that positive ele ments are squares of elements in R, and suppose M is a machine over R. Then the sequence (n0,xo),(n\,xi),...(nk,xk) € R x R" is a computation by M if and only if it satisfies (la"') for some « I , K 2 > . . . in R and (lb). If these conditions are satisfied then necessarily (nk, xk) € lv" x 5.
The proof is straightforward. We can characterize computations via polynomial sets of equations and (lb) even more generally. For example, if the field R has the property that every positive element is a bounded sum of squares (eg., by Lagrange, every positive rational number is the sum of 4 squares of rationals) we can replace every occurrence of «^_, above by £ > . i M?*_o7 where b is the given bound and u(*_ | ) ; are new variables for./« \,...,b*ndk = 1,2
1311 A^-COMPLETENESSr*KWRS!VE FUNCTIONS AND UNIVERSAL MACHINES
«
Over the ring of integers, we can replace (a") by
c f xt_,], -J2^-\)J-
' ) ( jc *-'i' + IZ u (*-')j
+ i
) f^ n *-i. jc *-i]i)- n *j =°
«*-l]l I **-|]l ♦ £«?*_,„ + I J U(»t-|.*t-|]l- ^ «?*-!,; I -"* J «*-l)l f**-lll - ^ «(*-!);- I 1 I 0 I 1*-|.*t-|]| +^2u^-t)j
=0
I -"* I =°
and Ck(gik){nk-i,Xk-i)-xk)
where tt(*_i); are new variables for j = 1
= 0
4 and k = 1,2,
4. Time T halting sets, equations, polynomials and computations. Let us consider now halting computations, i.e., those sequences (nk,xk) satisfying (la') or (la"), (lb), and with the added condition nT = N for some T < oo. Such a T will be called a halting time for the y in (lb) and the halting time for M to compute
nT = N,
0(xT) =
(O the output map). Let 1T = {v € 7|7*(v) < 7} be the time 7" halting set of M. Thus, &M = LUr, the union over T 6 Z + . Weremarkthat if y, y' € 7 agree on the first KT = max(Jtw> T + 2) coordinates, and y € 7r, then y' € 7r and 9uiy'),
1312 20
IENORE BLUM. MIKE SUUB AND STEVE SMALE
are not finite, nor are / and O polynomial. However, for each T, we can easily modify the time T halting equations to get afinitepolynomial system (modulo x) by eliminating allj coordinate equations from the second set in (la") and (lb'), for j>KT* m*x(kM,T + 2). We then have [y,w) € GM,T if and only if there is a a € {R x RKT)T*1 such that (y,o,w) satisfy the "modified" time T halting equations plus
[xrhj-x Wj
'\yj
for2j
Note, the only nonfiniteness comes in the last equations, and then only in a nonessential way. Over the reals, the time T halting equations are equivalent to a polyno mial system (just replace (la') by (la'") for k » l,...,T). By modifying this system as above we get a finite polynomial system. We can now use the following proposition to convert this system to a single polynomial equation of degree less than or equal to 4 which we call the time T halting polynomial equation for M. 1. Over the real numbers. (a) any system of polynomial equations is equivalent to a quadratic sys tem. (b) Any quadratic system is equivalent to a single equation of degree less than or equal to 4. PROPOSITION
Here equivalence means if one system has a solution then so has the other. For the proof of (a), consider the polynomial equation £ 0 OaXa = 0, a « (ai,...,a„), a, nonnegative integers. Let ta = x° be new variables. One has an equivalent system for the /„'» of the type t^p * tatp together with £ « tota ■ 0, a ranging over some set J. For the proof of (b), one just takes the sum of the squares. THEOREM. Let R » R and let M be a given machine over R. For each T € Z \ there is a polynomial fT: R*r x R * - . l of degree < 4 such that y € IT if and only if there is a z € R* such that fr(y\, • • •. y*r., 2) = 0.
Furthermore, s (and the number of monomials needed to express fr) is bounded by a polynomial in T, depending only on M. (If the input space is finite dimensional, we replace KT by /, the dimension of the input space.) PROOF. The existence of a single polynomial follows from the above dis cussion and Proposition 1. {2 encompasses all the variables in the "mod ified" time T halting equations, except in y, plus the variables needed to eliminate x as well as those needed to convert the system into a quadratic one.) To get the polynomial bound on the number of variables (and mono mials) we must count the number of variables and monomials appearing in the modified time T halting equations, the number of new variables (and equations) needed to get a quadratic system (and then the number of monomials needed to describe the resulting quartic equation). These
1313 *f-COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
21
counts are mainly affected by the number of monomials needed to de scribe the "fifth node" polynomials fik) for k = 1,...,7\ One thus uses the analyses given in §3 to get a requisite polynomial bound. Observe that the construction of the polynomial fj above is uniform in both T and M. We can restate the Theorem to reflect this, and in a somewhat different way. THEOREM'. There is a machine over R which on input (A/, T) (where M is a machine over R and T € Z+) outputs in time polynomial in lM (the "length ofM") and T, a polynomial fuj- R*r x R^"*-7"' — R of degree < 4 satisfying the following diagram
V = f^]T(0)
c »*• x Rpo«y(/«.n £ i R
7Jtf.7- = « 2 - , »,(F) c
R^^R*'
Here 7MJ is the time T halting set of M and nu iti are the projections of R*r x RP<»»y<'*-7"> and R°° onto R*T respectively. (If the input space is finite dimensional, replace KT by the dimension of 1M and ignore «2.) To make sense of Theorem', we have to define the "length of A/" and specify how machines M and polynomials / over R are to be represented in R°°. This will be done more precisely in the next sections. Polynomials will be specified by their "powerfree" representations (see §5) and machines by their programs (see §8). The "length of A/" will then be the length (as defined in §5) of the representation in R°° of the program for M. We now develop the relationship between the time T halting sets IT and the time T halting computations Tj. Here IV « TM.T is the subset of (R x 5 ) r + l consisting of ((no,JCo) ("r.-Xr)) satisfying, for some y € 7, the time T halting equations. (Note we automatically get that y € 7r and
0(XT) « *niy).) There are natural maps a, a' defined as follows:
7 7 .ir 7 -^7 r . o(y) - ((\J{y)),H(lJ(y)), H2(l,I(y)),...,HT(l,I(y))) a'((no,Xo),...,(r*T,XT))-y
and
where no * 1 and xo «I(y). Then a and a' are inverses to each other and provide a set-theoretical isomorphism between IT and IV. Moreover a' is continuous and even is the restriction of a linear map. But a in general is not continuous, since H is not (recall H involves a characteristic function). Now consider the map y t : Tj -» RT*X which is the restriction of the projection (R x 5)7"*1 - RT+1. In fact y.flV) C # r * ' C R™. Define 7: 7r -» If * as the composition yt ■ a. One can interpret y(y) as a halting path or as the computation path of y of length T in the directed graph of the machine M. So y(y) is the sequence of nodes traversed in
1314 23
UNORE BLUM. JUKE SHU1 AND STEVE SMALE
the computation fuiy). If y € Tf * , let Vr be the subset of IT such that y{y) s y. Then it is easy to see that a restricted to V, is a continuous map from Vy into TT. PROPOSITION 2. (a) Vrisa semialgebraic subset oflT, (basic, if the maps at computation nodes are polynomial), and fu restricted toVrisa rational map. (b) IT * Uy€Vr*' Vy ** semialgebraic. The V7 s are disjoint. (c) ft*/ * Ur>o^7" " fl countable union of semialgebraic sets. PROOF. Only (a) needs proof. By following the path y and noting the branches taken, one sees that Vr is defined by inequalities of the type
^(-ft1(ft,(/M)))].<0
and «t.(-fc J (fc I (/(y))))]. < 0
where the £*s are polynomial (or rational). If the g"s are rational, each inequality is replaced by a disjunction of polynomial inequalities by using the correspondence: p/q < 0 iff (p < 0 and q > 0) or (j> > 0 and q < 0). Similarly, by composing the computations along the path y, one can see that fu restricted to Vr is of the form OgJm(- • • gj2(gj,(l{y)))). If M is finite dimensional we are done. If M is infinite dimensional, we can, as in the modified T halting equations, use the bound KT on the number of "active" coordinates and variables in a computation of length T to get semialgebraic descriptions of V7, and polynomial (or rational) descriptions of fu restricted to Vr. Note that M with inputs restricted to IT is essentially a finite dimen sional machine. Moreover, this machine is equivalent to one without loops, a tree as in Smale. We remark that (b) gives an algebraic description of 7 r via an exponen tial (in 7") number of semialgebraic formulas. Compare this with the time T halting polynomial description of TT given by the previous theorems for the case A « R. As discussed in § 1, it follows from Proposition 2 that Julia sets in general are undecidable. There are other immediate applications. For example, we now easily get a lower bound on the halting time for computing the "greatest integer in." (See Example 5 in §1.) PROPOSITION 3. TM(X)> log(x) for
Suppose a machine M computes \x\ forxeR*. an unbounded set ofx.
Then
PROOF. Fix such a machine M. For each L € Z* let 71 be the maximum of the halting times Tu(x) for x < L, x € R+. (We can suppose TL < oo.) So for or < L, x € R* we have x £ 7TL, which is the union of at most 2TL sets VT, y of length TL. On the other hand, for each nonnegative integer /, there is an open set U of points in R such that for all x € U, ?H(X) * I. If / < L, there must be such an open set in Vr for some y of length TL. But fU restricted to Vj is polynomial (or rational). So, p M restricted to V7 must be identically equal to /. Thus, there must be at least L such sets V7. So 2 r i > L, Le., TL > togL.
1315 ^•COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
«
In a similar fashion, we get lower bounds for the Travelling Salesman Problem over R. (See Example 7 in §1.) Suppose M solves the TSP over R, i.e., for each mileage chart A over R, fu{A) is a tour of minimum distance. (Of course, we are supposing 7 = 9 = R°° and some specified representation of mileage charts and tours in R°°. For example, an n x n matrix A could be represented in R°° as n followed by a listing of the rows of A. A tour t of n cities could be represented as an ordering (/i,...,( n ) of the integers l , . . . , n where i\ = 1. Thus, t{ij) = jy+i for j « 1 n - 1 and t(i„) = 1 . ) For n € Z* let Tu(n) = max Tu{A), where A ranges over all n-city mileage charts over R. The lopological complexity of M, as a function of n, is the number of halting paths of length T»(n). PROPOSITION 4. Suppose M solves the TSP over R. Then the topological complexity of hi is at least (n - l)!/2. So
T W O * 10,11^2!. The proof is similar to above, noting that for each tour t, there is an open set of inputs U such that for all A € U, fu{A) = t or 7 (where 7 is the reverse of t). So, there must be such an open set in V7 for some y of length Tu(n). So, since fu restricted to V7 is polynomial (or rational), we have fu restricted to Vr identically t or t. The result now follows since there are (n - 1 )!/2 (unoriented) tours over n cities. The topological complexity of the TSP measures the branching neces sary and sufficient to solve all n-city instances. (It could be thought of as the "obstruction" to obtaining a rational formula solution.) This notion can be studied more generally. For example, a slightly more subtle argu ment shows the topological complexity of the Knapsack Problem over R is 2". (The Knapsack Problem is: Given {x\,...,x„) € R". Find a subset 5 C { 1 , . . . , ri) such that Y,,es x< " J • if o n e exits.) For lower bounds on the topological complexity of solving 1 variable polynomial equations, see Smale. REMARK. In the above two examples the machines compute functions that take values in a discrete set. We used open set arguments to get estimates on the number of components Vy of IT, thus getting our lower bounds. If these open set arguments were not applicable, we could still use estimates on the number of connected componenu of IT (see e.g., Milnor and Thorn) to get lower bounds. 5. Complexity theory of machines over R. A background reference for this section and the next is Garey-Johnson, which gives an account of the ideas of ^^-completeness of Cook and Karp. Here we are considering a version of this theory, more algebraic and especially over the real numbers. There are some differences in substance. Toward denning the property, "a machine M over a ring R is in class P" (polynomial time), we introduce the notion of size. The length of a nonzero element x € R" (n < oo) is defined as the largest it such that xk ft 0, where x = (x 1 ,x 2 ,...,jc t ,0,0,...). The length of 0 is 1. The size
1316 34
- M N O U NJUftMAS SHUB AND-STEVE SMALE
of x is to be the length of x plus the height of x where the height is yet to be defined. The height of x is the maximum of the height (x,) over all i, and it remains to say what the height is for an element of R. We only do this here for the cases R » Z the integers, Q the rational numbers and R the real numbers. If/?« Z, x € /?, then Aeigfa x «log(|x| + 1). If R « Q, x * 0/9 eR,p,q relatively prime integers, then height x is the maximum of height p, height q. Finally if R ■ R and x € R, then /w£/u x » 1. Note the difference in the case of x € Q C R depending on which field is considered. For some other rings R, the notion of height (e.g. Mazur) from algebraic numberfieldsis suggestive. The height over Z of an integer is essentially the number of bits. The height over R reflects the fact that as in scientific computation, the cost of multiplication is independent of the magnitude of a real number. PROPOSITION 1. Let M be a machine over R and for y el. let 7V(y) be the halting time as in the previous section. Then aiy) < kT^iy) where a(y) is the number of arithmetic operations and uses of 5th nodes used to compute ft/iy)- The constant k depends only on the machine M.
The proof is straightforward from the fact that M has only a finite number of nodes. Thus, there is a bound on the degree and the number of active variables and coordinates over all the computations at nodes. We now define the standard cost function Cyiy) of machine M on input y to be the product of the time, TM(y), and the maximum height, hM{y), occurring in the computation of fu(y). That is, hhtiy) m max height(xjt) where Hk(l,I(y)) » (nk,xk) some nk € V. where x£ is identical to xk on itsfirstku coordinates, 0 elsewhere. Note that for R - R, CM(y) ■ TM(y). For R * Z, CM(y) is polynomial^ related to the "bit complexity" of computing fuiy)Now say that a machine Ai over R is in class P {polynomial time) (or that the computable function fu is in class P over R) if there are constants c and q € Z* such that CM(y) < c(sizeO'))* aD>€7. Here and in what follows, often sense is made only in case height has been defined. This includes R - Z, R at least For R ■ Z, class P over Z is identical to the classical polynomial time class with respect to the bit complexity measure of cost Suppose Y C 7 is a space of admissible inputs. We say that the pair (M, Y) is in class P if C„(y) <
1317 Nr•COMPLETENESSJtECUJtSIVE FUNCTIONS AND UNIVERSAL MACHINES
23
An algorithm (or sometimes machine) which solves a decision problem (Y, Yy,,) over R is a pair (M, Y), where M is a machine with space of admissible inputs Y such that fu(y) « 1 (yes) or 0 (no) for all y e Y, and fu(y) * 1 if and only if y € Yy,,. Then we say that the decision problem (Y, Yy,,) is in class P if there is an algorithm in class P which solves it. PROBLEM 5.1. Would the class of decision problems in P over Z increase if the cost function were changed to Ty(y)1 (Certainly, the class of poly nomial time computable functions over Z would increase, eg., consider the function f(x) « |x||jr|.) We will say that a decision problem (Y,Yy„) over R is in class NP (nondeterministic polynomial time) if there are constants c and q € Z + and a machine M over R with 7* « 7 x 7\ where 7 « 7* « /?", n < oo and the space of admissible inputs for M is y x T, such that (a) the values of
(b) f * 0 \ / ) - 1 only if y € Yy,, and, (c) for each y € X m , there is a / € 7* such that fvCy,/) = 1 and C v O \ / ) < c(si2eCF))«. The / € t are the guesses as in the standard /^/'-completeness theory (see Garey-Johnson). It is easy to see that without loss of generality we can assume, in (c) that the size ( / ) < d size {y)9, for some fixed d € Z + . Again, if /? « Z, our class NP coincides with the classical one. PROPOSITION
2. NP D P over R.
The proof goes by using the machine for P and ignoring the guess. The following two examples illustrate our notion of NP over the real numbers R. PROPOSITION
3. The Traveling Salesman decision problem over R is in
NP. (See Example 7 in §1.) PROOF. We write a flow chart for the NP machine as follows (Figure 10): (U.A)./)
Htrc: kiR,A is ■ "miltaft chart", and y'iT, a guess
ZHAy)S*? Hert:DCAj')dtnotei the distance of tour y'
FIGURE 10
1318 26
LENORE BLUM. MIKE SHU* AND STEVUMALE
The proof now focuses on the subroutine, "Is / a tour?" Suppose mileage charts and tours are represented as in §4. Let n be the number of cities. To check i f / represents a tour of n cities just check if each integer ! , . . . , « appears once and only once in (y\,...,/m) and / , « 1. Now it can be seen that the subroutine can be implemented by a polynomial time algorithm. The second example is the 4-Feasibility problem, a certain feasibility problem for real algebraic varieties, which we write (F,Fyes). Here F consists of polynomials / : R" -» R of degree < 4, and / € Fyrs if there is some x € R" such that f{x) « 0. It remains to be specified how F c R°°. To do this we describe the powerfree representation: The polynomial / : R" — R of degree < 4 is powerfreely represented in R00 as (4,n) followed by a sequence of (Q,OO) where a « (ori, 02,03,04), a, € [0,...,n], a, < Q,+I and Oa € R. The pair (a,Oa) stands for the monomial a<,xaixa:xaixa4, with XQ =» 1 to allow for terms of degree less than 4. These (0,4,) are supposed ordered by the lexicographic order on the a. Thus f(x) = Z)o<*o-xoi-!CajJC«»a:iu- Note / can be considered as a polynomial on R°° which does not depend on x, for i > n. (For each degree d € Z + , it is clear how to generalize this description to get the powerfree representation in R°° of polynomials f: R" -»R of degree < d.) PROPOSITION
4. (F, Fy,,) is in NP over R.
The NP machine takes guesses for / : R" — R, points / in R°° and tests if f(y') m 0. Since degree / < 4, this evaluation is given by a polynomial time machine. PROBLEM 5.2. Does P » NP over R? 6. NP-compkttneas and the analogic to Cook's theorem wtr R. Inspired by the theory of NP completeness in computer science (see Garey-Johnson) we say that a decision problem (Y, Yy,,) is NP complete if it is in NP and: Given any decision problem (Y, Yy,,) in NP, there is a map y.Y — t with these properties: (a) w(y) € Yyn if and only if y € Yy,,. (b) iff » f M\Y for some machine M, and this machine is in class P. In other words w is polynomial time computable. This definition is meant over a ring R and a definition of height over R is needed. In particular, this is satisfied if R * Z in which case our definition checks with the classical definition. Also the case of R = R is included, in which case, we have a new definition of NP complete. This is the case of primary interest in what follows; but the case of real closed fields should be the same. MAIN THEOREM (ANALOGUE OF COOK'S THEOREM FOR R).
The
4-
Feasibility problem (F,Fy„) of the previous section is NP complete over R
1319 ^/■COMPLETENESS, .RECURSIVE FUMCTIOMS AMD UHWHtSAL MACHINES
37
The proof uses the machinery and results of the last three sections. Notefirstthat (F, Fyes) has already been proved to lie in NP, Proposition 4 of the previous section. Let M be the "nondeterministic machine" in the definition of NP for the problem (Y, Yy,,). In this definition there is the polynomial bound c(size(y))9 * T0(y). We must describe the map V- Y -» F with the requisite properties. Thus let y € Y and consider in the space (R x 5)7"*1, T = T0(y), the equations (a'") for k m 1 T. We adjoin to this set, the following (2) n0 - 1, /(>»,/) = xo. (Here / is a free variable of length KT-) (3) nT - N, and 0(xT) - 1 (yes). LEMMA. For y eY, this system of equations (la'"), for k - 1 and (3) has a solution if and only ify € Yy„.
T(2)
PROOF. One just has to trace through all the definitions. This system is equivalent to the time T halting equations of §4 (now with input space 7x7*) plus the equation 1 - X7-]]. Thus, it is easily modified to an equivalent finite polynomial system (we have already eliminated *). As in the Theorem in §4, we can convert this system to a single polyno mial equation / : R" — R of degree less than or equal to 4, to obtain using the powerfree representation, y: Y — F. By the previous lemma, yiy) is in Fyti if and only if y € Yy,,. It remains to see that y is a polynomial time computable map. For this one notes that the construction of w makes the computability clear. Moreover, since T « To(y) is a polynomial in the size of y, it is only left to see that the length of (the powerfree representation of) iy(y) is a polynomial in T. Here one uses the analyses given in §§3 and 4. COROLLARY. Any algorithm for the feasibility problem (e.g. from TarskiSeidenberg, see van den Dries, or the faster algorithms of Canny andRenegar) can be used to solve any problem in NP over R. If that algorithm is a poly nomial time algorithm, then any problem in NP over R can be solved in polynomial time.
This is a usual motivation for studying Af/'-completeness (Cook-Karp) see Garey-Johnson. Our theorem above raises many questions. Among them is: what are other ATP-complete problems over R? We only have very preliminary re sults in this direction as PROPOSITION 2. Fixing a degree d, let FL be the space ofall semialgebraic sets defined by polynomial constraints of degree d, powerfreely represented Let F'4ytt be the nonempty ones. Then (F'4, F^,) is NP-complete. PROOF.
The reduction is given by the inclusion map F — F'd.
3. Let F" be the set of all representations (powerfree) of polynomial systems of equations of the type utj - tk, and one equation E, € y '•■ -c. Let F'y\z be the feasible ones. Then (F", F£,) is NP-complete. PROPOSITION
PROOF.
This follows from the proof of Proposition 1.
1320 28
l£NORE BfcUM, 4UK£ SHUB AND STEVE SMALE
REMARK.. Finally note that the map / : R" -» R, / € F, gives a reduction to a one dimensional problem. Does 0 e Image ft (The image of f is an interval. Of course the description of the interval is not the standard one.) This presumably can be denned as an jVP-complete problem.
7. Computable functions, normal forms and partial recursive functions over R. We deal mainly with thefinitedimensional case. Suppose /, m < oo and let / : R1 -* Rm be a partial function (map) computable over R. We denote this by writing / € C^°°. Suppose M computes / . Added in proof. Without loss of generality we can assume M is finite dimensional.4 We may also assume M is in normal form. Using the (partial) computing endomorphism H:~N x S —» N x 5 for M we get a normal form description for / (and
1321 NF-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
19
(ii) the characteristic function x- Q* = R — {0,±1} where - l ifjr<0, 0 ifjr = 0, 1 ifjc>0, !
and closed under the operations of (a) composition, (b) juxtaposition, (c) primitiverecursionand (d) minimalization. For (a) we suppose / : R' -» Rm and g: Rm — R" are given partial functions. Then the composition g ■ f:R' — R" is defined by g f(x) = g{f(x)) where atf = / " ' « , . Next for (b), suppose / : R1 -» Rm< (/ = 1,.... A:) are given and that y is the natural isomorphism mapping Rm> x • • • x /?"»* onto Rm'+" *mi. (We often identify Rm> x x /?"" with JR"»i+-+»«» without writing y/-) Then the juxtaposition F = {fu...,fk): R' — /?""+-+->■ (of the /•) is given by F{x) = r(/i(Jf) /*(*)) for JC € Q f = nf.i « / • Note that juxtaposi tion of polynomial functions gives us the polynomial maps f:R' — Rm, (m * mi + • • • + m*) and thus we also get the general projections. For (c), we suppose we are given a partial endomorphism g: Rf -» /?'. Primitive recursion then defines a partial map G: Z-° x /?' -» /?' by G((U) »= x and <7(r + 1, JC) = *(G(f,x)) and Qc - {(n,x) € Z-° x /J'|n * 0 or n = / + l,(/,ar) € flo and G(/,x) € fl,}. So, for / € Z*°, (/(*,*) = g'(x) * g g(x) (g composed with itself t times applied to x). Finally, given F: Qf 2 Z-° *R' — R, the operation of minimalization (d) defines a partial function L: Rf — Z 2 0 where QL « {x € rt'|F(/, x) » 0 for some t € Z* 0 } and L(x) « min,(f(r,x) - 0) = min{/ € Z*°|F(f,x) » 0}. Now suppose M is finite dimensional. Then the computing endomor phism for A/ is given by H(n,x) « {fi(n,x{x)i)),gn{x)) where fi and g(n,x) ■ £„(x) are polynomial (rational, if J? is a field). (See §3.) Recall that the coefficients of ft and £ are in the ring generated by R and Q. Since R need not contain Q, some modifications are necessary. We first define a partial recursive function over R which for x,y € Z-°, y ? 0 gives the "greatest integer in" x/y: [x/y\ » min,(F(f,x,y) * 0) where F(t,x,y)mX{y(t+l)-x)-l. Now, for « € # * { ! JV} let
J€*
1322 30
LENORE BLUM. MMS£ SHUB AND STEVE SMALE
Let
?(y,o)= Y,
3
-0'^('')+([£^2i:^J+(ff+1)(1-ff))lI5-^^+C) »6B
(As before 2? = {branch nodes ofA/}.) Now let ~g{y, x) = ^2n€.-ya„(y)g„(x). All the above functions are partial recursive over R. Thus (using the basic functions, composition, juxtaposition and the above), H:RxS — RxS defined by 7l(y,x) = (P{y,x(x]i))>l(y>x) is partial recursive over R. And Thus H is a partial recursive endomorphism over R. So (using prim itive recursion and appropriate projections in addition to juxtaposition, composition and noting that / and O are basic), we see that h\ and h2 are partial recursive over R, each with domain Z-° x 7. It follows (using minimalization and the normal form description) that the input-output map
Construct M = A/(/l fk): Let 7 = R", O = v(Ofl x • ■ x Oft) and S = Rkm where m is the maximum dimension of all spaces occurring in each Mf. Then, for x e £2(/, fk) (see Figure 12). (c) PRIMITIVE RECURSION (BY LOOPING WITH COUNTDOWN). Suppose g: R1 — R' is computable via Mg. Let 7,, ~Og, "Sg be the corresponding input, outpuL and state spaces. We cojastruct M = MQ as follows: Let 7 = RxJt, 0 = Og and S = Rx7gxSgx Og.For z € S, let z = (zu z 2 ,2j, z4) where z, eR, z2€ 7g, z3 € 5 , and z4 € O,. Then for (t,x) eRx Rl, M looks like (Figure 13). (d) MINIMALIZATION (EVALUATE AND COUNT-UP). Finally, suppose F: ftF = Z 2 0 x R' — R is computable via_A/f and let_L(;c) = min,(.F(/,x) = 0). To construct M = Aft, let 7 = R',0 = R and 5 = 7 f x S f x 0 f . For z € S let z = (z_j_,z2,z3,z4) where zx € /?z2 € /?' (so (zi,z 2 ) e 7/-), z 3 € SF and z 4 € O. Schematically, M looks like (Figure 14).
1323 Af-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
Inputs as Mf
U(x) = {l4x),0.^.0)
r
Like node 1 of M(
-. !
Like nodes 2,...Jff-l on the first n( = dimSf) coordinates of Mf. Identity on remaining coordinates (f{x),0,...JO)
Like node N{ of U(
.r
,
qg(/(x)).O.....P) as if,
r
Like node 1 of if,
* —i Like nodes 2 - . J V f - l of Mf -x 1 igf(x))
Output*
Like node Ng of if,
FIGURE 11
Input*
/
I{x)-({IrSx),Q
0)
—5
I
Hi
(7r,(x),0
0))
m
*
((/,l(x),o,....o),a/,(x),o o),....c/A(x),o o))
^
((^(x),0
j
Q II
0),(^x),0.-,0)^_tfA(x),0
0))
-j
3
( as Af,, Output^
/
{((^(x).0
.
0),(/-2(x),0,...,0)^.,(^(x),0 * ^tbl^W^d))
FIGURE 12
1 0))
M
1324
.
32
LENORE BLUM, MIKE SHUB A N D STEVE SMALE
InputM
z «-/(*,x) = U,x,(0),(0)) (Branch)
OutputM
z2
^ ^ ^
»«-(ti,^Vti),(0))
Inputg
\
if
\
Like nodes 2... JV X -1 of Mg on z3. Identity on*!, *i, r 4
Like Mg I /
sir
/
s—Ui.ziAO), gUt)) 1 Output / * / z—izj.-1, z«,(0),(Q)) [ y
FIGURE 13
Inputs
a«-/tx)«(0,z.(0),0) I
Input* N.
I
\
Like nodes 2,...,Nr-I o{MF. (Identity on the first / +1 and the last coordinates.) * I Z*-Ui,22,(0),/'U 1 ,2 2 ))
\ \ \ Outputj.
\ \
Branch
Output*
I *7|
|z^(z 1 + l,z2,(0),))|
FIGURE 14
/
1325 ^-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
S3
Thus we have P^"° = CjJ00 i.e. thefinitedimensional partial recursiveJunc tions over R are the finite dimensional computable functions over R. THEOREM.
REMARK. In classical recursive function theory, the partial recursive functions are defined over the natural numbers Z*°, as the smallest class of partial functions / : (Z*0)' -» (Z*°)m (/, m < oo) containing the suc cessor, zero and projection functions, and closed under composition, jux taposition, "classical" primitive recursion (which we define below) and minimalization. It is clear, but hardly ever stated, that this definition ex tends naturally to the integers to define a class P which we shall call the "classical" partial recursive functions over Z. PROPOSITION.
P * P£°°.
To prove this it is sufficient to define the operation of "classical" prim itive recursion (cpr) and show it is equivalent to ours. DEFINITION. The operation of classical primitive recursion (cpr) asso ciates to given (partial) functions / : R1 — Rm and k: R x R' x Rm — * m , a (partial) function K:2£°xR< ~ Rm where K{0,x) « f(x) and K(t + \,x) = k{t,x,K(t,x)) with Q* »= {{n,x) € Z*° x R'\n - 0 and x € Q/ or n = / + 1 and (/,x) € Q* and (t,x,K(t,x)) € Q*}. To get the classical definition one replaces R by Z* (see e.g. Manin). Although cpr is seemingly more general than our definition of primitive recursion, we have LEMMA. Both definitions of primitive recursion are equivalent (in the presence of{\) and (a) and (b) in our definition of partial recursive function).
First assume cpr and suppose g: R1 — R' is a (partial) endomorphism. Let f:R'->R' be the identity, and k: R x R1 x R' -» R' be given by k(t,x,y) = g(y). Then (by induction on t) the function K given by cpr is the G stipulated by our definition (c). Conversely, suppose / and k are given as in the hypothesis of cpr and define (partial) g: R x R1 x Rm — R x R1 x Rm by g{t,x,y) « (f + l,x,k(t,x,y)). Let G: Z*° x R x R' x Rm - R x R' x Rm be the (partial) function prescribed by our definition of primitive recursion (c). Then letting K(t,x) - G{t,Q,x,f{x))\*. we get the function stipulated by cpr. To see this use induction on t. Hence, K{t,x) is gotten by the following composition: PROOF.
(t,x) -
{t,0,x,f{x))rG(t,0,x,f{x))
.-» G(«. 0,XJ(X))\K. prpj. oaif*
=
K(UX).
Thus, we have the above Proposition and the following COROLLARY. P - Z^°°, i.e. thefinitedimensional computable functions over Z are the "classical" partial recursive functions over Z. One can imagine extending the notion of partial recursive functions to the oo dimensional case.
1326 J4
IPIO«F
aura, MIKE aiua AND ITEVKJMALE
8. Existence of a universal machine over a ring. The first task is to de scribe a machine M over R, as an element x{M) of R°°. Thus "coding of Mm is actually a "program" for M, with M assumed in normal form. More precisely, x{M) is the sequence of labeled instructions, n « 1,2,3,..., N, \n,tn,fin,bn,gH), entries having the following meaning. The n at the be ginning refers to the node, and /tt » the type of the node. The symbol fiH stands for the next node, and bH is the lengthof the description of the computation g„ at node n. More specifics and constraints are: f„ « 1,2,3,4 or 5 corresponding to (1) input, (2) computation, (3) branch, (4) output and (5) fifth nodes respectively. Thus tn ■ 1 if and only if n = 1. Also, /„ = 4, if and only if n - N. If /„ = 3, then fiH is the pair {fi~{n), fi*{n)) and if t„ = 4, fit is omitted. If /„ # 2, then the pair [bH, g„) is omitted. If tn = 2 and
is of dimension k, we suppose gH is represented by k followed by the sequence of pairs of polynomials {pi,,qi), I * l,...,k where p'n,q'„, I ~ 1,.... k are given by some standard representation (see below for the one we are using here). At the beginning of x(M) we might wish to insert a 0 or 1 to indicate how the various instructions are to be interpreted (0 for finite dimensionally, 1 for infinite dimensionally). However, for simplicity we will assume all machines here are infinite dimensional This is not a serious restriction since, as can be easily seen, there is a uniform procedure to convert any finite dimensional machine to an equivalent infinite dimensional one. The following is the flow chart of the Universal Machine U. (See Figure IS.) If it takes as input the pair (*{M), u), u € QM, it will yield as output fw("). Note that enough information is given in x(M), so that given a node n, one knows where (n,/„,...) is in the sequence, and an appropriate subroutine using the fifth node will be able to access that instruction. The state space 3?r/ of V is presumed decomposed as (Z+ + Z + ) + (*°°) + (R + R + rt°°) + (rt°°) + (*°°). The x(M), supposed entered into the fourth of these five components, re mains fixed throughout the computation of the universal machine. The fifth component is to provide "work space" needed in some of the sub routines. The second component will hold the "current instruction of M;n the third will be used to simulate the "current state of M." In the flow chart we are aisnming various programming devices in each box that check the well formedness of the current state of U with respect to the operations to be performed. The universal polynomial evaluator is a subroutine which has input a polynomial (or rational) map g in some standard representation and a state z € J. The output is g(z) in 7.
1327 N f COMPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
Impute
/„(»(*),■) Sad component 3rd component 4th component
U.I.I,)
Id.i.*)-;^.
3}
Initialat Pint instruction of *(tf) Initial Mto of If (with input u) Program br U
Currtat labeled instruction of ti "Current state of it" mt »lM)
lij.z)
i , l O
. F » 5th node computation
Universal polynomial evaluation
FIGURE IS
The universal polynomial evaluator is described briefly as follows. First we write a routine which sends y m (1, 1 , H , J C I , . . . , * „ . . . ) into / =
Let and
( 1 , \,Xn,X\,...
,XH,...
).
f i ( i , l , * , . . . ) ■ ( i + 1.1.* - !)••••) ft(/,l,r,...)»
(1.1.*,...)• (See Figure 16.)
1328 36
UNORE BLUM. MIKE SHUB AND STEVE SMALE
FIGURE 16
A second routine, easily written sends (.... a, (,...) into (...,r°,...) for a € Z-°. More generally, writing a » (ai,...,ak), a, € Z-°, / = (/ tk) and f» « C " C ' * third subroutine takes ( . . . , a , / , . . . ) to (...,/",...). Let a polynomial / : / ? * - » / ? of degree be standardly rep resented as (&,, (a, o„)) lexicographically in a with £ a / < and with 4> € A. Then using the above, one easily constructs
(/,/)-.£«./•. a
Polynomial maps and rational maps are constructed similarly. A final remark on the use of the Sth node should be made. In the universal machine, the Sth node is used in (a) Universal polynomial evaluation, (b) To access a labeled instruction, (c) The Sth node computation. For (a) and (b) we may start (i,j) in Z* + Z + with (1,1). Therefore when the Sth node is used in (a) or (b), the first step is to start with a routine which stores the current (/, j), replace them by (1,1). Then after (a) or (b) is done, (1,1) is replaced by the original (/,;'). These routines are easily done. The existence of the universal machine may be used to construct R.E. subsets of R00 which are not decidabie by the usual Cantor style diagonal arguments—see e.g. Rogers. In particular Civ is such an R.E. undecidable set 9. Characteriziag RJL sets as oatyat sets aad pseodo-diophantiiie sets. The R.E. or halting sets over a ring /? are the domains of input-output functions of machines over the ring. The output sets are the image sets. For the ring of integers Z, the class of subsets of Z which are R.E. is the same as the class of output sets.
1330 LENORE BLUM. MIKE SHUB AND STEVE SMALE
38
a polynomial P € R[x\,...,xk] such that (xt,...,Xi) e fl iff there exist */,.,,...,x* in R such that P(xi,...,xhXt+l,...,xic) = 0. Over R, diophantine sets of reals are R.E., even decidable (by Tarski), but not conversely. As we have already seen, there exist R.E. sets of re als that are not decidable over R. But even more, it is rather simple to construct a decidable set over R with a countable number of connected components, for example {x € R|3/t € Z and \n - x\ < \}. Such sets can not be diophantine over R, for diophantine sets over R are semialgebraic and thus can have only finitely many components. Nevertheless, motivated by Davis-Putnam and Denef \, 2, for simple machines we may try to put a diophantine-like structure on the equations which describe the halting sets of machines. Since polynomial rings over any of our rings have many properties similar to the integers, we can hope for some measure of success. For a ring R and a finite (or countable) set of variables T, subsets S|,...,St c T and positive integers/|,...,/*, we say the set flc R[S\]'< x x /?[S*]/4 is pseudo-diophantine over R[T] iff there exist subsets SJUI , . . . , S,„ c T, positive integers 4+i,...,/ m , and a polynomial P over R[T], P: R[St]'' x • • • x R[Sk]'* x *[S k+ ,]''-' x
• x R[Sm]'" - R[T]
such that Q is the image of the projection on R[S\]'< x • •• x /?[5t]'' of P~]{0). We say a relation is pseudo-diophantine if its graph is. First we show that composition (or evaluation) is pseudo-diophantine. PROPOSITION 2. Let T = (*,,...,/„) and T = (t\,...,t'm). Then the relation K = FoGforKe R[T']S, F € R[T]S and G e R[T']n is pseudodiophantine over R[T u T']. (Here we are thinking of G as a polynomial mapping G: Rm -> R".) PROOF.
We claim there exists a n ^ x n matrix A over R[Tu T'] such
that (••)
F(T) - K(T') = A{T,T)(TG(T')) if and only if AT(D = F G(T').
Now for any monomial a//"1' • • f;"'*, there exist f, such that a,t"'' ■■■ t"hk - 0/g"t'[ ■ • £,"'' = £* = , fiMi, ~ #',)• Tim is easily seen by induction on the degree. For degree one it is clear. For degree greater than one we have
Thus there exists a n i x n matrix L over Z[7"u V] such that F(T) F G(T') = L(T,T')(T - G(T')). On the other hand if F(T) - K(T') = A(T, T')(T - G(T')) add and subtract F ■ G(T') to obtain F ■ G(T') - K(T') = (A + L)(T - G(T')). Substituting G(T') for T the right-hand side is identically zero, so F ■ G{V) = K{T).
1331 .\7'-C0MPLETENESS. RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
.19
Note, the above holds for rings with the property that any polynomial in many variables over R whose values are all zero, is the zero polynomial. Now we deal with an endomorphism. PROPOSITION 3. Let T = (t ,t„) and V = {t\,...,t'„) and R be a ring as above. Then the relation H = F G. H,F € R[T]S and G € R[T]n is pseudo-diophantine over R[T U T'\. {Here we are thinking of G as a polynomial endomorphism G: R" — Rn.) PROOF.
R[TuT']
We claim there exist K € R[T'Y and s x n matrices A,B over such that
(♦*)
F{T) - K{T') = A{T - G{T')),
{***)
H{T)-K{T')
=
B{T-T'),
if and only if H(T) = F ■ G(T). Moreover K{T) = F ■ G{T). There exist K, A satisfying (**) if and only if K{T) = F ■ G{T'). Now there exist B such that (* * *) if and only if the values of H and K agree when T = V. Thus H{T) = F ■ G(T). Similarly K{T) = F ■ G{T). Next we will show that dynamics is pseudo-diophantine. First we need two lemmas which we could derive from Denef 1 in a stronger form, but which we prove in our context. For P € 7J[t] let P4 be the derivative of P. LEMMA
1. The set {{P,P') € Z[t] x Z[t]} is pseudo-diophantine over
Z[t,s,t',s']. PROOF. Q = P' if and only if there exist H, K 6 Z[t, s] such that H{t, s) = P(t) + sQ{t) + s2K{t,s) and H{t,s) = P{t + s). The latter is a pseudodiophantine relation over Z[t,s, f'.j']. LEMMA
2. The set {{k,tk)}k^z>n
C Z x Z[t] is pseudo-diophantine over
Z[t,s,t',s']. PROOF.
Let g € Z[t]. Then g = tk if and only if tg' = kg and g{\) = 1.
Let G € Z[T]n, T = (ti,...,t„) be a polynomial endomor phism G: Z" -» Z" with Zariski dense image {i.e., only the zero polyno mial vanishes on it). Then the set {{k,Gk)}k€Z>n c Z x Z[T]" is pseudodiophantine over Z[S] for a finite set S D T. THEOREM.
Let T = (/,,_...,/„), T = (/', /'„), t = (/,,...,t„,s), P = {t\,...,t'„,s'). Let G = G x idj. Then we claim Gk = J for J e R[T]n if and only if there exist F,H € R[t]n, K e R[P]n and there exist n x {n + I) matrices A,B over R[fuP] such that PROOF.
(*) (*•) {***)
H{t)-F{t) = F{f)-K(P) = H{t)-K{P) =
skJ{T)-T A{f,P){f-G{P) B{t,P){t-P).
For suppose F, H,G, J, A,B satisfy («), (••), (* * «). Let F = £?„, / ( D J 1 with fd $ 0. Then F G = Y.fWT))? and since H = F G by {*»)
1332 40
LENORE BLUM. MIKE SHUB AND STEVE SUMJ.
and (* • *), d-\
sH-F = (ftlG{T))sd+l+Y!,M'G(T))-fM(T)s,+l-fo
=
skJ(T)-T.
1-0
Since fo ■ G(T) £ 0 (the image of G is Zariski dense), k = d + I. f0 = T and yi+i = fj-G which proves that J = Gk{T). In fact we have also shown
F = E*Jo' G's'. On the other hand, if J = Gk(x). Let F = £ G'(7>',
# =^
i*0
G^\T)s'.
<=0
Then (*) is satisfied and there exist A,B,K by Proposition 3. Now since {{k,sk)}keZ' c Z x Z[s] is pseudo-diophantine, by squaring and adding all the equations we are done. Note that this Theorem applies to complex analytic iterations. Given g(z) = z* + c, we may consider c as a variable and define G(c,z) = (c, z2 + c). More generally, consider a finite dimensional machine M with one branch node over a ring R (see Figure 18). t /
t*-g{t) I
/
(
A>0 >
Input t€R"
^
1
y
Replace t by git) where g-B.n-*B.n is a polynomial map
hsO t
Output
FIGURE 18
where A: R" -* R is given by A(r) = £ bit' and £ ■» (glt...,g„)
is given
For each nonzero coefficient c/,, of £ assign a variable C/,,; and each nonzero coefficient 6/ of A a variable 2?/. Suppose there are p new variables C/,, and q new variables £/. Let g, = £ C/,,V € Z[C, T] and A = £ 5,r' € Z[B, rj. Here C - (C/,,), r = (/ t„) and B = (B,). Define G € Z[C, r]xZ[C,r], G: R'xR" -+ R" x RH by G = Ij x g where £ = (&,... |„). Suppose that the image of G is Zariski dense. Then the set (k, Gk) is pseudo-diophantine. The element t - (t\,..., t„) is in Hu if and only if specializing CIJ to CIJ and B/ to */, we have hGk(t) < 0 for some k. Thus the set {k,h, Gk)
1333 ^-COMPLETENESS, RECURSIVE FUNCTIONS AND UNIVERSAL MACHINES
41
is pseudo-diophantine and to find the halting set of M, we need only specialize the coefficients and evaluate these polynomials. PROBLEM 9.2. Suppose Q c Z[T, U), r » ( / , / , ) , U = (K, um) is pseudo-diophantine. For each ueRm,\ei * a * - {x € JT\3p €tJ with />(x, u) < 0}. Then JTQJI is R~E< over U. Do all R.E sets over R arise this way? 10. Most Julia sets are ■wdcridaMe. In this section we investigate which Julia sets of rational maps g = p/q of the Riemann sphere (C) into itself can be R.E. over R. Recall briefly that the Julia sets J = Jt is the set of points z € C such that the family of iterates, ga, n > 0 of g fails to be normal on any neighborhoods of z. This set is, for degree g > 2, the same as the closure of the repelling periodic points of g. (See Brolin for this and other properties of Julia sets we shall be using.) In particular (i) J is closed and fully invariant for g, i.e. g(J) « J ■ g~l(J)(ii) For any relatively open set U c / there is a finite n > 0 such that g"(U) m J. (iii) If J has interior it is all of £. The fixed points ofg ore hyperbolic if the derivative of g has modulus different from one for eachfixedpoint, i.e., |^(z)| # 1 whenever g{z) « z. An R.E. Julia set of a rational map g:C — C is either, (a) Empty, and g is a rotation, or a constant, or, (b) a point; and g is fractional linear but not a rotation; or, (c) a real analytic arc, or, (d) a real analytic Jordan curve, or, (e) the whole sphere £. Moreover, if thefixedpoints ofg are hyperbolic then the arc in (c) is actually the arc of a round circle, the Jordan curve in (d) is a round circle and in this last case g is conjugate by an affine map to a Blaschke product » _ -w THEOREM.
Z — OO^Jli' Z Wi'h | f l ° 1 * 1 alUi N S I PROOF. From (ii) it follows that J is either connected or has uncountably many distinct components. Since J is the countable unions of semialgebraic sets each of which has finitely many components, / must be connected. If / has interior it is the whole sphere. If J is empty or a single point it follows fairly simply that degree g « 1 or 0. Rotations are easily seen to have empty Julia set and other fractional linear have one fixed point as their Julia set. Thus we are reduced to the case where J is the countable union of 1-dimensional semialgebraic sets, which we may assume to be closed. Now the Baire category theorem asserts that one of these semialgebraic sets has relative interior in / . Thus we may find a real analytic arc in it, / and a finite no > 0 such that *-(/)«/.
1334 42
LENORE BLUM. MIKE SHUB AND STEVE SMALE
g"a(I) may have as many as two ends and a finite number of branch Y or crossing points X. But if a Julia set has branching or crossing points they must be dense by (ii). Thus J has no branch or crossing points and is either an analytic arc or an analytic Jordan curve. This finishes the proof of parts (a)-(e) of the theorem. Now we may consider that we are in case (c) or (d) and the fixed points of g are all hyperbolic. The complement of J is then one or two fully invariant simply connected domains. By Sullivan's classification theorem these must be the basins of attraction of a hyperbolic fixed point and by Brolin (Theorem 9.1 and Lemma 9.1), the arc is the arc of a round circle, the circle is a round circle and the map in this case is conjugate to a Blaschke product. PROBLEM 10.1. If the Julia set of a rational map of C is a differentiable arc or difierentiable Jordan curve is it necessarily the arc of a round circle or a round circle respectively even without the hypothesis that all fixed points are hyperbolic? EXAMPLES, (a) If g has three attractive fixed points, Jt is not R.E. (b) If f(z) is a polynomial with at least 3 distinct roots and A/(r) = * - hz)lf{*) then JNf(x) is not R.E. PROOF. The basins of attraction permitted by (a)-(e) of the theorem are all of C or 1 or 2 simply connected domains. If g has three attractive fixed points, its basin has at least three distinct components. Nf{z) is the Newton iteration for the zeros of / , and has an attractive fixed point at each roots of / . PROBLEM 10.2. Classifying Julia sets as to their complexity is an im portant problem in complex analytic dynamics (see e.g., Blanchard). Can relative decidability be called into play? That is, define as in classical re cursive function theory, a (Julia) set A to be decidable in (Julia) set B if a machine with an additional node for deciding B (i.e., an "oracle") can be used to decide A. Do the resulting equivalence classes and hierarchies shed any light on the Julia set classification problem? 11. SMK final murks aad problems. There are a number of ways to further develop and modify the model presented in this paper. 1. For example, it would be of interest to explore an analogous theory for unordered fields such as C, fields of finite characteristic and valued fields (such as the p-adic fields Qf). A way to do this is to replace the branch nodes by branch nodes which distinguish between h{z) # 0 and h{z) « 0. For the case of valued fields, one could incorporate branching decisions based on the values of the valuation function. 2. We have already indicated how the theory of ATP-completeness could be developed forringsin general and have perhaps alluded to ways the NPcompleteness of the Feasibility problem might be generalized. We think this direction looks promising. 3. It would also seem natural to further develop ideas from recursive function theory such asfixedpoint theorems, reducibilities and hierarchies for machines over a ring R. 4. To develop a theory of probabilistic algorithms over R, one would naturally adjoin "coin tossing" nodes.
1336
44
* tgHOKMLUM. W U SHUB AND STEVE SMALE
is the degree and the a, are the coefficients of / given as pairs of real numbers. A recent paper, Kim, describes algorithms for the solution of the i all roots problem, Xf%. Algorithm 1 there, for example, can be seen to be given by a polynomial time machine. PROBLEM. What about other classical problems of numerical analysis? REFERENCES F. Abramson. Effective computation over the real numbers. Proceedings of the 12th An nual (IEEE) Symposium on Switching and Automata Theory, 1971, pp. 33-37. A. Aho, J. Hopcroft, and J. UUman, The design and analysis of computer algorithms. Addison-Wesley, Reading, Man., 1979. E. Becker, On the real spectrum of a ring and its application to semialgebraic geometry, Bull. Amer. Math. Soc. (N.S.) 15 (1986), 19-60. M. Ben-Or, Lower bounds for algebraic compulation trees. Proceedings I Sth ACM STOC, 1983. pp. 80-86. P. Bouchard, Complex analytic dynamics of the Riemann sphere. Bull. Amer. Math. Soc. (N^.) 11 (1984), 85-141. A. Borodin, Structured vs. general models in computational complexity. Logic and Algo rithmic, Monographie, no. 30, de L'ftitrignetnent Mathematique, Geneva, 1982, pp. 47-65. H. Brolin, Invariant sett under iteration ofrationalJunctions, Ark. Mat. 6 (1963) 103-144. J. Canny, Some algebraic and geometric computations in PSPACE, 20th Annual ACM Symposium on Theory of Computing, 1988, pp. 460-467. S. A. Cook, The complexity of theorem proving procedures, Proceedinp 3rd ACM STOC 1971, pp. 151-158. N. Cutland, Computability, Cambridge Univ. Press. Cambridge, 1980. M. Davis, Computability and unsohability, Dover, New York, 1982. M. Davis, Y. Matjjasevic and J. Robinson, Hilben s tenth Problem. Diophantine equa tions'. positive aspects ofa negative solution, Mathematical Development Arising from Hilben Problems. Pre*. Sympos. Pure Math., vol. 28, Amer. Math. Soc., Providence, R.I., 1976, pp. 323-378. M. Davis, and H. Puuum, Diophantine sets over polynomial rings. Ill, i. Math. 7(1963), 2S1-2S6. J. Denef, I, Diophantine sets over Z(7T, Proc. Amer. Math. Soc. 69 (1978), 148-150. J. Denef, 2. The Diophantine problem far polynomial rings andfieldsof rational functions, Trans. Amer. Math. Soc. 242 (1978), 391-399. B. C. Eaves and U. G. RothMum, A theory on extending algorithms for parametric prob lems, Department of O.R., Stanford Univ„ 1985. S. Eiienberg, Automata, languages, and machines, vol. A, Academic Press, New York. 1974. E. Engeter, Algorithmic approximations, J. Comput. System Sci. 5 (1971), 67-82. H. Friedman, Algorithmic procedures, generalised Turing algorithms, and elementary recursion theory. Logic Colloquium 1969, (R. O. Gandy and Yates, eds.) C.M.E., NorthHoUaad, Amsterdam, 1971, pp. 361-390. H. Friedman and K. JCo, Computational complexity ofrealfunctions, J. Tbeoret. Comput. Sci. »(1982), 323-352. N. Friedman, Some results on the effects of arithmetic comparisons, Proceedinp of the 13th Annual (IEEE) Symposium on Switching and Automata Theory, 1972, pp. 139-143. M. Garey and D. Johnson, Computers and intractability. Freeman, New York, 1979. D. Y. Grigorev and N. N. Vorobojov, Solving systems ofpolynomial inequalities in subexponential time, i. Symbolic Computation, 5 not. 1 and 2 (1988), 37-64. L. Harrington, M. Moriey, A Seedrov, and S. Simpson (eds.), Harvey Friedman s Re search on the Foundations of Mathematics. North-Holland, Amsterdam, 1985. G. T. Herman and S. D. hard, Computability over arbitraryfields,J. London Math. Soc. (2)2(1970). 73-79.
1337 fff-COMPLETENBK. MCUMIV8 FUNCTIONS ANDJINDSBISAL. MACHINES
45 .
H. J. Hoover, Feasibly constructive analysis, Ph.D. Thesis, Department of Computer Science, Univ. of Toronto, 1987. J. P. Jones and Y. V. Matijasevic, Registernmachine proof of the theorem on exponential diophamine representation of enumerable sets, J. Symbolic Logic 49 (1984), 818-829. M. H. Kim, Topological complexity of a rootfindingalgorithm, preprint. C. Kreitz and K. Weihrauch. Complexity theory on real numbers and functions. Lecture Notes in Computer Sci., 145, Theoretical Computer Science, (A. B. Cremers and H. P. Kreigel eds.), Springer-Verlag. Berlin and New York, 1982, pp 165-174. L. Lovasz. An algorithmic theory of numbers, graphs and complexity, CBMS-NSF Series 50, SI AM, Phil., Pa, 1986. M. Machley and P. Young, An introduction to the general theory of algorithms, NorthHolland, New York, 1978. X. I. Manin, A course in mathematical logic, Spnnger-Verlag, Berlin and New York, 1977; 1984 (2nd printing). Z. Manna, Mathematical theory of computation, McGraw-Hill, New York. 1974. B. Mazur, Arithmetic on curves. Bull. Amer. Math. Soc. (N.S.) 14 (1986), 207-259. J. Milnor, On the Betti-numbers of real varieties, Proc. Amer. Math. Soc. 15 (1964), 275-280. M. Minsky, Computation: Finite and infinite machines, Prentice-Hall, Englewood Cliffs. New Jersey, 1967. Y. N. Moschovakis, Foundations of the theory of algorithms. I, draft 1986. V. Ya Pan, Methods of computing values of polynomials, Russian Math. Surveys 21 no. 1 (1966), 105-136. M. B. Pour-El and I. Richards, Computability and noncomputability in classical analysis, Trans. Amer. Math. Soc. 275 (1983), 539-560. F. P. Preparata and M. I. Shamos. Computational geometry, Springer-Verlag. Berlin and New York, 1985. M. Rabin, 0.1, Computable algebra, general theory and theory of computablefields.Trans. Amer. Math. Soc. 95 (1960), 341-360. M. Rabin, 0.2, Proving simultaneous positivity of linear forms, J. Comput. and System Sci. 6(1972), 639-650. J. Renegar, A faster PSPACE algorithm for deciding the existential theory of the reals, Technical Report No. 792, School of Operations Research and Industrial Engineering, Cornell, April 1988. (Also, Proceedings, of the 29th Annual Symposium on Foundations of Computer Science, 1988, pp. 291-295.) H. Rogers, Jr., Theory' of recursive functions and effective computability, McGraw-Hill, New York, 1967. A. Shamir, Factoring numbers in O(logn) arithmetic steps. Inform. Process. Lett. S no. I (1979), 28-31. J. C. Shepherdson and H. E. Sturgis, Computability of recursive functions, J. Assoc. Comput. Mach. 10 (1963), 217-255. A. Schonhage, On the power of random access machines, Proc. 6th ICALP, Lecture Notes in Computer Science, no. 71, Springer-Verlag, Berlin and New York, 1979, pp. 520-529. S. Smale, On the topology of algorithms. 1, J. Complexity 3 (1987), 81-89. E. Sontag, Polynomial response maps. Lecture Notes in Control and Information Sciences, Springer-Veriag, Berlin and New York, 1979. M. Steele and A. Yao, Lower bounds for algebraic decision trees, J. Algorithms 3 (1982), 1-8. V. Strassen, Algebraische berechnungskomplexitat, Perspectives in Mathematics, (Basel), Birkhauser-Veriag, 1984. D. Sullivan, Quasi-conjormal homomorphisms and dynamics. Ill, IHES preprint. A. Tarski, A decision method for elementary algebra and geometry, 2nd ed., Berkeley and Los Angeles, 1951, vol. 1, 63 pp. R. Thorn, Sur I'homologie des variites algebriques reelles. Differential and Combinatorial Topology, Princeton Univ. Press, Princeton, New Jersey, 1965, pp. 255-265.
1338
46
-tZMORE BUm, HIKE SHOT MID STOTE SWXH
J. Tiuryn, /< smr>> o/;A? /oy/f of effective definitions. Lecture Notes in Computer Sci., 125, Logic of Programs, (E. Enteler. ed.), Springer-Verlag, Berlin New York, 1979, pp. 198245. J. F. Traub and H. Wozniakowiki, Complexity of linear programming, Oper. Res. Lett. Ino. 2, (1982). 59-62. L Valiant, Completeness classes in algebra, Proc 11th Ann ACM STOC, 1979, pp. 249261. L. van den Dries, Alfred Tarski's elimination theory for real closed fields, J. Symbolic Logic 1(1988), 7-19. J. von zur Gathen, Algebraic complexity theory. Technical Report No. 207/88, Depart ment of Computer Science, University of Toronto, January 1988, 37 pp. S. Winograd, Arithmetic complexity of computations, SIAM Regional Conf. Ser. Appl. Math. 33(1980), 93pp.
IMTEKNATIONAL COMPUTER SCIENCE INSTITUTE, 1947 CENTER STREET. BERKELEY. CAL
IFORNIA 94704 IBM T. J. WATSON RESEARCH CENTER, YORKTOWN HEKJHTS, NEW YORK 10598-0218 DEPARTMENT OF MATHEMATICS, UNIVERSITY OF CALIFORNIA, BERKELEY, CAUPORNIA
94720
1339 SUM R(«virw Vol i.". No. 2. pp. 211-220. June 1990
C 1990 Society for Industrial and Applied Mathematics
SOME REMARKS ON THE FOUNDATIONS OF NUMERICAL ANALYSIS* STEVE SMALEt Abstract. The problem of increasing the understanding of algorithms by considering the foundations of numerical analysis and computer science is considered. The schism between scientific computing and computer science is discussed from a theoretical perspective. These theoretical considerations have an intellectual importance when viewing the computer. In particular, the legitimacy and importance of models of machines that accept real numbers is considered. Key words, numerical analysis, foundations, complexity, algorithms AMS(MOS) subject classification. 65
Introduction. I would like to discuss a problem already raised by John von Neumann in 1951 ("The General and Logical Theory of Automata," in [von Neu mann, 1963]). Its relevance compels me to quote once more (i.e., [Smale, 1985]) from this article. Von Neumann wrote: The theory of automata, of the digital, all or none type, as discussed up to now, is certainly a chapter in formal logic. It would, therefore, seem that it will have to share this unattractive property of formal logic. It will have to be from the mathematical point of view, combinatorial rather than analytical. . . . a detailed, highly mathematical and more specifically analytical, theory of automata and of information is needed.
Von Neumann has anticipated what I like to call the conflict between scientific computation and computer science. Two subjects with many common goals have grown apart. On the theoretical side especially, this conflict seems broad and deep. For example, numerical analysts (I speak of numerical analysts as the theorists of scientific computation) have a disdain for Turing machines. On the other side, computer scientists often minimize the importance of calculus in the college curric ulum. These are just signs of a bigger schism. The following little table illustrates the contrast I am trying to make.
Mathematics Problems Goals Foundations Complexity "Machine"
Scientific Computation
Computer Science
Continuous Classical Practical, Immediate None Undeveloped None
Discrete Newer Long Range Developed Developed Turing
This display certainly oversimplifies the situation, so I will elaborate. Continuous mathematics implies the centrality of real numbers and includes calculus and differ ential equations. The discrete mathematics of computer science emphasizes logic, • Received by the editors October 5, 1989; accepted for publication October 5, 1989. This paper was presented as the 1989 John von Neumann Lecture at the SI AM 1989 National Meeting in San Diego. California, July 17-20. This work was supported in part by the National Science Foundation. t Mathematics Department, University of California, Berkeley, California 94720.
211
1340
212
STEVE SMALE
combinatorics, and number theory. "Newer problems" needs to be augmented by problems of number theory. I have felt the pressure of immediate goals when numerical analysts compliment a paper of mine, but urge me to implement the algorithms on useful problems. Theoretical computer science seems much farther from such concerns. Algorithms in numerical analysis are primarily a means to solve practical problems, while in computer science, algorithms are studied systematically in their own right. Later, problems of the foundations of numerical analysis will be addressed. For the moment note that algorithms are the main object of study in scientific computa tion, yet there is not a formal definition of algorithm. I am reminded of how the development of the definition of differentiate manifold was so important in the history of differential topology. In contrast to numerical analysis, the use of Turing machines gives the computer scientists a unifying concept of algorithm, well formalized. Thus complexity theory can speak of lower bounds of all algorithms without ambiguity. A suitable complexity theory for numerical analysis should at least possess theorems for the computing time of basic algorithms, deterministic or probabilistic, in terms of numbers, including desired precision and "size" of problem associated to the input. See the Appendix on this subject I hope the above communicates my view of numerical analysis as an eclectic subject with weak foundations. On the other hand, its achievements over many centuries, and especially since the revolution of the computer, have an undeniable greatness. I am reminded of the discipline, ordinary differentia] equations which, in 1960, was sometimes regarded as a bag of tricks. Topology and modern mathematics in general have since helped give unity to large portions of that subject around the development of dynamical systems. A major obstacle to reconciling scientific computation and computer science is the present view of the "machine," i.e., the digital computer. As long as the computer is simply seen as a finite or discrete object, it will be difficult to systematize scientific computation. The machine has a central role in theory due to the influence of Turing. I refer to the picture: (1) machines, (2) the derived notion of computable function defined as the input-output map of a machine, and (3) an algorithm as a computable function together with a class of problems it solves. Thus it is important to confront the question: What is the nature of the machine? Of course, we already have an important idealization of the digital computer in the Turing machine. However, the matter does not end there. Idealizations are not necessarily unique. As an example, for the material world we have idealizations of Newton, of Einstein, and of quantum mechanics. Each of these idealizations illustrates the truth of certain aspects of the physical world. And, of course, certain of these idealizations have been unified. Consider then the problem of reconciling the digital machine with the continuous mathematics of scientific computation. Newton faced an analogous problem in writing "Principia" (see, e.g., [Smale, 1988]). Newton's intended model of matter was based on real numbers for his calculus and geometry. Yet he believed, like other scientists of his day, that the universe was composed of atoms, indivisible and discrete. Eventually he resolved this paradox by taking a limit as the number of atoms became denser and denser. This could be justified by the large number of atoms in a material substance. I believe this little historical observation holds a lesson for us on understanding the computer. The high density of rational numbers available for use indicates that
1341 FOUNDATIONS OF NUMERICAL ANALYSIS
213
using real number inputs is not a bad idealization. The Newton story is a response to an objection I meet: "How can a digital computer input an arbitrary real number?" When asking numerical analysts for a formal definition of algorithm, I usually have not received a satisfactory answer. There was an exception. Arieh Iserles replied to my query with "a Fortran program.'' This makes a lot of sense. Recall that Fortran does take real number inputs. Blum, Shub, and Smale [BSS, 1989] attempt to meet the problem of giving a real number idealization of the machine. (See also [Blum, 1989].) The model described in that article may be seen as a mathematical idealization of the essentials of a Fortran program. Fortran is already formal but rather complex. Let me describe what we have done without much precision. Our model is more of a simple abstraction of aflowchartof a computer program than an innovative discovery. The following Newton-Method machine is a good illustration. Input x € R
I \f(x)\2
pr"l output. This example is a version of a program for using Newton's method naively to find approximately a zero of a real polynomial, f(x)« £o a,x'. There is a space of inputs R, the real numbers. There is an output space, in this case also R. Moreover, R plays the role of state space where the computation x' ■= x -f(x)/f'(x) takes place. There is a distinguished set ft« C R of inputs for which the machine halts. Thus flu consists of initial x such that Newton's method is defined iteratively up to an x' for which \f(x' )| < c S1M is the Halting set (unfortunately called R.E. set in [BSS]) of M (after Turing). An input-output map 4>»:fl*—»Ris defined by following the flow of the chart. It is easy to imagine extending the example to Newton's method in many variables. In general, afinite-dimensionalmachine will have an associated input space J, output space , and state space ^ , each some real Cartesian space. The nodes of the general machine do rational computations or are branch nodes (corresponding to "if, then" instruction) as in the example (besides an input node and output nodes). For each machine M there is a halting set QM C J and an input-output map 4>M'- QM —► 0. The formal definitions are carried out in [BSS]. The definition of a general, perhaps infinite-dimensional machine, has a slightly more technical flavor. But it is apparent how naturally the model follows and is close to actual programming. In fact, there is an equivalent definition in terms of a list of instructions. A computable map (over R), 4>:fi-»R', Q C R \ is a map which is the inputoutput map of some machine M. It can be checked that the composition of computable maps is a computable map; rational maps are computable, as well as the characteristic function of the set [x € R | x 2 0|.
1342
214
STEVE SMALE
Since the arithmetic operations on a computer are so primary, the computations on our model are restricted to rational functions (although eventually the power could be expanded to include wider classes of functions). Thus the theory of these machines has a pronounced algebraic flavor. The algebraic emphasis permits an extension of this theory of computation to the situation where the reals are replaced by a ring from some general class. If the ring R is not a field, then rational maps are replaced by polynomial maps. If/? is not ordered, then < is replaced by = in the branch nodes. Thus with machines over the integers Z, or machines over C, the complex numbers are defined. The generalization permits some unity with theoretical computer science. Com putable functions over / ? » Z coincide with the traditional computable functions (or partially recursive functions). This unity is deepened by passing to complexity theory as the following indicates. The celebrated and beautiful problem "does P = NPT" is central to complexity theory of computer science (see [Garey and Johnson, 1979]). The importance of U P=NPT' is based on Steve Cook's TW-compIeteness theorem [Cook, 1971] and subsequent work of Dick Karp [Karp, 1972]. The study of machines over C or over R yields a natural extension of the problem "P = NF!n to the complex numbers and to the reals. "Does P — NP over C? over R?" seems to be difficult, and because of our new AT'-completeness theorem, contains some of the importance of the original problem. The algebraic character of the machines I have described carries over to the problem "P - NP?", both in the new (over C) and classical setting (over Z). The AT'-completeness theorem over C of [BSS] (the proof there for the reals carries over simply to cover the complex case) asserts the existence of a universal problem among NP (decision) problems over C. This universal problem is essen tially Hilbert's Nullstellensatz, referred to as HN. The input is a set of polynomials P,, • ••,P m :C—»C. The output is one if the />, have a common zero; it is zero otherwise. This problem is in "NP over C" because there is a system of tests f € C" which correlates appropriately with the [P,| having a common zero. Namely, given input the Pi and f € O , test in polynomial time if P,(f) — 0 all /. For precision of all of these terms, see [BSS]. It is easier to show that a problem is in NP (non-deterministic polynomial time) in contrast to P (polynomial time) in the same sense that it is easier to check a proof of a theorem than to find a proof. Our W-completeness theorem asserts that any A^ problem over C can be reduced in polynomial time to HN. Thus an algorithm for HN yields an algorithm for any problem in NP over C. As a consequence, if HN is in P, then P = NP over C and conversely. There are algorithms over C for the Hilbert Nullstellensatz and there is a large literature on this subject. [Brownawell, 1987] gave a strong bound. See this paper as well as [Dube, 1989] and [Ierardi, 1989] for the mathematics and history of the problem. The existence of an algorithm for HN implies as a corollary to our TW-completeness theorem, that any problem in NP over C is decidable. This fact, in contrast to the traditional analogue, is not obvious. The continuum of guesses cannot be used to describe a decision procedure. The real version of an NP complete problem [BSS] is called 4-Feas. Here the input is a polynomial / : R"-+R of degree 2 4. The output is 1 i f / h a s a zero, 0 if
1343 FOUNDATIONS OF NUMERICAL ANALYSIS
215
not. The history of algorithms (over R) for this problem perhaps starts with [Tarski, 1951]. [Canny, 1988] has found good complexity bounds (still exponential) for 4Feas in a more general context. [Renegar, 1989] has a very general, systematic, and complete account with good bounds. The problems 2-Feas and 3-Feas are in P (have polynomial time algorithms), as shown in [Triesch, 1989]. Appendix. Round-off error, approximate solutions, and complexity theory. Since machines over R do exact arithmetic it is natural to ask, how could these machines be helpful in numerical analysis where round-off error is so important? Even using exact arithmetic, algorithms can usually only solve numerical problems approxi mately, to within "accuracy e." Complexity theory of computer science must be modified to deal with these problems if it is to be useful in numerical analysis. Largely following [BSS] we suggest a way of how machines over R may be used to give a basis for a complexity theory for numerical analysis. Computer scientists say that an algorithm defined by a machine M is tractable (or polynomial time, or in P) if the time T(y) = Time(j>) of computation associated to input y satisfies the bound (•)
T(y)£c(sizeyy
where the constants c and q depend only on M. Here time is the number of Turing machine operations and size is the number of bits. A problem is tractable if there is a tractable algorithm solving it. A way of translating (*) into a language of machines over R is to first let time T(y) be interpreted as the number of nodes traversed in the computation of 4u(y). The input space of an (infinite-dimensional) machine is the linear space J = R* of all sequences of real numbers of the form y=(yi,)'2,-
-,y„,0,0,■■■)
(the output space is the same?. Then s(y) = sizc(y) is taken as n. For the NPcompleteness theorem this interpretation of (•) means that M is in P over R. For purposes of numerical analysis it is important to formalize the notion of "problem" and to expand on the above definition of the size of an input. Toward this end, let the space of problem instances (or admissible inputs) be denoted by YC J. A problem is a subset X C Y x e ( the space R" of outputs). The idea is that for each y^-Y, there is a set of solutions 4>(y)e
1344
216
STEVE SMALE
Example 2. The problem is to find all roots of / Here Y is the same as in Example 1, but ff is C and * « | ( / f , , •••, W « Y x Cd \ f(z) = adr1(z - f,)|. (Strictly speaking, ( / f„ • • •, fc, 0,0, •••).) If it is restricted to the case a*« 1, then A' is canonically isomorphic to C, the space of (f,, •••,£/) and «•: A1—» y is the map defined by symmetric functions (sometimes called the Vieta map). Example 3. The n-variable versions of Examples 1 and 2. Here Y is some set of polynomial maps C—*C. The constructions of those examples can be extended to this case. We say that an algorithm solving a problem XCYx&isa machine over R (or its input-output map
where the norm measures the distance of ((c, y),
T(c,f)Sc(d\ogd+\logc\)<,
fePAD
is proved. Here T(e,f) could be taken as the number of nodes traversed in the computation of an approximate solution of/(f) «= 0. Essentially T(c,f) is the number of modified Newton iterates. That estimate suggests that (•) could be used as a criterion of tractability in numerical analysis provided size(.v) is replaced by size(>>) + | log e |. Indeed this suggestion is in therightdirection but more is needed. Let us return to the original Example 1 where the input space Y is the set of all polynomials / whose leading coefficient a^ is not zero. Define a function R — Rf of/by |4f|
/
This number R is used in a normalization to put /into PA 1). The preceding estimate becomes T(c,f)Zc(d\ogd+1
log r | +log/?)«.
The function R is the prototype of a real function W of inputs, I will call the weight. The weight of y is to measure the "difficulty of the problem instance y." Sometimes If could be taken as a norm, sometimes as a condition number. In writing [Smale, 1981], I teamed of an apparently general phenomenon that I am sure is familiar to numerical analysts. There is a certain set of ill-posed problems for which algorithms fail, and as a problem instance tends to these ill-posed problems,
1345 FOUNDATIONS OF NUMERICAL ANALYSIS
217
the time of computation tends to ». In this light the weight is to resemble the reciprocal of the distance to the set of ill-posed problems. A consequence of these various considerations and other examples is that the computer science tractability estimate (•) extends to numerical analysis in the form
TXcy)£cs{c,y)< (»NA) Hcy)**s(y)+1 log* | + \ogW(y). Here the weight function W, and even the spaces Y, X, have to be chosen with much thought, taking into account the particular numerical analysis setting, its limitations, and just what a good algorithm can be expected to accomplish. We might also consider Examples 1,2, and 3 from the point of view of quadratic convergence at a solution. It would be reasonable in this case to replace | log e \ by log | log c | and moreoever incorporate into the weight o f / a quantity going to » as / approached a polynomial with multiple roots. The work of [Renegar, 1987a, b] on Newton's method could no doubt be interpreted from this point of view. It would be fruitful to see how various other results in numerical analysis could be formalized via (*NA). For example, using [Kostlan, 1988], it seems that the power method of finding the largest eigenvalue could be dealt with. On the other hand, algorithms for approximating solutions of nonlinear partial differential equations give greater difficulty since a global analysis is required, not just an asymptotic analysis. Complexity theory of numerical analysis is not simply an academic exercise. Specifications of limits on inputs for which an algorithm is tractable become necessary. Ill-posed problems must be isolated. The general systemization would clarify and in the long run help improve the practice of scientific computation. Condition (*NA) needs to be supplemented by the condition of numerical stability in order that an algorithm qualify as tractable. The definition of numerical stability in the numerical analysis literature needs formalization. For example, [Conte and de Boor, 1980, p. 391] write: "Loosely speaking, we can say that a method is unstable if errors introduced into the calculations grow at an exponential rate as the computation proceeds." Very tentatively, I will try to capture that meaning with machines over R. In [BSS], to each machine M over R, there is an associated computing endomorphism HM: / x y i where A" is the set of nodes and 5* the state space R". Then the sequence of computations starting with input y is essentially: (ni,xi)=HM{n,.l,xl-i), «o=input node,
i= 1,2, • • • x0"y
n r - a n output node,
xT"
T=T(y).
Now say that 6 > 0 is an admissible roundoff error for (e, y) if for any se quence (/!,-, Xj) satisfying |(«,,x,) - HM(ii-u Xj-\)\ <6 and n 0 -input, x0~y, then \XT — 4>M(y)\ < * and nT is an output node (we could replace faiy) by 4>MU, y)). Let 6M(c, y) be the supremum of the admissible roundoff errors. Say that an algorithm defined by a machine M over R which solves a problem X C Y x 0 approximately is numerically stable if: (••)
6„(c,y)Scs(c,y)<
for some constants c, q depending only on M. Note the resemblance of (**) to («NA).
1346
218
STEVE SMALE
As an example, the results in [Kim, 1988b] imply the numerical stability of a global version of Newton's method for Example 1 above. She gives a good explicit estimate for 5M(c, y)The weight of a problem instance seems close to the condition number of the problem. Yet I have hesitated to use the same word for both objects. The condition number has a reasonably clear-cut definition in its own right and, under the right hypotheses, there might be a result asserting that the condition number could be taken as the weight. But it is necessary to be careful as the following example shows. L e t A ' C J ' x ^ b e a problem where X is smooth, and r(x) = y where *•: X—» Y is the associated map of the problem. The condition number of y may be defined as || Drix)'* | where Dx is the first derivative (or matrix of partial derivatives). In case of nonuniqueness, we could take the maximum over all x such that x(x) ■ y. See [Wilkinson, 1973] and [Demmel, 1987a, b] for a good background. Consider now Example 1 from this point of view. Polynomials with multiple roots are ill-conditioned (» condition number). Yet the algorithms mentioned above work reasonably fast on these ill-conditioned problems and so the corresponding weight does not need to reflect the condition number. In any case the condition number does play an important role in many aspects of the topics discussed here. For one thing, the condition number of a problem puts limits on the numerical stability of any algorithm solving that problem (roughly speaking by the chain rule). Probability considerations enter into our view of complexity theory in two ways. In thefirstway, it is supposed that the space of problem instances has a probability measure. A complexity estimate then might have the form 7Xe,tfSc(j<jO+|logc| + |log(l-M)l)« for a set of instances of measure /i. This is the point of view taken in [Smale, 1981] and [Renegar, 1987b]. This estimate is obtained by combining (*NA) with an estimate of the type: measure |^|
W{y)±A\*c
er-
The last inequality leads quickly to elimination theory of algebraic geometry and volume analyses of differential geometry. [Renegar, 1987b], [Demmel, 1987a], and [Ocneanu, 1989] have nice contributions on this. Demmel uses these ideas to help understand why poorly conditioned problems are unlikely. A particular case of the latter is the question, "can large linear systems be solved in the presence of machine roundoff error?" On one hand, the practice of recent decades has given one answer. On the other hand, it would be satisfying to have a theoretical analysis. In fact von Neumann studied this question. The log of the condition number of the matrix can be interpreted as a loss of precision defined by the corresponding linear system as discussed for example in [Smale, 1985] and [Blum and Shub 1986]. In my paper I posed the question of estimating the expected loss of precision. Recent papers that deal with this subject include [Demmel, 1987a], [Kostlan, 1988], [Edelman, 1988], and [Szarek, 1989]. The second way that probability plays a role is when the algorithm itself uses a guess. For example, a simple version of Newton's method chooses an initial point at random. This approach may be found in [Shub and Smale, 1985-1986], [Smale,
1347 FOUNDATIONS OF NUMERICAL ANALYSIS
219 1985]. and [Smale, 1987]. Here formalization in terms of machines over R could use a new kind of node which chooses a number at random. Before closing let me share some of my reservations about focusing on the notion of "the optimal algorithm." Of course, in general it is impossible to have a minimum of two functions simultaneously. Now consider the estimate (»NA) above which could be given in more detailed form as nctfSdsW'
+ c, |logr|«' + c3(log W(y))«>.
The constants c, q,, / = 1,2, 3 are functions of the machine or algorithm, and complexity estimates would aim to have these numbers as small as possible. But by the above remark we cannot expect even to minimize two of these numbers by the same algorithm. The situation is already met to a lesser extent in traditional computer science where problems posed by large ct are ignored in defining optimal algorithms. On the other hand, I believe that complexity lower bounds of various types have a big future in numerical analysis. In the body and in this Appendix, I have not tried to lay down what the foundations of numerical analysis should be. Rather I have presented some ideas which might improve the present situation in this regard.
REFERENCES The references below are the ones that I have referred to in the text. ID no way are they meant to be complete. They reflect strongly my own limited personal knowledge and perspectives. L. BLUM, Lectures on a theory of compulation and complexity over the reals {or an arbitrary ring), in preparation. 1989. L. BLUM AND M. SHUB, Evaluating rational functions: Infinite precision is finite cost and tractable on average. SIAM J. Comput.. 15 (1986). pp. 384-398. L. BLUM, M. SHUB. AND S. SMALE [BSS], On a theory ofcompulation and complexity over the real numbers: SP-completeness. recursivefunctions, and universal machines. Bull. Amer. Math. Soc., 21 (1989), pp. 1-46. W. D. BROWNAWELL. Bounds for the degrees in the Nullstellensatz, Ann. of Math., 126 (1987), pp. 577591. J. CANNY. Some algebraic and geometric computations, in /"-Space, Proceedings of the 20th Annual ACM Symposium on the Theory of Computing, 1988, pp. 460-467. S. CONTE AND C. DE BOOR, Elementary Sumerical Analysis, Third edition. McGraw-Hill. New York. 1980. S. COOK, The complexity of theorem proving procedures, in Proceedings of the Third Annual ACM Symposium on the Theory of Computing, 1971, pp. 151-158. J. DEMMEL, The geometry of ill-conditioning. J. Complexity, 3 (1987a), pp. 201-229. , On condition numbers and the distance to the nearest ill-posed problem. Numer. Math., 51(1987b), pp. 251-289. T. DUBE. Quantitative analysis of problems in computer algebra: Grobner bases and the Nullstellensatz, Ph.D. thesis, Courant Institute. New York University, New York, 1989. A. EDELMAN, Eigenvalues and condition numbers of random matrices. SIAM J. Matrix Anal. Appl., 9 (1988). pp. 543-560. M. GAREY AND D. JOHNSON, Computers and Intractability, Freeman, New York, 1979. D. IERARDI, Quantifier elimination in thefirst-ordertheory of algebraically closed fields. Proceedings of the 21st Annual ACM Symposium on the Theory of Computing. 1989, and Ph.D. thesis. Cornell University. Ithaca. NY. R. KARP, Reducibiltty among combinatorial problems, in Complexity of Computer Computations. R. Miller and J. Thatcher, eds., Plenum Press, New York, 1972, pp. 85-104. M. KIM, Topological complexity of a root finding algorithm. Bellcore, Morristown, NJ (1988a). preprint. . Error analysis and bit complexity: Polynomial root finding problem. Part I, Bellcore. Morristown. NJ (1988b), preprint.
1348
220
STEVE SMALE
E. KOSTLAN, Complexity theory of numerical linear algebra, J. Comput Appl. Math., 22 (1988), pp. 219— 230. A. OCNEANU, On the volume of tubular neighborhoods of algebraic varieties. Mathematics Dept, Pennsyl vania State University, March 1989, preprint. J. RENEGAK, On the worst-case arithmetic complexity ofapproximating zeros of polynomials, J. Complexity, 3 (1987a), pp. 90-113. , On the efficiency ofNewton's method in approximating all zeros ofa system ofcomplex polynomials, Math. Oper. Res., 12 (1987b), pp. 121-148. , On the computational complexity and geometry of the first-order theory of the reals. Parts 1, II, and III, Cornell University, Ithaca, NY, 1989, preprint. M. SHUB AND S. SMALE, Computational complexity: On the geometry ofpolynomials and a theory of cost, I, Ann. Sci. Ecoie Norm. Sup (4), 18 (1985), pp. 107-142. , Computational complexity. On the geometry of polynomials and a theory of cost, II, SLAM J. Comput., IS (1986), pp. 145-161. S. SMALE, The fundamental theorem of algebra and complexity theory. Bull Amer. Math. Soc. (N.S.), 4 (1981), pp. 1-36. , On the efficiency ofalgorithms of analysis, Bull. Amer. Math. Soc., 13 (1985), pp. 87-121. , Algorithms for solving equations. Proceedings of International Congress of Mathematicians, Berkeley, CA, 1986, Amer. Math. Soc., 1987. , The Newtonian contribution to our understanding of the computer, in Newton's Dream, Marcia Stayer, ed., McCill-Queens University Press, Kingston, Canada 1988. S. SZAREK, Condition numbers of random matrices, IHES, Bures sur Yvette, France, May 1989, preprint A. TARSKI, A Decision Methodfor Elementary Algebra and Geometry, University of California Press, 1951. E. TRIESCH, A note on a theorem of Blum, Shub, and Smale, RWTH, Aachen F.R.G., 1989. J. VON NEUMANN, Collected Works, Vol. V, A. Taub, ed., MacMillan, New York. 1963. J. WILKINSON, Rounding Errors in Algebraic Processes, Prentice-Hall, Englewood Cliffs, NJ, 1973.
1349 60
Theory of Computation
The reason for a theory of computation, for me in particular, comes from an attempt to understand algorithms in a more systematic way. The notion of algorithm is very old in mathematics; it goes back a couple of thousand years. Mathematicians have talked about algorithms for a long time, but it was not until Godel that they tried to formalize the notion of algorithm. In Godel's incompleteness theorem one saw for the first time the limitations of computations or the need to study more clearly what could be done. To do so, one has to establish more explicitly what an algorithm is, and I think that this became clearer in the way that Turing interpreted Godel. So let us stop for a moment and look more closely at Turing and his achievements. I think it is fair to say that he laid down the first theory of computation. Perhaps I will be more specific later about what Turing's notion of computation was. We can take as set of inputs the integers Z, so in the Turing abstraction the input is some integer; perhaps not every integer is allowed but only those that were eventually called the halting set ftM of the machine M. This is the domain of computation of the machine M. Given an integer in fiM, we feed it to the machine M and obtain as output another integer.
Z D=i H"M
There is some kind of mechanism here, described by Turing, which I will later on formal ize in my own way. Turing gave different versions of the input set; for instance, finite sequences of zeroes and ones. Thus Godel's incompleteness theorem can be stated in the following way. THEOREM. There is some set S C Z which is definable in terms of a finite number of polynomial conditions and is not decidable. This is Godel's incompleteness theorem as formulated by Turing. Not decidable means that there is no computable function over Z which is 1 on S and 0 out of S. In other
1350 61
words, 5 is decidable if its characteristic function is computable by some machine. Thus, Godel's incompleteness theorem asserts that there exists a set S which is very definable mathematically, yet is not decidable. This is in some sense the beginning of the theory of computation, which shows the limits of decidability. Eventually, from this evolved a theory for present-day computers. It is from this formulation that it evolved into one of the foundations of computer science. Even a very refined theory of computer science is developed from this: This is complexity theory, which 1 could say today lies in the center of theoretical computer science, specially after the work of Cook [3] and Karp [7]. They made use of the notion of speed of computation; now the question is not whether a set is decidable, but whether it is decidable in a time that can be affordable by present-day machines, or whether it is a "crackable" problem. The fundamental question is
PjLNP? This is a very famous conjecture and it is the most important new problem in mathematics in the last half of this century; it is only 20 years old. To me it is the most beautiful new problem in mathematics. Very hard to solve, a very fine notion coming from this theory of Cook and Karp. So we have a very active subject in this area but there is something that is missing. I have talked about the need for a notion of definable algorithm, yet the algorithms mathe maticians have used for a couple of thousand years at least do not fit into this framework. The algorithms we are talking about have to do with real numbers, and specially since the time of Newton they have had to do with differential equations, nonlinear systems, etc. We see the notion of the real numbers R is central. Newton's method to me is a paradigm of a great classical algorithm like the procedure of the Greeks for finding square roots, and it does not fit naturally into this framework, because the framework is quite discrete and to fit Newton's method into it requires de stroying geometric concepts. One can do this in a very cumbersome way —I find this a very destructive way— to deal with the algorithms of continuous mathematics with Turing machines. Indeed, some work has now been done to adapt the Turing machine framework to deal with real numbers. Let me mention two such attempts. One of them is recursive analysis, which initially was worked out by Ker-I-Ko and Harvey Friedman [5] and the main name connected with it is Marian Pour-El. She worked with Richards [8] in developing a kind of real number analysis based on Turing machines. There has been very extensive work on this which deals with partial differential equations, and the way to deal with real numbers in this context is to consider a real number s defined by its decimal expansion s = 1.2378... . A real number is computable in this sense if there exists a Turing machine which says that the first digit is 1, the second 2, the third 3, and so on, with the decimal point in the appropriate place. So a computable real number is given by a Turing machine. These work with computable real numbers and eventually provide a very successful theory. Similarly there is the notion of interval arithmetic from R. E. Moore (see [1]). In some ways it is close to the work of Pour-El but in a quite different direction. Thus, the foundations are probably being laid for a theory of computation over the real numbers.
1351
62 Now, continuing from a very different point of view, I will devise a notion of computa tion taking the real numbers as something axiomatically given. So, a real number to me is something not given by a decimal expansion; a real number is just like a real number that we use in geometry or analysis, something which is given by its properties and not by its decimal expansion. Eventually, I will talk about a notion of computability over the real numbers which takes this point of view. There one thinks of inputing a real number not as its decimal expansion but as an abstract entity in its own right. Some mathematicians and computer scientists have trouble with the idea that a ma chine takes as input an arbitrary real number. I wrote a paper [13] on precisely this point, saying that here one idealizes, as in physics Newton idealized the atomistic uni verse —making it a continuum— in order to use differential equations. One can idealize the machine itself by conceiving it as allowing an arbitrary real number as input, but I am not going to argue about this point today. In a preliminary phase, I was concerned with the problem of root finding for polyno mials for many years. In that process I faced the kind of objects known as tome machines. It is not a theory of computation, but just a preliminary. We take as input now the coefficients {ao, a i , . . . , a^} of a complex polynomial / of degree d, and we think of it over the real numbers, i.e., each a, is given by its real and imaginary parts. Thus, we think of this as the input and we describe the computation in the language of flowcharts. Then comes a box describing the computations. We replace a by g{a), where g is a rational function. This vector of numbers [ao, a i , . . .,aj] can be considered as a state and in this step this state is transformed by a rational function. Then we can put down and answer the question of whether some coordinate of the state is less than or equal to 0. Depending on the outcome of this comparison, we continue along the corresponding branch.
{a\,---,an}
'' a - » g(a)
'
output approximate zero
So we go down the tree in this way, and eventually we may output the approximate
1352
63 zeroes { { , , . . . ,£d] of / up to some e. In fact this is a good way to express an algorithm for solving this equation. We next ask: How about these nodes which branch? To what extent are they necessary? The topological complexity of a problem in general is the minimum number of these branch nodes for any machine which solves the problem. I am not being completely precise about what resolving a problem means. One can imagine this example as a prototype of the general situation of a machine that solves problems. In particular, for this problem of finding the zeroes of polynomials we have the notion of its topological complexity. And the theorem [12] is as follows: THEOREM. to log d.
The topological complexity of the root finding problem is greater than or equal
So the topological complexity increases with the degree, and the proof of this theorem actually is not so easy; it uses the cohomology of the braid group worked by Puchs in Russia [6]. Subsequently, Vasiliev [14] extended this bound to ~ d. So the answer eventually emerged that the topological complexity grows linearly with d and this is a sharp bound. One notes that the Turing machine framework could never deal with this way of looking at all possible algorithms, even in this limited class. There is no useful way of thinking about it in terms of Turing machines, whereas using this kind of tree we were able to give necessary conditions on all algorithms, what we call lower bound theorems. There is some early work dealing with this kind of tree, but this is the first time we have obtained topological complexity results using algebraic topology. Then, shortly after this, we did a joint work with Lenore Blum and Mike Shub [2] and developed this into a complete theory of computation over the reals by allowing loops. We certainly increased the computational power of tame machines by allowing loops to give a notion of computation in general. This situation is reflected in the next picture.
computation
branch
output
1353 64
So here we have also • an input space R , • an output space R , • and also a state space S = W for things happening inside the machine. If /, k, j are finite, this essentially defines a machine. We can take here an oriented graph where the nodes are computation nodes given by rational functions, branch nodes given by inequalities, and input and output defined inside accordingly. At each computation node there is a single output, a single branch going out of the node. A decision node has two. Remember that the number of input branches is arbitrary except for the input node where nothing comes in and there is a branch going out, and an output node, where nothing comes out. This gives a theory of computation for finitely dimensional input and output spaces motivated directly by the flowcharts used in scientific computation. Yet the full theory will have to allow /, k, j = 00 and we will have to have a little more technical process to access far out coordinates, but this is the idea. This model gives an algebraic flavour to the process of computation. We defined this not only over the real numbers but over any ordered ring, eventually any ring. In particular, if we take the ring to be Z, the input space to be a subset of Z, the output space again Z, and the state space Z°°, we obtain Turing theory, and so this extends the Turing theory of computation. We can now say that a Turing computable function is one which is given on some Q of the machine, a domain of inputs, by following the flow of the machine and doing what is said at each node. {admissible inputs] = ft*/ —^* R And this essentially is a complete picture of what we mean by computable function. A function $ M defined by a machine going from the admissible inputs or the halting set of the machine to the output set. And it is precisely equivalent —or practically so— to the notion of Turing computable in the case when the ring is the ring of integers. We have developed for this model notions of computability itself; we have, for instance, shown the existence of universal machines. We have a complexity theory and the problem U P ^ NP ?" is also defined over these rings; for example, the theory for R or C possesses universal or NP-complete problems just as in the case of Cook and Karp. An TVP-complete problem over C (a machine over C is like one over R except that the branch nodes just ask "^ 0 ?") is the following: Does a system of quadratic polynomials have a zero ? The idea of the reduction is to have more polynomials than varYables. So it is an open question whether there is a machine that can decide in polynomial time if there is such a zero. All this is written very carefully in our paper. In Barcelona, Felipe Cucker [4] gave an analog for the real numbers of the arithmeti cal hierarchy of classical recursion theory. There have been developments in different
1354 65 directions in the theory of computation of these machines from the point of view both of complexity theory and of computability. There has also been a lot of controversy and criticism. Let me deal with one main point, making some comments on two sharp critiques by Pour-El and Moore. Our theory of computation is very different from their two theories. In a way this is more or less the basis of their criticism. It has to do with the branching
/
x>0
^
,
We branch according to whether one of the coordinates of the state space is greater than or equal to 0. This is in some respects one of the most controversial elements in the kind of machines we have, because the question is that an actual machine cannot do this. Given a number, for example 0.0000...00... , it may or may not have a one after that eventually. If it never has a one, and we input it to an actual machine, we can never decide this question. If it does have a one, we wait long enough and we can decide it. So we have a problem here when branching at > 0, or equivalently at = 0, and this is the focus of attack of both Pour-El and Moore. Let me give an example here. Both of their theories of computation lead to a notion of computable function which is continuous. Every computable function here is continuous. Even in a strong sense: They have to be constructively continuous. Now the clue to this lies in the philosophy of thinking about t h e real numbers as abstractly given, and choosing the idealization of the right machines. For example, in scientific computation this is the kind of computation carried out traditionally by algo rithms like Newton's method. One does test if something is > 0, then do this, if not do something else. Moreover, the need for these branchings is given by our earlier results on topological complexity. Topological complexity states that if one wants to find zeroes of polynomials then one has to branch, and the number of branchings in the machine is given approxi mately by the degree. Even to approximately solve the fundamental theorem of algebra one needs to branch. So 1 would imply that these two theories of computation do not lead even to an approximate solution of the fundamental theorem of algebra. Here I would refer to a letter I received from Moore a year ago. I do not intend to dwell here on my opinion that numerical analysis and scientifical computing have weak foundations. Moore is the main developer of interval arithmetic and he wrote that "There are foundations for scientifical computation. More than 2000 papers and dozens of books. I invite you to read all of Aberth's book" [lj; it is a book that Moore even sent to me. He said "It will open your eyes to a whole new world." So I opened the book —actually a few weeks ago— and read on page 34 of the book (called Precise Numerical Analysis) "The problem of deciding whether two computable real numbers are equal is therefore a computational problem one should avoid." But problem 3.1 on this book reads: "Given
1355 66
two numbers a, b decide whether a = 6." Later, on page 62, Aberth says: "Solve the problem 6.1: Find k decimals for the real and imaginary parts of the zeroes of a polynomial of positive degree." But the answer to this solvable problem —the fundamental theorem of algebra— needs to pass through d versions of this single problem that "one should avoid." Marian Pour-El very kindly sent me a review that she has given of our paper in the Journal of Symbolic Logic [9], in which she says that it is a very good, highly developed theory of computability over the reals. In the review she confirms that in her theory she can only produce continuous functions and so she cannot solve the fundamental theorem of algebra, not even approximately. What I will now do is to pass on to something which relates to this problem of A'in completeness if only a little indirectly. This is work done jointly with Mike Shub in the last few months. It is an example of an algorithm which fits into our framework. But it is a simple algorithm, so the fact that it is an algorithm in our strict sense is secondary. It is the problem of the complexity analysis of Bezout's theorem. Let me say a little bit about what this is. The situation we look at is as follows: We have a polynomial system / : Cn -
Cn
of n polynomials in n variables of degrees dlt..., dn respectively. One wants to find an algorithm and analize its speed for solving the equation /(--) = 0; not to produce a solution but to analize how much time it takes. The idea is to make complexity analysis on this. The work done so far on polynomial equation solving can be summarized by dividing it into two parts: one is Newton's method as the basic algorithm —it is essentially the method used by the Greeks for finding square roots— and the other method is elimination theory, a very algebraic method; it works over arbitrary fields. In the first one we are using some kind of norm or metric, so it is metric-oriented. It is the method of choice of numerical analysts. The second is probably the method that would be chosen by a computer scientist. My own inclination is to the first side. Numerical analysts have a better focus on the problem. They do not have a complexity theory or any kind of foundation, but they have a better instinct about how to solve this problem. In any case, what we use is some global version of Newton's method to solve Bezout's theorem. That is, we use Newton's method to follow a path in the space of polynomials. All these are homogeneized. It is more elegant to think entirely in terms of homoge neous coordinates and projections Vw
= {/ : C» -» C")
and H(i) = {/ : C n + 1 -> C n homogeneous of degree
d},
with Thus we are going to work in projective space, and the main thing we will analize is a projective version of Newton's method which is due to M. Shub [11]. The previous work
1356
67
has been done mostly in one variable on this problem, using Newton's method to solve this; I spent many years doing that. Jim Renegar [10] has some extension to n variables. What we want to do here is give a very conceptual process. All the ideas are lying there; we can try to understand the best, most elegant ways of looking at algorithms to solve this. In this way perhaps eventually we will be able to see more clearly the problem "P ^ NP ?" over the real numbers or the complexes. I hope it will eventually shed some light on the practical problem UP ^ NP ?" Given the space H.(i) we consider a function /o for which we know the zeroes. The zeroes of f0 could be given by a set of intersections in a grid so we get a set of equally spaced zeroes
This could be the initial element in Ti,(d) for which we know the answer, and we simply homotope that back /i=i/ + (l-0/o, where t goes from 0 to 1. Let us denote by T the curve /, in the space H^y We try to trace the zeroes, which we know for / 0 , to give the answer for / j . This seems to be a very good method that has been used for the last decade or two. It embodies some kind of global Newton's method. The idea is to consider some sequence t, and apply Newton's method to the function / t , + 1 , starting from some approximation Xi of the solutions for fti
N^JXi)
= Xi+1
where, here, Nf(X) stands for applying Newton's method to solve / starting with the initial guess X. In the projective space P „(R) we can see the paths given by the solutions A', of /, and the algorithm provides a sequence of points X, following this path very closely. So, what kind of results can we expect here? What kind of things may we prove? Here is the main theorem. The question is how many iterative steps are necessary, how many ti, in such a way that we can follow this path very closely, and our result is THEOREM.
The number of steps is bounded above by
where a is a universal constant (given by a set of equations which can be solved itself by Newton's method) which is approximately 1/16, D is max{di,..., dn] and p is the distance from the arc joining f0 and fx to the discriminant variety.
1357 68
It should be recalled that the discriminant variety is the subset of "H(d) of all singular polynomial systems. It is the variety of polynomial systems which are degenerate at some zero. And this is an algebraic variety that we shall call E. The theorem says that what is crucial are not the coefficients of / . They do not even enter. In fact, not even the dimension comes directly here; this is even dimension-free. But what is crucial here is the distance p between T and the discriminant variety E. This is the crucial factor —the only factor— in estimating the complexity for finding the zeroes of a polynomial system. Now we have to make a little caveat here because we are not finding the zeroes of every polynomial system. There may be a continuum of zeroes and then we cannot do this. So we have to put some kind of condition, let us say to solve / + e where e is a small polynomial. This is the thing we solve. We cannot find the solution of arbitrary polynomial systems; there may be a continuum of solutions, but for some deformation we can find the zeroes in a very exact sense. Now, the great problem to me is: To what extent is the term D3^ necessary? While we have no proof of this, we suspect that the D itself could be eliminated from the formula. For a polynomial in one variable, this is the fundamental theorem of algebra, and we show that we can take off the 3/2 to get D. This is what we have done in the last months; the proof is in handwritten form. Since last week we believe that we can eliminate the D in the one variable case, but this uses the theory of Schlicht functions, which is only available for one variable. There is no theory of Bieberbach conjecture for more than one variable. If it is true, if it is D-hee, if we can do this, then one can find for example one zero of a polynomial in one variable in a universal number of steps, say one hundred.
References [1] 0. Aberth, Precise Numerical Analysis, Brown Publishers, Dubuque, Iowa, 1988. [2] L. Blum, M. Shub and S. Smale, On a theory of computation and complexity over the real numbers: ^^-completeness, recursive functions and universal machines, Bull. Amer. Math. Soc. (N.S.) 21 (1989), no. 1, 1-46. [3] S. A. Cook, The complexity of theorem-proving procedures, Proceedings 3rd ACM STOC (1983), 80-86. [4] F. Cucker, The arithmetical hierarchy over the reals, to appear in J. Logic Comput. [5] H. Friedman and K. Ko, Computational complexity of real functions, Theorei. Comput. Sci. 20 (1986), 323-352. [6] D. Fuchs, Cohomologies of the braid group mod 2, Functional Anal. Appl. 4 (1970), 143-151. [7] R. Karp, Reducibility among combinatorial problems, in Complexity of Computer Computations, R Miller and J. Thatcher (eds), Plenum Press, New York, 1972, 85-104. [8] M. B. Pour-El and I. Richards, Computability and noncomputability in classical analysis, Trans. Amer. Math. Soc. 275 (1983), 539-560. [9] M. Pour-El, Review of [2], to appear in J. Symbolic Logic.
1358 69 [10] J. Renegar, On the efficiency of Newton's method in approximating all the zeroes of a system of complex polynomials, Math. Oper. Res. 12 (1987), 121-148. [11] M. Shub, Some remarks on Bezout's theorem and complexity theory, to appear in Proceedings of the Smaleftst, M. Hirsch, i. Marsden and M. Shub (eds.). [12] S. Smale, On the topology of algorithms I, /. Complexity 3 (1987), 81-89. [13] S. Smale, Some remarks on the foundations of numerical analysis, SIAM Rev. 32 (1990), no. 2, 211-220. [14] V. Vasiliev, Cohomology of the braid group and the complexity of algorithms, to appear in Pro ceedings of the Smalefest, M. Hirsch, J. Marsden and M. Shub (eds.).
Stephen Smale Mathematics Department University of California Berkeley, California 94720 USA
TVanscribed from the videotape of the talk by Felipe Cucker, Francesc Rossello and Alvaro Vinacua; revised by the author.
1359 JOURNAL OF THE AMERICAN MATHEMATICAL SOCIETY Volume 6. Number 2, April 1993
COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS MICHAEL SHUB AND STEVE SMALE
TABLE OF CONTENTS
Chapter I: The Main Result and Structure of the Proof 1-1: Introduction 1-2: Complexity of path following in Banach spaces 1-3: Complexity for polynomial systems in terms of the condition number ft 1-4: Complexity in terms of the distance to the discriminant variety I Chapter II: The Abstract Theory II-1 Point estimates II-2 The domination theorem II-3 Robustness Chapter III: Reduction to the Analysis of the Condition Number III-1: The higher derivative estimate III-2: Projective Newton method Chapter IV: Characterizing the Condition Number IV-1: The projective case, fi = l/p IV-2: Bounds on zeros and the affine case CHAPTER I: THE MAIN RESULT AND THE STRUCTURE OF ITS PROOF 1-1. INTRODUCTION
This paper represents a step in the general program of establishing principles for solving nonlinear systems of equations efficiently. Let W,d) be the vector space of all homogeneous polynomial systems / : n+l _> c" of degree d = (d{, ... , dn) (so that degree f( = d,). c Received by the editors November 4, 1991 and in revised form June 8, 1992. 1991 Mathematics Subject Classification. Primary 65H10. Both authors received partial support from the National Science Foundation. Shub was visiting the International Computer Science Institute and the Mathematics Department of the University of California, Berkeley, during much of this research. ©199} American Mathematical Society 089*0347/93 JI.00 + J.2S per par
1360 460
MICHAEL SHUB AND STEVE SMALE
Consider the computational problem, given / e W(d), solve /(£) = 0. What does this mean? A reasonable answer is: exhibit x e C"+1 such that x is an approximate zero of / restricted to Nx, f/Nx : Nx -> C" where Nx = x + {y € C+l/[y, x) = 0} and ( , ) is the Hermitian inner product on C"+1 . See (*) of Theorem 1 of the next section for the precise definition of approxi mate zero. In particular, Newton's method for f/Nx , starting at x, converges quadratically to some C € Nx with f(Q = 0, and e relative accuracy is ob tained with log | log £ | further steps. Nx is the Hermitian orthogonal complement to the vector x e C" +l through x, and can be thought of as the tangent space to complex projective n-space. In order to apply Newton's method, f is restricted to an ^-dimensional subspace. The choice of Nx is natural from the projective space point of view and optimizes some of our estimates. That x is an approximate zero of f/Nx is invariant under scaling; i.e., kx is an approximate zero of f/N^ , X ^ 0, if x is an approximate zero of f/Nx . Thus we say x is an approximate zero of / in the projective sense with associated actual zero £. A curve F : [0, 1] - Z[d) x C n+1 satisfying /,(£,) = 0, F, = (/,, £,) is called a homotopy-path. An important computational problem is to produce from the input (ft) and an approximation z0 of C0. a sequence zjt i = I, ... , k, which fits (Ct) in the following sense. Each z( is an approximate zero of f( in the projective sense with associated actual zero C, , 0 = f0 < • • < /,._, < tt < ■ • • < tk = 1. Projective Newton's method proposed by Shub [24] yields z. from z(._, by applying Newton's method to / fN . The problem we deal with here is how i
l-1
small can k be taken to obtain z,, ... , zk fitting (£,). The answer we demonstrate is that the controlling factor is the distance along P(%[d)), p{F), in the corresponding projective spaces, of the curve Ft in W(d) x C"+l to the discriminant locus l! (an irreducible algebraic variety of ill-posed problems). The other factors are the length L of the curve (ft), a modest contribution from the degree of / and a small constant. More precisely: Main Theorem. Let Ft = (ft, C,) be a homotopy-path in 2f,d) x C" + l , and let z0 satisfy
Let
l*,-M
= 035
C
'
C LDi/2 ? , , C2 = 8.35.. 2 P(F)2 then I projective Nevton steps are enough to produce z , , . . . , zt fitting (£,). Here D = max(rf(). />
The number of variables n does not enter directly into the complexity /, but f (F)
1361 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
461
The problem of finding the starting point can be dealt with by choosing a universal f(d) e %^d), or from aspects of the particular problem. The invariant p(F) needs to be studied from a geometric probability point of view. Part II is devoted to these problems. For the problem offittingall the solution curves of given f(, we find similar results, using the distance in 2f(d) to the discriminant locus Z. This will lead to the complexity of algorithms forfindingall the zeros of a given system / € &!d.. Moreover, we provide a similar complexity analysis for the (nonhomogeneous) general polynomial system / : C" -» C" using a more traditional (yet global) form of Newton's method. One novel feature in our development is unitary invariance. For example, if U is a unitary transformation then x is an approximate zero of / in the projective sense iff U(x) is an approximate zero of f o U~ in the projective sense. In general, one would not expect the particular coordinate representation of the polynomial system to be reflected in the basic invariants of the theory. The distances of solutions to each other seem basic. Our way of dealing with this is by using a fully unitarily invariant theory. This has an added feature of forcing a more elegant development of the mathematics. Our proof of the main theorem puts into a very general setting theorems of Eckart-Young, Houth and Demmel on the condition number and the reciprocal of the distance to the algebraic variety of ill-posed problems. The most important work on this problem previously is that of Renegar [20]. That paper was very helpful to us. Very roughly, work on algorithms for Bezout's problem can be divided into two distinct schools. One is algebraic, represented for example by Brownawell [2], Grigoriev [9], Heintz [10], Canny [3], Renegar [21], and Ierardi [12], and a second more numerical analysis approach represented here. The algebraic algorithms tend to be less numerically stable (see, e.g., Morgan [18]). The convergence and practice of path-following algorithms (or homotopy methods) may be seen for example in Allgower-Georg [1], Garcia-Zangwill [7], Hirsch-Smale [11], Keller [13], Li-Sauer-Yorke [16], Morgan [17], Wright [34], and Zulehner [35]. One variable complexity results on these algorithms is in Shub-Smale [25, 26] and Smale [27, 28]. The proof of the Main Result is quite long. The rest of Chapter I is devoted to giving the structure of this proof by displaying some intermediate theorems. In fact, in these theorems there is an effort to isolate some main concepts. We would like to thank Matt Grayson for his calculations of some of our constants. 1-2. COMPLEXITY OF PATH FOLLOWING IN BANACH SPACES
Here we state some general results on complexity which are valid in a wide setting, yet form the framework of the main proofs on the complexity of Be zout's Problem. These ideas revolve about an invariant a(f, x) proposed in Smale [29, 30]. Subsequent work of Royden [23], Wang-Han [32] sharp ened and broadened these first results, and that work is incorporated into the present treatment, Theorems 1 and 2 below. Moreover, robustness results of the
1362 462
MICHAEL SHUB AND STEVE SMALE
Q-theory, only suggested in Smale [30] and in Renegar-Shub [22], are formu lated in Theorems 3 and 4. Besides applications to polynomial systems, these theorems may be used to analyze complexity of linear programming algorithms. It turns out profitable to give a complete demonstration of both the old and new results of this a-theory together. Thus the theorems stated in this section are proved in Chapter II. For some motivation and broader perspective one can see Smale [30] as well as the other references. Throughout this section and Chapter II, E and F denote Banach spaces. In the main applications they are both C m , m-dimensional complex Carte sian space (or subspaces) with a norm defined by the standard Hermitian inner product. We consider analytic maps / : Dr(x0) -» F where x0 e E and Dr(x0) = {x € E| ||x - x0|| < r}. For x e Dr(x0) let Df(x): E -» F denote the derivative of / at x (see Lang [ 15] for our way of doing calculus). If Df(x) is not an isomorphism all the following a, P, y are oo (or not defined). Otherwise define for / : Dr(xQ) -» F and x e Dr(x0) \\Df(x)-if(x)\\,
fi(f,x)
=
y(f,x)
= sup
a(f,x)
=
Df(xylDkf(x)
k>\
^
*!
0(f,x)y(f,x).
Newton's method (when defined) constructs a sequence of points xl, x2, in Dr{x0) by the formula xn=xn_l-Df(xn_l)-if(xn_i),
n = \,...
.
We also write Nf(x) = x-Df(x)-lf(x)
and xH = Nf(x„_x).
Thus /*(/,*„) = \\xn+l-xn\\. We also frequently use these quantities: (l+a)-x/(l+«)2-8a T(a)=1± 1
t
_ for 0 < a < 3 - 2 \ / 2 ~ . 1 7 1 5
and a 0 = i(13-3vT7)~.157671. Theorem 1. Let f : Dr(x0) - F be analytic, 0 = 0(f, x0), y - y(f, x0), « = fiy and suppose r > ^ . Then if a
<•>
K+.-^H
2"-l
11*1 " *oll
1363 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
Moreover ||C - x0|| < T-f, and ||C - JC, || < ^
463
.
A point xQ e E is called an approximate zero (of f) if (*) is satisfied. In that case C is called the associated (true) zero. The following is an easy consequence of Theorem 1. Corollary. Let f : E -* ¥ be analytic, x0 e E satisfy a - a(f, x0) < a0 and have associated zero C • Then the nth Newton iterate zn of z0 is within e of C provided n > (log | log $&\) + 1. Remark. The \ in Theorem 1 may be replaced by any A, 0 < X < 1, with aQ redefined. See §11-1. Consider as an example the following family of real-valued functions (which have a universal quality as we will see):
Let a = 0y satisfy (a + l) 2 - 8a > 0 or equivalently 0 < Q < 3 - 2\/2. Then hg has two distinct real positive roots at (a+l)±y/(a+l)2-Sa 4y
T(Q)
y 2
2
Moreover d hB Y(t)/dt > 0 as long as 0 < t < £ . Thus Newton's method starting at 0 converges to the smaller root since by convexity the Newton se quence is monotone. Let tn = Nh (/„_,) where tQ = 0. Theorem 2 (Domination Theorem). Let f : Dr(x0) -> ¥ be analytic, fi = /?(/. x0), y = y(f, x0), a - fiy and suppose r > ^ and a < « 0 . These values of p ,y define hfi , and the sequence tn . Then K-^-ill<'„-V>.
« = i.2,...,
where xn is the Newton sequence of f starting at xQ. The Domination Theorem yields the last sentence of Theorem 1 as follows.
K-Cli <£lK + 1 -xj|
n=0 oo
x
' ,
\\xx - en < E K + . - J ^ E ' - . -'« = ^
.
- fi ■ D
We next deal with the question, "How does
1364 464
MICHAEL SHUB AND STEVE SMALE
Proposition 1. Let f: Dr(x0) -* F be analytic. For x e Dr(xQ), we have
«(/.*) <" (/ '*°¥(»))( V ) + " where u = y(f, x0)\\x0 - x\\ and u < 1 - ^ . Proposition 1 plays a role in the proof of the following result which permits repeated applications of Newton's method. Theorem 3. There are universal constants a about .08019667,
a about .02207
with this property. Let f : Dr(Q -» F be analytic with y = y{f, C) < ? (some constant), 0 = fi(f, (,) < 'j, and r > *f. Suppose x € Dr(Q satisfies \\x - Cll £ f and C, is the associated zero of C • Then ||x, - f, || < | , where x,=Nf(x). Theorem 3 can be readily seen to have global implications. A homotopy ft : E -» F is a family of analytic maps 0 < t < 1 with the induced map [0, 1] x E —> F continuous. An associated path is a continuous map [0, 1]-»E, / -♦ C,. satisfying for each / e [0, 1], (a) yj(C,) = 0 and (b) the derivative Df(C,) : E -» F is an isomorphism. Sometimes we call {/,, (,,} a homotopypath. The central algorithm in this paper (as used in Smale [28]) is designed to follow a path associated to a homotopy and works this way. To a subdivision T = Co = 0, / , , . . . , tk = 1}, tt.,< tM , \T\ = k , x0 with |JCO - C0II < &. define inductively by Newton's method (*)
x ( = A^(jf,_,),
i=l,...,k.
If \tj - tj-t\ and ||x0 - C0|| are small enough, then ||x( - C, II is small for all i = 1 , . . . , k . More precisely we will say that Newton's method follows the homotopy-path {ft, C,}, relative to T and 8 provided x( of (*) is well denned, a(ft , x() < a 0 , and C, is the associated actual zero to the approximate zero x of / , ii = 1 , . . . , \T\. i
Note that in this case the number of Newton steps to reach an ^approxima tion of the zero d of /, is given by |r| + l o g l o g ^ ey where Q = Q(/, , x{), y = y(/,, x,). The central complexity measure is |7"|. The most important example of a homotopy is a linear homotopy: /, = tfx + (1 - t)f0, f0, /, : E -♦ F. Often in this case a zero (or all the zeros) of f0 is known. Theorem 4. Let F = {/f, C,} be a homotopy with an associated path as above. Let A = {, k an integer, and y > 0 be such that £(/,., C,) < f and y(f,>, C,) < ? if\t'-t\
1365 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
465
Let \\z0 - t0|| < | . Then if T = {0, A, 2A, ... , kA = 1}, Newton's method follows F. In fact
The proof of Theorem 4 from Theorem 3 is practically apparent. One uses only the appropriate continuity of the associated zero which comes from The orem 1. Theorem 4 indicates the importance of estimating p(ff., £,), y(ft., £,) and much of the rest of this paper is devoted to just that.
1-3. COMPLEXITY FOR POLYNOMIAL SYSTEMS IN TERMS OF THE CONDITION NUMBER, H
We start this section with some general background on polynomial systems. It turns out that the affine (usual) and projective developments shed light on each other, and in fact we treat, in part, the affine problem in a homogeneous context. In both cases a representation of the unitary group plays an important role in our study. Subsequently, the condition number ft = n(f, x) for / : C" -» C", x e C" , is introduced. This is a modified version of the simple ||£>/(jr)~'||. Our condition number must assume a more technical definition for several reasons, mainly related to natural scalings. The algorithms to follow a path of a homotopy are modeled on that of Theo rem 4 (previous section) and we are able to estimate the appropriate a invari ants in terms of the degree and the condition number of the homotopy. The passage from Theorem 3 to Theorem 4 in the previous section gives the under lying idea of how we then obtain complexity results. As usual, this section is on the overall structure with full proofs in Chapter III. We turn to describing spaces of polynomial systems together with a unitarily invariant metric. This metric while natural and used in the theory of group representations (Stein-Weiss [31]) is not traditional in the numerical analysis literature of equations. It has been suggested by Kostlan [14] and seems to be well suited to purposes of complexity, and corresponding estimates appear to be more elegant. Unitary invariance plays a central role in our approach to complexity. We use &.d. to denote the linear space of all polynomial systems f :Cn — C" , f = (f{,... , f„), each f. a polynomial of n-variables of degree < di, and rf = (,,...,„),rf,.> 0 . Let ftj-. be the homogeneous counterpart. That is, / e 2t?.d. is a map n+1 C -» c" of the form / = ( / , , . . . , / „ ) where each ft is a homogeneous polynomial of degree exactly di. We suppose 0 e X"(d^ so that 3T^ is a linear space. Note that there is a natural linear isomorphism : &.., -» &.. given by
1366 466
MICHAEL SHUB AND STEVE SMALE
homogenization as follows. Let / = ( / , , . . , / „ ) € ^d),
fi(2l,...,2n)=
J2
so that
^
M
where za = z? •••*"« and |«| = £ a , . Then
v=i
For homogeneous polynomials g, f: C n+I -♦ C of degree d, let V
\a\=d a
a
where f(z) = ^aaz , g(z) - ^2baz . uct on J^d.. Simply write, for f,g€
"•
'
This induces a Hermitian inner prod ^d),
(f,g) = J2^i'8ih,i
Proposition 1 (Kostlan [14]). Let the unitary group act on C n+I in the canonical way and on Sff.d. by the induced representation. Then ( ) on &!d) is invariant. In other words (/«"' ,gu~l) = (f,g) for all f,ge X[d), and u : C n+1 — C"+1 unitary. Of course ||/w"'|| = ||/||. By the isomorphism : &d) -»ff.d) we obtain an induced Hermitian struc ture on l?,d). We will denote the corresponding norm on / by simply ||/|| for each &.ds and &d). We sometimes use the same symbol for / e &d) and If E is a linear space over C, let P(E) denote the corresponding projective space of lines through the origin in E. So P(E) = (E - 0)/C*, C* = C - 0. A Hermitian structure on E induces a Riemannian structure on P(E) (of con stant curvature). Thus we have canonical metrics on P{^,d)), P(^j)) • Some times for / e Jf(ds, the same symbol will denote the corresponding element of We define, for
xeC+i, NUH(JC) = { I ; € C " + , , ( I ; , J C ) = 0}
and an affine subspace Nx = x + Null(x)cC n + l . Let e0 = ( l , 0 , . . . , 0 ) e C n + 1 . We will use the notation A(j>() to mean the diagonal matrix whose ith ele ment is y{. We are ready to define the condition number fi(f, x) of / e &d) at jc e C" . The idea is to take fx to be ||D/(x)"'||, e.g., as in Wilkinson [33], but it
1367 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
467
is important to take into account the special polynomial nature of / . We also want to make n compatible with homogenization. Moreover, for sharper estimates on complexity it is convenient to have a further factor of d} . Thus define M,x) = \\f\\\\Df{x)-xUidl)hx\ffl)\\ or 1, whichever is larger, and where ||x||, = ( £ " x2t + \y . If / is treated as an element of &.d. this is equivalent to /*(/. x) = H/ll \\Df\N ( x r V / W ' " ' ) or 1 '0
where x = (x0, ... , xn), x0=l. We will use the two versions interchangeably. The projective version of the condition number is: For / e %^d) and x e
c +l ,
W whichever is larger. For feX[d), let
*> = ll'll \\Df\Nx(x)-1A(d}\\x\\d-l)\\
or 1
/"iiivir^' ^ iX) .!*(£wprt«!.
For fe&>y>, ri(f,x) is the same except ||jr|| is replaced by ||x||,. Let / e &d), x eC" . Appropriate versions of /?(/, x), y(f, x) are Rlf o(f x)
r\-£±LA
* ' -^xjr' y0(f,x) = y(f,x)\\x\\r Thus a(f,x) = fi0(f,x)y0(f,x). The projective case goes as follows. For / 6 ^d),
x e Cn+ ,
P0(f\NM,x)
=
fi{f\Nx,x)/\\x\\,
?0(f\sx,x)
=
y(f\N\x)\\x\\.
These definitions are invariant under scalings of / and x, so they make sense on P(JTd) and / > (C n+1 ). Proposition 2. For f e &>{d), xeC
,
P0(f,x)
xeCn+l, ^{f\Nx,x)<^ny{f,x)ri{f,x).
The proof of Proposition 2 is obtained by putting together the definitions with \\Ab\\ < | | 4 | \\b\\. In Chapter III we will show
1368 468
MICHAEL SHUB AND STEVE SMALE
Proposition 3.
( /
V0(f\N,, x) < ^
'
2
X)D
^ ,
/ € ^ , jrec"'.
For / , # € ^ , let dp(f, g) = minAeC ^ ^ . dp(f, g) is independent of scaling of both / and g and defines a distance function on P(.%[ds) ■ To see that this function is in fact a metric we compare it to the standard metric on P{%[d)). There is up to scaling a unique unitarily invariant Hermitian metric on P{%',d)). One way to get the existence of one is simply to restrict the Hermitian structure on X".d. to Null(/) at each / . This Hermitian structure defines a unique Riemannian structure and a distance dR(f, g) for / , g e P{%[d)) • Proposition 4. dp{f, g) = sindR(f, g) for f,g€
P{X[d)) ■
Proof. dPKJ P{f, S>#) = min
'
*C
ll/H
11/11
V
||/|| 2 ||S|| 2 /
by expanding the norm in the numerator as a Hermitian product. Now |2 \ 1/2
V
2
2
imi 2 IISIIV ii*ii / II/II
smarccos'^'^'
mm
so we have only to see that dR(f, g) = arccos l ^ j ^ . To see this last, use unitary invanance and the uniqueness of unitary structure on C up to iso morphism. We may assume that dR and arccos Wji'iM are defined on CP( 1) corresponding to C spanned by / and g. Moreover in affine coordinates we may assume / = (1, 0), g = (1, x0) where the metric (see Mumford [19]) is ,2_
dxdx l+\x\2
(xdx)(xdx) (l + \x\2)2
_
dxdx (1 + W 2 ) 2 '
Now integrate on the path (1, tx0) for 0 < t < 1. / ^ — ? dt - arctan |jcn| = arccos ( r-r^ ) 2 2 ol 'o l + / | x 0 | V(1 + |* 0 I 2 ) I / 2 / which verifies the formula in this case. Note that to see that dp is a metric it is enough to note that s\n(A + B) < sin A + sinB for 0
1369 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
469
Proposition 5. Let f, g 6 &d), C € C + 1 . Then
w
"^■^,->vf(/.,);TO,/.ofor
c#o
ay /o/i£ ay /«£ denominators are positive. Remark. Di/2 in (b) may be omitted by a different proof. For a while now we restrict ourselves to the affine case. So in §1-2, E and F both become C" and a homotopy ft: C" -» C" is a (continuous) curve in &d. Then £, is a curve in C" with f^Q = 0 and Dft(C,): C" — C" nonsingular. Theorem 1. Suppose C0 e [0, 1] ana1 {/,,£,} « a homotopy-path in ^>{d) x C", / , / ' € [ 0 , 1], fi = fi(f,,Ct) with
Let
Di/2
/ 1 + C0\
£(/,<, C,)?
is about .11629. We give the short proof here (assuming the previous propositions). Choose in Propositions 2 and 3, f = f,>, and x = C, ■ Then
fiif,-, C,) < v(fe, C,)»f(^. £,)• Lemma 1 (easy).
»(/,<, £,)<<W,<>/,) = */>• By Proposition 5 ^ " C ^ ( l - D < , ) so we obtain using the lemma,
«*■<■> W d s ^ ) -
1370 470
MICHAEL SHUB AND STEVE SMALE
Lemma 2. 1+A, ^ 1 + C0 l l-D '%p~ l-C0Proof. Using the hypothesis on A^ and the fact that D > 2, p. > 1, one sees that pDi/2&p < C0 . Then A^> < C0 and the lemma is proved. The estimate on /? and y of Theorem 1 now follows by making the appro priate substitutions. We now will give an estimate of the number of steps of the algorithm of §1-2 described right after Theorem 3. For a homotopy-path F = {ft, C,} in &(d) x C" , define L = L(F) to be the length in the metric dp of the curve ft, 0 < / < 1. Define the condition number /* = //(F)=max/i(/,,C ( ). Note that n takes different meanings in different contexts. Theorem 2. Let F = {ft,C,}
be a homotopy-path in &>(d) x C" . Let Ir .
LZ>3/2
2
k> — t - p . . a Then k Newton steps are sufficient to follow the path C,. [0 < t < 1 ], in the sense of §1-2. Again we give the short proof which follows from Theorem 1, but first note: Remark. For the main case of a linear homotopy ft = tf + (1 - t)f0, recall that L is less than or equal to the diameter of projective space, which is 1. Proof. One may choose t( so that
dP(f,rfti,_,)<§.
i = l,...,*.
Then since £ < a'/D ' p., the hypotheses of Theorem 1 are satisfied where C0 = a. The conclusions of Theorem 1 put us into the situation of Theorem 3 of §1-2 and this finishes the proof of Theorem 2. The next theorem is an important step in analyzing the projective version of Newton's method. Theorem 3. There exist numbers a ~ .07364, u ■ ~ .0203... with the i/2 following. Suppose f e W(d), y > D and x € A^ satisfies (a) f/(/,C)^ p r o j .(/,C) . (c) y0(f,O
IIC'II
<
"proj.
1371 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
471
where x = N^N (x), x e N(. and k£ is the associated zero of x for f\N , some k e C. Remark 1. If we take J> = £>3/2/z/2, n = ^ proj ( / , f) then (c) is automatically satisfied by Proposition 3. Remark 2. That / e ^d) is not crucial to the proof. ft homogeneous complex analytic of degree dt with large enough radius of convergence around £ suffices with J> > max(l, Di/2), D - max;=1 n \d^ . Note also that the expression rjfi of (a) does not involve ||/|| and is denned for all such / . A homotopy ft : C n+ -» C" in the homogeneous case is a curve in the space %[d). 0 < t < 1. An associated path C, is a curve in C n+I satisfying f,(C,) = 0. Let f ^ = ^Proj.(f) = m a x V>j.(/< • f«) • = U > CJ, be the condition number of the homotopy-path F. We will now prove the main theorem, §1, with p(F) replaced by -4?y. In the next section, the main theorem asserts these quantities are equal and so that result will finish the proof of the main theorem. Let C, be the first positive root of
for D = 2, (i = 1 and let C2 = ^-. Then C2 = 8.35 . . . . Let 1 . /i(\+A)DV2 A = C2DlTTr—r '2n2 a n d J>' = 2(l-Di'2A»)' Choose tt, i = 0 , . . . , k = [£], such that for s € [/., r,+ l ] , dF(fs,f,) < A. Here [x] is the smallest integer > x. Let n' = sup ie( , , , n{fs, C,) • Then * \~D"2An by Proposition 5 and for s € [/,, tj+l], y,f
C
)< Ml+A)^2
=
by Proposition 3. This gives condition (c) of Theorem 3. We now check con ditions (a) and (b). For (a), note that tj(fs, C,) < A for s e [/,., / (+l ] by Lemma 1. Also AD,/ < 1 H(l+A)DV2 /i(l+A) VM il2 2 ~ C2D n 2(1 - DxllAn) (1 - D " 2 ^ ) _C, (l+C,/D3/V)2 "
2(1-CJDrf
-
apr J
°-
1372 472
MICHAEL SHUB AND STEVE SMALE
by the definition of C,. Thus rj < aproj /yfi. For (b), by hypothesis,
11*0 - y < r p(F)
- ' tf'1 - ? '
uyi
Apply Theorem 3 inductively to obtain
K-y < «
proj-
IK,, II
for all /'. Finally from Proposition 2 of §111-2
a <
K ( C
W W *W <.Q24, ^(«proj.)2
the last using a pocket calculator. Here Q, stands for a of /. at x, . Thus x. is an approximate zero of / and so certainly x. is. Theorem 1 of §1-2 applies to yield log log estimates, i.e.
where C,
here means the associated root of x, in Nv . a 1-4. COMPLEXITY IN TERMS OF THE DISTANCE TO THE DISCRIMINANT VARIETY I
The goal of this section is to replace the condition number in the estimates of the previous section by the distance to the discriminant variety. To that end, we extend an idea going back to Eckart and Young [6] and developed especially by Demmel [4]. Consider the product space %{d) x C n+I with quotient ^{d) x Pn where Pn = / > (C n+l ) is n-dimensional complex projective space. Let
v = {(f,z)ejr(d)xpn\f(z) = 0}. We may consider the projection of V onto the second factor V -» Pn as a vector space bundle with fiber over z given by Vz = {f e &,d) \ f(z) = 0}. The associated bundle n2 : V -* Pn with fiber P(VZ) is a smooth algebraic hypersurface V c P(%[d)) x Pn (see Shub [24]). Let l! be the algebraic hypersurface in V given as the set of ( / , <£) € V such that Df(Z): Cn+1 -> C" is singular (i.e., of rank less than n). While we have considered V as a bundle n2: V -» Pn we may also consider the projection n : V - . P{^!d)) on the first factor.
1373 473
COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
Remark. Z' may alternately be described as the set of critical points of n . That is, ( / , x) 6 V is in Z' precisely when Dn{f, x): T(f X)(V) -» Tf{P{X[d))) is singular. The image under n, TT(Z') = Z c P(%[d)). is the algebraic hypersurface of polynomial systems with a degenerate zero (i.e. / e Z if and only if there is some C € Pn with /(C) = 0 and Df(Q) singular). This variety Z c P(^!j\) is called the discriminant variety. It is familiar in the affine one variable case as the set of all polynomials with nonvanishing leading coefficient having a multiple root. The map n : V -> P(&d)) is an n-dimensional generalization (homogenized) of the well-known map taking roots of a polynomial onto the coefficients by the symmetric functions (in one variable; sometimes the "Vieta map"). The variety Z has played an important role in recent complexity analysis of polynomial zero finding since it consists of "ill-posed problems". For one variable Newton method see especially Smale [27, 28], Shub-Smale [25, 26]. For the many variable case see Renegar [20]. On the algebraic side, a similar situation prevails; see Canny [3], Heintz [10], Ierardi [12], and Renegar [21]. In both cases, however, there is a subvariety of the discriminant locus of a more seriously ill-posed system which contains an infinite number of zeros. An underlying theme in much of this literature is the idea that the condition number is bounded by the reciprocal of the distance to Z. This theme also comes from numerical analysis (see for example Demmel [4, 5]) even more explicitly. We sharpen and develop that idea here with Theorem 1 below. Our account continues with one version of a result seen in undergraduate numerical analysis texts. Let \\A\\F be the Frobenius norm of a matrix A € Jt(«), the set of all n x n matrices. Thus
Mllf = d K / ) l / 2 . Let S c Jf (n) be the subset of singular matrices and let dF{A, 5) be the distance from A to S in the Frobenius norm. Proposition 1 (Eckart and Young [6]).
The proof is in Golub and Van Loan [8]. Here ||/4 -1 1| refers (as always here) to the usual operator norm induced from the Hermitian structure on C" . Next we define a function p on V which represents the distance to the discriminant variety. For ( / , x) € V, take p(f, x) as the distance in the fiber Vx of it : V — Pn over x of ( / , x) to l ' n K x . Recall that this fiber is the projectified subspace {/ e ^(d) \ f(x) = 0} of ^{d), and the distance is computed in the dp metric. Thus p is ultimately defined by our unitarily invariant norm on M^d). Theorem 1. Let f e *£,,, x e C + l , f(x) = 0. Then ^(f,
x) = ^
.
1374 474
MICHAEL SHUB AND STEVE SMALE
On one hand Proposition 1 is used to prove Theorem 1; on the other hand it is a special case of Theorem 1. In the case of one variable polynomials, there is the work of Hough and Demmel [4] giving upper and lower bounds for the condition number of / at x in terms of a version of our p(f, x)~ . In the passage from Proposition 1 to Theorem 1 we use heavily unitary invariance. Unitary invariance has already played a role in the proof of Proposition 3 of 1-3 and continues to do so throughout many of our proofs. In more detail the unitary group U(n + 1) acts on &,d) x C n+I by sending ( / , z) to (/u~ l , UZ) for u G U(n+l). This action induces actions on ^ x i ^ leaving V invariant and on P(^d)) x P„ leaving V invariant as well as Z' C V invariant. As a corollary to Theorem 1 we have immediately that, in the main complex ity results of §1-3, we may replace nVHi (f, x) by -^j-^ . Next we give a result which corresponds to Theorem 1 with the x eliminated. Quite simply for f e Jf(d), let »proi(f)=
m ax x
*W(/,x),
/(JC)-0
p(f) = min
p(f,x).
Then /iproj (/) may be thought of as the condition number of / . Corollary. Let f e Sff(d). Then
/W(/)
=
~pjf}-
Remark. It is easily seen that p(f) > dp(f, Z) so that //(/) < l/dp{f, I ) . We now proceed to an analysis of the condition number in the affine case. The situation here is more complicated. Define l'0 = V n Z^ where Zoo = {(f,z)€P(2'{d))xPn\z0 = 0}. Thus ( / , z) € IQ means that / has z as a zero at oo. Let Z0 = niX0), n : V -► P(^d)) ■ We may consider 1^ as contained in ^d) via O - 1 : ^d) -» &>(d) (abusing notation). This way Z0 consists of all polynomial systems / : C" -» C" with the property that the highest order homogeneous parts of ft have a common nontrivial zero. It was observed in Hirsch-Smale [11] that if / £ Z 0 , then / is proper. From the construction Z'0 and hence Z0 are varieties and in fact irreducible (compare Shub [24]) hypersurfaces in V , P(^d) respectively. The following proposition gives a bound on the zeros of a polynomial map / : C" - C" . Proposition 2. Let f e &(d), xtC"
and f(x) = 0. Then
\M\ - ( 1 i r i , ^ ' / 2 ^ ' A K ' / V i i ^ ' / 2 i i / i i
1375 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
Let pQ(f) = d(f, Z 0 )/||/||. Thus (1 + n l*,.|2)1/2 < Dl/2/p0(f)
475
if f{x) =
0. The following gives us the affine condition numbers in terms of the projective ones. Theorem 2. Let f e &>(d), <J e C" . Then M/,«r)
x£ C" with f(x) = 0. Then DM2
H(f, x) <
PQ(f)p(f> * ) ' D1'2 Po(f)Pif)' The corollary uses both Theorems 1, 2 and the proposition. This yields another version of Theorems 1 and 2 of the previous section. We summarize this section by reviewing what must be proved in Chapter IV. These results are Proposition 1, Theorem 1, Proposition 2 and Theorem 2. CHAPTER II: THE ABSTRACT THEORY II-1. POINT ESTIMATES
We prove the results stated in §1-2. The proof of thefirstpart of Theorem 1 of §1-2 is given. We proceed directly with a general result which is rather technical sounding. It is used in proving all of the theorems of §1-2. Use the basic notation of §1-2 and besides let i//{c, «)= 1 - 2 ( c + \)u + (c+ \)u2. Proposition 1. Let f : Dr{z) -» F be an analytic map and z e Dr(z). 0 = fi(f, z), fi' = /?(/, z) and c,S>0 satisfy
¥>m-xffm
, = 2(3
k\ Ifu = \\z - z'\\8 and if/{c,u)>0,
Moreover, if c = -^-^,
i.e., u < v V + c/(c + l), then
S' = -fa , then
Wizr^A:')^^-^ Finally, if K = 08, K = p'd', then
, = 2>3
Let
1376 476
MICHAEL SHUB AND STEVE SMALE
Note that the K estimate is a consequence of the /?' estimate and the defi nitions of S', K , K . We write down the special case of Proposition 1 for c = 1 and 6 = y(f, z). Let y(u) = yt(\,u). Proposition 2. Let f: Dr(z) -» F be analytic, z € Dr(z) with ip(u) > 0 where u = \\z' - z\\y(f, z). Then
fi(f, *') < ^ f ytf ,f
z')
<
(0 " M W > z ) + Hz' " z») •
y(/»z)
_'x «, ( 1 - M ) a ( / , Z) + K
This proves Proposition 1 of §1-2. We now prove Proposition 1. Lemma 1. Let A, B : E -» F be bounded linear maps with A invertible such that \\A~lB -1\\ < c < 1. Then B is invertible and \\B~lA\\ < ^ . Proof. Using the series ^ = 1 + x + x2 + ■ • • for ||x|| < 1, A~lB = I (I-A~XB) is invertible. So B is invertible and ||£~'/l|| = \\{A~xB)~l\\ < ^ . The following very easy lemma is left to the reader to prove. Lemma 2.
If w(c u) > 0, then c((^L) 2 - 1) < 1. Lemma 3. With notations and hypotheses of Proposition 1 (1) Df{z) is invertible, (2) \\Df(z')< (1 - u?lv{c, u), \\Df(z'YxlDf(z)\\ < l k (3) \\Dr (z')D f{z')/k\\\
about z is
So
mm-'DfW) - ,n < E t||0/(2)",p n'nu- - xi* S^t^-'llr'-rll'-1 *-2
1377 477
COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
This bound is less than 1 by Lemma 2 and thus Lemma 1 applies to yield \\Df{z')-XDf{z)\\ < l-c((r^-l) By Lemma 2 we obtain Lemma 3 parts (1) and (2). For part (3) of Lemma 3 we use part (2) as follows. Df(zTlDkf(z') k\
<
\\DAz')-lDf(z)\\ 1=0
^ (l-uf
^ (k + l ! .*+/-!.. /
^ —,(c T/
<" ' ") fa
c(l-u)2
Df{z)-xDkMf{z),j ~k\ "
t
,,„ !/!
CO
{Z
^ ~Z)
„/
\\z - z\\
. , v ( f c + /)! ,
" V(c,u)° fa k\l\ < g(»-")2j*-' (J-)k+l - y/{c,u) \l-u)
< c ( S ) - y/(c, u) \\ -u)
k-\
«i>rW(*')ii< y^ii*'-*IIProof. \\Df{z)-xf{z')\\ <
< low1™+*' - *i+±m*rgmu._ „. But
*=2
u K* f{z \\Df{z)--lD k\
iii*'-*n*<(f;rf*-,Bz'-xB*-,)ii*'-*ii ^*-2 ^
c
u
' II '
I
To prove (b) note that in this case ||Z>/(z) ' f(z) + z - z\\ = 0. For (a) we have \\Df(z)-lf(z')\\
1378 478
MICHAEL SHUB AND STEVE SMALE
This proves Lemma 4. Proposition 1 follows, fi' = ||Z>/-(z'r7(z')ll < \\Df(z'rlDf(z)\\ ^ (1 -u)
(D
., ,
* kd) ('
..
\\Df{z)-Xf{z')\\
cu ,. ,
+ l|z z +
,,\
- " r^ - z"j l|z
by Lemma 3 and the above. Proposition 3. Under the conditions of Proposition 1, let z = zThen
Df(z)~xf{z).
fi'<-£-rK(l-K). V{C,K)
Moreover take S = y and c = 1, to obtain V <
V(a)O-a)'
2 a <
T
W(a)2 as in Smale [29]. Note that for z-Df(z)-lf(z)
z' = we have u = K and
\\Df{z)-Xf{2)\\
fi(z') =
<\\Df(z')-lDf(z)\\\\Df(z)-lf(z'[ < 5
(1-")2 cu , I|Z (c,u) ( 1 ) ¥ U
„
Z|1
CK(I-K)
-
y/(c,K)
using Proposition 1 and Lemma 4(b). Now recall that for 5 = y we may take c = 1 and K = u = a. Thus the inequality for y follows from Proposition 2, the inequality for /?' from the above both by substitution and the inequality for a by multiplying the inequalities for / and /?'. This proves Proposition 3. We next make a change of variables from c to a as follows: CK
0<<7<
(l-K)2 i
*
i
CK
7, 4
i
0-KV
We are continuing the use of notation from Proposition 1.
1379 COMPLEXITY OF BEZOUT'S THEOREM. I: GEOMETRIC ASPECTS
479
Proposition 4. Under the hypotheses of Proposition 3,
(a) *'
^'<(T^)
2
.
Proof. Observe that V , = 1 - 2(7 + OK. (1-K)2
Then (b) follows from Proposition 3 and an easy substitution, and (a) follows from (b) since K = fi'S' and d' < j ^ by Proposition 1 recalling that u = K . It remains to confirm (c). Let y/ = /{c, K) . Then / O =
CK C CK r^ < (1-K')2 ~ V V
1 (1-K')2
But (1-K)2
<
( 1 - K )
V(I-K')
2
(1-K)2
_
2
(1 - 2(c +
V-CK
1)K
+ K2)
1 -2oc/(l
-K)2
1 -
2<X
proving Proposition 4 from Proposition 3. We now suppose the hypotheses of Theorem 1 of §1-2. Proposition 5. For 0 < a < \ , let 0 < X < 1 satisfy er = A/( 1 + X)2. Then Pk<X2k-lfi0,
* = 0,1,2
For the proof of Proposition 5, we use the following lemma. Lemma 5. Let f?k = /?(/, z t ) awd a te f/ie jth iterate under a -* (rfjo) a0 = a. Then
(b)
of
(T^-)
1J
2*
1
Since ri/=o ^ = ^ > Proposition 5 is a consequence of the lemma. More over, (a) is a consequence of Proposition 4(b). Note that ,2
a
,2
2 ~ Vl-2/1/(1+A) 2A/(1+A)V/
It follows that
,2 J
**<
( 1 + A2' \)2
X2
"(1+A ■ ,2N2 '
1380 480
MICHAEL SHUB AND STEVE SMALE
But since (l-2ak)(l+/l2i)2> 1 we have l-2ak*X proving (b) of the lemma. Take c = 1. Then K = a and a ff
2
k 2
"(l-a) "(l+A) "
Choose A = j , s o a = a0 = j(13 - 3vT7) and Proposition 5 implies the first part of Theorem 1 of §1-2. Recall that the rest of Theorem 1 follows from Theorem 2. The next section is devoted to a proof of that theorem. II-2. THE DOMINATION THEOREM
For the proof of the domination theorem (Theorem 2 of §1-2), we will use some lemmas. The first concerns the monotonicity of the functions defining the inequalities in Proposition 4 of the previous section. Let v, , a)\ = -—z aK K(K , v ' 1 - 2a + OK
Lemma 1. Suppose 0 < KX < K2 < 1, 0 < ax < a2 < \ and 0
(b) (c) (d)
Then
K(OX,KX)
B(Bx,Kx,oi)ox), S(ax)<S(a2), S(ax)
< $ , 0 < K{ax,
KX) and
0 < B{BX
,KX,OX).
Proof. For (a), (b), (c), compute the derivatives of K, B, S with respect to the appropriate variables and note that they are nonnegative. The proof of (d) is straightforward. Next introduce the functions h
e K „W = P ~ ' + {l~K) fi,K,a\ if
g-T^Vj _ K(
K
We have purposefully not simplified the expression for ease of manipulation. By direct calculus we prove Lemma 2. (a) Dhf K a(t) = - 1 + i i ^ ( 7 ( - l + 1/(1 - It)2). (b) g ^ ^
= ii—V-o
KV
Ul,
fori>2,
1381 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
481
(c) hfiKO{P) = a{\-K)P, (d) Dhfi[Ka(0) = -{\-2a + Ka). Let T{0,K,a) and
BJ(P,K,O)
=
(B{fi,Kto),K{o,K),S{o))
be the 0 component of the ./illiterate TJ of T.
Lemma 3.
w/iere (/?', *', a') = T{P,K,O). Proof. We prove this algebraic identity as follows: the Taylor series for h fi,K,o(S + P) »S:
i>2 v 1
"^
using the previous lemma. V « . , ( ' + *) _
as ~ ( KS V"' + Ko)£'A(l-K)fi)
K(i-2o v
'—
~ * -"
— ' ,>2
.2
K(1 - 2a + KO) (1 -
jfas)
Comparing terms finishes the proof. Newton's method has the following basic property. Proposidon 1. Let L be a linear automorphism of V, A : F -» E an affine isomorphism, U c E, and f:U->E. Then NL.f.A =
A-lNfA.
Let Tr(ft) denote the translation by b. Lemma 4. Tr(-XV(/*, K, a ) W v >=o ' /or n > 1. Proof. For n = 1
J r ( g ^ . K,a)) = N v o '
_V«,,-Tr(^)
,
1382
482
MICHAEL SHUB AND STEVE SMALE
The first equality follows from the last proposition and the second equality from Lemma 3. Now induction finishes the proof. Recall that in the setting of the domination theorem we have 2
vt h* *(0 = fi-t + -r1P,JK '
>"
a = h„ n „ (0 where a = P>a>a.
\ - y t
=•
a
'
n_a\2
and tn is the nth iterate of 0 by Newton's method so tn = N%
(0).
Lemma 5. For n > 1 t„-tn^=Bn-\p,a,aa). Proof. For n = 1, this is obvious. Since t0 = 0, induction gives us at stage n - 1 that n-l
(*)
*,,_,= j V t f , a , ff„). >=o
It follows from Lemma 4 that
V.„.., (0) = X,..,{?:BJ(0>«. o ) - E ^ . « . O; *
7=0
7=0
substituting (*) gives
where the last equality is the definition of tn . Now B definition the 0 component of T"~l(B, a, aa) so B"-\fi,
a, ff«) = V ' w . . . o ( ° )
= f
(0, a, oa) is by
- "'"-■■
Proof of the Domination Theorem (Theorem 2, §1-2). By Proposition 4, §11-1, Lemma 1 and induction it follows that fi(f.\.l)
for«>l.
K - * „ - J = W . *„-.) by definition and Bn~\fi, by Lemma 5; thus \\xn - xn_x\\
a, oa) = tn - r„_,
II-3. ROBUSTNESS
Here we give the proof of Theorem 3 of §1-2. Toward the proof of Theorem 3 consider the function a{t) = T(0 - t with x(t) = (1 + / - \/{\ + t)2 - 8r)/4 as usual.
1383 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
483
Lemma 1. The map a is a differentiable homeomorphism, a : [0, t0] -»[0, u0] where t0 = 3-Vi, uQ = ^ ^ , a'{t) > 0, t e (0, tQ). Proof. It is easy to check that a(0) = 0, a(t0) = u0, 4
W l - 6 f + f2
)
a'{0) = 0, a'(t0) = oo. Thus a ' : [0, M0] — [0, t0] is well defined and differentiable on the interior. Lemma 2. Suppose o > 0. The function 2 ^ = T(*")~to is monotone increasing for 0<s< *=j&. It is sufficient to show that this function has a positive derivative. We leave this as an exercise for the reader. Define t{l-u) + u s(t,u) = ^—-^— • We remind the reader that y/{u) = 2u - 4«+ 1. Lemma 3. The functions ^ " j
and * £ ^ are monotone increasing for 0 < u <
Once again check the positivity of the derivative. Lemma 4. s{t, u) is monotone increasing in t and u for 0 < t < 1 and
0 0 for small u>0. Definition. Let a be the first positive zero of a'(u) and a = a(S), fi = 0.02207..., Q = 0.08019667... . Let G(y, u) =
y(B)fl_M)
for 0 < y and 0 < u < 1 - ^ . Let 5 = s(a, it) and
G = G(y,a). Lemma 6. As in Theorem 3 (§1-2) let a = py, p = p{f', 0 . y = y ( / , 0 • *-*' a/50 a x = 0 ^ . ^ = P(f, x), yx = y{f, x). Suppose a
a(aj 7,
£(£) -
G•
1384 484
MICHAEL SHUB AND STEVE SMALE
Proof. a(<*x) ^aiWJy.u))
'■^•«sH^((|-«"+*-aWr=is
by Proposition 2 of §11-1 again. This is /(1-K)2
" V ¥(u) by the hypotheses, and is
{1-U)U\
1
"•" v(u) )v{U)(l-tt)
/(l-a)2 (i-o)a\ < I i—7-^r-a + -—r-f-
l , ...
'
=s(a, 8) = S
by Lemma 3. By Lemma 2 a(fixG) < a{5) and thus a(ax)/yx < a{5)/G. Proof of Theorem 3 (§1-2). By Theorem 1 (§1-2) » .. o(a) a(5) . c, - {,|| < ~ V ^ < - ^ by Umma 6 Q
= - by the definition of 5 and G. ? CHAPTER III: REDUCTION TO THE ANALYSIS OF THE CONDITION NUMBER Here we give the proofs of the statements in §1-3. Some references back to that section are inevitable. III-l.
THE HIGHER DERIVATIVE ESTIMATE
Our main goal of this section is to prove the estimate on y of Proposition 3 of §1-3 and Proposition S of that section. We start with Proposition 1. Let f : C n+1 -» C be a homogeneous polynomial of degree d. Then l/(*)l < ll/ll W forallxeC"*1. +l Proof. Let x e C" and y — (\\x\\, 0 , . . . , 0). Take a unitary automorphism U : C"+l «-- satisfying U~xy = x. Then !/(*)! \\x\\d
=
\fU-\Ux)\ \\x\\d
_ \g{y)\ \\y\\d
where g = fU~l = E < A * a and ||f|| = \\f\\ by Proposition 1 of §1-3. We have l*fr>l - '^.o
<>ll|x||rf
=
\b
| < 11*11 = 11/11 D
1385 COMPLEXITY OF BEZOITTS THEOREM. I: GEOMETRIC ASPECTS
485
Proposition 2. / / / 6 X(d), then ||A(||x|rV(x)|| < 11/11 • Proof. From the previous proposition we know that
iwr^wisii/ji.
'=i
n.
Just square both sides and sum over <. Proposition 3. Let f be a homogeneous polynomial of degree d. Then \\Dkf(x)(wl,..., for all
JC.U;,.
wk)\\
1)||/|| I M l ' - V , || • \\wk\\
€C"+1.
The proof uses two lemmas. Lemma 1. Let U : Cn+l -» C + 1 be a unitary automorphism, f : C" +l -» C a homogeneous polynomial and x,w€ C n + 1 . Then D(fo U~l)(U(x))(Uw) = Df(x)(w). The lemma follows from the chain rule, D{fo U~l)U(x) =
Df(x)U~\
Apply this to Uw . Next let 9~d be the space of homogeneous polynomials C"+l -» C and / e ^ , w e C" +1 . Then Df(x)(w) is a polynomial of degree d - 1 in x and can thus be considered as an element say Df(w) of ^ _ , . Lemma 2.
\\Df{w)\\^
Here the subscripts on the norms are temporary. It is sufficient to prove Lemma 2 for ||w|| = 1 by scaling and then for w = e0 = (1, 0 0) by choosing unitary U with Uw = e0 and using Lemma 1 together with the unitary invariance of the norm. Then since DAx)(w)= ^aoVoao"'<'-^ a
we have
a a 0 *0
a
This proves Lemma 2.
1386 486
MICHAEL SHUB AND STEVE SMALE
Note considering D f(wl, applied twice 2
w2) as a polynomial in x we have from Lemma 2
\\D f{w,, w2)\\^_2 < d(d - ijiiyvjKH \\w2\\
and similarly by induction: \\Dknwl,..., u ^ i i ^
< d(d -1) • • • (d - k+OH/II^KII
• • • \\wk\\.
Now apply Proposition 1 to obtain the assertion of Proposition 3. Lemma 3. Let d > k > 2 be positive integers. Then maxX k
d(d-l)--(d-k (d(d-l)--(d-k
>. I
+ +
l/{k l) l)V l)\ -
d^k\
)
is at k = 2. Proof. Observe ' * - ' / , / _ nK\-'ld-k /*-i (YT(d-i)Y< \LL i+l )
>
k+\
1=1
for 2 < k
terms in the product is bigger than j^t
/(^r + 1 '>' / ( *"'Y ^'/2't""(nt1'(^))1/(t-),, ,/J (*+l)!
Thus (-!)•••(-* + l ) i ; * - i 1
l/2 Jt!
j
is a decreasing function of k. Lemma 4. Let f: C"+1 -» C be a homogeneous polynomial of degree d. Then ( \\Dkf(x)(wl,...,wk)\\ l 2 \d > \\x\\d-kk\\\f\\\\wl\\---\\wk\\
y/<*-'>
rf^V-l)
)
for every k > 1. This follows from Proposition 3 and Lemma 3. Recall from the introduction that D = max(rf(). Theorem 1. Let f 6 W{d) and x e C n+ . Then (\\A(\\x\\d--kd-l/2)-'Dkf(x)\\\W-»
\
k\\\f\\
j
Proof of Theorem 1. By the definition of || ||, (\\A(\\x\\<-kdli>2ylDkf(x)\\\ •/(*-') /
pW(D-l)
~
/
2
D^
-
up'f+xn
2•
\2\ '/2'*-"
1387 487
COMPLEXITY OF BEZOUTS THEOREM. 1: GEOMETRIC ASPECTS
and by Lemma 4 ?" 2 r,f _ n \ * - ' i i n i \ 2 \ ' / 2 ( * - ' )
^
<
^
-. D
We next prove Proposition 3 of §1-3. For / : C" - C" , x&C" , 7 0 (/. x) = y(f, JOIMI,
\Df(x)-lDkf(x) max it!
11*111
Z>/(jC)-|A(rf//2)A(||JC||f'-')A(
!/(*-!)
k>l
< max fi{f, x)
\l(k-\)D
k>\
2
by Theorem 1. The last is less than fi(f, x)Di/2/2 since fi{f, x) > 1 . This proves the first part. The proof of the second part is essentially the same. Lemma 5. Let A, B : E -» F be bounded linear maps of Banach spaces where B is invertible, and \\A - B\\ \\B~l\\ < 1. Then A is invertible and \\A~l\\ < \\B-l\\/(\-\\B-l\\\\A-B\\). Proof. \\I - AB'l\\ < \\A - B|| ||B"'|| so, by Lemma 1, §11-1, ■ i„ ^
||^"ll<.
1
„, \-\\A-
B\\ \\B
___, „ A-\„ ^ „ „ - i „ „ „ A-\
,„ and | M " | I < P
IIIIJM II-
We next prove Proposition 5a of §1-3. For X # 0, X s C,
nig. o = *(**. o = ii(M~,/2iicif"',~l))0(A*)i/v (or'iPsii 'o
\md-l/2u\\'l'~l))D/]N U2
l
X)
i-wd- K\\- <- )Dif-ig)\H
'0
({»"'IIPSII
c:)\\\\(Md;,/2\K\\-l"-'l))Df\N (or'ii
by Lemma 5 as long as the denominator is positive. Thus ^'C)-l-||AK,/2)(/-A^)||^fi which follows from Proposition 3, and
We apply the last inequality to that X for which df,(f,g) = ^jjf^ which by hypothesis makes the denominator positive. The proof of 5(b) is the same replacing u by unmi and N. by Null..
1388 488
MICHAEL SHUB AND STEVE SMALE
III-2. ANALYSIS OF PROJECTIVE NEWTON METHOD We give the proof of Theorem 3 of §1-3. Part of this proof is very similar to that of Theorem 3 of §1-2 in §11-3, where we use the same notation. For 0 < a < a„, 0 < u < 1 - ^ , let a(u, a) = (*(",<*)) ( i
-f—
=R(u,a)
S(a,u)
where R(U, a) =
5—* J 5 . 1 - u((2u - u2)/(l - u)2 + 2a)(l - « ) 2 M " ) For each u let a(u) be the maximum of a such that if/(u){l - u) Recall that a(t) = r(f) - / is defined in §11-3. Let a . be the maximum of a(u) (over u) and u that a(«) = a ^ . Lemma 1. Both a approximately
■ and u-
are defined uniquely and are positive. Moreover
a,™ = 07364.... proj.
the least u such
u-
'
= .0203....
proj.
The proof is very close to the proof of Theorem 3 of §1-2 in §11-3. One must check the monotonicity and boundary conditions of ft. We leave the details to the reader. Proposition 1. Let f e X[d), x, C e C" +l with x e NQ where # c = C + Null { , Nx = x + Null,, Nullc = {v G C + 1 | (v, C) = 0} etc. Let r0 = u
= W/U
• <)•
Then
\Wf{x)\-NxDf{x)\N\\
< K where
20/2
K =
(1+rj) 1 - r0((2 - «)«/(! - M)2 + Dnn)(\ -
u)2/y(u)
where ft = /iplX)J (/\ C). 1 = f ( / \ 0 and as long as the denominator remains positive. Proof. For the proof we use a series of lemmas. Lemma 2. Let L : C n+I - C" have rank n. Suppose C n + ' = V e Vx as unitary direct sum where V has dimension n and Vx dimension 1. With respect to this splitting write L(x + y) = Ax + By, A = L/Vn, B = L/Vl. Let Wn be an n-dimensional subspace of C n+1 which is given as the graph of a linear map a : V" -» Vx, tV" =
{(x,o(x))\xeVn}.
1389 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
489
If A is invertible and §A~X Bo^ < 1 then L \ Wn is also invertible and
ii(Li^r^ii<(1+ll12)l/2l-||/T'fl<x|| Proof. We wish to solve the equation (L\ Wn)~lA(v) = x + a(x) where x e Vn or (L\W")(x + a(x)) = A(v) or yet Ax + Ba(x) = A(v). Inverting A we have x + A~ Box = v. This last equation can be solved for x by (1 - t)~x = 1 + / + t2 + ••• , x = (I + A~xBo)-\v) and ||x|| < . . ^ i ^ J H . Finally since \\x + o(x)\\ < (1 + IM|2)I/2H*II, multiplying gives ||x + (x(x)[|< [1+}°}\1
l-WA-'BaW
Jv\\
Lemma 3. Let x e Null(C)+C- Then Null, is the graph of a : Null(C) — C(fl£if) where a{w) = -^"jmj"^ • ( # « * C(ii£i[) means the subspace generated by ^ .) Proof. Given u> e Null{ we want to find ff(to) such that (w+o(w)-fa , x) = 0. Solving for a(w) (o(tV)-^r
Xj =-{W , X)
but now (»(w)jj|j|. * ) = ^ j p «f. x-C) + (H,C)) = a{w)\\Q and , x - C> since x - { 6 Null, and u; e Null c . TiCmma 4.
\\(Df\N(x))-,Df(x)x\\
S'-pjj
< W - *)*(/. *)*>•
/Voo/. Df(x)x = A(rf.)/(x) is Euler's identity and IKA/l*
(x)y,A(di)f(x)\\ \\x\\
< IKAfl^rV/'W'"')!! ||A(J,.)|| \\A(d-],2\\xf-)f(x)
1390 490
MICHAEL SHUB AND STEVE SMALE
Lemma 5. Let x € N(. Then
\\(Df\N((X))-lDf(xK\\ < i i t i i l i ^ ^ c r , CM/, c)/>+£^*f). Proof. \\{Df\N((x))~lDAxK\\
< \\(Df\N((x))-lDf\N((Q\\
\\(Df\N((Q)-lDf(x)i:\\.
The first term of the product < ^~^1 by Lemma 3, §11-1. The second term satisfies
I W L (0) Df(x)C\\ = *=0
< li WU c (or'AAOCII +
IICII J
> + »)(y(C)iiJc - f
ID*
k=\
which by Lemma 4 and summation of the series is less than or equal to
IKIl(iW(/. <W> UD + ^T^T " *)• Now multiply the two estimates together. Proof of Proposition 1. We apply Lemma 2 with (w, x - Q a{w) = -■
IICII ' B = Df(x);_c_ 'IICII' ^ = D/| W{ (x).
Thus ||
l
B\\\\a\\<
IICII
(1 ^( ^.(/,0,(/,C» i^), < 1 - v(u) DV r p r o j w ' " " " ' " + {l-u? 0 where the last inequality is a hypothesis of Proposition 1. □ Proposition 2. As in Proposition 1, let fi, = PQ(f\N., C). ^ = A 0 (/U • ■*) *'cThen ( 1 - « ) ( ( ! - « ) ^ { + r 0 )||;|| PX
v(«)(i-«)iicir ^((l-lQftg-HQ rt_ <
V(") 2
1391 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
491
This proposition is a consequence of Proposition 2 of §11-1, and the previous proposition. Proposition 3. Let x € N^ where XC' is the zero of f\N associated to x by Newton's method on f\N and where x is one Newton iterate of x. (We suppose a(f\N , x) = ax < aQ.) Then ! r
II*'-C'II^/TK)-« V yY0(x) 0(X) )J\\
°
IIC
y0(y0(x)/
First we prove a lemma. Lemma 6. Let £', x e Nx. Let n(x') be the radial projection of x into the space TV,;. Then (a) n(X)
X
(x',S')X-
K>f-(i'-x',Z'-x)
(b) / / | | x - { ' | | > | | a c ' - « ' | | then
IW/)-^|<||x'-^|(l+V5J!^). Proof, (a) Since ({' - x , x) = 0,
IK'll 2 -(£'-*', {'-*> = (*',{') so to prove (a) it suffices to prove <||i'|| x'/(x', £') - £', <j') = 0 which is im mediate. b) IK'll2
\\n{x)-Z\\ =
| (Z'-x',Z'-x) -x) \\\Z'f-(i'-x',{' (i'-x',i'-x) ,
-M
>• ii ^
= \\x-?\\(l
IK'-*llll*'ll\ \(x'
Now we note that i
_/.
,_/
;.
2\{x , O l > \(x', O + «T, x')| = | ll-x'H2 + |K'||2 - ||x - i'll'l
1392 492
MICHAEL SHUB AND STEVE SMALE
which follows from expanding (x/ - {rl , xI -
f IK'~*ll 11**11 V <
2
4|K'-*llVlf 2
v \(x',i')\ ) -|iu'ii + i n - i i y - { W 4iK'-xii2(iix|i2 + iix-yn 2 )
(||x||2 + ||JC - x'||2 + ||x||2 + ||x - O l 2 - ||JC' - «'||2)2
< 4ii^
/
-xii 2 dixii 2 -nix-yii 2 )
(2||x||2 + ||x-x'|| 2 ) 2 4|K , -x|| 2 J\\t!-x\\\\ 2 2 (2||x||- + ||x-*T) - v I W I ; • _»_ MJC-jc'ii substituting above yields "
IW*V^II
J !
^ ) .
Proof of Proposition 3. Let n(x') be the radial projection of x on Nx(^ . Then ||x,-C,H IIC'II
=
||»(x,)-AC,|| /.I
IIAC'n
>
and
l|x-AC'||<^, Wx'-K'WK^'01* yx by Theorem 1, §1-2. Therefore !!«(*')-*f II < * ( « , ) - « , (x , ^ T K ) \ IIAC'II " IIAC'lly, V y J W by Lemma 6 and since P('|| > ||x||, we are done. After these preliminary results we go directly to the proof of Theorem 3 1-2 So suppose t] = >/(/, Q < apni jyn as in the hypotheses of the theorem. (§1-3). Let » = «(MproJ.'aP«»J>
We will show (a) Qc ^ ° W . (b) K < * ( « p r o j , Q p r o j ) = « ,
(c) ax < a, (d) Jl^Jly < a(a)(l + vlT(d))^(Mpf0j.)(l - «proj.)/K
1393 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
493
First note that (d) with the definition of a • , upmi yields Theorem 3. Next (a) is a consequence of the bounds on r\, y0 and Proposition 2 of §1-3. Here is the argument for (b). Observe (i) ro? ^ "proj. bV hypothesis and y > 1 so r0 < «proj , (ii) u = r0y0 < rQy < «proj , (iii) raDntf < ^ • D^IBL = *>
^ ' —
y
y
„2
proj. proj. —
proj. proj.
by the hypotheses. Finally note that ft(u, a) is monotone in u and a, as long as the denominator doesn't vanish. Part (c) is a consequence of Proposition 2 and (a), (b). For (d) we have by Proposition 3 that
IICII
" >W v
r* I
By Proposition 2
y <_^0L_J!£li 7 0'-K«)(l-")IICII so by (ii) and the monotonicity of K , ¥(u)l\-u) and the hypothesis that y^ < ? we have 7 x
°
A 1 - M ~ ~\I/(U • )( • ) * proj. proj.'
Let ff(v\
proj
=
-
P"" 7 ""' y
Then by Lemma 2 of §11-3 and the assumption that
y0x ~
m?)
ax
•
By Proposition 2 A .r.rm - * ( 1 " " ) ( 0 " " ) / ? < +
rp)
*(Qproj'
U
™)?
Then ^ < o^_/J> by Proposition 2 of §1-3 and the hypothesis that n nptoi < aproj/J». Also by hypothesis r0y < wproj . Thus by Lemma 3 of §11-3 and (b) fi0xH{y)
a(a)
y0x
-m?Y
= &.
1394 494
MICHAEL SHUB AND STEVE SMALE
As|i>l, U>
~ w(u • )(1 -u
•)
and we have a(ax)
(**)
^
?K(aproi.>«proj.)
"
Now we consider the term *{ax)ly0x ■ By Lemma 2 of §11-3 x{bs)/s is also monotone increasing in s and T(/) is monotone in t by Lemma 1 of §11-3. Thus as above *K)
<
T(q)
since ? > Z)1/2 > 1, ic(aproj , «proj) > 1 and viu^l denominator is > 1. Hence *(<*x)/y0x < i(a) • And (***)
(l + y/2^)<{l
- «proj) < 1 the
+ y/2x(&)).
Multiplying (**) and (* * *) and substituting in (*) finishes the proof of (d) and hence the theorem. CHAPTER IV: CHARACTERIZING THE CONDITION NUMBER IV-1. THE PROJECTIVE CASE H =
\jp
In this section we prove Theorem 1 of §1-4. We begin with the same notation and a preliminary proposition. Given two /i-dimensional complex vector spaces Vx, V2 with Hermitian structures and a linear map A : Vl -» V2 we define the Frobenius norm of A , ^A\\F as \\M\\F where M is a matrix representation of A with respect to any orthonormal bases of K, and V2. By the following standard lemma, \\A\\F is well defined. * Lemma 1. Let A,VvV2benxn = \\A\\F.
matrices with Vv V2 unitary. Then H^/l^ll,,
Let 0 / x € C" +l . Let Lx(Cn+l, C") be the subspace of linear maps vanish ing at x. Let 5fx c &[d) be the subspace of maps / = (/, fn) of the form d x d l fi{z) = ({z,x) >- l(x,x) >- )Li and L = (L,, ... , L„) e Lx(Cn+*, C"). Let +I £>, : %[d) - L(C" ,C") be the derivative / - DJ. Recall Vx = {/ e **(rf) I /(•*) = 0} • F ° r / e Kj, Dxf{x) = 0 since / is constantly zero on the ray through x. Thus we may consider Dx : Vx -> Lx(Cn+l, C"). Let Gx = {feVx\Dxf-0). \\A\\F is the same as the Hilbert-Schmidt norm (trace(/4*/4))'
1395 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
495
Proposition 1. (a) Vx in the Hermitian direct sum 2Cx®Gx. (b) For h€J?x, \\h\\ = ||A(rf- ,/2 ||j f |r (rf '- |) )DA(x)| Nllll J| f . For the proof of this proposition we prove two lemmas. Let u : C n+ -+ C n+1 be a unitary transformation and u : Vx -* Vux the induced isometry «(/) = / o » " . Lemma 2. Let M : C " + 1 ->C n+1 be unitary. Then (a) « ( ^ ) = ^ , (b) *(GX) = GUX. Proof, (a) Let / e &x so (x, x) ■ +I
with L = ( L , , . . . , L„) 6 L X (C" , C"). But then Lou - 1 = ( L , o i r ' , ... , L„o « " ' ) a „ ( C " + ' , C " ) and / O K - ' H / , . ! ! " 1 / , . « ■ ' ) where -i
.
(u_1z,jr)'_1
_i
^—nrr-'
V " (> 7
z =1
^°" ( )
z =
(z,«jc) rf '" 1 . .
T7rr(
LoM
-i.
)/ z -
(x, jt) < («x, ux) • This shows U(JSTX) c ^ , but as «"' = « " ' , «"'(^ u x ) c -% and M(-%) = ux
(b) By the chain rule, D(fou~l(u(x))) iff D/(x) = 0.
= 0 iff Df(x)ou~i
= 0 which holds
Lemma3. Lrt L e L x (C n+1 , C") anrf f(z) = (/,(z), ... , /„(*)) »W»
(b) H/ll = IIA^'Vir^V/WlNunJlF/Voo/. (a) /(z) = A((z,x) < ''- , /^.^) < ''" l )i-(^) so
= 0 + L(v). (b) Let u be the unitary transformation mapping x to ||x||e 0 . By Lemma 2 / o «-' = * = (hl,...,h„) where
and where n
-I V~'=£*■>*,-
1396 496
MICHAEL SHUB AND STEVE SMALE
Thus
111/1
-"•'"'■-sSdfe
= \\A(d;l/2\\x\\-{d-l))Lou-l\UM
\\F
= IIA^-'^lljflf^-'^LI^JI, by Lemma 1. Proof of Proposition 1. (a) Dx:Vx-> LX(C"+I, C") is linear. The kernel is Gx by definition, and Dx : Sfx -» L x (C n + l , C") is an isomorphism by Lemma 3(a). Thus Vx is the direct sum of J?x and Gx. We need only check they are orthogonal. As Sf^ = ^ and G^ = Gx for A € C, X ^ 0 it is sufficient to do this for 11*11 = 1 and by Lemma 2 for x = e0. In this case 5ft = (fx,..., fn) and n
If g = (*,, ■ ■ ■, g„) and g(eQ) = 0 then gt(z) = T,aijzizd0'~l+'£la aaza where a = (a0, a,, ... , an) and a 0 < d( — 2. Then D
(Eaa^)(^0) = 0 * o
'
where a 0 < di, — 2. Thus for Dg(e0) = 0 all the a, = 0. This establishes the orthogonality. 1(b) is Lemma 3(b). Proof of Theorem 1. First we prove that H^U', x) > pJx^: Let (g, x) € I* be such that p(f, x) = dp(f, g) = Hjj^H and let f - g = h. First we claim that h € Sfx. By Proposition 1 h = h~ + hG ,
h~ €&,
hG e G ,
and \\h\\ > \\hj?J\ with equality iff h = h# . Since DxhG = 0,D(g + hG )(x) = Dg(x) and ( / + hG ,x)eZ but dp(f'g + hG ) < \\h<,\\/\\f\\. Thus' \\h\\ = W
X
JT
X
\\h# || and h = h& . That £ € I' n Vx means that Dx(f -h) is singular or A(d-l/2\\x\\-{d--l))(Dx(f-h)) is singular. It follows that rff(A«-,/2||x|r("'-,))Z)/(x)|Nul^5) < ||A(rf-,/2||x|rW'-")D/.(x)|NuH ||, = ||/t||
1397 COMPLEXITY OF BEZOI/TS THEOREM. I: GEOMETRIC ASPECTS
497
by Proposition 1. By Proposition 1 of §1-4
ii(A(«/f''Vir
A, = «z,x)''-7<*.*»£,(z). By Lemma 3(a) Dh\mu
= B\UM = B and / - h e Z' n Vx . By Proposition 1
IWI = ±«nd , ( / , * ) < ft = ^ = j ^ . IV-2. BOUNDS ON ZEROS AND THE AFFINE CASE
We first prove Proposition 2 of §1-4. It follows immediately from Theorem 1. Let fe^{d) and x ? 0 e C + 1 with f(x) = 0. Then
rf(/,i0)<j^|iiA(rf;/2)/n. For the proof we first construct a perturbation H e ^ r f ) . Let Ht(z) = fj(x)({z, x)d'/(x, x)d') be the /th coordinate of / / where /) is the "highest order homogeneous part" of ft. Precisely f{z) = ft{z)\z _0 or yet ft consists of the sum of monomials of ft which do not contain z 0 . Note that we have immediately that H(x) = f(x) so that f-HzI^. The theorem is thus a consequence of Lemma 1. Under the hypotheses of Theorem 1
ll",H * {j^lK'^HProof of Lemma 1. Using unitary invariance of the norm it is easy to see that 11
'"
\\x\\d- IWIV ^ \*o\ \g(x)\ \\x\\ \\x\\d-1
\\x\\"-1
)
where zQg(z) = f(z) - ft(z) and degree g is ,- - 1. Thus by Proposition 1 of §111-1, ||ff,|| < (|x 0 |/||x||)||*||. Thus for Lemma 1 and Theorem 1 it is sufficient to prove
1398 498
MICHAEL SHUB AND STEVE SMALE
Lemma 2. ||^||2 <
dt\\ff.
Note that all the terms of ft{x) = £ I Q I = ( / aax
*(*) = E <w> „.)■*"• M=«/,.-i
Therefore "*"
' a a 0+ l
2_
a. I
(d.-l)\
and
',t>- £ iv J'"*T""'' |a|-rf,-l
'"
so ||*||2
x '• <^"+1 ~* ^u^x ^
tne
orthogonal projections. Then
K*<(*)ii _ /j_ jc_\ II^WII MKH'IWI, and it^x) is orthogonal to Null, nNull^. Proof, n^x) = x- fy$f . Now if w e Null, nNull,, then (x, w) and (<J, w) are both zero, so {n((x), w) = 0 and TT{(JC) is orthogonal to Null^nNull^. Let v = ic((x). , \ (v>x)
"<{v) = v-j^?x-
Note that (v, x) = {v, v) since v = TT,(JC) . Thus nx{v) = vV r=
71x V \r +
"
(\\v\\2/\\x\\2)x.
\\V\\\
"X\r
and |2
,. . , , 2
...
,,2
1 5 ^ , , _ N T , J!£z^L
Nr
iwr ^*^
2
by Pythagoras
iwr by the definition of v.
\\t\n\x\\ +1 Lemma 4. L r t ^ , < f e C + l such JUC/I r/;a/ (x, 0 / 0. Let jr{ : C — Null^ te the orthogonal projection. Then ||(jr{ | Null x )"'|| = (gJgf .
1399 COMPLEXITY OF BEZOUTS THEOREM. I: GEOMETRIC ASPECTS
499
Proof. Let u,, ... , «„_, be an orthonormal basis of Null^nNull,. Then u, vn, nxi/\\nxi\\ is an orthonormal basis of Nullx and vt, ... ,vn, it(nxt is an orthonormal basis of Null,. Let
Then
IM-(D'/)" ! .
s'w-tv,^,^,
and
1=1
by the previous lemma. This is less than or equal to ffJ^fdMI) with equality if all a, = 0 for / = 1 «,an+I#0. D Proposition 1. Let A : C n+1 -» C" be linear. Suppose A{{) = 0 and A | Null{ is invertible. Let x e C"+1 5wcA f/iaf (x, £) ^ 0. FACTI ^ | Null^ w invertible and MA | Nuig-'y < M M | | ( y j ! N u U { ) - ' | | . /Voo/. ||M Null,) -1 1| = ||(/J | Null { )o R{ | NullJ-'H < ||(JT€ | Nullje)_l||||(-4 | Null c ) _, || and the previous lemma finishes the proof. Theorem 2 of §1-4 now follows from Proposition 1. Let A=
A(d-*/2)A(\\t\\~{drl))Df(0
and x = e0. Then ||JC|| = 1 and {x,
577-591. 3. J. Canny, Generalized characteristic polynomials. J. Symb. Comput. 9 (1990). 241-250. 4. J. Demmel, On condition numbers and the distance to the nearest ill-posed problem, Numer. Math. 51 (1987), 251-289.
1400 MICHAEL SHUB AND STEVE SMALE
S. 6. 7. 8. 9. 10. 11. 12.
13. 14. 15. 16. 17. 18. 19. 20. 21. 22. 23. 24.
25. 26. 27. 28. 29.
, The probability that a numerical analysis problem is difficult, Math. Comp. SO (1988), 449-480. C. Eckart and G. Young, The approximation of one matrix by another of lower rank, Psychometrika 1 (1936), 211-218. C. Garcia and W. Zangwill, Pathways to solutions,fixedpoints, and equilibria. Prentice Hall, Englewood Cliffs, NJ, 1981. G. Golub and C. Van Loan, Matrix computations, Johns Hopkins Univ. Press, Baltimore, MD, 1989. D. Grigoriev, Computational complexity in polynomial algebra, Proc. Internat. Congr. Math. (Berkeley, 1986), vol. 1, 2, Amer. Math. Soc., Providence, RI, 1987, 1452-1460. J. Heintz, Definability and fast quantifier elimination in algebraically closedfields,Theoret. Comput. Sci. 24 (1983), 239-278. M. Hirsch and S. Smale, On algorithms for solving f(x) = 0, Comm. Pure Appl. Math. 32 (1979), 281-312. D. Ierardi, Quantifier elimination in thefirst-ordertheory of algebraically closedfields,Proc. 21st Annual ACM Sympos. on the Theory of Computing and Ph.D. Thesis, Cornell Uni versity (Computer Science). H. Keller, Global homotopic and Newton methods, Recent Advances in Numerical Analysis, Academic Press, New York, 1978, pp. 73-94. Eric Kostlan, Random polynomials and the statistical fundamental theorem of algebra, Preprint, Univ. of Hawaii (1987). S. Lang, Real analysis, Addison-Wesley, Reading, Mass., 1983. T. Li, T. Sauer, and Yorke, Numerical solution of a class of deficient polynomial systems, SIAM J. Numer. Anal. 24 (1987), 435-451. A. Morgan, Solving polynomial systems using continuation for scientific and engineering problems, Prentice-Hall, Englewood Cliffs, NJ, 1987. , Polynomial continuation and its relationship to the symbolic reduction ofpolynomial systems, Preprint (1990). D. Mumford, Algebraic geometry. I, Complex projective varieties, Springer-Verlag, New York, 1976. J. Renegar, On the efficiency of Newton's Method in approximating all zeros of systems of complex polynomials, Math. Oper. Res. 12 (1987), 121-148. , On the worst case arithmetic complexity of approximating zeros of systems of poly nomials, SIAM J. Comp. 18 (1989), 350-370. J. Renegar and M. Shub, Unified complexity analysis for Newton LP Methods, Math. Programming (to appear). H. Royden, Newton's method, Preprint (1986). M. Shub, Some remarks on Bezout's theorem and complexity theory, From Topology to Computation, Proc. of the Smalefest (M. Hirsch, J. Marsden, and M. Shub, eds.) (to ap pear). M. Shub and S. Smale, Computational complexity: On the geometry of polynomials and a theory of cost. I, Ann. Sci. Ecole Norm. Sup. (4) 18 (1985), 107-142. , Computational complexity: On the geometry ofpolynomials and a theory of cost. II, SIAM J. Comput. IS (1986), 145-161. S. Smale, The fundamental theorem of algebra and complexity theory. Bull. Amer. Math. Soc. (N.S.)4(1981), 1-36. , On the efficiency of algorithms of analysis, Bull. Amer. Math. Soc. (N.S.) 13 (1985), 87-121. , Newton's method estimatesfromdata at one point. The Merging of Disciplines: New Directions in Pure, Applied and Computation Math (Ewing, R., Gross, K. and Martin, C. eds.), Springer-Verlag, New York, 1986.
1401 COMPLEXITY OF BEZOUT'S THEOREM. I: GEOMETRIC ASPECTS 30.
501
, Algorithms for solving equations, Proc. Internal. Congr. Math. (Berkeley 1986), vol. 1, Amer. Math. Soc., Providence, RI, 1987, 172-195.
31. E. Stein and G. Weiss, Introduction to Fourier analysis on Euclidean spaces, Princeton Univ. Press, Princeton, NJ, 1971. 32. Xinghua Wang and DanFu Han, Precise point estimates on continuous complexity theory, Preprint, Hangzhou Univ. (1986). 33. J. Wilkinson, Rounding errors in algebraic processes. Prentice Hall, Englewood Cliffs, NJ, 1963. 34. A. Wright, Finding all solutions to a system of polynomial equations. Math. Comp. 44 (1985),
125-133. 35. W. Zulehner, A simple homotopy method for determining all isolated solutions to polynomial systems, Math. Comp. 50 (1988), 167-177. IBM, T. J. WATSON RESEARCH CENTER, 32-2, YORKTOWN HEIGHTS, NEW YORK 10598-0218
E-mail address: [email protected] DEPARTMENT OF MATHEMATICS, UNIVERSITY OF CALIFORNIA AT BERKELEY, BERKELEY, CALI
FORNIA 94720
E-mail address: [email protected]
1402
COMPLEXITY OF BEZOUT'S THEOREM II VOLUMES AND PROBABILITIES M. Shub- S.SmaleA b s t r a c t . In this paper we study volume estimates in the space of systems of n homegeneous polynomial equations of fixed degrees di with respect to a natural Hermitian structure on the space of such systems in variant under the action of the unitary group. We show that the average number of real roots of real systems is V1!7 where V = fj[ d, is the Bezout number. We estimate the volume of the subspace of badly conditioned problems and show that volume is bounded by a small degree polynomial in n, N and V times the reciprocal of the condition number to the fourth power. Here N is the dimension of the space of systems.
Section 1. I n t r o d u c t i o n . This paper can be read independently of Shub-Smale hereafter referred to as [I], but is closely related to it. Here we confine ourselves to homoge neous polynomials and projective spaces although some extensions to the affine case may be dealt with as in [I]. The paper [I] can be read for background and more references. First consider a real polynomial system / : K n + 1 —► R" so that f(z) = (/i(*o. • • .,*n), • • .,/n(*o. • • -.*n)) and each /,• is a homogeneous polynomial of degree d, > 0. Let Tif^ be the linear space of all such / where d — {du ..., d„) (permitting fi = 0). There is a natural inner product on "Hfd-. invariant under the induced ac tion of the orthogonal group 0(n+1) acting on R n + 1 (so that (foO, g°0) = (/,g) for O € 0(n+1)). See Section 2 for this. This inner product defines a Riemannian structure and volume element on the corresponding projective space Pi'Hfj,)) of lines (Fubini-Study). In turn this volume element defines a probability measure so that the following makes sense. "Some of this work was carried out when Shub was visiting the Berkeley Math De partment for 2 months in 1992, "Supported partially by NSF funds
1403 268
M. SHUB , S. SMALE
Theorem A. The average number of real zeros in P„(R) of f 6 12
isV '
P(Hfd))
where V = Yl"=ldi.
According to Bezout's theorem, the (average) number of zeros over C is V. Eric Kostlan-1991 proved this result earlier in the case that all of the d,'s are the same. In this paper Kostlan has a similar result for underdetermined systems, or gives average volumes of real varieties. Moreover, Kostlan-1987 suggested using an orthogonally (or unitarily over C) in variant metric in these complexity matters. Such a metric was already used in the theory of group representations and harmonic analysis (see e.g. Stein-Weiss). In 1943, M. Kac had found in one variable, the expected number of real roots to be asymptotic to % log d using the traditional measure. Next consider the complex case, / : C + l —► C", with each coordinate function /,• a homogeneous polynomial of degree
vw = {(/,<) e P(nw) x P(cr+1) | /(c) = o}. This complex non-singular subvariety of codimension n plays a central role in our work. Inspired by Wilkinson we define the condition number \i : V^ — R + U oo by
M/.0 = IP/(Ol^lA(^/2|Ki|'-l)||||/||.
Here A(j/,) is the diagonal matrix with j/,- as the (»', i) entry and Nc = {w6P(C+i)\(wX)
= 0}-
The careful reader will have noted our customary practice of identifying ob jects in linear spaces and their quotient projective spaces. But appropriate homogenization gives sense to our definitions as fi(f>0 above. If £>/(OIJV< is singular then /i(/, C) = oo. In [I] n was defined and called ^ p r o j . Example. Let e0 = (1,0, . . . , 0 ) and /,(z) = z$i~lzi, ,
i = l , . . . , n . We
,.1/2
claim that /*(/, eo) = ( £ " = 1 3-) D 1 ' 2 , D = maxjc/j, and thus (/, eo) is a very well conditioned set of pairs varying over d. Observe Neo = {(0,u1
u„),v,- G C} and that £>/(e0)|w.o is repre
sented by the identity matrix. Moreover it is checked that ||/|| = (5Z"=i 37) This yields our statement. Typical numerical algorithms follow paths in V^ and loss of precision due to roundoff errors can be controlled by upper bounds on the corre sponding condition number /i. Moreover ft plays a primary role in the (exact arithmetic) complexity analysis of such algorithms (see [I]). Thus it is important to understand the probability distribution of \i. We will show:
1404 COMPLEXITY OF BEZOUT THEOREM
269
Theorem B. If n > 1, and D > 1, then the probability that / i ( / , ( ) > po is less than K^. Mote precisely Vol{(/,C)€^)l/i(/,0>/io} VolV(<0 " Here N is the dimension
KnN
/ig-
ofTi^y
The constant K is & universal constant less than 25. The case n = 1 will be dealt with in Theorem D. Note that as a consequence of Theorem B, we see that most (/, C) in V(i) are well conditioned. The bound is independent of V and a low polynomial in n and N. Numerical analysis has a useful tradition of relating the condition num ber to the distance p of the nearest ill-posed problem (see e.g. EckartYoung, Demmel). In our setting the ill-conditioned pairs E' C V(d) are described by ^ = {(/,<) € V(d) | £>/(C)k is singular}. It can be shown that E' is a non-singular hypersurface in V^ by a transversality argument. For C € C* + l let V( = {/ € "H(d) I /(C) = 0} and V( = P(V(). Then Vc can be naturally identified with x f l(Q C V(i) where *3 : V(rf) —» P„ is the restriction of the projection P(W(j)) x P n —► P„, and Pn=P(C+1). Define p(/,C) to be the distance in V<; of (/,C) to E'n V^. This projective space distance is taken for convenience to be dp = sin dn where dR is the Riemannian (or Fubini-Study) distance. The diameter of projective space is one. A main result of [I] is: Condition Number Theorem.
A sketch of the proof is given in Section 2. Thus Theorem B may be interpreted as giving distribution estimates of pairs (/, C) close to ill-conditioned ones. We now pass to the more subtle situation corresponding to all roots of a given / € P(tt()). Define the condition number (t(f) of / € P(7i^), H : PCH(d)) — R + U oo by
/!(/)= max/i(/,0-
1405 270
M. SHUB , S. SMALE
Example, n = 1, ft(z) = ZQ — zf. The zeros of ft are (1,?), whre q is a d"1 root of unity. Take typically q = 1. Then ||(|| = V2, { = (1,1) and \\fd\\ = V2. Also N( = {(v,-v),v G Q , ||Z>/d(C)|*j|| = J3 a n d i l fo »°ws that/i(/d) = ^ . Note that for this series fd the condition number grows exponentially in d, and so is extremely ill-conditioned. Yet this ft and its variations with exponential behavior are used in numerical analysis as a model starting polynomial whose roots are known. We will prove the existence of wellconditioned sequences {gi) (even for general n > 1), yet are unable to exhibit them even for the 1-variable case. In a further account we hope to develop this subject which in the case n = 1, is intimately connected to elliptic capacity and transfinite diameter (see Tsuji). Defining p(f) = mii^ / ( o = 0 p(/, C) we have: Corollary of t h e Condition N u m b e r T h e o r e m .
"U)'iry T h e o r e m C. Let N, = {f € P(H(d)) | />(/) < p). n > 1, D> 1 Vol/V, Vol P(nw)
Then for p < ^ - ,
p*n2(n + l)(N - l)(N - 2)V " 4
where N = dimW(«i). The closest previous results in the direction of Theorem 3 are due to Renegar. Remark. The exponent 4 in Theorem C is somewhat surprising. Since E C P('W(rf)) is a hypersurface, dimensional considerations say that the volume of the tube of radius p around E should vary like p2. Problem. Is d(f, E) = 0(p2f))? The theorem says that on the average it is. There is one exceptional case where d = ( 1 , . . . , 1) in which case E C Ptfi(i,\ i)) is of complex codimension 2. In this exceptional case the p* could be expected. If we consider x2 - lexy - f{x, y) / € W(j), x2 - Itxy + e V S E but
Pif) = 0(e). Corollary. For n > 1 and each d = (di,...,d„) such that p.(fa)
<
there is fa € P(H(d))
(»'{n+w-w-7Tvy/\
This is a consequence of Theorem 3 and the Corollary of the Condition Number Theorem.
1406 COMPLEXITY OF BEZOUT THEOREM
271
As we have suggested, it is an interesting open problem to exhibit such /(,/). These could serve as better starting points of numerical algorithms than those currently used. The probabilistic estimates given in Theorems B and C, together with the results of [I] have implications on the efficiency of algorithms for solving non-linear systems. We hope to develop this point in a future paper. In Theorem D we give the results of Theorem B and C for the case n = 1. Theorem D. For n = 1 and d > 1 Vol{(/.0€Vw|/i(/,Q>/io>l}^ VolK(<0
i)
h)
Vol N Vo\P{fi(i))
-
d{1
~
(1
-
p2)
'~'(1
+ (d
~
l)p2))
for0
Section 2. Background. The following diagram motivates the more abstract treatment of the next section. P(H(d)) x P«
U tfl
P(nw)
/
V
*2
\
Pn
The map *! : V^ —» P{H(d)) is a branched covering, branched along E' so that on V^) — » f ' ( £ ) , where E = *i(E'), *i is a covering map. The fiber w f ' ( / ) , / € P(H(d)), f $. E consists of V points, corresponding to the zeros of / . Moreover, *i : V^y —» P„ is a fiber map with fiber T^ X (C) = V( over
Cefl.. The Unitary Group U(n + 1) is the group of linear automorphisms of C n + 1 preserving the standard Hermitian inner product. This induces an action x —► ux, u g U(n + 1) on P„ as well as the action / —» / « _ 1 on P(?iw). Moreover U{n + 1) acts on P(7iw) x Pn by ( / , r ) - ( / V 1 , ux) leaving V^) invariant. The fiber map * 2 : V(d) -* P„ is also invariant under these actions. The invariant norm on 7i{d) is given by ||/|| 2 = E , E a l a a I2 (
*
)
\ < * l i • • • i <*n /
where/^(z) = T
a\,za,a
= (ori,.. .,a„) is a multi-index and ( ' \ai,...,a„J is the multi-nomial coefficient, _ ?>'—:.
)
1407 272
M. SHUB , S. SMALE
Here is a sketch of the proof of the condition number theorem. The condition number \i : l^j —► R is invariant under U(n + 1) using the chain rule (as in Lemma 1 of Section III-l of [I]). The variety E' C V(<<) is unitarily invariant as well as the distance in V(. This implies that p : V^) —► R is invariant under U(n + 1). Given (/, x) € V{i) pick u€U{n+ 1) with ux = e0 = (1,0 0). Then P{f,x) = p(/« - 1 ,«o), and
_
- = ———.
Thus it is sufficient to prove the condition number theorm for x — eo. For / S H(d)y write / = ( / , , . . . , /„)
(*)
/<(*) =
If / € Vi0 then the a, are all zero, and note thus that K = {/ € ft^ | fi(z) = a,-*0'} is the orthogonal space to Ve<). Let L(d) = {/ g T^^j | fi(z) = z*'~lEaijZj}, J(d) = the orthogonal complement of L(d) in V,t, and 3r : Veo —» L(d) the projection. Note that Null(e0) = {u S C" + 1 | (u,e 0 ) = 0} is {(0,ui,.. .,«„) € +1 C } and can be identified with C . In this way, Df(eo) may be regarded as a linear map from C* —<■ V and as a matrix. This is the matrix (a<;-) when / has the form of («): Let M(n) be the space of n x n matrices endowed with the Frobenius norm ||A/|| 2 = E|m,-y|2. Recall that A(y,) denotes the diagonal matrix with entries ( y i , . . . ,$/„). —1/2
Lemma. The map L(d) —► M(n) sending f —♦ A(rf, ' )D/(e 0 ) is a norm preserving line&r isomorphism. The proof follows from the definition of the norm on H^ 0 L(d). Now suppose H/ll = 1. Then /i(/,e 0 ) = ||£>/(eo) _ 1 AdJ / 2 || using the operator norm. By the theorem of Eckardt-Young (see [I]), /*(/, eo) = d(A(
>} \ completing the sketch
Remark 1. In [I] it was part of the definition that p(f,Q > 1. In fact it is a consequence of the original definition that /i(/, () > y/n. One uses the condition number theorem and the fact that p < y/n for any any matrix on the unit sphere in M(n). Remark 2. There is a real version of this section where the unitary group is replaced by the orthogonal group, 0{n + 1), acting on R n + l , etc.
1408
273
COMPLEXITY OF BEZOUT THEOREM
We end this section by stating some standard facts that we use in our proofs. The volume Vo^S""1) of the unit (n - l)-sphere S " - 1 is rfziT) using the gamma function. Let Pn{lR.) be real projective space of dimension n and P„ complex projective space of (complex) dimension n. Then Vol/>„(«)= 5 Vol(S") Vol/'r, = ^ : V o l ( 5 2 n + 1 ) . We also use the following integral formula: K
J0
'
2r(*±i + 9 + i)
Section 3. Some General Integral Formulae. Let M, N be (real) compact Riemannian manifolds and V a compact submanifold of the product M x N with dim V = dim M. Suppose that the restriction ir^ : V -+ N of the projection M x N —♦ N is a locally trivial fibration. Let Vy = *il(y)Let x be a regular value of *i : V —► M, the restriction of the projection MxN —> M. Define A(x, y) : Ty(N) —► TX(M) to be the linear map whose graph is the orthogonal complement to TVy (x, y) in TV(x,y). Let U be an open subset of V and # ( x ) be the number of points in TJ"'(X) f~l U.
Theorem 1. /
#{x)dM=f
JT£T,U
det(A'(x,y)A(x,y))l':ldVydN.
I JrfJV,rtU
Proof. f JnU
#(x)
/
\det(D*i)\-l—dVydN.
JNJV,C\U
jvy^2
Here Njx2 = |det(D*-2 | T(Vy)*-)\ is the normal Jacobian and T(Vy)x is the orthogonal complement of TVy in TV(x, y). Here the first equality is a version of the usual change of variable formula for integrals. The second equality is the coarea formula (Morgan). By Sard's Theorem it is sufficient to show that (♦)
| d e t ( Z ? T 1 ) | - i - = (det>lM) 1 / 2
where both sides are evaluated at (x,y) € M x N and x is a regular value ofxj :V — M. Let Hi and H? be finite dimensional real vector spaces (complex vector spaces) with inner product (Hermitian product). Let A : Hi —► H? be linear and define the graph of A as T(A) = {(x,i4(x)) | x € Hi). Let T : T(A) —► Hi be the restriction of the projection. Let T(A) inherit the inner product structure of the product.
1409
274
M. SHUB , S. SMALE
L e m m a 1. I det r\ =
det(/ + A M ) ' / 2
where A' is the adjoint of A. Proof. Let I . 1 : Hi —♦ H\ x Hi be the map x —► (x,Ax).
There is an
orthogonal (unitary in the complex case) automorphism ) of H\ such that
'(((j)'U)n-(i) and hence | det I
V;
ToI
| = det(/ + A'A)1'2.
Also Det* =
Detj --r\
U)
since
= / . These two equalities give the proof.
Lemma 2. Let B(x, y) : TM(x) —- TN(y) be the linear map whose graph T(B(x,y)) is TV(x,y) for a regular value x of*i. Then det(I + B'{x,y)B(x,y))
= det(I +
A-i'(x,y)A-1(x,y)).
Proof. TVy(x,y) is contained in the TM(x) factor of TM(x) x TN(y). As TV(x,y) is the orthogonal direct sum of TVy(x,y) and T(A(x,y)), it follows that TV(x, y) is the graph of the linear map B : TM(x) —► TN(y) which is zero on TVy(x, y) and A~l(x, y) on (DTI(X, y))(T(A(x, y))) which in turn is orthogonal to TVy(x, y) in TM(x). Consequently det(7 + B'(x, y)B(x, y)) = det(7 + A~l-(x, y)A~l(x, »)). Now we return to the proof of («) and the theorem. By Lemmas 1 and 2, using A = A(x,y), |det(D,r 1 )(*,y)| = det(/ + , 4 - 1 * A - 1 ) - 1 / 2 . By Lemma 1 and the definition of A, — t -=det(I Njviix.y)
+
A'A)l'\
Thus |det(DTi(i,y))|
Nj*i{x,y)
The next lemma finishes the proof.
det^+yl-'M"1)1^-
1410 COMPLEXITY OF BEZOUT THEOREM
275
Lemma 3. Let A : V —* W be a linear isomorphism of finite dimensional vector spaces with inner product, then det(I + A'A) _ det(/ + A - > M - i ) - d e t ( A i 4 ) Proof. Multiply numerator and denominator by det(/ + AM) det(I + A'~M-1)
det(AA')
det(AA')det{I + A'A) ~ det(AA')det(I + A—iA-1) _ det(AA')det(I + A'A) det(AA' + I)
But then det(AA' + I) - det(I + A'A) since AA' and A'A have the same eigenvalues. Let U, V, M, Ti, *i etc. be as in Theorem 2. Theorem 2. det(I+A'A)1/2dVydN.
Voli/= / / JN Jv,nu Proof.
Vo\U= f \dV= f I Ju
-1—dVydN
JN Jv,r\u
= ( l det(I + JN Jv,nu
Hjxi
A'A)l"DVydN
by Lemma 1 and the definition of A. Remark. The complex versions of Theorems 1 and 2 are true with the same proof. Then the exponents of i on det(/ + A' A)1'2, (det A'A)1!"1 musts be removed as we pass to the real determinant. That is |real determinant! = |determinant| 2 for a complex matrix.
Section 4. Integration Formulae in V^y We use the notations introduced in Section 1, 2 and specialize the com plex versions of Theorems 1 and 2 of Section 3 (cf. the remark at the end of Section 3) with M = P(Ti(,i)) and N = P„. Section 2 provides the background for the following.
1411 276
M. SHUB , S. SMALE
T h e o r e m 1. Let U be an open set in V(d) which is unitarily invariant e0 = (1,0,.. .,0) S Pn and # ( / ) be the number of points in f f l ( / ) n (; (i.e. the number of zeros of f in U). Then (a) Volt/ = VolP(n)JV < o n t ; det(/+ Df(e'0)Df(eo)) (b) /„ # ( / ) = VolP(n°)/v. 0 ni/
det(Df(e0yDf(e0)).
We will use: Proposition, det A'(f, x)A(f, x) is invariant under the unitary group act ing on V(i), and (*)
A'(f, e0)A(f,e0)
= Df(e0y
Df(e0).
Postponing the proof of the proposition for the moment, we will prove Theorem 1(b). Use Theorem 1 of the previous section and both parts of this proposition to obtain
/
#(/)= /
/
det(D/(eorD/(e0))
= VolP„/ det(D/(e„r.D/(e 0 )). Jv.0nu The proof of Theorem 1(a) is similar. Since *2 : V(j) —► P„ is unitarily invariant, so is the orthogonal comple ment of TVx(f,x) in TVw(f,x) where Vt = T J ^ I ) . Then by definition A(f, x) transforms by unitary compositions and thus det A'(f, x)A(f,x) is unitarily invariant. This proves the first part of the proposition. Working in the corresponding vector spaces, write as in Section 2, 7i(t) = K + L{d) + J{d). Also T(Vw)(f,e0)
= {(h,w) | Df(e0)w
= h(e0)}
so that A : C" — K C T/(P(W(,i))) is characterized by Aw - h e K and /»(eo) = Df(eo)(w). Here we are using the notation of Section 2 and the definition of A = A(f, eo). Then A(w)i
= hi =
^aijWjZ^ i
recalling Df(e0)(w) = £ \ a,;tuy and the Hermitian structure on K. Then A'A = Df(e0)'Df(eo).
1412 COMPLEXITY OF BEZOUT THEOREM
277
Section 5. Proof of T h e o r e m A. We start with: Proposition 1. l e t x : S[ —► Dk be orthogonal projection of the unit sphere S[ C R ' + 1 on the unit disk Dk C Rk a. subsp&ce o/"R', 0 < it < /. Let 4 :£>*—► R be continuous and U C /?* be open then ( (
= (
^(^)(l-||x||J)1^1^(r)VolS}-t.
JDk
For the proof we use the following lemma. Lemma. The normal Jacobian ofx : S' — Dk at x € S' i s ( l - | | T ( x ) | | 2 ) 1 / 2 . Proof. Let ||x(x)|| = r. Let Sk~l be the sphere of radius r about 0 in R*. Then x - l ( S * ~ ' ) = Sk~l x S'j^j which leaves only one normal direction to T - 1 ( T ( X ) ) , namely the one which maps to the day in Dk. Thus the problem reduces to Sl C R 2 and the norm of the projection is easily seen to be V I - r 2 . Proof of Proposition 1.
f (+ox)X{x-\U))= f JS"
*
ox
Jx-HU) (V)
(4°x){x-\x))-
1
^voisi^/a-iixii2)1^1^*) Ju = VolS{-i/
X(U)(l-\\x\\2)i=^^x).
JD*
Let cf = ( < / , , . . . , d „ ) , V = xfdi. Then
Ad =
fp(H't))#(f) Vol/»(««,)
is by definition the average number of real roots. We will prove that Ad = V1'2. To prove this we will show
1413 278
M. SHUB , S. SMALE
Lemma 1. A4 = Z>1/2G(n) where G is a function ofn. Apply the lemma to the case d = (di,..., dn) = ( 1 , . . . , 1). Then clearly Ai - 1 and V = 1. Therefore G(n) is identically 1, and Ad = V1'*. Thus it is sufficient to prove Lemma 1. For this we use the real analogues of results of the previous sections, with orthogonal invariance replacing unitary invariance. Thus the real version of Theorem 1(b) of Section 4 gives (with U = P C ^ ) ) : Lemma 2. Ai =
\^Lr°>^°i^''
Here e0 = ( 1 , 0 , . . . , 0), Df(e0) : R" — R n is the derivative and Df{e0)' its adjoint. Moreover Se„ is the unit sphere in V* = {/ g n*^ | f(eQ) = 0} and VCo = P(V*) has ± the volume of S«0. Let N = dimW^ so that dimS«0 = N — n - 1. Lemma 3.
/ .5. 0
(det D/( e o )*D/( e o )) 1 / 2 =
-SI 2
V^-jj^Hin) l
\T)
where H(n) depends only on n. Since \o\P{7ifd))
= ^ Vo\(SN~l)
= r^rpTj, Lemmas 2 and 3 yield that At V1*7 VolP„(R)//'(r») proving Lemma 1. Thus it remains to prove only Lemma 3. As in Section 2 let L(d) be the linear subspace of / € V,* of the form /< = x i'~l Z)?=i a'ixi an<* T : V* ~' L(d) the natural projection. Then J(d) consists of polynomial systems with no terms of the form z0' ~l, z*' ~l Ea,;- ZJ in the i'* coordinate. Next we may identify L(d) with the space of n x n matrices A = (a<;-) as above. Here L(d) comes endowed with the Hermitian invariant norm from V'*. Let M*(n) denote the space o f n x n real matrices endowed with the Frobenius norm. By Proposition 1,
i
f / € v. ( ,(DetD/( e o)*D/(eo)) 1 / 2 =
,A€L(n)(DetA'A)l'2(l J11 IMU
- P U 2 ) * " * ^ Vol
1414 COMPLEXITY OF BEZOUT THEOREM where Vol is the volume of SN~n
" " " ' or
? r
7„
279
L.\■
l
Now use the change of variables A~ A = M £ M i ( n ) , A = A(d t I / 2 ) noting ||.4|| L ( n ) = ||A~M||ju„( n ). Thus this last integral is: Vol LMm{n)(DH((&M)'&M))l»(l J
- ||A/|| 2 )
N-*'-*-l
\\M\\<\
or yet since det A 2 = V V o I P 1 ' 2 LM.W{V*{M'M)YI\\ J ||Af||
- | | M | |
2
) *
= J
^ .
Use polar coordinates to obtain that this last is
^ f y S ^ T / ' d - ' - 2 ) * ^ - 1 dr / l|M|l=r (det M-Mf'\ 1
(
j 12
1 J°
J
MZM(n)
n3+ , l
nutfmi=r(detM'M) ' = r ' - /||Af||=i(detA/*A/)1/2andwherer',J"1 scales the volume element and r" the (det M'M)1/*.
The last follows from the identity
and the substitution r2 — s (compare Section 2). Putting these together yields Lemma 3 and finishes the proof of Theorem A.
Section 6. Proof of T h e o r e m C. Our proof of Theorem C depends heavily on the following simple corol lary of a result of E d e l m a n - 1992a T h e o r e m . For n > 2, 0 < p< - j ^ ,
4.A.S)
1415 280
M. SHUB , S. SMALE
Here M(n) is the space o f n x n complex matrices with the Frobenius norm, 5 is the set of singular matrices in M(n) and d is the distance in M(n). We use Proposition 1 of Section 5, just as in the last section, but now we are working over C so that the real dimension of M(n) is 2n 2 etc. The argument of the last section yields (where N = complex dimension ~H(i)) I
det(D/(eo)-D/(e 0 )) =
JjtN,nS.0
V L€M(n)
det(M'M)(l
- WMWY"1'-"''
Vol
Jd(M,S)
where Vol = V o l S 2 ' " " " _ n ) _ 1 and S is the set of singular matrices. Use polar coordinates to evaluate the integral on the right as (following Sec tion 5) JMZMW det(ATAO(l - H M H 2 ) " - " ' - " - 1 HW||
/ " r ' V ' - ' O - r 2 ) " - " ' - " - 1 L€M{n) JO
detM' M.
J \\M\\ I|M|| = =l
Now apply Edelman's Theorem to obtain the upper bound for both sides as (for n > 1): V o l ( 5 - ' - l ) £ [l r - J « » - ' ( l - r r 4 J0
n
'-
n
-lr-Ur
n l T ^ n +V r ( n 2 + n - 2)
T " ' p* n 2 r ( n 2 ) r ( n + 2) T(n 2 -I- n - 2)T(N - n 2 - n) ~ T(n 2 ) 4 T(n 2 + n - 2) T(N-2) Now we put this information together to obtain:
Vol P„ 1 f VolP(« ( < 0 )2x JjZN,r>S.0
det(K/( e o )*D/(eo))
by Section 4, Theorem 1(b). Continuing, this is less than (by the above calculation) T"' n y p V o l S 2 " * 1 1 r(r» 2 ) 4 V o l S 2 " - ^ *
^s-n'-n^-i
W - " 2 ~ ")r(" 2 )F(n + 2) r(AT-2)
1416 COMPLEXITY OF BEZOUT THEOREM
281
which by a further easy calculation turns out to be ^ D n 2 ( r » + 1)(N - 1)(AT - 2). 4 Remark. Theorem C and Theorem B might be improved especially for rea sonable ranges of p and /JO- For example, Edelman's result which we have used above actually says that
/
I
' 3 n2_^A ^\n-r„i K
AiM(n) det(A'A)
=
IMII=i 4-*,S)<,
»,»+,-, r ( r + l ) r ( r(n^)r(n+i)r(» + 2) n-r)r(n + r-l)r(n J
UA,
+ 2-r)
^
where Vol = Vol(S 2n _ 1 ) and we have.somewhat crudely estimated by disregarding the (1 — nX) factors and then maximize the terms of the sum at r = n — 1 and finally multiplying by n. Section 7. P r o o f of T h e o r e m B . For the proof we use several lemmas and a theorem of Alan Edelman. L e m m a 1. For y € [0,1], (1 - y)k > 1 - ky. Proof. It is true for y — 0. The rest follows from a comparison of the derivatives. Lemma 2. For x > 0 and n € N 1 - (1 - min(l, nx))-'-1
< n(n' - l)x.
Proof. l-(l-min(l,nx))nJ-1 < 1 — (1 — (n7 — 1) min(l, nx)) by Lemma 1 < n(n 2 - l)x. Lemma 3. For 0 < x < 1 D, n e N
( l+
£)" (,-,)»
Proof.
,nD/n
(Thy =^'+'7+-">D,n^i+^-
1417 282
M. SHUB , S. SMALE
Lemma 4. Let M be an n x n matrix with eigenvalues Xlt...,X„ Det(/ + M) = n ( l + A,).
then
Lemma 5. Let A be an n x n matrix which is a wealc dilation, i.e.
IW»)II>IMI
V.jiO
«€R n .
Let M be a matrix and 0 < Ai < • • - < A„ the eigenva/ues of M'M 0 < ^i < • • • < A'n the eigenvaiues of M' A'AM. Then /i, > A,- V »'. Proo/. The eigenvalues of N'N for any n x n matrix iV are the principal axes of the elipse hich is the image of the unit sphere by N. The image of the unit sphere by AM is outside the image of the unit sphere by M since A is a. weak dilation. Lemma 6. Let M be an n x n matrix then Det(7 + Af*A(d,)Af) < (l + 5M^.y waere D = maxcf;. Proof. A(D 1 / 2 )A(
< Det(/ + M'A(D)iW) = Det(7 +
DM'M).
< Xn < Pn be the eigenvalues of M'M. D
S i
Det(7 + A M ' M ) = 0 ( 1 + *>) < (l + ^ )
bv the
Then
inequality of the
geometric and arithmetic means, and the last equals fl +
"n » j
by
defintion of the Frobenius norm. Theorem (Edelman). The voiume of N( n S 2 " 3 - 1 is (1 - (1 -minU.nC 2 ))"'- 1 ) VolS 2 "'" 1 where S 2 n
-1
is the unit sphere in M(n) with the Frobenius norm.
Proof. Actually this follows immediately from Corollary 3.2—Edelman1992 and the fact that Amin = (miim.^i ||i4(v)||)2 = C2Now we pass to the proof of Theorem B. First assume D > 1. By taking Np = U and applying Theorem 1(a) of Section 4, this amounts to estimating:
Q(P
'
}
Jv
det(I + Df(eo)'Df(eo))
■
Here Q M ) = ^ } j p . Just as in Section 5, we obtain: Q( W
., _ /M € ^n).HM||< 1 .^M,5)< < ,^( / + ^ ^ ) M ) ( 1 - | l A f H 2 ) W - n ' — ' '
JMsfiUnUMKi
det
( / + M-(A
1418 COMPLEXITY OF BEZOUT THEOREM
283
Apply Lemmas 3 and 6 to obtain an upper bound for the numerator. Use the fact that det(/+ M' A(dt)M) > 1 to bound the denominator from below. This leads to
< f^^s^-im2)*-"'-"-1 1-* WO-IOT)*-"'---
,p d) Q Q{p,)
Using polar coordinates ^ J01(l-rT-"'-"-1-°r'"'-'Vol(jy,/rn5»"'-i) /01(l-r2)'v-*',-'-1r2n,-1Vol(52n'-1) Apply Edelman's Theorem and Lemma 2
n, ~ - ^ } -i)/,'(i-ry-'-'-v^r Qip.a) < w ' '"
; J01(l-r2)"-"'-'->r2'>3-'dr
.
T(N - n) r(AT - n2 - n - D) n r(yV-n-D-l) ' r(//-n2-n)
2 "P
We conclude the proof with Lemma 7. Let n > 1, D > 1. Then ( 7 7 3 ^ = ^ 0 )
< K.
2
Proof. Let P = N — n . Then the estimate is
or
\
N - n 2 - t» - £ ) /
The worst case is seen to be = (1,2) and TV — n 2 — n — Z) = l.
1419 284
M. SHUB , S. SMALE
Finally, the condition number theorem allows to pass from the above estimates to Theorem B. The case where D = 1 is simpler /iY,nS?-'-det(1 +
Q{p d)
' -
M M
'
)
/ s J . , d e t ( l + A/-M)
where S, 2 "'" 1 is the unit sphere in M(n). Now 1 < det(l + M'M) < (l + ±)»
g,,,fl<,M^!<^.,) Vol(S 2n " ' )
~
by Edelman's theorem and lemma 2.
Section 8. Proof of T h e o r e m D . For the proof of Theorem D we will use a simple lemma on indefinite integrals which can be checked by differentiating. Lemma, i) / ( l +
Voiifa
~Q{p'd)-
/ J € C i | j | < 1 (i + ^l a )(i-l*l J )-- 2 -
Now use polar coordinates and Lemma i). For the proof of Theorem D ii). Remark that by Theorem 1 (b) WolP(H,d)) = Vo\P(n) I
det(D/(e0)'D/(e0))
so for n = 1 as in the proof of Theorem C Vol VolP(W (-) ) VolP(l)
d(A-D 1 L i c. id\z x | » ( i _ w »)--» J\'\«>
Now use polar coordinates and lemma ii) to finish the proof.
1420 COMPLEXITY OF BEZOUT THEOREM
285
REFERENCES
Demmel, J., On condition numbers and the distance to the nearest ill-posed problem, Numerische Math. 51 (1987), 251-289. Eckart, C. and Young, G., The approximation of one matrix by another of lower rank, Psychometrika 1 (1936), 211-218. Edelman, A., On the distribution of a scaled condition number, Math, of Comp. 58 (1992), 185-190. Edelman, A., On the moments of the determinant of a uniformly distributed complex matrix, in preparation (1992a). Kac, M., On the average number of real roots of a random algebraic equa tion, Bull. Amer. Math. Soc. 49 (1943), 314-320. Kostlan, Eric, Random polynomials and the statistical fundamental theorem of algebra, Preprint, Univ. of Hawaii (1987). Kostlan, E., On the distribution of the roots of random polynomials, to ap pear in "From Topology to Computation" Proceedings of the Smalefest, Hirsch, M., Marsden, J. and Shub, M. (Eds) (1991). Morgan, F., Geometric Measure Theory, A Beginners Guide, Academic Press, 1988. Renegar, J., On the efficiency of Newton's Method in approximating all zeros of systems of complex polynomials, Math, of Operations Research 12 (1987), 121-148. Shub, M. and Smale, S., Complexity of Bezout's Theorem I, Geometric Aspects, Preprint (referred to as [I]) (1991). Stein, E. and Weiss, G., Introduction to Fourier Analysis on Euclidean Spaces, Princeton Univ. Press, Princeton, NJ, 1971. Tsuji, M., Potential Theory in Modern Function Theory, Maruzen Co. Ltd., Tokyo, 1959. Wilkinson, J., Rounding Errors in Algebraic Processes, Prentice Hall, Englewood Cliffs, NJ.
Michael Shub IBM, T.J. Watson Research Center, Yorktown Heights, NY 10598 USA email: SHUB @ WATSON.IBM.COM
Steve Smale Math Department, University of California, Berkeley, CA 94720 USA email: SMALE @ MATH.BERKELEY.EDU
1421 JOURNAL OF COMPLEXITY 9 , 4 - 1 4 (1993)
Complexity of Bezout's Theorem III. Condition Number and Packing MICHAEL SHUB A N D STEVE SMALE* IBM T. J. Watson Research Center, Yorklown Heights, New York 10598-0218 Received September 23, 1992 DEDICATED TO JOSEPH F. TRAUB ON THE OCCASION OF HIS 60TH BIRTHDAY
1.
INTRODUCTION
The best known example of a condition number is that for the problem of solving a linear system Ax = b. The condition number is KA = ||i4|| ||i4~'||, where an operator norm on the matrix A is used. It measures the sensitivity of output error relative to input error. Since K ^ I , log KA is positive and measure the "loss of precision." For finding a root of a polynomial/, or solving/(O = 0, the condition number has been taken as W/-t = 1/|/'({)| as in Demmel (1987) (see also the earlier accounts Wilkinson (1963) and Wozniakowski (1977)). Again it measures sensitivity relative to input error. This condition number lacks certain naturality properties as invariance under the trans formations/-* A/, or i -*■ a£ for X, a E C - 0. Moreover, W/it does not satisfy Wf,r s 1. For these and other reasons we are motivated to define a new notion of condition number /*(/, £) of a polynomial at £ e C. Our emphasis is on homogeneous polynomials but the ideas apply in general. See Bez I and Bez II (Shub and Smale, 1992) for a general account of how /x(/, £) plays a role in the complexity theory of solving polynomial systems. Our present account is mainly self-contained. Let 9d be the space of complex polynomials of one variable of degree less than or equal to d, and 9€j homogeneous polynomials of exact degree * Partly supported by NSF funds. 4 0885-064X/93 $5.00 CnpvriKtil *> I W hy Academic Pre«. Inc
1422 COMPLEXITY OF BEZOUT'S THEOREM, 111
5
d in 2 variables. There is a natural isomorphism 9d-* dXd,f—* g, given by g(x, y) = Ej( a , J C V where f(x) = 1^ a,*'. The projective space of lines through the origin of C2 is denoted by 9(C2). The zeros of g E dXd are naturally considered as points of P(C2). Now we must define the derivative of g G aXd at a point u = (z, H») G C2 so that it can be inverted, and makes sense in projective space. This cannot be done purely algebraically. Let (u, v) denote the standard Hermitian inner product between vectors u, i) 6 C!, and ||u|| = (u, u)m be the associated norm on C2. The complex numbers supposed to be imbedded in C2 by z -» (z, I) G C2 and we write ||z||, = ||(z, 1)|| = (1 +
|z|2) 1/2
A model for the tangent space at u G P(C) is N„ = {v G C2 | <«, u) = 0}.
The projective derivative of g G 7Hd at u G C2 will be the restriction of Dg(u): C2 -* C to N„ C C2. Write this as DPg(u): Nu -* C. (Recall Dg(u) (a, b) = {dg/dz)(u)(a) + (dg/dw)(u)b if u = (z, *).) The reciprocal of the magnitude of DFglu) is the key ingredient of our notion of condition number. LEMMA 1. Let g E 3Cd be the homogenization off E 9d and £ be a simple root off. Then
\D,g(i, D| = Mh\f'(0\. Proof. We may write/in the form/(z) = ad Ilf (z - £,), Ci = £. Then g(z, w) = ad Il| (z - iiw) and/'(C) = od Hi (z - £,•)• For the proof we may suppose ad = 1 by rescaling.The line N ({1) C C2 is one dimensional, contains the unit vector (1, -{Vllflli ar>d so
£«•'>-<£«•'> n « - w + « n « - o lltll.1
WrKtW-lfi^j^2
1
lid.
(-2
= (l + UI2)|/'(Olj^ji; = IICll.l/'(0|. Before we are able to define our condition number, a norm must be defined on !*£,/.
1423 SHUB AND SMALE
The group of linear automorphisms of C2 which preserve the inner product is called the unitary group, L/(2). Thus if u, v G C2, a G U(2), then (au, av) = (u, v). As in Kostlan (1987), Bez 1, Bez II, we replace the customary norm on %tj, defined by ||g||2 = 2 |a,|2, for g(x, y) = 2 aix'y1'', by a weighted version to make it unitarily invariant. Thus define
M-CEkPft)"')"
sex,
where (?) is the binomial coefficient. Thus ||#||2 = (g, g) where (h, g) = 2 bjOi (f)"1 if h(z) = 2o bJzJwd~J. The unitary group 1/(2) induces an action on Hi by sending g into the composition g « a - 1 for a G (/(2). It can be shown that the Hermitian structure and hence the norm || || on cKd is unitarily invariant in the sense (h° a'1, get'*) = (h, g). If/G 9d, define ||/|| as equal to ||g|| where g is the homogenization of/. Finally, we define the condition number for g E 3f,/, and u G C2 by w
*'
j
|o^(«)|
The d"2 is a convenient normalization factor which eventually makes a certain formula more elegant. Note the following properties of fi(g, «): (i) fifrg, u) = nig, u) all X * 0 G C; (ii) n(g, ku) = fi(g, u) all X * 0 G C; (iii) M ( £ « _ I , ««) = fig, u) all a G 1/(2); (iv) /i(g, M) > 1 (= a>foru a double root of g). (i) insures that y. is defined on P(^Cj), the projective space of lines in lid. (ii) implies that \L is defined for u in AC2) and (iii) that n is invariant under the unitary action on %/ x C2. (iv) implies that its log or loss of precision is non-negative. Let g G Mj, u = (z, w) with g(u) = 0. Differentiating this equation yields u = -/>,*(«)"' (v 2 a,z'w*-'). o ' where u E Nu and the a, can be considered as perturbations in the coeffi cients of g. It follows that (tig, u) reflects the sensitivity of the solution u to an error in g. In fact, using unitary invariance it is not difficult to see that n(g, u) = dm sup ||w||, it = (a0, . . . , ad)
1424 COMPLEXITY OF BEZOUT S THEOREM, III
7
where u is considered as tangent to PiC1) at a and a is tangent to PCXd) a t / with the induced Hermitian structures. In Bez 1 and Bez II, fi(g, z) is denned generally for systems g: C"+l -»• C" of homogeneous polynomials, with properties generalizing (i>—
and so is °° if g has a double root. If/ E 9d, fi(f) = max /i(/, {)• {
/(O-o EXAMPLE. For each d = 1, 2, . . . and fixed a > 0, let a polynomial g e. 3€rf be defined by
gAz, w) = zd - adwd. A zero of gd is £„ = (a, 1). One can use Lemma I to show ,
„,
dm(i + a2)w"2V2(l + aM),/2
By unitary invariance iiigd) - fi(gd, &)• For each a>0, itigd) grows exponentially fast as a function of d. Thus one could say that {fd} is a poorly conditioned family of polynomials. This has the following consequences. Versions of gj are typically used to initiate homotopy methods for solving a polynomial equation, or a polynomial system of equations. See Bez I for the literature on this. The speed of the algorithm can be proved to depend on the conditioning of the homotopy (Bez I). But if the homotopy is poorly conditioned at time = 0, it is certainly poorly conditioned. This raises the question (to which this note is addressed), "does there exist a well-conditioned family {gd)d-\x... gd E VCd, and if so can one find it?" An answer to thefirstproblem is provided by Bez II. Use the standard measure on the projective space of lines through 0 in %*, POCd), normal ized to have total volume 1.
1425
8
SHUB AND SMALE
Then it follows easily from part (ii) of Theorem D of Bez II that for 0 < m < 1 there is a subset Sm C P(Xd) of measure 1 - m such that for g &Sm
In particular there is a subset S of />(%/) of volume 1/2 such that for n(g) < d
all d a 2.
This suggests Main Problem. Find explicitly a family {gj}, gd e 3Crf with (4.gd) £ rf. In fact it seems difficult even tofindsuch gd with (i(g) ^ 4 any fixed q (as ^ = 100). What does it mean "to find explicitly?" There are different levels of interpretation, along the lines of "giving a handy description." More formally, describe a polynomial time machine (as in Blum, Shub, and Smale, 1989) to output such gd as a function of d. The next section will relate the main problem to the Fekete theory of Transfinite Diameter, in its elliptic version in Tsuji (1959). This translates the problem to a kind of packing problem on S2: Find a set of d points on S2, no two of which are very close together. This also brings our problem into good contact with problems of finding electrostatic equilibria on S2, of Tomography, and distributions of points on the sphere as in Lubotzky, Phillips, and Sarnak (1986).
2.
ELLIPTIC FEKETE POLYNOMIALS ARE WELL CONDITIONED
Let S2 be the sphere of radius 1/2 centered at (0, 1/2) in C x R = R\ If z e C, let z be the point in S2 obtained from (z, 1) by stereographic projection from (0,0). See Hille I (1962) for LEMMA 2.
For any Z\, Zi&
C,
"*' " f 2 " = (i + k,|2j,/2(i + W 2 )" 2 f
1
Here || || is the R3 norm.
(z, 1)
' 1 + l*P"
1426 COMPLEXITY OF BEZOUT S THEOREM, 111
9
One may consider S2 as a model of P(C2), and || || a variant of the standard distance in P(C2). In this way, some of this paper extends to n variables. Let Vd be the maximum value of f[
Ik - JCH /
over all Jt,, . . .
,xdES2.
The elliptic Fekete points (of order d) are a (/-tuple (xt,. . . , xj) which realizes this maximum. By compactness, Vj and such U,,. . . , xd) exist. By a rotation of S2, we may suppose no JC, is (0, 0) and so there are z\, . . . , id 6 C such that z, = JC,, / = 1, . . . , d. Then for each d, Fj(z) nf (z - z,) is an elliptic Fekete polynomial. Hide II (1962) has an account of the original Fekete theory and the elliptic version we use here may be found in Tsuji (1959). The sequence 8 = Vj<*"l) is a decreasing function of d with limit as d—► °°, the elliptic transfinite diameter, T.D., of S2 according to Tsuji (1959). This number is XlVe. A main result of the Fekete-Tsuji (see Tsuji, 1959, p. 93) theory is the identification of T.D. with the elliptic capacity, and an elliptic Chebyshev constant, ail proved for the general case of the points constrained to lie in some region of S2. F o r / € tyj with zeros Zi, • • • , zj, define a continuous function/: S2->Uby
fW = fl Ik " tilMoreover, for a zero ( 6 C off, let
AW =
Ik -
With this notation, we will prove in Section 3. PROPOSITION 1.
1/2
"■/««)
S
n^iiz*-^
Here we are using the usual L2 norm for functions on S2. PROPOSITION 2.
M/,o.^ni4 f (n t
1427
10
SHUB AND SMALE
Our main result is an immediate consequence of these two proposi tions. For any polynomial f: C —» C of degree d with zeros z\, . . . , Zd, i one of the z,, THEOREM.
(A) /*(/> = iiif, o < y/d(d + i) vdiuk
3.
PROOF OF PROPOSITIONS 1 AND 2
First we give the proof of Proposition 1. Let/, z,, £ be as in Section 2 with £ = Zj say. First note:
n II** - an = n II** - ziWub*
(*)
*
Thus by maximality of Vd, for z = z,,
/<(*> n iia - an * vd k
and in fact for any z this is true for the same reason. Using (*) again we obtain
MO
1428 COMPLEXITY OF BEZOUT S THEOREM, 111
11
(«W"M"^".
fi(0 Since Vol(S ) = it. This yields Proposition 1. For the preparation of the proof of Proposition 2, we prove some lemmas. 2
LEMMA
3. Let Si = {(z, w) e C2| |z|2 + \w\2 = 1}. Then
f | z |»k|* = 2 ^ m
+ 1)F(/
+ °
/Voo/. As in Bez 11, / j ) | Z | > r = 27r/ / ) l |z|«(l-|zP)' = =
2TT2
J ' (*2)*(1 - s2)1 2sds
, m + on/ + i) rot + / + 2) •
Let (f, g)' = (Ss'fg)112 be the L2 Hermitian structure on 3fd. (/«, #«>' = (/> #)' for a unitary transformation, so we know by the uniqueness of unitanly invariant inner products (see Kostlan, 1987) that (/, g)' and (f, g) differ by a multiplicative constant. LEMMA
4. For f, g GJfj
= 2TT2
w + \w(d - 1 + i) T( + 2)
2TT2
//!( -
+ 1 \ 2ir
!
Q!\
/
2
+ I
1429 12
SHUB AND SMALE
Let 7: Si -* S2 be the composition (z, w) -»((z/w), 1) -+ z/vf if w 4 0 and 7(1, 0) = (0, 0) £ S2, where S 2 is represented as in Section 2. Let z e 9d have zeros zi, . . . , Zj, and g e 2td, be the homogenization of/. Then LEMMA 5.
|*fc. H-)|=/(7(z,H'))n||zJ 1 . Proo/
Use Lemma 2. Then
nikii./(Az,H-)) = n^-dikii, 7
;
BtV
||
V M i (i + l*M2)l/2"yl1'
= n \z - zjw\ = \gu, w)\. j
J
6. The map 7: 5 —► S2 has as fibers, unit circles, and the normal Jacobian of J is 1. (// is the Hopf map.) Thus LEMMA
L *J=2ir L f for a continuous function ip: S2 —* R. The last sentence is essentially the Fubini Theorem or the co-area formula. For the proof first note that 7(z, w) = w(z, w ) e J ! C C x B = R' since z/w - (z/w, 1)/1 + \zlw\2 = (wz, ww) using Lemma 2. Then J(p»a, pieb) = p-»b(piea, p»b) = b(a, b). So 7 respects the S1 action. Differentiating 7 at (a, b) on the tangent vector (u, v) gives v(a, b) + b(u, v). Applying this to the normal to (a, b), (-b, a) gives a(a, b) + b(-b, a) = (a2 - b2, ab + ba). The length squared of this vector is I. ■ LEMMA
7.
With f, g as above $g\\o = \\f\\u V2~7r 11, Mi-
Proof. Use Lemma 5 to see 1/2
ifb - n Mi (/„ to*. "»2) •
1430 COMPLEXITY OF BEZOUT S THEOREM, 111
13
Let us now proceed to the proof of Proposition 2. Let / £ 9d and £ be a zero as in Section 2. From the definition of condition number and Lemma 1, we have
dm\\f\\\\ar ' ° ~ \\Cl\f'(0\ '
(fn M/
By Lemma 2, with z
, Zj the zeros off, and £ = Zj some j ,
\f(o\ = n ic - *i = (n ic - tii) war n NI. = MiM\\rU\\z,h. I*j
Thus W"2|l
Using Lemma 4 this becomes
ML 0 =
Vd(dTT)
\\8y
Now apply Lemma 7 to obtain
/*", 0 =
VSwTljVSr ll/llt.ILWI, ^ Mb HI n„, ik/ii, W(d + 1) II/IL * ,/2 /<(£)'
proving Proposition 2. REFERENCES BLUM, L., SHUB, M., AND SMALE, S. (1989), On a theory of computation and complexity
over the real numbers: NP-completencss, recursive functions and universal machines, Bull. Amer. Math. Soc. 21, 1-46. DEMMEL, J. (1987), On condition numbers and the distance to the nearest ill-posed problem, Numer. Math. 51, 251-289.
1431
14
SHUB AND SMALE
HILLE, E. (1962), "Analytic Function Theory, 1 and II," Ginn, Boston. KOSTLAN, ERIC (1987), Random polynomials and the statistical fundamental theorem of algebra, preprint, Univ. of Hawaii. LUBOTZKY, A., PHILLIPS, R., AND SARNAK, P. (1986), Hecke operators and distributing
points on the sphere I, Comm. Pure Appl. Math. VXXXIX, 149-186. SHUB, M., AND SMALE, S. (1991), Complexity of Bezout's Theorem I: Geometric aspects, J. Amer. Math. Soc, to appear. SHUB, M., AND SMALE, S. (1992), Complexity of Bezout's Theorem II: Volumes and proba bilities, in "Computational Algebraic Geometry," (F. Eyssette and A. Galligo, Eds.) Birkhauser, to appear. TSUJI, M. (1959), "Potential Theory in Modern Function Theory," Maruzen Co., Ltd., Tokyo. WILKINSON, J. (1963), "Rounding Errors in Algebraic Processes," Prentice-Hall, Englewood Cliffs, N.J. WOZNIAKOWSKI, H. (1977), Numerical stability for solving non-linear equations, Numer. Math. Tl, 373-390.
1432 SIAMJ.NUMOLANAL Vol. 33, No. 1. pp. I2S-14S. Fetratfy 1996
© 1996 Society for Industrial and Applied MadMratici
COMPLEXITY OF BEZOUT'S THEOREM IV: PROBABILITY OF SUCCESS; EXTENSIONS* MICHAEL SHUB* AND STEVE SMALE* Abstract. We estimate the probability that a given number of projective Newton steps applied to a linear homotopy of a system of n homogeneous polynomial equations in n + 1 complex variables of fixed degrees will find all the roots of die system. We also extend the framework of our analysis to cover the classical implicit function theorem and revisit the condition number in this context Further complexity theory is developed. Key words. Bezout's theorem, complexity, path following, homotopy methods, integral geometry, unitary group AMS subject classifications. 65,68. S3,58
1. Introduction. 1.1. Bezout's theorem revisited. Let / : C + 1 -*■ C be a system of homogeneous polynomials / = (/i fn), fcgfi=dt, i = l n. The linear space of such/is denoted by H(
where D = max,-(,-), V = [T^, dt and N is the dimension of7iw. Remark 1.1. An approximate zero is denned in Bez I, also see Theorem 1.4 below, without recourse to an arbitrary accuracy c. Newton's method starting at an approximate zero converges quadratically, immediately to an associated actual zero. Remark 1.2. Projective Newton steps are defined in Shub [1993], and in Bez I one can see the detailed full algorithm. Remark 1.3. See Bez I, II, HI for background. This paper may be read largely indepen dently from Bez I, n, m, although it uses some of those results. In particular, Theorem 1.1 uses the main theorem of Bez I and Theorem C of Bez II. Remark 1.4. For the case n = 1, the number of steps is j ^ . Remark 1.5. Renegar [1987a] is an important predecessor to this paper. His result specialized to the case n = 1 has a factor ■£;?■ In Smale [The Fundamental Theorem of Algebra and Complexity Theory, 1981],thereisasimilarresult (n = 1) with a ry^r, but that 'Received by the editors June 1,1993; accepted for publication (in revised form) April 17,1994. ♦IBM, T. J. Watson Research Center, Yorktown Heights, NY 10398-0218 ( s h u b g w a t s o n . ibm. com). This research was supported in part by NSF grants. 'Department of Mathematics, City University of Hong Kong, Tat Owe Avenue. Kowloon, Hong Kong ( m a s r a a l e g c i t y u . e d u . hk). This research was supported in part by NSF grants. ■Bez I. Bez n, Bez m . Bez V refer, respectively, to the four Shub and Smale [1993] and (1994] references.
128
1433 COMPLEXITY OF BEZOUT'S THEOREM IV
129
paper has no theory for systems. There are much better results dealing with n = 1, and no probability er. For example, see Shub and Smale [1985], [1986], Kim and Sutherland [1991], Renegar [987b], Neff [1993], and especially Pan [1987]. Remark 1.6. The constant c in Theorem 1.1 and Remark 1.4 can be estimated from the explicit constants of Bez I and is not very large. Remark 1.7. The proof of Theorem 1.1 is given in §2 below. Remark 1.8. Let us elaborate on a. The space Ttyj is given a unitarily invariant Hermitian inner product which induces a Riemannian metric and probability measure on the associated projective space V(H(d))- Then given a, 0 < a < 1, there is a set in V{H^)) of measure a such that for / in that set, the bound of Theorem 1.1 holds. Remark 1.9. What is g of Theorem 1.1? Our theory asserts the existence of such a g, but it is not known how to find it In fact that is the main problem of Bez III (even for n = 1). Besides the references in Bez III, see also Conway and Sloane [1988], especially §2.3 and the references there. Remark 1.10. For finding one root of / € H(
OG(ao) = —(ao, yo)
_.3F -T-(«O. yo) da
is given explicitly, and so is its norm, the condition number. Example 1. Let Vd = [(ao aj) = a € R"+1) represent the space of polynomials of degree d and define F : Vd x R ->■ R by F(a, y) = £ o a,y' • Then /*(/. r) bounds the infinitesimal change in the solution of / ( f ) = 0 as a function of an infinitesimal change in the coefficients (see e.g. Wilkinson [1963], Demmel [1987], [1988], Bez I, n, m). Example 2. Generalize Example 1 to systems of polynomials / : R" -*■ R" (or / : C" -»■ C"). Example 3. Let T be a linear subspace of Vd over C and F : T x C - > C b e F ( / , r) = / ( f ) . Defining the condition number of these sparse systems is a great convenience; if Vj is replaced by a linear subspace T c Vd, the "sparse case," only infinitesimal changes in T are taken into account in the condition number, say M T ( / . f )• Then nr(.f,r) < /*(/, r ) and for certain T, fir may be much smaller than fi(f, ?). Now see the discussion after the proposition of §1.4.
1434
130
MICHAEL SHUB AND STEVE SMALE
Example 4. The special case of Example 3,T = [f eV \ f(x) = xd - a] defines the condition number for the
x HC"* 1 ) | / ( f ) = 0}.
n(f, f) = ||DC(/)||.
/*(/.?) = Il/ll IID/(?)|^ I A(||f||*- 1 )||. Here Df(r)\N( : N( -*■ C is the derivativerestrictedto ty = {v € C"+1 | (v, r) = 0}, and Afllfll* -1 ) is the diagonal matrix Diag(||f H*-1 Hfll*"1). Compare Bez I, II. +1 Uf€Nfc H(d), f 6 ty c C , then .. i n
H\\T((V{C*'))
11/ II K«)
=
ll/llww' U liellflC"'
hence n may be thought of as a relative condition number as in Wozniakowski [1977]. In general, condition numbers for homogeneous problems will have natural relative condition number interpretations. In Bez I, n, m we normalized n using factors of d/ /2 . That is, M — ( / . S) = Il/ll IDM)\~N) A t f / ^ A d l f ||*-')||. Henceforth we will call that the normalized condition number. (In Bez I, this was called /Xpnj(/, f).) The normalization gave the condition number theorem,restatedbelow, a shorter form.
1435 COMPLEXITY OF BEZOUTS THEOREM IV
131
Now one may describe n in the case of sparse systems of homogeneous polynomials as well, as in Example 3. This permits the sharpening of various results of our previous papers in the sparse case, as will become evident in what follows. Remark 1.13. The condition matrix has an interpretation in economics as the matrix of comparative statics. For example, it dictates how infinitesimal equilibrium allocations change as a function of infinitesimal endowment charges (see e.g. Smale [Global Analysis and Economics, 1981]). In the situation of Example 5, we restate the condition number theorem proved in Bez I and Bez n. Let E' C V (see Example 5) be the ill-posed set, i.e., (/, f) € II' if and only if f is a degenerate root of / (or the derivative D/(f) : C"+I -► C has rank < n). The map n2: V -> V(Cn+l) hasfiberV( over r e P(C"+1), given by vc = {/ € vww)
i no = o) = JTJ-'G).
Thus as a subspace of V(H(d)), V( has an induced metric d(. The condition number theorem gives a formula for //(/, f) in terms of thisfiberdistance. THEOREM 1.2 (condition number theorem (Bez L IL III)). For (/, r) € V C V{U{d)) x 7>(C+1) 11
||AW/ /2 )/llsin^((A(^ /2 )/.f).2:'nV f )'
'
A corollary to the condition number theorem computes /i. (/) = max < /i (/, f) in terms of min d( in the obvious way,
**(/) =
11/11 || A ( < 0 / | | minc s i n 4 ( ( A ( < 0 / , {), £' n V{)
Remark 1.14. If instead of thefiberdistance rff((A(
Next, we consider the condition number for the eigenvector, eigenvalue problem and show how itfitsinto the proceeding picture and how a coresponding condition number theorem holds. For background see Wilkinson [1984], Demmel [1987], [1988]. Let M(n) be the space of all n x n complex matrices and V be the subvariety of Min) x P(C")xC defined by V = {(M, v, X) € M(n)
x V(C)
XC|MD =
kv).
The tangent space to V at (M, v, X) is defined by (M, v,X) € TM(M(n)) x TV(.V(.C)) x C satisfying (just differentiate Mv = Xv) (M - XI)v + (M - XI)v = 0, (v, v) = 0. If M is a regular value of the restriction x\ : V -+ M(n) of the projection, then v and X are each linear functions of M. These functions are the condition matrices, say Ki(M, v, X)M = 0, K2(M, v, X)M = X.
1436 MICHAEL SHUB AND STEVE SMALE
132
Multiplying (A/, A) by a scalar c and leaving vfixedwe see that KdcM, v, cX) = -KdM, v. X), K2(cM, v, cX) = tf2(W, v, X). The Hennitian structure on P ( C ) is the usual one, so for Mi,U2 6 V 1 , ( « 1 . « 2 > =
("1."2>C" {V, V)&
We take trace(A5*) = (A, 5) as a Hennitian structure on M(n). Here 5 ' is the Henniuan transpose of B. The induced norm is the Frobenius norm ||A||2 = Y.ij k»y I2- "H* condition numbers are then defined from the condition matrices as usual, say C,(Af, v, X) = \\Kt(M, v. A.)II, i = 1,2. Define the "ill-posed" variety £ ' by E' = {(A/, u. A) € V | rank(X/ - M)k < n - 1, some integer * > 0}. Consider V
M(n)
P ( C ) x C,
where ;ri and ;r2 are therestrictionsto V of the natural projections from the product space, and E = it\ (£')• On V - TTJ"1 (£), ii\ is an n-fold covering map. ThefiberV„,>, = n^1 0>, A), of JT2 is an affine subspace of M (n) of codimension n. Let dVt>. be induced metric on VV,A. THEOREM 1.3 (second condition number theorem). For (M, v. A) e V (1)
CdM, v, A) = [dv,MM, v, X), £' n V^)]" 1 ,
(
IIAfll2
V'2
WU),rnv,/') • As in the (first) condition number theorem one has an obvious corollary for C,(M) by taking the maximum over (u, X) 6 jrf'(A/). The proof of the second condition number theorem is in §3. The formulas for C\ and C2 first appeared in Wilkinson [1972] and Demmel [1987]. 13. Moore-Penrose, Newton, and complexity. Newton's method can be generalized to search for zeros of maps / : R" -*■ Rm, n > m, using the Moore-Penrose inverse of the derivative (as in Allgower and Georg [1993]). We recall the definition of the Moore-Penrose inverse A* of a surjective linear map A : V -*■ W, where V, W arefinite-dimensionalvector spaces with inner products. A* : W -*■ V is simply the inverse of A restricted to (ker A) 1 . It may also be described as the unique linear map A* : W -*■ V satisfying AA* = /, and A^A is the orthogonal projection onto (ker A) x . Note that Af = A*(AA')-\ A' the adjoint Now we will define Newton's method for / : R" -+• R™ ( / could as well be defined on adoniainofR",orfromC"toC m ). Let AT/ : R" -»• R" be Nf(x) = x- D/OOVto
1437 COMPLEXITY OF BEZOUT'S THEOREM IV
133
and if xo is a given "starting point" in R", x, = Nf (*,_i). Note that Nf is well denned at x provided Df(x) is surjective. Moreover if m = n, Nf is the usual Newton method. If Df(x) is surjective then x is a fixed point of Nf if and only if / ( x ) = 0. PROPOSITION 1.2. Suppose 0 is a regular value of f : R" -► R". For £ e /"'(()), let W{ = {x € R" | Nkf(x) -+ t as k -»• oo}. By Nf we mean the kth iterate of Nf. Then (a) the union over ? € /~'(0) of Wj is a neighborhood of / _ 1 ( 0 ) , (b) W{ intersected with a small neighborhood of /~'(0) is a cell varying continuously in r, and (c) DNf(r) restricted to kerZ)/(f ) x « z*ro. 77»e tangent space ofWf at t is the orthogonal complement f o r f ( / - ' ( 0 ) ) = kerD/(f). This extends the usual basin of attraction theory from the case m—n. Proof. The existence and continuity of W£ are contained in Theorems 5.1 and 5.5 of Hirsch, Pugh, and Shub [1977]. The fact that the union fills a neighborhood follows from a simple degree argument. We will make the size of the small neighborhood more precise in §5, Proposition 5.2. We now give the complexity theory for using this method for finding zeros of analytic / : R" -»• R™, generalizing Smale [1987], Bez I. Define forx € R", 0 ( / , x) = IID/OOVOOII (or oo if Df(x) is not surjective),
,„-£>|
r±r Y(f< *) = max 1 Df(xy ———1| (or oo if Df(x) is not surjective), anda(/,x) = ^(/,x)y(/,x). THEOREM 1.4. There is a universal constant oo approximately | with this property: if f,x = xo, are as above with o ( / , x ) < ao, then all the Newton iterates x\,xi,... are defined, converge to t e R" with / ( f ) = 0, andfor allk>\ (3)
Hxi+1-xt||
Hxj-xoll.
A point xo E R" is called an approximate zero of / if (3) is satisfied. Then r is called the associated zero. The proof of Theorem 1.4 is in §4. Imagine an operation (an ideal operation) which produces from an approximate zero the associated actual zero. This could be done in the Blum-Shub-Smale [1989] model of compu tation with a "6th type of node" for example. Given an approximate zero as in Theorem 1.4, iterating Newton a fixed number of steps gives a desired final accuracy. Since convergence is immediately quadratic and noroundoffis assumed, it is reasonable to assume the exact answer is computed. We will assume in our complexity estimates below that such an operation exists. Justification is based partly on the gain in conceptual simplicity of the results and on simplicity in the arguments. Moreover the results in Bez I give a mathematical justification. Using the robust or theory of Bez I (Theorem 3, §§1-2 and JJ-3), one can bypass the use of this 6th node, at a cost of more technical work. Using these arguments it seems likely that the use of the 6th node here could be eliminated, obtaining estimates with slightly worse constants. Let us see how one can use Theorem 1.4 to get global complexity results. Consider / , : R" -»• R™, y, € Rm, a homotopy and path, respectively, for 0 < t < 1 and let ft € R" satisfy /o(?o) = yo- Define A,j= Observe A,,, = 0.
max a ( / , - - > , s x ) .
1438
134
MICHAEL SHUB AND STEVE SMALE
Hypothesis. Suppose A,,,- < oto whenever \t -t'\ < A = j,ka positive integer. COROLLARY 1.1 (of Theorem 1.4). Let f,,y,,r0beas above and satisfy the hypothesis. Then a number of steps (6th node) sufficient to solve f\ (f i) = y\ is k of the hypothesis. The proof is immediate. Let to = 0, t, = f<_i + A, so <*(/,,, ?,,_,) < c*o. Then the 6th node yields ?,, from ?,,_, starting from r0 with /(?,,) = y,,. One may use Moore-Penrose in place of projective Newton forfindingroots of / 6 Hy). In projective Newton, at z 6 C"+1 one restricts Df(z) : C + 1 -+ C" to the orthogonal space Nz = {v 6 C + I , (v, z) ss 0}. In Moore-Penrose this is simply replaced by the orthogonal space of ker Df(z). For (/, t) e V - £', N( and the orthogonal space of ker Df(r) coincide, but this is not the case in general. The main theorem of Bez I estimates the number of steps needed in following J, from Jo, where /,«,) = 0, and /, is a curve in Wy). The same estimate can be proved for the Moore-Penrose version. THEOREM 1.5. Let F, = (/,, ?,) be a homotopy path in Hy) x C + 1 (so /,(?,) = 0), and fo satisfy /(fo) = 0. Then k = CLD^n1 (the greatest integer in) Moore-Penrose steps are sufficient to produce {,,, r^ f,k with f* = 1 and so /i (f i) = 0. Here, C is a modest universal constant, /x = max, Mno™ (ft, rt) is the normalized condition number and L is the length of the curve /, in V(Hy)). We use the 6th node and our proof (see §4) does use a couple of results from Bez I, but is much shorter with the concepts more clearly exposed than the proof of the main theorem in Bez I. In other words, the number of steps required to follow the homotopy successfully grows as the square of the largest condition number of the zero J, of any polynomial /, along the homotopy path. ConsiderthecomplexityoftheproblemoffbllowingthecurveF~l(0),where F : R"+1 -»• R" has zero as a regular value. Here the algorithm has Moore-Penrose as one ingredient of a predictor-corrector method just as in Allgower and Georg [1993]. THEOREM 1.6. The number of predictor-corrector steps sufficient to follow an arc A of F"'(0), F : R"+I -»• R" as above, is CyL, where L is the length of A, C a constant (not more than 20), and y = maxx€A y(F, x). Theorem 1.6 yields a way of dealing with the problem of producing a complexity theory for zero-finding of real polynomial systems. The difficulty here is that the set of ill-posed problems has codimension 1 so that paths will generally have to meet that set. Now one could use the parameter t of the homotopy for the extra variable. If one wants to follow a zero of /, : R" -»• R \ just define F : R"+1 -»• R" by F(t,x) = f,(x), where R"+1 = R x R \ Generically 0 will a regular value of F even when the path /, meets the set of ill-posed problems, so mat Theorem 1.6 applies. It would be interesting to see this idea carried out to obtain explicit bounds. Theorem 1.6 is proved in §4. The proceeding results extend to Riemanrrian manifolds provided with a computation of the exponential map. The following result has a different version in Theorem 1.8 of §1.4. THEOREM 1.7. Let F : R" ->• R" have zero as a regular value and define y = maxje/-i(0) y(F,z). Then there is a universal constant C so that ifthe distance d(z', F"'(0)) < £, then z' is an approximate zero. For the proof see §4. 1.4. Complexity and condition number. Approximate zeros were defined without ref erence to any actual zero. The corresponding zero was then derived. We now define XQ as an approximate zero of the second kind (as in Smale [1987]) for / : R" -*■ R" provided there is some t g R \ f(r) = 0, and
1439
COMPLEXITY OF BEZOUTS THEOREM IV
\\*k-t;\\<(\)
ll*o-fll.
135
* = i.2
** = x*_i - D/ta-ir'/ta-i). THEOREM
1.8 (Smale 1986). L« / : R" -»• R", x, f e R\ /(?) = 0, and 3-V7 ll*-flly(/.f)<—2^-
77ien x is an approximate zero of the second kind. We will suppose that our "6th node" has the power of producing the actual zero from an approximate zero of the second kind. Now consider the setting of the implicit function theorem F : R* x R" ->■ R", where F may also be defined on some domain or over C Suppose that t -»• (a(t), f ( 0 ) € R* x R", with F(a(f), {(/)) = 0, is a curve, 0 < t < 1, such that |^(a(/), f (f)) is nonsingular for all t. The idea is that a{t) is given explicitly (the input of a problem) and f (f) is given implicitly. Suppose that f (0) is also given and that we want to find ?(1). This is a general setting for path-following algorithms. For our algorithm and complexity result, it is convenient to write ?, = £(») and F,(x) = F(a(r), x), so that F,(f,) = 0,0 < / < 1. The algorithm (slightly idealized) depends on a partition to = 0, t, < ti+l, tt = \, i = 0 , 1 , . . . , k, and the condition is that ?,, is an approximate zero of the second kind for F,(+l and £,,„.,. Thus it "6th node" operations are sufficient to produce f i and so A: is the main ingredient in the complexity. We may estimate k thus by Theorem 1.8. Accordingly, the required condition to implement the above procedure is
Let A ; + 1 = ||f„+l - f,,||, r. = y(.F,M,t,^).
(4)
So
* = Ef;= c E A ^
where c = T^-J-, is sufficient. This yields the following theorem. THEOREM 1.9. Suppose that F,(?,) is as described above, y = max, y(F,, {,), and L is the length of the curve t -*■ f,. Then given fo. a number of steps ("6th node") sufficient to reach f « TZTRLY-
Use (4) and that Yl A/K < y J3 A, < yL. The number of Newton steps without using the 6th node could probably be estimated at about three times the above, using the robust a theory of Bez I. Then one would obtain an approximate zero of f\ instead of f i. Theorem 1.9, while dealing with an idealized algorithm, is nice because of its extreme simplicity in statement and proof. It displays the main complexity ingredients. The condition (4) is sharper. For example if y(F„ ?,) is monotone, the complexity is bounded by/p'llf/llyr^. In the main theorem of Bez I, the condition number fi (/, f) rums up as the main ingredient. It is quite natural to ask why, since (i(f ?) is an infinitesimal invariant reflecting other aspects of computation. We are now in a position to deal with that question.
1440 136
MICHAEL SHUB AND STEVE SMALE
Consider the environment of the previous theorem. One of the two complexity ingredients is L, the length of the curve £,. This curve is only implicitly given, and hence L is also. It would be preferable to replace L by a more direct invariant of the input curve a(t). Here F,(;t) = F(a(f), ?(/)) = F(.a(t), G(a(/))), and G : U -> »" is defined on a{t). Let La be the length of thecurvea(f). Recall that the condition number n (a, ?). w i m G(a) = f at (a, f), is ||DG(a)|| and so the condition number n of (a(f), f (0) is max, /x(a(f), f (f)). PROPOSITION 1.3. L <
/xiV
Proof.
L= f Him = /" \\G(a(t))'\\dt Jo
Jo
= /" ||£>G(a(0)a'(f)||df < [ l|Z>G(a(r))|| ||a'(0ll^ Jo Jo : M / \\a(t)\\dt = (iLtt Jo Proposition 1.3 and Theorem 1.9 yield the estimate 2 (5)
complexity k <
=y/ila. 3-V7 The last results give some way for taking advantage of sparsity in complexity estimates. In Bez I we found that for full systems, i.e., allowing all coefficients to be nonzero for systems of polynomials / : C" -*■ C" of given degree {d\ dH), the condition number it squared was decisive. One factor of fi1 came from an estimate on y. The second factor of fi could be interpreted as the tt in (5). But now in the theory just preceding, the sparse case leads to the condition number fi of (S) which could be much smaller than the n of Bez I. We end section 1.4 with a couple of comments. (1) The condition matrix itself leads to an algorithm of predictor-corrector type, and its complexity may be estimated as above. (2) Results here may be extended to maps of Riemannian manifolds using the exponential map extensively. Remark 1.15. The unitarily invariant norm on H^) used in Bez I, II, IE, and this paper is described in detail by Weyl [1932] with his focus explicitly on unitary invariance. Stein and Weiss [1971] use the same norm (and corresponding inner product), but don't discuss its unitary invariance. As mentioned in Bez I, Eric Kostlan brought this approach to our attention. 2. Proof of Theorem 1.1. The proof uses the main results of Bez I and Bez C. First we follow Bez I to obtain Theorem 2.1. We use some of the notation of Bez I. For example, we write AH>n>j(/< O for Mnorm(/. f )> ^ normalized condition number, and extend the definition to (/, f), where / ( f ) is not necessarily zero by the formula of §1, Example 5. Let B p (/, s) be the dp ball of radius s around / € "H^) - {0}, where dp is the chordal metric BP(f, s) = {g e Hw | g # 0 and <*,(/, g) < s}. THEOREM 2.1. Given Ci > 1, 3Ci > 1 with the following property: if g e 7 i w , g(x) = 0, and fijmjig, x) < oo, then x may be continued to a zero x{f)for all f in Bp(g, s), where s =
1 C,D»/*Mkf(j.x)
1441 COMPLEXITY OF BEZOTJTS THEOREM IV
137
and / ■ W / , * ( / ) ) < C2/iproj(«, x).
Proof. We prove three preliminary propositions and a lemma. Here we recall Proposition 5(b) of Bez I, §1-3. PROPOSITION 2.1. Let /, g € H(d). Then
as long as the denominator remains positive. LEMMA 2.1. Let f,g e W((i). L«r $(x) = 0. 77I«I (a) &(/.*)<
(b)>t)(/,*)
(c)«(/.x) < i M ^ j ( / ^ ) ^ ( / . g ) £ » 3 / 2 .
Proof. From Propositions 2,3, and Lemma 1 of §1-3 of Bez I A>(/. *) < * W / . *)«»(/. *) < / W / , x)dp(f g), which proves (a), (b) is Proposition 3 of Bez I §1-3 and (c) results from multiplying (a) and(b). PROPOSITION 2.2. There is a constant K\ > 0 such that, if g € H(d) and g(x) = 0, then x may be continued to a zero x(f) of f for all f € Bp(g, s), where s = Ki//j^(g, x)D3/2. Moreover given constants K2, K3, K+, K$ > 0, K\ may be chosen small enough such that (a) /v>j(/. * ) < 0 + * 2 ) / W * - *)> (b) ^igpsl < 2K3/nvmj(g,x)D3'2, (.d) Mfx) < K3/nTni(g,x)D3'2, (e) » ( / . x) < ^ / W * . *)*>3/2Froo/. By Proposition 2.1 (a)
lW/
'* ) -
.
A
/ \
< MmiCff. ^) ( f r f j - ) ^ 'V>J^. *>0 + *a) for ^1 small enough (d) /%(/. x) < Mproj(/, x)dp(f, g) by Lemma 2.1(a) _/V»j(*^)^1_is:JM2n).(g;t)Z)3/2
for /kj small enough.
1442 MICHAEL SHUB AND STEVE SMALE
138 (e)
H>(/. x) < -^p«.j(/, x)DV2 by Lemma 2.1(b)
< \««
({±£W,||x|| Thus by Smale [1986] (referred to hereafter as P. E.), or Bez I,
2(^)^,1x11 l|x(/)-x||< Hw(g,x)DW (c)
(\wn\D-1 < /B*(/)-XII
y D-l
* (*£?♦•) <«»(!$)* choose K\ sufficiently small so that e \T=r^i ' < 1 + K4. (f) llx(/)-x|| 2K3 fl + K2\
._ 3 / 2
= tf3(l + K2) by (b) and (e) and we may choose £3, K2 sufficiently small so that £3(1 + #2) < KsPROPOSmON 2.3. Let f e W(
w/iew r0 = lffi,
u = r0yo(/ | Nx, x), a/uf (l+r|)'/»
1 - ro (§E$ + D<W/. *)»(/• *)) W as /ong a s O < « < 1 — 2 positive.
am
* r° 's
sma
^ enough so that the denominator of' K remains
1443 COMPLEXITY OF BEZOUT'S THEOREM IV
139
Proof. / W / . y ) = ll/ll \\(Df, I Nully)-1(A(
x||(D/, |NuH t r l A(
«)\22
< K-
\\M\)
/||.,||\0-1 ,, JM\~ '
n
by Proposition 1 of Bez I, §111-2, Lemma 3(2) of Bez I, §11-1 (initially from P.E.), the definition of Aiproj. and the fact that the norm of a diagonal matrix is the largest norm of its entries. Proof of Theorem 2.1. Choose K2, K3, K*, K5, K6, Kn, K% > 0 so that (1 + K2){\ + tf4)(l + * 6 )(1 + * 7 ) < C2> (1 + KlVft/V - K3((-0^)Kt + l ) ^ j | | f ) < (1 + Kn), ^ £ | f <1 + K6,and0
< 1 + K6,
lK")
|7-T
< 1 + *4,
11*11
and A W / , x) < (1 + tf2)/v>j(£. *)• Thus /V»j(/> * ( / » <(> + *7><1 + «6)d + Ktin^ii, x)(\ + AT«) D < C2/V)j(s. *)• COROLLARY 2.1 (of Theorem 2.1). Let L be a great circle in Hyj and Ne the special neighborhood ofL in H^ as in Bez I. Then there is a universal constant c so that ifL n Np ^ 0, then Vol(L l~l N2fi) > jj&. The "Vol" is the measure of a subset of the circle and a great circle is just the intersection of a real 2-dimensional linear subspace with unit sphere. By Theorem 2.1, there is a constant C\ such that n(f) = (i(f, y,) < 2/j.(g, xt) < 2n(g) for all / with dp(f, g) < c,lli{lniDi/i- Therefore, if n(g) < ± then n(f) < ± for all / such that
d,(f,g)<
1 « ( $ » »
V ClD3/2
'
Now let / € L n Np, i.e., /*(/) > ± and g € L D (N2 -£fcm. Hence if / € L n Np, then the interval of ^ length 4 c p3/i around / in L is contained in N2p. Since rfp = sind* this interval has Riemannian length greater than C4^i/S and so its "volume" is greater than jfe. D Let 5 be the unit sphere in Euclidean space E of some dimension and £ be the space of great circles of 5. The orthogonal group O of E acts isometrically on the product S x C by {x, L) -»• {Ox, OL) for O € O. The subspace V = {(x, L ) e J x £ | x e t ) is invariant under this action, and O acts transitively on V. For* € S, let £, = {L e C \ x € Z.).
1444 140
MICHAEL SHUB AND STEVE SMALE
2.4. LetU c S,W cCbe open sets. Then (a) there isanx € 5 such that PROPOSITION
\o\{Cx n W) VolW Vol Cx ~ Vol£' (b)/ £££ Vol(WnL) = 2 i £ & ^ . Here L0 is just a standard great circle and Voll0 = 2 jr. For the proof of the proposition let ;n : V -* S, 7t2 '• V -*■ C be the restrictions of the projections and W = jrf'W, W = n^W. First we prove Proposition 2.4(b). Let xo € 5 and £o = £x»- Let W-^i and iV/2 be the nonnal Jacobians of the maps TT\ and 7t2, respectively (from the coarea formula as in Bez II). These are constants by the orthogonal invanance. We use X to denote the characteristic function. LEMMA 2.2.
f I
XQA'
n jrf 1 !) = (jfy
Vol(W)Vol(£o).
Proof. Jc h-'L Nh
Jv
= I 4T [ W 7s "^l Jvx
n V,) = -^-Vol(W)Vol(£o), NJ\
where V, is the fiber over x of 7t\, proving the lemma. In the lemma consider the special case U = 5, W = V to obtain that NJ2 _ VolQC)Vol(Lo) # / , ~ Vol(£o)Vol(5)' Observe that /"
XQX n JT,-1 L) = voi(W n L).
Putting this and our evaluation of ^f into the lemma yields Proposition 2.4(b). Now for part (a) of the proposition: the coarea formula gives
Lxm=JsikLx(w'n^
and
Lxm=L Thus
w2 f
1445 COMPLEXITY OF BEZOUTS THEOREM IV
141
by the computation of ffi above. Therefore Vol W
minxes fv, X{W n Vx)
Vol£ > -
Vol£o _ Vol (£ x fl W) Vol £ x
for some particular x. This proves Proposition 2.4(a). We now apply Proposition 2.4 to the case where E = H^), with £, S defined for this specialization. Then for g € 5 (i.e., g € W^, || g II = 1), Cg is the space of all great circles in H(d) containing g. Define £ , . , = [L € £ , | L n iV„ # 0}. THEOREM 2.2. 77icre « a constant c > 0 5«c/i that for each d = (d\ g € Hy) of norm 1 and ^ ° 1 , £ ; " < cp2n2(n + 1)(AT - l)(tf -
VolC,
d„), fAere is a
2)D3/2V.
For the proof we use Proposition 2.4 taking U = N2p and W = {L e £ | L n N^ ^ 0}. Using Corollary 2.1, Vol W - ^ < /" Vol(L n N2„) < /" Vol(L n N2p), which by Proposition 2.4(b) is Vol £ Vol LpVol N2P Vol 5 By Proposition 2.4(a) there is x = g so Vol(£ x fl WQ < / c / P 2 \ " ' Vol LpVol AT2p Vol£ x ~ VD3/2/ Vol 5 Now use Theorem C of Bez IL It yields (for/i > 1) ^ £ ^ < 4 p V ( « + l)(N - 1)(JV - 2)P. Vol 5 Joining this with the previous estimate yields Theorem 2.2 (note that £ x n W = £p,,). The case of n = 1 is implied by Theorem D of Bez II. Proof of Theorem 1.1. Fix g as in Theorem 2.2. Given / € P(H{d))* the homotopy 0 < t < 1, /, = tf + (1 - t)g has length less than or equal to one. Let p{f, g) = sup p {/,} n Nfi = 0. Then by the main theorem of Bez I, 3 ( y » projective Newton steps are sufficient to find all the approximate zeros of / , with c2 around 10. Now by Theorem 2.2 the probability that />(/, g) > s is greater than or equal to the probability that the great circle through / n N, = 0 > 1 - cs2n2(n + 1)(JV - \)(N - 7)DV2V. Thus the probability that &£— steps suffice is greater than or equal to 1 — cs2n2(ri + \)(N - 1)(N - DD^V and setting t = cs2n2(n + \)(N - \)(N - 2)D3'2V we find that with probability 1 - 1 , vt
1446 MICHAEL SHUB AND STEVE SMALE
142
3. Proof of the second condition number theorem. Note that VvA. is the set of A/having i; as a A eigenvector and that if M € £' n VvX, then X is a multiple eigenvalue. The unitary group U(n) acts on P(Cn) in the natural way and acts on M(n) by sending A -+ UAU~l for U € U(n). Moreover if (A/, v, A) € V, Mv = kv, so UMU~x(Uv) = k{Uv). Thus V is invariant under the product action of U(n). Since n\ : V -*■ M(n), m : V -* P(Cn) x C both commute with the action of U(n), it follows that Kt(M, v, k), i = 1,2 are also invariant under the action of U(n). Thus Kt(M,v,k),i = 1,2 only depends on the linear map M \ v1, where v1 is the Hermitian complement of u in C , and not on a particular basis. Let M = JT„IM | vx and M\ = nvM \ v1, where nv± and nv are the orthogonal projections of C" onto v x and v, respectively. Since U(n) acts transitively on P(C) we may assume for proof that i; = (1,0,..., 0) and thus that (M, v, k) e Mv.x has the form M = (o i*')- *" 8 e n e r a l w e assume ||v|| = 1. LEMMA 3.1. (a) Ki(M, v, k)M = (A/„_i - ^ r ' ^ M W , (b)K2(M,u,A)Af = ^ , where y satisfies (M* — kl)y = 0, M* the adjoint ofM. Proof. The equations (6)
(kl - M)v + (kl - M)i> = 0,
(7)
(v, v) = 0
define (A/, v, A) e r (M . vA) V c r w (X(n)) x TV(P(C)) x C. Applying jr„x to (6) gives Jtv+M(v) = Jtv±(kl — M)v, 1
and since v € v by (7) W„J.M(U) = (A./ — A/)0 proving (a). For (b) note that (y, (kl - M)u) = 0 for all u € C . Take the inner product of (6) with y to obtain (y, (A/ — M)v) = 0 so that A=H: L = K2(M, v,k). {y.v) LEMMA 3.2. (a) ||jr;,(M,w.A)|| = ||(A/B_, - M)-% (b) \\K2(Miv, A)|| < (1 + ||Af||2ll(H.-i - Wr'H 2 ) 1 / 2 . Proof- Assume M = (£ £'), w = (1,0 0) and write M = (^ ^ )• The11 f o r (a)> ii = 0 for M of the form (^ ^') since the eigenvector v is constant for these perturbations. Thus K\{M\, v, k)M factors through projections on M2 and ||Af2|| = ||7r„j.Af2(t;)||. For (b) y may be taken as (8)
y =v+
(kln.,-M)-l(\f'-kI)v
for Image(M • - kl) c Ker(M - A/) x c vx. Applying (A/* - kl) to both sides of (8) gives (AT - A/)y = (AT - kl)v + (AT - A/)(A/„_, - A/)*"1 (AT - kl)v = (AT - kl)v - (A/* - kl)v = 0, using (AT - A/)(A/„_i - A/)- 1 = -/„_,. Now from (8)
llyll2 < l + ||(A/ - J O - Y I I W - *')f II2 l l + IKAZ-A/r-'fHA/rvll2 < l + ||(A/-Af)|| 2 ||A/|| 2
1447 COMPLEXITY OF BEZOUTS THEOREM IV
143
using ||A*|| = ||A||. Also from (8) it follows that {y, v) = (v, v) = 1. Thus \\K2{M,v,X)M\\ =
(y. Mv) < \\y\\ \\M\\ (y,v)
using Lemma 3.1, Cauchy-Schwarz, and the above. The estimate of ||>'||2finishesthe proof of(b). For the proof of the theorem we also require the following proposition. _1 PROPOSITION 3.1 (Eckart and Young [1936]). ||A || = 3^755, where A € M{n), S C M(n) is the set of singular matrices, and df is the Frobenius distance. Proof of Theorem 1.3. Assume M = (* ^ ' ) , v = (1,0 0). Then it is simple to see that the closest matrix in V„,A n E' is N = (* ^'), where N is the closest matrix in M(n -1) to M such that A./„_i — N is singular. That is, dv,MM, v, X), E' n Vv.k) = dF(M, N) =
dF{XIn.,-M,S)
= iia/B.1-M)-1ir1 by Proposition 3.1. Here S is the set of singular (« — 1) x (n — 1) matrices. Now Lemma 3.2 finishes the proof of the theorem. 4. The proofs for §13. The proof of Theorem 1.4 is adapted from Smale [1986] (P.E.). LEMMA 4.1. Let f : R" -* RM, z,z' € R", and u = \\z' - zl|y(/,z). Let V, =
kerD/Ot)-1- andnx : W -+ Vx be the orthogonal projection. Suppose u < 1 — ^ . Then \\Df{z)^Df(z')-nz\\<(-^\
(9)
-1<1,
(10)
l|D/(zK'ZV(z)M < - ^ T r " i/(u) where ir(u) = 2u2 — 4« + 1, a-") 2 (11) II D/(z')fD/(z)II < Vr(u)
Proof. For (9) expand Df(z') by the Taylor series Dfiz')
= Df(z)
+
j:^(z'-zrK
Apply Z)/(z) t , noting D/(z) f D/(z) = 71z to obtain ||D/(z)tD/(z') - JTJI < £ * < / - ' - 1 = ( - j - ^ - ) - 1 (compare P.E.). Next note that IIO/(z)|;.'D/(z')lv; - h.\\ <
\\Df(z^Df(z')-7tz\\,
1448 MICHAEL SHUB AND STEVE SMALE
144
so as in P.E., (10) follows. Here Iv. : Vz -*■ Vc is the identity. Finally (11) is a consequence of the fact that for any surjective linear A : Rn+1 -*■ R" and hypeiplane V c R"+1, v € R\ l | A | ^ A i w | | < IAIV'WII.
Using the lemma, the proof of Theorem 1.4 follows just as in P.E. We now prove Theorem 1.5. This uses Moore-Penrose (i.e., Newton's method extended by Moore-Penrose) to follow a curve (/,, f,) in W(rf) x C"+1 with /,(?,) = 0, using x'M = Nf fa). Using the 6th node we may take *, = f,,, each i = 0,1 k provided that a(fi, &,) < a0 for f, < / < ti+\. By Bez I, especially the higher derivative estimate, ,i?Z> 3 / 2
«(/,. ?») < /*«»»(/„ U) - L 2~• where r) < dp(ft, ftl) = dp is as in Bez I and dp(f, g) = sind(/, g), d(f, g) the Riemannian distance in P(H^). Use Proposition 5 in Bez I, §1-3 to obtain ,t
*\*
Promt ftn fft)0 + 4>)
^ Jt(l+^) -l-D>/2
, M = maXAtnonn(/f.?f)-
r
To apply Theorem 1.4, we thus need Hd+dp)
\2 DW
^\-D**dpn)
2
dp < Cfo.
There is a universal constant c which makes dp < . ^ sufficient Thus we obtain the complexity * = ^Cl^^L. D Proof of Theorem 1.6. Let zero be a regular value of F : R"+1 -» R \ u € F _1 (0), u € r„(F-1(0)),and ||u|| = 1. Suppose that (12)
or(F, u + hu) < a0.
Then the predictor-corrector algorithm with the 6th node produces u ->■ u + hu -► u 6 F"l(0), where u' is the zero associated to the approximate zero u + hu. The estimate (12) has two parts, the estimate for y(F, u + hu) and for fi(F, u + hit). One uses Lemma 2c of P.E., extended by Moore-Penrose, to see that y{F, u + hu) < cy, for hy < C\ -1
andforalli/€AcF (0). The Taylor expansion of F about u yields F(«+A«) = 2 ^
71
since thefirsttwo terms are zero. Compose with DF(u + hu)* to obtain p(r, u + nu) < 2^ 75
•
1449 COMPLEXITY OF BEZOUT'S THEOREM IV
145
Here we used DF(u)DF(u)1 = /. Equation (11) of Lemma 4.1 applies to yield \\DF(u + hu)^DF(u)\\ <
(1 " * y . ) , hy < 1 - ^ . 1r(hy) 2
Thus we obtainfi(F,u + hu) < ch2y and so a(F, u + hu) < c{hy)2 for hy < c'. Thus (12) is satisfied provided that hy is less than an easily estimated constant. Tofinishthe proof of Theorem 1.6 we need to relate A to the complexity. It is sufficient to show that || H' - «|| > ch where u' e F~' (0) is obtained by the predictor-corrector step, and c, as usual, is a new universal constant In factitissufficientto show that ||(u+hii)-u'|| < c\h,c\ a small universal constant Since w+/tuis an approximate zero, \\u+hu—u'\\ < 20(F,u+hu) and (f(F, u + hu) < c{hy)h. But hy can be assumed to be universally small as above. D Proof of Theorem 1.7. LEMMA 4.2. With the setting of the theorem, let u = \\z - z'\\y(F, z), where F(z) = 0. Thenif*{u)>0,a(F,z')<^. Proof. Use Proposition 2 of Bez I, §0-1 extended to Moore-Penrose as before. Then
a(F,z')< (1 -t^ Z) + "Vr(«)2
But a(F, z) = 0 since F(z) = 0, proving the lemma. There is a constant c such that if u < c, jfo
\\DNf(y)- | , W I < < 1 ^ ( - ! _ - , ) +
2« ^(«) 2
as long as V(«) > 0, i.e., u < 1 - &. Proof. DNf(y) - DNj(x) = D(Dp o /)(y) - D(Df* o /)(*) = ZHD/OOVOO - Df{y?Df{y) - D(D/(y)t)/(x) + D/(jt)tZ>/W = Z W O O V t v ) + Df{x)^Df(x) A
B
We will prove in the lemmas below that (13)
l|A|| <
2u lK«)2
and
(M)
hB|l
'
,).
Df{y)*Df{y).
1450 146 146
MICHAEL SHUB AND STEVE SMALE
Recall that Df(x)^Df(x) is an orthogonal projection on ker Df(x)x and Df(y)^Df(y) is an orthogonal projection on kerD/(>>)J-. First we estimate the norm of a : ktr Df(x) -*■ kerD/X*)-1 such that ker Df(y) = graph(or) = {(V,
0:ktvDf-*H
the zero map and Df(y) = (D\, Dz) in these coordinates. LEMMA 5.1. |M| < ^ ( ^ - 1). Proof, a = -D;lDi
so ||
(10)ofLemma4.1,and flC-!D,|| = Wto* £ ? * / ( g V ' " ' U ^ Tnb " 1LEMMA5.2. I « £ c HixHibegivenasthegraphofo : #i -* # 2 . then\\nE-nHlII < ||ff||, where JT£ anditn are orthogonal projections on E and H, respectively. Proof. Let A : V -*• Hi x Hi. Then orthogonal projection (Image A) is given by A{A*A)~X A*, A* the adjoint of A. Thus *£
_ / 7tHl -y
(/+
+
-/
CT.a)-i
(/ + ffV)-'ff 1 ffv)- ff'
(7 (/ +
and the norm of the matrix as an operator is easily seen to be less than or equal to |CT ||. Lemmas 5.1 and 5.2 prove (14) using that UT£J. - w^ll = ||(/ - nE) - (/ - 7rW|)|| = ll*£ - nHl II. Now we turn to (13) D{Df{y)<)f{y) = -DfHy)(D2f(y)Df(y)' + D2f{yT){Df(y)Df{yr)-xf{y) + D2f(y)'(Df(y)Df(yrrlf(,y) t = -Z)/ (y)(P 2 /(y)D/'Cy))(D/(y)D/(y)*)- 1 /( > ) Df\y)Df{y))(D1f(ynDf{y)Dny)')-lf{y)).
+ (/ C
H
G is a projection so || G || = 1. We will prove || £|| < - ^ and ||tf|| < jfc. LEMMA 5.3. ||£|| <<*(/, y). Proof. \\E\\ < \\Df(y^D2f(y)\\\\Df'(Df(y)Df(y)'r1f(yn < H / . y ) W , y ) = «(/,y). The notation D 2 /O0* has the following interpretation: D2f{y)(u,) for fixed u is linear. EPfiyY means the adjoint of this linear map. LEMMA 5.4. ||D 2 /(y)*(0/(y)D/(y)')- I /(y)|| < a ( / , y ) . Proof. \\D2f(ynDf(y)Df(yyrlf(y)\\<\\D2f(ynDf(y)Df(y)')-lDf(y)\\ x||D/(y)*(D/(y)D/(yr)-'/(y)|| using that D/(y)Z)/(y)*(D/(y)D/(y)*)- 1 = Id < y ( / > ) £ ( / . y) = <*(/,y) using that ||A*|| = ||A|| for linear A. D Now the proof of Theorem 1.7 gives a(/, y) < ^ y , proving (13) and Proposition 5.1.
1451
COMPLEXITY OF BEZOUTS THEOREM IV
147
We make part of Proposition 1.2 more precise for analytic / : E -»• F, where E, F, are Hilbert spaces. Suppose 0 is a regular value for / and r e /~ l (0). Write E = kerD/(f) © kerD/(f )■"-. Let S, = j + {(x, y) € ker £>/(?) © ker JD/($) X | ||x|| > ||y||}. Let £ = ? + {(x.y) e kerD/(C)®kerD/($) x I II*|| < lly||}. Let f = (ft, ft) with respect to ker £>/(£) © ker Df($)L, and let Br(x) denote the ball of radius r centered at x. PROPOSITION 5.2. There is a universal constant c > 0 such that for analytic f and f as above and y = y (/, r),
(a) v . /-'(O) n a*(o = nB>0^/(5. n »j(0). (b) V is the graph ofC function av : Bi(Ji) -► kerD/(?) x am/ ||DCT„(?I + x)|| < 3||x||y, («) " & * = {(*.y) e **(f) I N}ix,y) e «*({) Vn > 0 and Nnf(x,y) -+ r as (d) W'c5loc ij die graph of a C1 function
<jw:BL(r2)^knDf(r), Dow(r2)=0and
sup HZ>
Proo/. To see V as a graph over B*(?i) restrict / to (f) + x) x kerD/(O x and apply formula (10) of Lemma 4.1 to deduce that a ( / | (ft + x) x kerD/(f ) x , fi + x) < jfo. Now choose u small enough so that this quantity is less than ao. NOW Theorem 5.1 of Hirsch, Pugh, and Shub [1977], the use of a bump function, and remarks on center manifolds finishes the rest of the proof. The estimate of || Dav (f i + x) || follows from Lemma 1. Note added in proof. Remark 1.1 is accomplished in Bez V. REFERENCES E. ALLGOWER AND K. GEORC, Continuation and path following, Acta Numerica (1993), pp. 1-64. L. BLUM, M. SHUB, AND S. SMALE, On a theory of computation and complexity over the real numbers: NPcompteteness recursive functions and universal machines. Bull. Amer. Math. Soc., 21 (1989), pp. 1-46. J. CONWAY AND N. SLOANE, Sphere Packings. Lattices and Groups. Springer. New York, 1988. J. DEMMEL, On condition numbers and the distance to the nearest ill-posed problem, Numer. Math., 51 (1987), pp. 251-289. , The probability that a numerical analysis problem is difficult. Math. Comp., 50 (1988), pp. 449-480. C. ECKART AND G. YOUNO, The approximation of one matrix another of lower rank, Psychometrika, 1 (1936), pp. 211-218. A. EDELMAN, E. KOSTLAN, AND M. SHUB, HOW Many Eigenvalues of a Random Matrix Are Real?, J. Amer. Math. Soc., 7 (1994), pp. 247-267. M. HIRSCH. C. PUGH, AND M. SHUB, Invariant Manifolds, Springer Lecture Notes # 583, Springer, New York. 1977. M. KIM AND S. SUTHERLAND, Families of Parallel Root-finding Algorithms, SUNY Stonybrook. Institute of Mathe matical Science, Stonybrook, NY, 1991, preprint. A. NEFF, Specified precision polynomial root isolation is in NC, J. Computer System Sci., 48 (1994), pp. 429—463. V. PAN, Sequential and parallel complexity of approximate evaluation and polynomial zeros. Comput. Math. Appl., 14 (1987), pp. 591-622. J. RENEGAR, On the efficiency of Newton's Method in approximating all zeros of systems of complex polynomials, Math. Oper. Res., 12 (1987a). , On the worst case arithmetic complexity of approximating zeros of polynomials. J. Complexity. 3 (1987b), pp. 90-113. M. SHUB, Some remarks on Bezout's Theorem and complexity theory, in From Topology to Computation, M. Hirsch, J. Marsden, and M. Shub, eds.. Proceedings of the SmalefesL 1993, to appear. M. SHUB AND S. SMALE, Computational complexity: On the geometry of polynomials and a theory of cost, I, Ann. Sci. Ecole Norm. Sup. (4), 18 (1985). pp. 107-142. , Computational complexity: On the geometry of polynomials and a theory of cost U, SIAM J. Comput., 15 (1986), pp. 145-161. , Complexity of Bezout's Theorem I: Geometric aspects. J. Amer. Math. Soc., 6 (1993), pp. 459-501.
1452
148
MICHAEL SHUB AND STEVE SMALE
M. SHUB AND S. SMALE, Complexity ofBezout's Theorem II: Volumes and probabilities, in Computational Algebraic Geometry, F. Eyssene and A. Galligo, eds.. Progress in Mathematics, Vol. 109, Birichauser, Cambridge, MA, 1993, pp. 2*7-285. , Complexity ofBezout's Theorem UJ: Condition number and packing, J. Complexity, 9 (1993), pp. 4-14. , Complexity ofBezout's Theorem V: Polynomial Time, Theoret. Comput Sci., 133 (1994), pp. 141-164. S. SMALE, Global analysis and economics, in Handbook of Math. Economics, Vol. 1, Ch. 8, K. Arrow and M. Intrilligator, eds., North-HoUaad, New York, 1981. , The fundamental theorem of algebra and complexity theory. Bull. Airier. Math. Soc. (N.S.), 4 (1981). pp. 1-36. , Newton's Method estimates from data at one point, in The Merging of Disciplines: New Directions in Pure, Applied and Computation Math, R. Ewing, K. Gross, and C. Martin, eds.. Springer, New York, 1986. , Algorithmsfor solving equations. Proceedings of the International Congress of Mathematics, Berkeley 1986, American Mathematical Society, Providence, RI. 1987. E STEIN AND C. WEISS, Introduction to Fourier Analysis on Euclidean Spaces, Princeton University, Press, Princeton, NJ, 1971. H. WEYL, The Theory of Groups and Quantum Mechanics, Dover, New York, 1932. I. WILKINSON, Rounding Errors in Algebraic Processes, Prentice-Han. Englewood Cliffs, NJ, 1963. , Sensitivity of Eigenvalues 25, Utilities Mathematica, 1984, pp. 5-76. , Note on matrices with a very ill-conditioned eigenproNem, Numer. Math., 19 (1972), pp. 197-198. H. WOZNIAKOWSKL Numerical stability for solving non-linear equations, Numer. Math., 27 (1977), pp. 373-390.
1453 Theoretical Computer Science 133 (1994) 141-164 Elsevier
141
Complexity of Bezout's theorem V: Polynomial time M.Shub* IBM TJ. Watson Research Center. P.O. Box 218. Yorktown Heights. NY 10598. USA
S. Smale* Department of Mathematics, University of California. Berkeley. Berkeley. CA 94720. USA
Abstract Shub, M. and S. Smale, Complexity of Bezout's theorem V: Polynomial time, Theoretical Computer Science 133 (1994) 141-164. We show that there are algorithms which find an approximate zero of a system of polynomial equations and which function in polynomial time on the average. The number of arithmetic operations is cN", where N is the input size and c a universal constant.
1. Introduction The main goal of this paper is to show that the problem of finding approximately a zero of a polynomial system of equations can be solved in polynomial time, on the average. The number of arithmetic operations is bounded by cN*, where N is the number of input variables and c is a universal constant. Let us be more precise. For d=(di d,) eachrf,a positive integer, let Jf W) be the linear space of all maps f:C"*1 ->C",f=(fu ..,/„), where each/-is a homogeneous polynomial of degree dt. The notion of an approximate zero z in projective space P(C"+1) of/has been defined in [11,12,14,6] and below. It means that Newton's method converges quadratically, immediately, to an actual zero ( of / starting from z. Given an approximate zero, an e approximation of an actual zero can be obtained with a further log | log e | number of steps. Correspondence to: M. Shub, IBM TJ. Watson Research Cent., P.O, Box 218, Yorktown Heights, NY 10598. USA. Email: shub(oiwatson.ibm.com. smalefgimath.berkeley.edu. • Supported in part by NSF, IBM. and CRM (Barcelona). 03O4-3975/93/SO6.00 © 1994—Elsevier Science B.V. All rights reserved SSDI 0 3 0 4 - 3 9 7 5 ( 9 4 ) E 0 0 0 6 0 - V
1454 142
M. Shub, S. Smalt
A probability measure on the projective space (of lines) P(Jf{i)) was developed in [5,11] and "average" below refers to that measure. Let N = dimension Jf(d) as a complex vector space. Main Theorem. Fixing d, the average number of arithmetic operations to find an approximate zero offeP(JV{i)) is less than cN*, c a universal constant, unless n<4 or some d, = 1. Remark. If n<4 or some rf,= 1, we get cN5 The result is also valid in the non-homogeneous case/:C"-»C. The import of the Main Theorem can be understood especially clearly in the case of quadratic polynomials. Thus consider the case
1455 Complexity of Bezout's theorem V: Polynomial time
143
The algorithm of the Main Theorem, developed in [11, 12, 14], is a homotopy method, with steps based on a version (projective) of Newton's method. There is a weak spot in its present use in that the existence of a start system-zero pair (g, () is proved, but not constructively. Thus the algorithms depends on d=(d1,...,d„), and even on a probability of failure a. It is not uniform in the sense of [1] in d and a. In Section 2 an obvious candidate for (g, 0 is given. If our (highly likely) conjecture stated there is true, then the uniformity of the algorithms is achieved. The Main Theorem has the following generalization, which includes the case studied in [14]. We say that zx z, are / (distinct) approximate zeros of/eP(Jf M) ) if they converge under iteration of (projective) Newton's method to / distinct roots Ci C/ of/ Generalized Main Theorem. Fixing d, the average number of arithmetic operations to find 2 > / ^ 1, where Q = fj". i dt is the Bezout number, approximate zeros offeP(Jf(d)) is less than cl2N*. c a universal constant, unless n<4 or some d{= 1 in which case cl2N5 suffice.
2. Main theorem, weak version Let yt (jj oe as in Section 1 and we suppose it is endowed with the Hermitian inner product in [5,11] invariant under the unitary group U(n+ 1). Then S(Jfid]) denotes the unit sphere in Jf {i) and V-{{A<>)eS(jrU))xP(C'*l)\m*0}. Then let 1" be the set of singular points of the restriction it,: V->S(J?ltl)) (i.e. ( / £ ) such that the derivative Dtix: Tf_;( V)->Tf(S(Jf ((0)) is singular). Compare all this with the similar notions for Kof [11,12,14]. In fact, one has the fibration V-* V with fibres SO(2) induced by the fibration S(Jf M ,)-P(^ ( 1 „); thus the vertical distance to I' of [11] defines a similar vertical distance p (same notation) and neighborhood NP(I') of £' in V. Let ¥, be the space of great circles of S(Jftd)) which contain geS(JiTilt)). Let 4i:S(Jf,,(,)-{±g}-*if,, i/'(/) = Z./ be the map which sends /into the unique great circle containing / and g. For such an/we may define L / = 5tf 1 (L / ) <= V. If LfnZ*=
1456 144
M. Shub. S. Smalt
Our Hermitian structure on Jr*U) induces natural Ricmannian metrics and prob ability measures on S(Jr*W)), P{Jf(d)) and if,; see [12,14]. Moreover with these measures, the natural maps S{Jf(i))-*P{#«>), ^•■S(JfU))-{±g}-*
n
Proposition 12. We have n3D3^cN unless some d{= 1 or n^A in which case there is a slightly weaker estimate. The proof is left to the reader (use (DZm)^N). Using this proposition and [11] for the case n = 1, one can use Theorem 2.1 to get the estimate: a^cp2N3. Theorem 2.1 has a ready interpretation in terms of the condition number of/ t*,.df)=
maX
(k.s)*L(f,a,0
/*»orm(M-
Here see [11,12,14] for fiBOfm(h,z) as well as the condition number theorem '<*•*)'
71—i' (Kz)eV, (or?). MnormC».Z)
Let Sn.p-{feS(jrw)-{±g}\LV,g,i:)nN,(?)
+ 9}.
Theorem 23. Given p>0, there is a gsS(Jf?ld)) such that 9
Vol5(JfM))
"
p
Note that Theorem 2.1 is a consequence of Theorem 2.3, since the left-hand side of Theorem 2.3 is an average over the zeros £ of g. For at least one £ one gets less than
1457 Complexity of Bezout's theorem V: Polynomial time
145
the average. Hence there exist a pair (g,£)eV such that g»
V
°!, S , - { -'
^cp2N2n3D3'2
proving Theorem 2.1. Conjecture 2.4. The pair (g, f) of Theorem 2.1 given by 0<(z) = Zo' -1 z{, i= 1,..., n, amf f = eo = (l,0, ...,0) makes the conclusion of Theorem 2.1 true. The truth of this conjecture would make our algorithms more constructive and in fact algorithms in the sense of [1] with input (d, a,f). Remark. Let S,,„ = \J ^ S,,{.„. So
Vol(S,.p)<Xv<>lS,.(.,. In this way it is seen that Theorem 2.3 is a sharp form of Theorem 2 of [14]. In fact that suggests a proof. Theorem 2.3 is proved in the next section following [14], but with a multiplicity function taken into account Next we use the Main Theorem of [11] and Theorem 2.1 to obtain a weak version of our main result. Main Theorem (Weak version). Let be given a probability of failure a, 0
1458 146
M. Shub, S. Smale
Let M2p: S(Ji?«,,)-► 2 + be defined as follows. M2p(f) is the number of roots of / i n 2fi(Z') (perhaps oo) or more properly, the cardinality o{Ttil(f)r\N2l>(Z'). Here we are following the notation of Section 2. N
Theorem 3.1. For any p > 0
(•^M)) JS( VolS(jr„,)Js (jr
M2p(f)^cp*n3N2S.
This is a sharper version of Theorem C of [12], but the same proof works. Define for (g, f )e V, if,.;,p c if, as follows: Suppose that
lnt=0
(L,2cS(jr w ,)).
Then itf * (L)-»L is a ©-fold covering map and rtf l (L)—Jifl (—g) consists of © open arcs in V. Let A,, ( denote the arc among them that contains £. Then define:
i?,.<., = {Z:sif,M,. { nN,(£')*0}. Lemma 3.2. With notation as above, VolS(jf W) )
"
Vol i f ,
The proof follows from the fact that 4>:S(Jf(d))—{±g}-*3ft ability measures and that
preserves the prob
Lemma 33. For any p > 0 , d = {d^, ...,d„) there is a geS(Jt?id)) such that I { ., ( 0 .oVol ■?,,;,, / c p ^ V 1 VolS 1 3/ Vol i f , AJ> V VolS(Jfu\) Js<jf,„) M) )J S(jrMI , (Here VolS'=271.)
Note that Lemmas 3.2 and 3.3, and Theorem 3.1 give the proof of Theorem 2.3. One has to just check that the constants come out correctly. Thus it remains to prove Lemma 3.3. For this we sharpen Propositions 4(a) and 4(b) of [14] as follows. Let S£ denote the space of great circles in S(Jtf(
M {f)
,.L » «V6lX
\ Jf*L
M2p (/)•
1459 Complexity of Bezout's theorem V: Polynomial lime
147
(b) Moreover, 1 Vol if
M
* J/
VolS 1
(/)
« t -' ~vois(.*»
M2fi(f). SVC,.,)
The proof follows so closely that of Propositions 4a, b of [14] that we leave it to the reader. Corollary 3.5. For any d,p there is a geS(Jff(i]) such that — i i M. 51 Vol. &* J * . J t
in< —
VolS 1
Vo1 s
(^wi) J S(jrwi)
Let i,:3e9->Z* be defined as follows. T P ( I ) is the number of A,^ meeting N ,{£'). Then from the definitions I
V0l * , . ; . , = tsir.
f
UD-
From the corollary of Theorem 1 of [14] we have
Therefore, we obtain for any geS{ Jf
{i)),
J \ - I
Vol if,
^Volif,V^3'2
M2p(f). x.J
The corollary of Proposition 3.4 now finishes the proof.
□
4. Integral geometry The goal here is to estimate the volume of certain real algebraic sets. The arguments go back to Crofton and [19], but we use a modern form closer to [12,14].' The following theorem illustrates what we are doing. Theorem 4.1. Let M c / > (R') be a real algebraic variety, given by the vanishing of real homogeneous equations with its complexification having dimension m and degree 5, over C. Then the m-dimensional volume of M is less than or equal to 6 Volf(R* + 1 ). The aftine version can be dealt with by the same methods and in fact is in [15] for the one-dimensional case in R2. Let V<=. S(Jf(i)yx P(C" + 1 ) be as in Section 2 with restrictions of the projections denoted by n,: K-»S(Jr*M)), * 2 : V-P{C"+l). Let L be a great circle in S(Jr*((l,), with 1
Added in proof. See also [18].
1460 148
M. Shub, S. Smale
Lr\I=0. Then AJ"l(L) is a one-dimensional (real) submanifold in V. Also A2(jtj"'(L)) is a curve B in P(C" + 1 ). Theorem 4.2. 77ie /en^rti o/"B is less than or equal to 221. We sketch some basic results on integration, especially Fubini's theorem, in a Riemannian manifold setting (the Coarea Formula, see [7]). Suppose F: X —■ Y is a surjective map from a Riemannian manifold X to a Rieman nian manifold Y, and suppose the derivative DF(x): TX(X)-* Ts (x) (Y) is surjective for almost all xeX. The horizontal subspace Hx of TX(X) is defined as the orthogonal complement to ker Df(x). The horizontal derivative of F at X is the restriction of DF(x) to Hx. The Normal Jacobian N,F(x) is the absolute value of the determinant of the horizontal derivative, defined almost everywhere on X. Example 43. Suppose a compact Lie group G acts transitively and isometrically on a manifold 5. Fixing s0eS, the normal Jacobian of the map G-*S, g-*gs0 is a constant. More generally the following result is seen easily. Proposition 4.4. Let F:X-*Ybe a map ofRiemannian manifolds, equivariant under the action of a compact group G ofisometries ofX and Y. IfG acts transitively on X then the normal Jacobian is a constant. Fubini's theorem takes the following form. Coarea Formula. Let F:X—Ybe a map ofRiemannian manifolds satisfying the above surjectivity conditions. Then for
Here the usual integrability conditions of Fubini's theorem are supposed. Next suppose that G is a compact Lie group acting transitively and isometrically on the manifold S. Let N be a submanifold of S such that the subgroup IN of G leaving N invariant acts transitively on N. Thus the quotient space GN = G/Ifl = {gN\geG} represents the various images of N under applications of elements of G. Our application will be to the case S is real projective space P(R'), G is the orthogonal group 0(1) and N is P(Rk*1) considered as inbedded in P(R') as a coordinate k subspace. In this case GN can be identified with the Grassmannian Gk of fc-dimensional linear subspaces of P(R').
1461 149
Complexity of Bezoul s theorem V: Polynomial time
Returning to the general setting let W<=.GNxS be the submanifold W~{{gN,s)\segN}. Let p , : W—GN, p2: W-*S be the restrictions of the natural projections. The following proposition can be easily proved. Proposition 4.5. The above W is indeed a submanifold, the product action ofGonGNxS leaves W invariant, and acts isometrically and transitively on W. Moreover, p t and p2 are equivariant under G. Corollary to Propositions 4.4 and 4.5. The normal Jacobians ofpi and p2 are constant. Since G acts on 5 it acts also on the tangent bundle T(S) by the derivative. It also acts on the associated bundle Gm(T(S)) with fiber, all m planes through the origin in T,(S). We say that the action of G on S is m-transitive if this last action is transitive. Note that in our application, m-transitivity is satisfied for all the relevant integers m. Proposition 4.6. Let M be an m-dimensional submanifold ofS. Suppose that G as above is also m-transitive on S. Let M = pf l (Af) <= W. Then the restriction p2\jz:M-*GN has normal Jacobian a constant c (possibly not defined everywhere) depending only on G, S, N and m. Proof. Define an associated bundle E(T(W)) = E over W of T(W) as follows. Let we W: we will define the fiber £„ by £w-{Dp 2 (w)- I (L)<=r w |L6C.(r M ,.,(S))}. Then the induced action of G on E is transitive by our hypothesis of m-transitivity. Let Hw be the orthogonal space to kerDp,(w) in Ew. Then NJpi\u(w) is the determi nant of the restriction of the derivative of p, to Hw. By our transitivity we are finished, noting also that the surjectivity of the derivative holds everywhere if at one point □ Theorem 4.7. Let M be as above. Then c VolM = Vol W0,
\o\(gNnM), gNeC,
where c is the constant of Proposition 4.6 and rV0=p2l(s0)for
some fixed s0eS.
The volumes are of course in the appropriate dimensions. Proof. Apply the Coarea Formula and Proposition 4.5 to p2 restricted to M to obtain VolAf = VolMVolWV
1462
150
M. Shut, S. Smalt
Next apply the same argument to pi restricted to M to obtain VolM=
\o\(gNr\M)
.
1 \NJpi\M
gNeGH
By Proposition 4.6, and by elimination of Vol M we obtain the result.
□
Returning to our special case of projective spaces recall that Gk denotes the space of fc-linear subspaces of P(R')Theorem 4.8. Let M c P(R') be an m-submanifold. Then
Proof. It suffices to prove that c Vol W0
VoIP(R" +1 ) 1 VolP(R" + *-' +2 )VolG»
Apply Theorem 4.7 to Af = P(R" + l ). So VolP(R" + 1 )=
c
Volffoji LeCk
Vol(Z,nP(R" +i ))-
The theorem follows, noting that Vol(Z.nP(R" +l ))«=VolP(R" + '-' +2 ).
D
Proof of Theorem 4.1. Let Afc <= P(C') be the complexification of M c P(R') of the theorem. So M-Af c nP(R'), P(R')<= P(C'). The real dimension of M is less than or equal to m and we suppose that it is m. The generic (/—m — 1) linear subspace in P(R') meets M transversally and in at most 5 points since its complexification can meet M c in at most S points of transversal intersection. Thus Vol(Z.nMX<5Vol(G l _ M _ 1 ). Since VolP(R')= 1, and we are finished by Theorem 4.8.
□
Proof of Theorem 4.1 Real projective 2n+l space P(R 2( " +l) ) fibers over P(C" + 1 ) with S 1 fibres, by the isometric action of the unit complex numbers mod± 1. Let ^:P(R 2 ( " + l , )-P(C" + l ) be thisfibration.Denote by A, q ~l (B), where B is as in Theorem 4.2. Note that A is a surface. Lemma 4.9. (a) The length of B equals {l/n) Area A. (b) A generic (2n— 1) linear subspace o/P(R 2, " +1) ) meets A in at most Q2 points.
1463 Complexity of Bezout's theorem V: Polynomial time
151
Proof of Theorem 4.2 (continued). We first show that Theorem 4.2 follows from Lemma 4.9. We use Theorem 4.8 just as in the proof of Theorem 4.1. Thus, Arca
^^VoT7^
and so length B^(VolP(R 3 )/jt)D 2 proving Theorem 4.2.
□
So it remains to prove Lemma 4.9. Part (a) of Lemma 4.9 is again a (rather simple) case of the coarea formula. Now consider (b). The idea is to lift the setting to P(R 2 " + 2 ) and then complexity. Observe that LciP(jf^), so * r 1 L c I x ? ( C ' + 1 ) c P ( j C w ) x ? ( C , + 1 ) and of l course it^ L<=. V. The following diagram helps: ^; 1 7tf 1 (Z.) = ( Z . x P ( R 2 " + 2 ) ) n ^ - ^ — P ( R 2 " + 2 ) J i f ' d ) -(LxP(Cm*l))nV
— ^ — P(C" + 1 )
Here q , : L x P ( R 2 " + 2 ) - Z . x ? ( C " + 1 ) is induced by q and V=qll(V). Thus A = n2q* ' Jtf'L. Now I may be described as £0i+(s-t)g 2 for particular gi,gi6JF(d), t,5eR and thus L may be identified with P(R 2 ) c P(C 2 ). Then qZ' jtf' L = X may be defined by the In real homogeneous equations iRe/+(s-r)Re0=«O,
tlmf+{s-t)\mg=0. These equations are homogeneous of degree 1 in (s, t) and (d,,..., d„) in (x,, yi,) where Zj=Xj + J^lyj. Complexifiying these equations in s,t,Xj,yj, we obtain a variety A"c in P(C 2 ) x P(C 2( "* ")• The generic In-1 linear subspace K c P(R 2 " + 2 ) has the property that the complexification P ( C 2 ) x K c meets Xc in at most Q1 points of transversal intersections (via elementary intersection theory). This yields the upper bound Q1 for the number of real intersection points, finishing the proof of the Lemma 4.9 (and hence Theorem 4.2). □
5. Some approximate zero Theory In this section we do some of the a-theory and approximate zero theory of [15] and [11] in the projective Newton setting. See also [6]. We take a slightly different perspective than [11], but still relying on it in part. In this account we do not attempt to get the best constants.
1464 M. Shub. S. Smale
152
Let /:C" + l -»C",/=(/i, ...,/,), where/ is analytic and homogeneous of degree dt, t.g.fe*w. Recall that the projective Newton method is a map Nf:C"*1—C*\ defined by ^-^-(0/W!N,.,,)'7W
and an induced map Nf:P(C*1)-*P(C*1) which we have denoted by the same letter. We also sometimes identify x and its equivalence class in P(C"*1). Also Nf is obviously not defined everywhere, so the above represents a certain abuse of notation. A point x e C * ' or P{C*l) will be called an approximate zero2 for /if the sequence x, denned by x=x 0 and Nf (x,-)=xt * 1 is defined for all natural numbers i, and there is a zero ( of/such that <M*iiCX(i) 2 ' -1 <M*o.C)- Here
»||x |,
*U,x)-fi0(f,x)y0(f,x\ where /?0, y0, ct are taken as infinite if (Df(x)\NMx)~i does not exist Note. We have not restricted Dkf(x) to Null x as in [11]. This change does not effect the theorems of [11]; they still hold with this y0. If C is a simple zero of/ i.e. (Z)/(C) INUII ;)" l exists, then (is afixedpoint of Nf. We begin by showing that Nf contracts discs of a certain size in P(C+1), centered at {, towards {. Theorem 5.1. There are constants c>0 and i2»>0 such that given fas above, and x,Ce/»(C"+l)wir/i /(C) = 0, <Wx,C)
tfen HDN/(x)||
In [IS] there are approximate zeros of the first and second kind. Our approximate zeros here correspond to the second kind.
1465 Complexity of Bezout's theorem V: Polynomial time
153
Let Br(x) denote the closed ball of radius r around x. Corollary 5.2. //><min(l,«i»/yo(/ 0), then N / : B r ( C ) - i > ( C + 1 ) is well defined and Nf(Br(Z))c: B,(£). It is a contraction with contraction constant cryQ(f,!i), so N / (B r (C))c2MC),r'=cr 2 y 0 (/t). Proof. Observe that BP(C) is convex. By Theorem 5.1 the length of the image of a curve of length L in B,({) is at most cry0{f, C)L. This establishes the assertions on the contraction constant. Applying the contraction estimate to the straight lines from £ to x, X6BP(C), shows that A7/(B,(C))cMC), r' = cr 2 y 0 (/f). As cry 0 (/,C)
O
Corollary 5.4. Iff (0=0, <**(*,£)< 1 and
°
Theorem 5.1 follows immediately from the next two propositions which are of independent interest. Proposition 5.5. Let f be as above. For \\DN,(x)\\^2a(f,x). Proof. Let £ x = x + Nullx, £*,,„ = N/(x) + Null N r (x).
xeP(C+l)
1466 M. Shub. S. Smale
154
We use Ex and EN/ix) as charts for P(C+1) at x and Nf(x), respectively, in the obvious way. □ Let ueNullx. i)N/(x)(r)-it(iy(x)|Nu,IJ-,D,l)/(x)|NllUx(r)(I>/(x)|Nll„,)-1/(x), where n is the orthogonal projection onto the null space of Nf(x). Thus \\DNf(x)v\\c."^2y(f,x)P(f,x)\\v\\c". Recall that N^xjex+Nullx so iiN/(*)ic"«>n*iic"'Now |DAr,(x).| r . / l - W :...,-
\\DNf(x)v\\c»> |N/W|c.„
JDN,(x)v\\c-*^^ II x H e - 1 = 2B(/x)||p||r.?(C-')-
JMIcllxllc-1 d
For the next proposition we use a simple geometry lemma. Lemma 5.6. Let x,yeC" +l -{0}. If dR(x,.y)
and
u=———-y0(f,g)
also differ at most by a multiple of tan(l).
1467 Complexity of Bezout's theorem V: Polynomial ime
155
Now apply Proposition 2, Section III-2 of [11] to conclude that 2a(/, x)< and \{/{u) are close to one for u small so we are done. Here we have been using the notation of [11]. £ 2(K : U/I/'(U) 2 ); K = K{U)
Next we prove a version of the a-theorem. Definition 5.8. Let y0(f x) = max(y 0 (/ x), 1) and &(f, x) = y0(f x)0o(/> x). Theorem 5.9. (Projective a-theorem). There is an &pni>0 such that if &{f x)
i„-c,
(•)
Now if 6t(fx) is small so is P0(fx) a n d s 0 i s l l ^ - C I I / l l x H b y t n e Domination Theorem again. Thus ||x-CI|/||x|| and iR(x,C) differ by a multiplicative constant close to one and dR(x,£) is also small. Since C is a zero, kerZ)/(0 _ 1 =Null(Q and ll(D/(C)| N u I i ; )- 1 D/'(;)| N u l u |Ki. Hence in Proposition 2, Section III-2 of [11], K may be taken as 1 and ro(/,C)<
Vo(/*) ^(T(a(/,x))(l-T(a(/,x))))'
1468 M. Shub. S. Smale
156
Substituting in (*) 4t(x,C)Vo(/„CKca(/x) for cc(fx) small enough. Finally, a(f,x)
so if doroj is small enough we are
6. The bomotopy The goal of this section is to give the proof of Theorem 6.1 below. Throughout this section we suppose that (./",, C) is a curve in V-t', O^t^l. Except for Proposition 6.2 and Lemma 6.3, we assume moreover that f, can be represented asf, = tf+(l-t)g for some f,geS(3^w)). Let y oe an upper bound for 1 andyo(/C,),04t
r,
tk = l
with *o = Co>
Xi^tyjxj.,),
i=l,...,k
well-defined for each i and xf is an approximate zero offu, with associated zero {„. Also k^c^^fHl
+ D^LX
where L is the length of the curve C,. Moreover, (as we will see) :• can be easily calculated atU-i-
Recall that Nft is given by projective Newton's method. Towards the proof we have the following proposition. Proposition 6.2. There exist universal constants a*, u* with the following property. Suppose t, f + Are[0,l],
x r e?(C" + 1 )
satisfy.
Mt(x„t,Ku*, no(.f,;X,)^a*
/ o r a / / r ' e [ t , t + Ar],
then for all r'e[t, t+Ar], ^„(x', C,K«*, where x' = Nfi (x,) and x, is an approximate zero off,- with associated zero £,-. Moreover, given any positive constant K we may take u* < Ka*. In fact, K = 1/48 in what follows.
1469 Complexity of Bezout's theorem V: Polynomial time
157
Proof. As long as a* <<2proj, x, is an approximate zero for/- from Theorem 5.9. That C, is the associated zero is a simple continuity argument. As usual there is a constant K{ close to one such that, d*{N/r(x,),xt)*Klh(f,;Xl)*Kl**/1. So from Proposition 5.3. <*»(*, Cr)<2JC,«V? and by Corollary 5.2
«WNy:(*,U.-Kc[—i-J v
(2K,«*) 2 c
c the constant of Corollary 5.3. Choose a*, u* so that 2K\c(a*)2
D
Lemma 63. There are universal constants K>0, u+m > 0 with the following property. If feP(JTU)),f (Q=0 and <Mx.Oyo(/C)<«... 4(x,C)
x)^KnMm(f, {)•
For the proof we may take xeC+Nullf, Following [11] and the notation there, tinm(f,x)=\\f\\HDf(x)\Uanx)-1Mdln)A(\\x\\d')\\ < II / IIII (Df(x)|NBll,)" > D/(x)|Nyll{llll (Df(x)|Nu!lc)"' £>/(£) |Null<||
tfy(C)iN.uc)"^tf«/2)^(iici*)iM((~)*) -A»-(/C)l(D/WlM.u,)~ l D/l(x)li..ii;l
xii(£!/'wiNulu)-,iy(oiN».Kii
<(i)") Let r 0 = ||x-CII/||CII and u = r0yQ(f, 0 so r0 and <*„(*, 0 differ by a multiplicative constant close to 1, and the same for u and d^(x,!i)y0(f, g). By Proposition 1, Section III-2 of [11] IKD/WlN.ii«)"1jyMlN.iicll<-
d+rg)1'2 f(2-u)u\
1470 M. Shub, S. Smaie
158
By Lemma 3(2), Section II-1 of [11] I W M I N U . . ; ) - ' D/(C)INU.K I <
^ p
and these quantities are both bounded as soon as dK(x, C)yo(/ 9) is small enough. Finally
so (l|x||/||CII)d' is also bounded and we are done. Now let ^,=
HDf,(xt)U»J'1((f-9)M)\\,
B^UDfAxM^J-iDV-gHxMProposition 6.4. (a) B,^2fiumn(ft, x,).
(b)r?o(/,-,x,Kr,
where AtW, + p0(f„x,) ITAfB,
H
as long as Af < l/B,. Proof. We may assume ||x,|| = 1. Then B.tZUDf.MUuJ-11|
\\D(f-g)(x,)\\
$H(f„x,)(\\Df(x,)\\ + \\Dg(x,)\\) ^2n{f„x,) using Proposition 2, Section III-l of [11]. Finally, ^(^.^Xl^onn^.X,)
proving (a). Since Po(fr,x,)=Ht'-t)(Dft(x,)\Nunx,)-l(f-g)(x,)
+ (DfAx,)\suiiJ'1f(xt)\\.
Proposition 6.4(b) follows from the following lemma. Lemma 63. (a)
1 UDfAx,)\HMJ-lDf,(x,)\NuU„^Y-\t'-t\^B,f
(b) MDfAx,)\Kmn»)-lDf,{x,)\KMx,Huin>l
a
1471 Complexity of Bezout's theorem V: Polynomial time
159
Proof. £!/;(x,)|Null;(, = Z)/;(x,)|NullXl + (f,-r)Z>(/-3)(A:,)lNuux1, so (Df,(x,)\NMx,)'1DfAx,)
=I+
(t'-t)(Df,(x,)\NMXl)-lD(f-g)(x,).
Now the minimum norm UDfAx,)\HuiiJ'lDf,{x,)\max,\
1 UDf,(X,)\NM«)- DfAx,)\llM«\ l
1
l + |f-r|*,' which proves Lemma 6.5(b). Part (a) follows from the additional fact that if | | / - / l _ 1 S | | < c < l then | | B _ 1 / l | | < l / ( l - c ) . D Define M = ^ » . c ( / ) = : m a X / ' n o r m ( / t . Cr)r
Choose j = Di!1(i. This is permissible by Proposition 3, Section 1-3 of [11]. Set a**=min(a*,u**), a*,u** of Proposition 6.2, Lemma 6.3, respectively. Also Ar = | r ' - t | . Proposition 6.6. With notation as above, there exists a universal constant c as follows. Given t with V4(Xr-C,K"*, then there is a Ar such that ^(AfKa**/? and At>- or else d R (;„ C,+A,)>ca**/7In fact Ar is easily computed as will be seen. Also there is an obvious adjustment to make in case r + Ar>l. Proof. If P*lAl-l/2A <«**/? let Af = 1/2B,. Then by Proposition 6.4,
P(f,;X,)Z*"/?.
1472 160
M. Shub, S. Smalt
Moreover by the same proposition, B, <2/inorm(/„ x,). Then by Lemma 6.3 we obtain Ar=—>2B, 4^norm(/„xt) 4K/inonn(y;, £,)' Otherwise let At be the solution of /T(At) = a**/y. Then Ar < 1/2B, and it only remains to show that
Lemma 6.7. l/nier r/te conditions of Proposition 6.6, (a)
*(*„«.*«(*,•. C - X j g ^ .
(b)
/Jo(/,.x,)<~.
(c) (d)
(c)
la** 0-(At)>-j—, 6 y 1 a** «*,(*„ x , . ) £ _ _ d«(C„C,)>iC. 24 y
Proof. Since (e) gives our proof of Proposition 6.6 it only remains to prove Lemma 6.7. The first part of (a) is in the hypothesis and since (Proposition 6.4) /?<>(/,•, x,K/J + (A0=a**/7\ Proposition 6.2 yields the second part of (a). Use (a) (first part), that p0(f„ t,)=0 and Proposition 2, Section II-I of [11] to easily obtain (b). For (c) we argue as follows. Recall ArB,«$}. Since 1-AlB,
y
using (b), la**
*W>2-f
3 a**
-^o(/,C,)>g — •
Therefore, *
l
'
1 + ArB,
3\,8 8 / f
6 ?
1473 Complexity of Bezout's theorem V: Polynomial time
161
We obtain (d) using Proposition 6.4, /?<,(/•■ x,)>/?~ (At) and (c). This uses the definition of fj0 as the Newton vector and the exponential map. Finally, (e) follows from
Z,)-dK(xr,
C)
□
Proof of Theorem 6.1. We use Proposition 6.2 and 6.6. Let r 0 =0, r.— tj-j+Ar according to Proposition 6.6. So at each step At satisfies one of the alternatives in(b). Since
l4(t,„C„-i)^, / » - l W A We get the result. D
7. The main theorem The goal of this section is to prove the Main Theorem of Section 1. To this end we first prove two theorems on the number of projective Newton steps sufficient to find a zero. Theorem 7.1. Fix d = (di dK) and a probability offailure a, 0 < a < l . Then there exists (g,{)eV such that the number k ofprojective Newton steps, starting from (g,ZX sufficient to find an approximate zero of input /eS(Jf*W)) is cN3
1
(or cN*/(o1-') if some d,= 1 or n<4). Thus the set of/where the algorithm fails to produce such an approximate zero in k steps has probability measure less than a. For the proof we have the following result. Proposition 12. Fix (d) = (di,...,dl>) O ^ e ^ l are given. Then 2
and suppose (g,{)eV, feS(JfU))-{±g}
and
Kci^iftf-WD*
steps ofprojective Newton's method are sufficient to produce an approximate zero off. In Section 2 we have defined /*#.<(/) and an arc £(/, g, C) in K<= S(Jf U) ) x Let I be the length of A2(L(f, g, ())•
P(C*1).
1474 162
M. Shub, S. Smale
Lemma 73. (a) L^W2. (b) L*nnt.<{f). Part (a) of the lemma is a consequence of Theorem 4.2. Part (b) is a projective space version of the Proposition, Section ID of [14]. The proof is the same noting that n(h, z)^.u„orm(h,z) for all (h,z)e V, and that the length of a great circle in S(Jf(d)) is 2TI. Returning to the proof of Proposition 7.2, we have from Theorem 6.1 that
cti^(f)LDw steps of projective Newton suffice. Hence by the lemma (b),
cftMiff-'L'D312 steps suffice. Finally, from Lemma 7.3(a), %,(,(/)2-^aD3'2 steps suffice. This proves Proposition 7.2
□
Proof of Theorem 7.1. In Proposition 7.2 take e=2/log® so that 22' is a universal constant. By Proposition 7.2 we need to show there exists (g, C)e V such that the function Hu.i)(f?~l®:uI>3'2 is bounded above by cN 3 /(ff 1-t ) for a subset offeS(Jf(d)) of probability measure at least 1 — a. Solve the equation a=cp2N2n3D312 for p. Apply Theorem 2.1, for this p and the condition number theorem to conclude the existence of (#,() such that cN2n3D312
aw/)) a<1 -"< ,"-. for all/in a subset of probability measure at least 1 — a. Now Proposition 7.2, Section 2 and a little arithmetic finish the proof. □ Theorem 7.4. The average number of projective Newton steps sufficient to find an approximate zero offeS(Jf?(d)) is less than or equal to c(\og@)N3. (clog i^N* if some
1475 Complexity of Bezout's theorem V: Polynomial time
163
More formally let parameters of our quasi-algorithm (,•, Ci)e^ be given by Theorem 7.1, with probability of failure aL= 1/2', i' = 1,2.3,4. Let K(a) = cNi/(al~') (or cN*/(oll-l)) if some dt=\ or n^A). For input/ i'= 1, do K(ff.) projective Newton steps for the homotopy {\ — t)gt + tf starting at Co as in Theorem 6.1. If XKia.. is an approximate zero of/by the alpha test, Theorem 5.9, halt and output If not set ; = i'+ 1 and repeat (some/). The average number of steps of this algorithms is less than or equal to '1
£
ff/l
K(<7,))
<x,K(ff,Kc"
K(o)da.
The first inequality follows by summing a geometric series. For the second note that ot-Wi—Oj-11, i= 1,... with a 0 = 1, that K is monotone on the interval [ffj, <Xi_i] and K(ai)/K{ai.1) = 2n~t) so the Riemann sum I
X k,-
K(
Finally, JJ K(ff)da = clog^/V 3 (or slog^/V* if some d, = l or n<4). See [10] for more arguments of this sort. Proof of the Main Theorem. To prove the Main Theorem, we need only make the passage from the number of Newton steps to the number of arithmetic operations. This argument uses well-known facts from numerical analysis about the number of arithmetic operations needed for approximations, for solving linear problems, etc. We omit the details. Proof of Generalized Main Theorem. We sketch some of the changes necessary for the proof. Theorem 2.1 has the following version which also follows from Theorem 2.3. Fixing geS(Jf{i)) and Ci.---.CieP(C + 1 ). ' distinct zeros of q. Let Oi = <Ji(p,g, Ci.---.Ci) 0<
there is a g<=S(JPw) and
Now we can apply Proposition 7.2 to each of the / homotopies starting at (g,^), i = l , . . . , / as in the proof of Theorem 7.1. To prove an /-zero version (change an approximate zero to / approximate zeros and /c^c/V 3 /(o- 1- ') to fc^c/v'3/2/(o'1",))one factor of / is for the probability estimate the other because we follow / homotopies. The /-zero version of Theorem 7.4 follows similarly.
1476 164
M. Shub. S. Smale
References [lj L. Blum, M. Shub and S. Smale, On a theory of computation and complexity over the real numbers: NP-completeness, Recursive functions and Universal Machines, Bull. Amer. Math. Soc. 21 (1989) 1-46. [2] F. Cucker. M. Shub and S. Smale, Separation of Complexity classes in Koiran's weak model, preprint, 1993. [3] M.-H. Kim, Error analysis and bit complexity: polynomial root finding problem. Part I, preprint (Bellcore, Morristown, NJ, 1988). [4] P. Koiran, A weak version of the Blum, Shub and Smale model, 34th Found, of Comp. Sci. 1993, 486-495. [5] E Kostlan, Random polynomials and the statistical fundamental theorem of algebra, preprint, Univ. of Hawaii, 1987. [6] G. Malajovich-Munoz, On the complexity of path-following Newton algorithms for solving systems of polynomial equations with integer coefficients, Ph.D. Thesis, University of California at Berkeley, 1993. [7] F. Morgan, Geometric Measure Theory, A Beginners Guide (Academic Press, New York, 1988). [8] O. Priest, On properties of floating point arithmetics: numerical stability and cost of accurate computations, Ph.D. Thesis, University of California at Berkeley, 1992. [9] M. Shub, On the work of Steve Smale on the theory of computation, in: M. Hirsch, J. Marsden and i4. Shub, eoX From Topology to Computation: Proceedings of the Smalefest (Springer, Berlin, 1993) :>1-301. [10] i. Shub and S. Smale, Computational complexity, on the geometry of polynomials and a theory of -ost: Part II, SUM J. Comput. 15 (1986) 145-161. [11] ivi. Shub and S. Smale, Complexity of Bezout's theorem I: Geometric aspects J. Amer. Math. Soc. 6, .1993) 459-501. [12] M. Shub and S. Smale, Complexity ofBezout's theorem II: Volumes and probabilities, in: F. Eyssette and A. Galligo, eds^ Computational Algebraic Geometry, Progress in Mathematics, Vol. 109 iBirkhauser, Basel 1993) 267-285. [13] M. Shub and S. Smale, Complexity of Bezout's theorem III: Condition number and packing, / Complexity 9 (1993) 4-14. [14] M. Shub and S. Smale, Complexity and Bezout's theorem IV: Probability of Success, Extensions, SI AM J. Sumer. Anal* to appear. [15] S. Smale, The fundamental theorem of algebra and complexity theory, Bull. Amer. Math. Soc. (N.S.) 4 (1981) 1-36. [16] S. Smale, Algorithms for solving equations, Proc. Internal. Congr. Mathematicians, Berkeley, CA (American Mathematical Society, Providence, RI1986) 172-195. [17] S. Smale, Some remarks on the foundations of numerical analysis, SI AM Rev. 32 (1990) 211-220.
References added in proof [18] R. Howard, 77i« Kinematic Formula in Riemannian Homogeneous Spaces, Memoirs of the AMS, Vol. 509 (AMS, Providence, RL 1993). [19] LA. Sautalo, Integral Geometry and Geometric Probability (Addison-Wesley, Reading, MA, 1976).
1477
32 The Godel Incompleteness Theorem and Decidability over a Ring* LENORE BLUM AND STEVE SMALE
Here we give an exposition of Godel's result in an algebraic setting and also a formulation (and essentially an answer) to Penrose's problem. The notions of computability and decidability over a ring R underly our point of view. Godel's Theorem follows from the Main Theorem: There is a definable undecidable set over Z. By way of contrast, Tarski's Theorem asserts that every definable set over the reals or any real closed field R is decidable over R. We show a converse to this result: Any sufficiently infinite ordered field with this latter property is necessarily real closed.
1. Introduction Godel showed [Godel, 1931] that given any reasonable (consistent and effec tive) theory of arithmetic, there are true assertions about the natural numbers that are not theorems in that theory.1 This "incompleteness theorem" ended Hilbert's program of formalizing mathematics and is rightfully regarded as the most important result in the foundations of mathematics in this century. Now the concept of undecidability of a set plays an important role in understanding Godel's work. On the other hand, the question of the un decidability of the Mandelbrot set has been raised by Roger Penrose [Penrose, 1989]. Penrose acknowledges the difficulty of formulating his ques tion because "decidability" has customarily only dealt with countable sets, not sets of real or complex numbers. Here we give an exposition of Godel's result for mathematicians without background in logic and also a formulation (and essentially an answer) to
* Partially supported by NSF grants and (thefirstauthor) by the Letts-Villard Chair at Mills College. 1 This formulation of Godel's Theorem, with consistency in place of co-consistency, is due to [Rosser, 1936]. 321
1478
322
L. Blum and S. Smale
Penrose's problem. The notions of computability and decidability over a ring R as developed in [BSS]2 underly our point of view. A generalization of Godel's theorem and various intermediate assertions may be formulated over an arbitrary ring or field R. If R = Z, the integers, the specialization is essentially the original theorem. Our proof of it is valid in the case R is an algebraic number ring (i.e., a finite algebraic extension ring of Z) or number field (a finite algebraic extension field of Q). Using R = R, the real numbers, the undecidability of the Mandelbrot set is dealt with. Suppose that R is a commutative ring or field (perhaps ordered), which contains Z, and Rk is the cartesian product, viewed as a ^-dimensional vector space (or module) over R. A set S c Rk is decidable over R if its characteristic function X: Rk -> {0,1}, *(x) = 1 if and only if x e S is computable over R in the sense of [BSS]. A set S c Rk is definable over R if it is of the form S«={(y„...,y»)eJl*|3x1Vx23x3...VxII such that (y„...,y»,x,
x„)eY}
k+n 3
for some constructible (or semi-algebraic) set Y in R . A natural question is: [D]
Is a definable set over R necessarily decidable over R?
If R = Z, the answer is no, and this is the backbone of the Godel Incom pleteness Theorem. This may be interpreted (Section 2) as asserting there is a "polynomially defined" set of assertions over Z,4 and that there is no way of deciding which are the true ones. The proof is in Sections 3 and 4; the Godel Theorem is proved in Section 5. If R = R, the real numbers the answer is yes. Tarski's fundamental decid ability result [Tarski, 1951] for the case of the reals may be interpreted as this assertion. These results extend to certain other rings andfieldsas well. Julia Robinson extended Godel's result. Robinson did not have the notion of decidability over a ring and put her theorems in the context of "decidable rings" (or fields). A ring R is decidable if the set of "first-order" sentences true in R is decidable in the traditional sense (of Tuling et al., or decidable over Z in our sense). Her way of showing a certain ring R is undecidable is to reduce the problem to Godel's Theorem by defining Z in R. We use her algebraic results [Robinson, 1959, 1962, 1965] on the definability of Z to answer [D] negatively for finite extensions of Z and Q. 2
We use [BSS] to denote our main reference [Blum, Snub, Smale, 1989]. That is, Y is a finite union of finite intersections of sets defined by polynomial equations and inequations (and inequalities, if R has an order) over R. (See Section 2.) 4 By this, we mean a set of sentences of the form {3x1Vx23xi...VxHP(z,x1,...,xm) = 0 } , . Zk where P is some polynomial over Z. 3
1479
32. The Godel Incompleteness Theorem
323
On the positive side of [D], Tarski's work was done in the generality of real closed fields. His work also implies that [D] has an affirmative answer for algebraically closed fields. Tarski's arguments are a developed form of elimination theory, which in a sharp complexity-theoretic form can be seen in [Renegar, 1989]. What about the answer to [D] for the remainingringsandfield?Let us say that R has Property D, just in case every definable set over R is decidable over R. We give some results (Section 6) in the direction of showing Tarski's exam ples exhausts all fields (sufficiently infinite) satisfying Property D. Here we follow [Maclntyre, 1971], and [Maclntyre, McKenna, van den Dries, 1983]. Indeed, our results could be considered an infinitary version of theirs. Note that Property D, as stated, is a nonuniform property. That is, each decidable set might have a different decision procedure. However, an immediate corol lary to these results (for sufficiently infinite fields) is the uniformity of Prop erty D once it is satisfied. We would like to acknowledge helpful insights gleamed from conversa tions with Michel Herman, Adrien Douady and others concerning the math ematics underlying the undecidability of the Mandelbrot set. The first author would like to acknowledge helpful discussions with George Bergman, Lou van den Dries, Leo Harrington, and Simon Kochen; the personal influence of Julia Robinson is deeply felt. [Friedman, Mansfield, 1988] and [Michaux, 1990] have results related to what we are doing here.
2. Background We give here some background on decidability and definability over a ring R. Then we state our Main Theorem. Let R be a commutative ring or field, which contains Z, and for the mo ment is ordered. Let Rk be the direct sum of R with itself k times. If k = oo, then x e Rk is a vector (x^Xj.-.-.x,,,...) with x,, = 0 for sufficiently large n. In this section, we may take k finite. The notion of a computable function over R, is taken from [BSS] where fiM c Rk, the domain of
1480
324
L. Blum and S. Smale
say that the set of admissible inputs of the corresponding machine is Y) In that case, if Y' c Y, then S n Y' is decidable relative to Y' over R. Next a very brief review of definability over R is given. First suppose R = Z. A subset S of Z* is called definable over Z if there is a polynomial P in n + k variables with integer coefficients such that S = {y = (y„...,y t )eZ*|3x 1 Vx 2 3x 3 ...Vx 11 P(y 1 ,...,y 4 ,x 1 ,...,xJ = 0}. The "defining" formula 3x,Vx23x3... VxJ)P(y1,...,yk,x, x„) = 0 contains an alternating sequence of quantifiers which could begin or end with either 3 or V. In this expression, ylt..., yk are called free variables (and x,, ..., x„ bound variables). If we replace each free variable yt by an integer y,, the expression becomes a sentence perhaps true, perhaps false in Z. (If there are no free variables, the formula is already a sentence.) Note, we have been using bold letters to denote elements (constants) or vectors and nonbold letters to denote variables. Henceforth, in the sequel we shall use nonbold letters for both; the intended meaning should be clear from context. Quantifiers 3z or Vz with a new variable z may be added at any place in a formula without changing the set S. Thus the assumption of an alternating set of quantifiers places no restriction on the notion of definability. More over, unot V" is logically equivalent to "3 not." Similarly, "'not 3" is equivalent to "V not." Thus, negations of formulas could be incorporated using P(y,x,,...,xj#0. But over Z, we have P # 0 if and only if 3z13z23z33z4((P - (1 + z\ + z\ + z\ + z\)) x (P + (1 + z\ + z\ + z\ + z\)) = 0). This assertion follows from the next lemma. Suppose R is an ordered ring orfield.Then: Lemma. // P(x) and Q(x) are polynomials in n variables over R, and x e R", then: (a) P{x) # 0 if and only if -P(x) > 0 or P(x) > 0. (b) IfR = Z: P(x) > 0 if and only if 3z u ..., z4 e Z 4
such that P(x) = £ zf + 1
(Lagrange).
i-i
(c) P(x) = 0 or Q(x) = 0 if and only if P(x)Q(x) = 0. So over Z, negations of formulas are equivalent to formulas of the specified
1481
32. The Godel Incompleteness Theorem
325
type.5 In the sequel, the following equivalences will also be useful: (d) P{x) = 0 and Q(x) = 0 if and only if P2(x) + Q2(x) = 0. (e) Q(x) }0if and only if - Q{x) >0or Q(x) = 0. Now for definability over a general ordered ring or field R, we must modify our definition to incorporate semi-algebraic sets: A basic semi-algebraic set X a R" (over R) is defined as the set of x = (x j , . . . , xn) satisfying basic conditions of the type P(*i
* J = 0,
Qi(xu...,xH)>0,
i=l,...,m,
where P and Qi are polynomials over R. A set X c R" is semi-algebraic (over R) if it is generated by basic semi-algebraic sets using a finite process taking unions (i.e., "or's" of basic conditions), intersections ("and's"), and comple ments ("not's"). From the above lemma, and by adding and substituting new variables, and by de Morgan's laws, it easily follows that: Proposition 1. Every semi-algebraic set X c R" can be expressed as finite union of basic semi-algebraic sets (and conversely). Intersections are eliminated using (d) and complements using (a) and (e). Definition. A set S <= Rk is definable over R if there exists a semi-algebraic set X c Rk x R' such that S = {(yu...,yk)|3x1Vx23x3...Vx11
such that (yu...,yk,xl,...,xl>)e
X}.
Note that the image of a definable set in Rk x R' under the projection Rk x R' -»R k is a definable set. Indeed, definable sets over R are precisely those sets that are "derivable" from semi-algebraic sets by means of a finite se quence of projections and complements. A (defining) formula over R for the above S is 3x^X23x 3 ...Vx 4 (>' 1 ,..., yk,x1,...,x„), where
or
...
or
(j* = o&e'1>o&---&e;Hi>o), the P's and Q's being polynomials over R (in the variables
yi,...,yk,xlt...,
9 Note here, and in the sequel, we are implicitly moving quantifiers to the front of formulas, changing variables as necessary to avoid clashes and to maintain logical equivalence.
1482
326
L. Blum and S. Smale
x„) that describe the basic semi-algebraic pieces of X in a union as given by Proposition 1. Definitions for free/bound variables and sentences over R can be given as above for Z. We remark that for ordered rings, our notion of definability over R is equivalent to the classical notion of definability given in "first-order" logic over the language containing the mathematical primitives +, x, = as well as > and constants from R. That is, the sets defined are the same. For consistency we need: Proposition 2. / / R = Z, our two notions of definable coincide. But this follows from Proposition 1 and the lemma because we can elimi nate occurrences of > using (b), "and's" using (d), and "or's" using (c). Both the concepts of decidability and definability over a ring can be de veloped without >, retaining #0. For decidability, one uses machines with branch nodes of which divide according to h(x) = 0 versus h(x) # 0. For definability, one omits all constructions requiring >. Semi-algebraic sets are replaced by what are usually called constructible sets, finite unions of sets satisfying basic conditions of the type Pi(x1,...,xH) = 0,
i=l,...,k,
e(*i,...,*„)*o. Similarly, a basic constructible formula is of the type P, = 0 «&•••& Pk = 0 & Q # 0 and a formula over R is a finite disjunction of basic ones. The following results apply to rings and field with an order, or without. Main Theorem. Supposes R is a ring (of "algebraic integers") which is a finite extension of Z, or a field which is a finite extension of the rationals Q. Suppose k < oo. Then there is a set S a Rk which is definable over R, yet not decidable over R. The Main Theorem may be interpreted as saying that there is a reasonable family of sentences over R, but there is no way of deciding which are true in R. It is an immediate consequence of Propositions A and B below. Suppose R is as in the Main Theorem. Then: Proposition A. For k <, oo, there is a halting set S c Rk (of a machine) over R which is not decidable over R (i.e., S is a "semidecidable" undecidable set over R). Proposition A will be proved in Section 3. Proposition B. For k < oo, any halting set S <= Rk over R is definable over R. Proposition B will be proved in Section 4.
1483
32. The Godel Incompleteness Theorem
327
Remark. We remark that over the reals R, Proposition A is true (see [BSS]). But by [Tarski, 1951], the Main Theorem must fail over R for each k < oo, and, thus, so must Proposition B. See Section 6 for more discussion and results along these lines.
3. Decidability: Proof of Proposition A Proposition A is proved in part in [BSS] (see Proposition 2 below). Friedman and Mansfield have a complete proof [Friedman, Mansfield, 1988]. But to keep our paper accessible to nonlogicians, we indicate in this section a proof of Proposition A. Say that subsets S, c Rk, S2 => R' are computably isomorphic over R if there is a bijection / : Si -»S2 such that /, f~l are computable over R. Proposition 1. / / R is as in the Main Theorem, there is a computable isomorphism f.R^Z'cR' over R where n is the degree of the extension. PROOF. Let R = Z[w,,...,w„] and for x = £" =1 X(H>; let f(x) = (*,,...,xj. Clearly, f~l is computable over R. Now list the elements of Z" with norm nondecreasing. A comparison machine using this list shows that / is computable. Similarly, if R is a finite extension of Q of degree n, then R is computably isomorphic to Q" over R. To finish the proof, note that if R is any extension field of Q, tnen there is a bijection g: Z -» Q with g, g~l computable over R.
The next step in our development (see [BSS]) is to associate to each ma chine M over a ring R, a point
1484
328
L. Blum and S. Smale
Propositions 2 and 3 now combine to yield Proposition A. As remarked earlier, it is shown in [BSS] that Proposition A holds for R = R (although Proposition 3 does not). Problem. Find k, R such that Proposition A fails, i.e., such that every halting set S c Rk is decidable over R. (See also [Friedman, Mansfield, 1988].)
4. Definability: Proof of Proposition B Suppose S
z0,...,zTe
Rk+1, xeR" such that (1), (2), and (3) are
However, this expression for QM does not make QM definable over R be cause the number of equations and variables is not afixedfinitenumber. It depends on T. So we need more. Generalized Godel Sequencing Lemma. Let R be as in the Main Theorem and k <, oo. There is a map a: N x Rk -* Rk such that (a) given a0, ...,ame Rk,3ue Rk such that a(i, u) = a,,i = 0,..., m, and (b) ifk
1485
32. The Godel Incompleteness Theorem
329
We now combine the register equations of a given finite-dimensional ma chine M as above with the function a of the Godel lemma to show that ilM is definable over R. Let F(y) denote the following "formula": 3reZ+3u3xVieZ [If 0 < i <, T, then <x(0, u) = (1,1(y)), H{a{i - 1, u)) = a(i,«), and a(T,u) = (N,x)l Here Rk of the Godel Sequencing Lemma is identified with R x S. Now on the one hand, flM = {y\F{y) is true in R}. This follows from the properties of the register equations and of a(i,u) of the Godel lemma. On the other hand, the above expression for F(y), and the definability of H and a, and Z (also Z + and N) in R [Robinson, 1959, 1962, 1965], gives us the definability required in Proposition B.6 It remains to consider the Godel lemma. First observe that if the Godel lemma is true for k = 1, then it is true for general k. This can be seen as follows. Suppose a is the map of the Godel lemma for k = 1. Then let
{a(i,ul),a(i,u2),...),
where u = (u 1 ( u 2 ,...). Now let k = 1. For R = Z, we have the original Godel lemma [Godel, 1931]] which is derived from the Chinese Remainder Theorem. The general case now follows by noting the isomorphisms given in the proofs of Proposi tions 1 and 3 in Section 3 with k < oo are definable over R. Problem. For which commutative rings R is the following true? If Z is defin able in R, then for each k < oo, any halting set S <= Rk over R is definable over R (and conversely).
5. Incompleteness In this section, the Godel Incompleteness Theorem is formulated and proved using the Main Theorem. First fix a ring or field R, with or without order. Let l.R be the set of all firstorder sentences over R, i.e., the set of sentences in the first-order language with mathematical primitives + , x , = (possibly > ) and constants from R. (See,
' Here we are also using the logical equivalence of the expressions "if P then Q" and 'not P or Qr
1486
330
L. Blum and S. Smale
for example, [Cohen, I960].)7 The sentences over R described in Section 2 are first-order sentences and can be considered in "normal form" because each first-order sentence is equivalent to one of these. One may consider ER as a subset of R°° by an appropriate natural coding. For example, one can use pairs (yi,y\), (yi,y'i), ••• where either yf or y[ is zero. If yt # 0, then it stands for a logical symbol, not in R normally, but thus coded by an element of R. If yk - 0, then y[ e R plays its role as an element of R, as, for example, the coefficient of a polynomial used in the sentence. One may use the codings of polynomials and rational functions of [BSS], for example. It makes sense to talk about the subset of sentences TR <= Z R that are true in R. For each sentence a e l.R, either a e TR or not a (the negation of IT) e TR (completeness), but not both (consistency). A set of axioms Y over R is simply a subset Y c TR. Of course, Y could be empty. Given a set of axioms Y, we will define the derived set YD, Y c YDcz TR, via (finite application of) a set of rules of inference. Then YD is the body of theorems, or theory, generated by the axioms in Y. The rules of inference, which are convenient for our purposes are Rules A-G (pp. 9-11 of [Cohen, I960]). As an example we state: Rule B. If A and A-* B are sentences, then so is B. Thus, if A and A -* B both belong to Y, YD must include B. A similar situation prevails with the other rules. Recall [BSS] that S <=/?*, k ^ oo, is an output set over R if there is a computable function q>: £iM -» Rk over R with
1487
32. The Godel Incompleteness Theorem F((x 1 ,...,x„),
a underfined,
331
il((
Any element in YD is in some Yp for some finite Y' c Y. So by the above lemma, the range of F is YD and F is computable if
6. Decidability over R and Property D In the following, R will denote R= (a commutative ring or field of characteris tic 0) or R < (an ordered ring or field). L will denote the corresponding (firstorder language L= with mathematical primitives {=, + , x ,0,1} or L< with the additional primitive < In case R is a field, we will also assume the primitive -H is included. Without loss of generality, we may assume L has constants for each integer or each rational as appropriate. In general, if C <= R, then L c will denote the corresponding first-order language allowing addi tional constants from C. Now let I c denote the set of sentences in L c , and Tc the subset true in R. Recall that in the classical setting, a ring or field R is said to be decidable iff T 0 (viewed as a subset of Z°°, as in Section 5) is decidable relative to Z 0
1488
332
L. Blum and S. Smale
over Z. This is what is meant for example when one says that the reals are decidable. But often one really has more. Thus, in the case of the reals, we see that TR< (<=R°°) is decidable over R<, by a machine with parameters from Z. This follows immediately from Tarski's theorem that R admits uniform elimi nation of quantifiers.8 Thus, we are motivated to extend the notion of decidability over a ring. Definition. R is (strongly) decidable over R iff TR is decidable relative to I * over R (by a machine with parameters from Z). If R is decidable over R, then clearly the Main Theorem fails for each k < oo, i.e., R has Property D: Definition. R has Property D iff for each He < oo, any X c J ! 1 definable over R is also decidable over R. Property D is a weak form of elimination of quantifiers: Proposition \. If R has Property D, then for each LR-formula q>R(x!,...,xk), there is a finite set C c R and an (effectively) countable set of quantifier-free Lc-formulas ^l(x^,..., xk) such that S = \J S,, where S, Sj are the sets defined by q>R, ipl, respectively. Furthermore, we can assume without loss of generality the ifrl to be basic semi-algebraic (or basic constructive) formulas. ProMem. Can we assume that C is just the set of constants occurring in (p„? If so, we shall say R has strong Property D. What is the relationship between the above notions of decidability? Theorem. Suppose R is a field of infinite transcendence over the rationals. Then the following are equivalent: 1. R admits uniform elimination of quantifiers. 2. R is strongly decidable over R. 3. R has strongly Property D. 4. a. R is an algebraically closed field in case R = R_. b. R is a real closed field in case R — RK. And these imply: 5. R is decidable. 8
Definition (Classical). R admits elimination of quantifiers iff for each L-formula cp(x,,..., xk) there is a quantifier-free L-formula ^(x,,..., xk) such that q> and ip define the same subset of Rk. The elimination is uniform iff the map from cptoip is computa ble over Z.
1489
32. The Godel Incompleteness Theorem
333
Furthermore, if we add the condition "R is dense in its real closure in case R = R<"9 we can add to the above list of equivalences: 2.' R is decidable over R. 3.' R has Property D. Remark 1. The stipulations of uniformity in 1 and effectiveness (of the count able decomposition of definable sets into semi-algebraic sets) as implied by 3 (or 3') are not necessary. These will follow from a simple analysis of the proof of the theorem. Remark 2. We note that, in general, the notion of decidability is weaker than decidability over R (or Property D). For example, R= is decidable, but be cause R= is not algebraically closed, the theorem implies it cannot be decida ble over R,. It is easy to see that 1 implies 2 which implies 3; 4 implies 1 by Tarski's theorems [Tarski, 1951]. Clearly, 2 implies 5, and 2 implies 2' implies 3'. Thus, we are left to show that under the appropriate hypotheses, strong Property D or Property D imply 4, i.e., the closedness conditions. Here we are inspired by, and indeed closely follow, the proof in [Maclntyre, McKenna, van den Dries, 1983] of the converse of Tarski's theorems. We start with some preliminaries. In the following, R will be a field (of characteristic 0), and n > 1. Fix n. Let Poly„(n) denote the space of monk degree n polynomials in one variable over R. There is a natural identification, Poly^(n) 2; R", where PROOF OF THEOREM.
r0 + • • • + r^z'-1
+ z " ~ r = ( r 0 , . . . , r ^ ) e R*.
Let S c R " correspond to the set of polynomials in PolyR(«) not solvable in R. Note that S = {r e R"\t(r,z) has no root in R}, where f(r,z) = r0 + ■■■ + r,,.^" -1 + z". Thus, S is defined by the L-formula Vzf(x 0 ,...,x B _ 1 ,z) =£ 0. So if R has Property D, then, by the Proposition 1, S = \JjfJSj, where J is countable and Sj = {r e R"\P/(r) = 0,i=\,...,mj,
QJ(r) # 0} in case R = R=,
or Sj = {re R"\PJ(r) = 0, Q{{r) > 0, i = 1,..., ro,} in case R = R<. Here the P's and Q's are polynomials over Q(C), where C is a finite subset of R. 9
R is dense in its real closure R means: For each r and e > 0 in R, there is an r' in R such that \r — r'\ < e.
1490
334
L. Blum and S. Smale
Proposition 2. Suppose R is of infinite transcendence. If R has an algebraic extension of degree n, then for each finitely generated subfield K a R, S can not be covered by a countable union of proper algebraic varieties Vj in R" defined over K (i.e., with Vj= {re R"\P/(r) = 0, i = 1, ..., kj}, the P's being nonzero polynomials with coefficients from K). Modulo Proposition 2, we proceed with our proof. We first consider the case R = R= and suppose R has property D. Suppose R is not algebraically closed, i.e., for some n, R has an algebraic extension of degree n. So by Proposition 2 (and noting S # R") we may assume for at least one ;', Sj= {re R"\QJ(r) # 0} * 0 , and so Q} # 0. Also S; <= S. So for r e R", whenever Q*(r) # 0, then f(r, z) = 0 has no solution in R. This contradicts Lemma 1. For each nonzero polynomial Q: R" - » R , there is an reR" such that Q(r) # 0 and f(r, z) has all its solutions in R. To see this let a: R" -*■ R" be the polynomial map given by the elementary symmetric functions alt a2,..., aH. a is an algebraically independent map. So for Q # 0, and because R is infinite, there is a = (a,,..., a„)e R" such that Q{a(a)) # 0. Letting r = a(a) we have Q(r) # 0 and f(r,z) = a,(a) + a2z h CTn(a)z"-1 + z" = Y["i-i ( r — a i)- Thus, r has the required properties. Now we consider the case R = R< and suppose R has Property D. Suppose first that some odd degree polynomial has no solution in R. This implies, as above, that for some odd n > 1 and j in J, we may assume Sj ={re R"\Q{(r) > 0, i = 1,...,ro,-}# 0 . Let
J/? [Q,
if R is dense in R in case c € Q m .
If \ji(c, r) is true in R, then \j/(c, r) is true in R, and so either trivially in the first case, or by the transfer property for real closed fields in the second case, i/f(c,r*) is true in F, some r* e F". Now, f(r*,z) = (z - £)(a0 + a,z + ••• +
1491
32. The Godel Incompleteness Theorem
335
a^iz"'1), some £, a0,..., a,,., e F. For £', ^o. •••» a «-i 6 ^> define r' e F" by f(r',z) = (z —
Remark 5. If X has property A, and X <= S <= R", then S has property A. 10
We are grateful to George Bergman for pointing us to this paper.
1492
336
L. Blum and S. Smale
Lemma 4. Suppose F: R" -» R" is an algebraically independent polynomial map (i.e., for all polynomials g: R" -*■ R, if gF = 0 then g = 0). Then if X c R" has property A, then so does F(X). PROOF.
Let Y = F(X) and consider the following diagram: XcR"
Y = F(X) c R" —^— R First suppose {Yj}Jej covers Y. Let X} = F _1 (^). Then {X}}}tJ covers X. Now suppose Yj is a variety in R' defined by a (finite) set G, of polynomials over K c /?. Then ^ = {re K"|F(r) e Y,} = {re R"|f/F(r) = 0, all g e G,} is a variety in R* defined by the (finite) set of polynomials {gF\g e G,} over K(at,...,am). Here a,,..., am are the constants occurring in F. So if {Yj}jtJ is a countable cover of Y by varieties over a finitely gene rated subfield K, then {X}}}9j is a countable cover of X by varieties over K(al,...,am). UX has property A, then X} = R", for some j e J. So for all r e K", gF(r) = 0, all 0 e Gj. Thus, because K is infinite, gF = 0, all g e G,. But F is alge braically independent. Therefore, g = 0, all g e G,. Therefore, >} = J?" and so y has property A. This proves Lemma 4. Now let R be the algebraic closure of R, and so R" can be considered the space of {sequences of all) roots of elements of Poly^(n). The Viete map n: R" (roots) -»R* (^ Poly«(n)) is given by *«)~ri(2-W
/or{ = « „ . . . , £ . ) e l l " .
We consider the following diagram: Poly„(n -l)x
R"~R'
x R" -
ProJ
PolyR(n - 1)
Eval
R" (Space of roots) (Viete map)
- - ^ — — K" a Poly*(n)
(Space of degree n—\ polys//?)
(Space of monk degree /»polys//?) Here PoiyR(n — 1) is the space o/ univariate degree (n — 1) polynomials over R which is naturally associated with /?" via r0 + • • • + r,,.!y _1 <-► r = 0o> • • •, r.,-1) e K". And for r e K", a = (a,,..., a,) e R', Eval is given by Eval(r, a), = r0 H /•„_, a,""1, F = it • Eval and F.(r) = F(r, a). Thus,
h
1493
32. The Godel Incompleteness Theorem Remark 6. {Eval(r,ot),=1 associated with Fa{r).
337
„} is the set of all roots of the monic polynomial
Let RJ = {a e R"|{a,}(=1 „ are distinct conjugates of degree n over R}. So, RJ # 0 just in case R has an algebraic extension of degree n. If £ e RJ, then 7t(iS) is the minimal polynomial of £, over R. Lemma 5. Suppose ae R%. Then (a) Fa: R" -> R" and (b) F„ is algebraically independent. PROOF. See
[Maclntyre, McKenna, van den Dries, 1983].
Recall X = {re /?"|rf ¥= 0, some i > 0} and S = {r e R"\f(r,z) has no roots inR}. Corollary 1. // a e R;, r/ien F„(X) c S. PROOF. Suppose txe Rl and r e /?". Then by Lemma 5(a), F„{r) e R". If, in addition, r e X, then Eval(r, a), ^ R, i = 1,..., n. So by Remark 6, /•",(/•) e S.
Corollary 2. Suppose the degree of transcendence of R is infinite. If RJ # 0 . t/ien S /ias property A. PROOF. By Lemma 3 and Remark 4, X has property A. Let a e RJ. Then, by Lemmas 5(b) and 4, F„(X) has property A. So by Corollary 1 and Remark 5, S has property A.
Thus, we have proved Proposition 2 and hence the theorem. Remark 7. Using a lemma of [Michaux, 1990], we need only assume as hypothesis in the theorem, in case R = R= (likewise, in case R = R<), that R be a commutative ring without zero divisors (an ordered commutative ring) of infinite transcendence over Q.
7. On the Undecidability of the Mandelbrot Set over R For a heuristic discussion of the problem of the decidability of the Mandelbrot set M, one can see [Penrose, 1989]. For the mathematics of M, see [Douady, Hubbard 1984-1985]. A well-known and presumably difficult conjecture is: The boundary of M has Hausdorff dimension 2. On the other hand, it seems much easier from the work of Douady, Hubbard, Misiurewicz, Tan Lei, and others that the follow ing holds.
1494
338
L. Blum and S. Smale
Weak Conjecture. The boundary of M has Hausdorff dimension greater than 1. Proposition. / / the Weak Conjecture is true, then the Mandelbrot set is undecidable over R. PROOF. Suppose the contrary. Then M is the halting set of some machine over R and, hence, is the countable union of basic semi-algebraic sets (see [BSS]). M is closed, and the closure of a basic semi-algebraic set is a semi-algebraic set. Thus, we may suppose
M = 1=1 0 S„ where each S, c C = R 2 is a closed semi-algebraic set. For each i, we claim: dim(3M n S,) < 1, where dM is the boundary of M and dim is the Hausdorff dimension. If dim Si < 1, this is immediate. However, if dimS, > 1, then dimS, = 2 because S, is a closed semi-algebraic set. So dimdS,- <, 1. Next note that the interior of S, must be contained in the interior of M. Therefore, dM n Sj is contained in dM n S, and has dimension ^ 1. So now we have 00
8M = \J (dM n Ss) has dimension ^ 1, i=l
contradicting the Weak Conjecture. Remark. It is easy to show that the complement of M is a halting set, so once more we have (provisionally) an example of an undecidable "semi-decidable" set. Added in Proof. At the Smalefest in Berkeley (August 1990), Dennis Sullivan has shown us what appears to be a direct proof that the Mandelbrot set is not the countable union of semi-algebraic sets.
References Amitsur, S.A., "Algebras over Infinite Fields," Proc. AMS American Math. 7 (1956), 35-48. Blum, L., Shub, M., and S. Smale, "On a Theory of Computation and Complexity over the Real Numbers: NP-Completeless, Recursive Functions and Universal Machines," Bull. AMS 21 (1989), 1-46. Cohen, P., Set Theory and the Continuum Hypothesis, Benjamin, New York, 1960. Davis, M., Computability and Unsolvability, Dover, New York, 1982. Douady, A. and J. Hubbard, "Etude Dynamique des Polynomes Complex, I, 84-20, 1984 and II, 85-40.1985, Publ. Math. d'Orsay, Univ. de Paris-Sud, Dept. de Math. Orsay, France.
1495
32. The Godel Incompleteness Theorem
339
Friedman, H., and R. Mansfield, "Algorithmic Procedures," preprint, Penn State, 1988. Godel, K., "Uber formal Unentscheidbare Satze der Principia Mathematica und Verwandter Systeme, I," Monatsh. Math. Phys. 38 (1931), 173-198. Maclntyre, A., "On wr Categorical Fields," Fund. Math. 7 (1971), 1-25. Maclntyre, A., McKenna, K., and L. van den Dries, "Elimination of Quantifiers in Algebraic Structures," Adv. Math. 47, (1983), 74-87. Michaux, C, "Ordered Rings over which Output Sets Are Recursively Enumerable Sets," preprint, Universite de Mons, Belgium, 1990. Penrose, R., The Emperor's New Mind, Oxford University Press, Oxford, 1989. Renegar, J., "On the Computational Complexity and Geometry of the First-Order Theory of the Reals," Part I, II, III, preprint, Cornell University, 1989. Robinson, J., "The Undecidability of Algebraic Rings and Fields," Proc. AMS 10 (1959), 950-957. Robinson, J., "On the Decision Problem for Algebraic Rings," Studies in Mathemati cal Analysis and Related Topics, Gilberson et al., ed., University Press, Stanford, 1962, pp. 297-304. Robinson, J., "The Decision Problem for Fields," Symposium on the Theory of Models, North-Holland, Amsterdam, 1965, pp. 299-311. Rosser, B., "Extensions of Some Theorems of Godel and Church," J. Symb. Logic I (1936), 87-91. Tarski, A., A Decision Method for Elementary Algebra and Geometry, University of California Press, San Francisco, 1951.
1496 Theoretical Computer Science 133 Elsevier
(1994)
3-14
3
Separation of complexity classes in Koiran's weak model F. Cucker* Universitat Pompeu Fabra , Balmes 132, 08008 Barcelona, Spain
M. Shub** IBM T.J. Watson Research Center , PO Box 704, Yorktown Heights, NY 10598, USA
S. Smale** Department of Mathematics . University of California, Berkeley, CA 94720, USA
Abstract
Cucker, F., M. Shub and S. Smale, Separation of complexity classes in Koiran's weak model, Theoretical Computer Science 133 (1994) 3-14. We continue the study of complexity classes over the weak model introduced by P. Koiran. In particular we provide several separations of complexity classes, the most remarkable being the strict inclusion of P in NP. Other separations concern classes defined by weak polynomial time over parallel or alternating machines as well as over nondeterministic machines whose guesses are required to be 0 or 1.
1. Introduction Very recently Pascal Koiran introduced in [10] a model of computation that comes from a modification of the cost notion of the real Turing machine of [2]. This new model - that following Koiran will be called weak - drops the unit cost assumption for the arithmetical operations and only allows a "moderate use of multiplication" [11]. The main result of [10] states that when restricted to Boolean inputs the class of sets decided by these machines in polynomial time coincides with P/poly. As a consequence, if P=NP in the weak model then the Boolean polynomial hierarchy collapses at the second level. Correspondence to: F. Cucker, Universitat Pompeu Fabra , Balmes 132, 08008 Barcelona, Spain. ` Partially supported by DGICyT PB 92/0498/C02/01, and the ESPRIT BRA Program of the EC under contracts no. 7141 and 6546, projects ALCOM II and PROMotion. ** Work done at the Centre de Recerca Matematica . Supported in part by NSF grants. 0304-3975 /94/$07.00 0 1994- Elsevier Science By. All rights reserved SSDI0304-3975(94)00069-U
1497 4
F. Cucker et at.
In the present paper we continue the study of the computational power of the weak model. In particular, several separations between complexity classes for that model are proved, the most important one being P # N P . In fact, it is shown that NP W (the subscript stands for "weak") strictly contains its subclass NP WD consisting of those sets that can be decided using binary guesses. Note that since P w <=■ NP WD the above mentioned separation holds. The problem of whether P w = NP W D remains open and we provide two kinds of partial answers to it. On the one hand, in Section 3 and following the line of ideas of [10] we show that the above equality would imply the collapse of the polynomial hierarchy at its second level, a fact seen as unlikely in complexity theory. On the other hand, we prove in Section 5 that if we restrict our attention to machines that branch only on equality tests, we can prove that the forementioned equality does not hold. This is done by showing that a well-known problem (the Knapsack problem) belongs to NP W D and cannot be solved in determin istic weak polynomial time. Finally, in Section 4, we consider the alternating variation of the weak model, and we give a doubly exponential lower bound for the parallel time needed to decide problems solvable in polynomial alternating time.
2. The weak model In the following we shall denote the direct sum © J° R by R°°. Also, we define the size |x| of an element xeUx as the largest i such that its ith coordinate xf is different from zero. We shall denote by I the subset {0,1} c R and - following the custom in Complexity theory - by I* the set of all finite strings over I. Note that there is a natural inclusion I* c; R°° and that the membership of a point in R" to I* can be algebraically expressed by n equations of the form X(X —1) = 0. Also, we shall consider real Turing machines over R°° as they were defined in [2] but in a normal form that requires that every computational node performs a single arithmetic operation. This requirement does not modify the running time of the machine up to a constant factor. Let M be a real Turing machine whose running time is bounded by t(n), and let ocj,... ,a t be its real constants. For any input size n, the machine M determines an algebraic computation tree 7"M,„ with depth t(n). At an arithmetic node v of this tree a value is assigned to a variable z corresponding to an arithmetical operation on some previously computed values. This value z can be expressed as/»(xi,... ,x„, ax, ... ,ak) where f, is a rational function with rational coefficients and (xi,... ,x„) is the input. These rational functions are used to define the running time in the weak model. In the next definition, we shall understand by the height of a rational number p/q its bit length i.e. Llog(|p|+l)+ (log(||)J. Definition 1. The cost of any arithmetic node v is defined to be the maximum of deg (_/",) and the maximum height of the coefficients of/,, while the cost of any other node is 1. For any x e R x of size n the weak running time of M onx is defined to be the
1498 Separation of complexity classes in Koiran's weak model
5
sum of the costs of the nodes along its computational path in TMt„. The (weak) running time of M is the function associating to every n the maximum over all xeR°° of size n of the running time of M on x. The classes P w and NP W of weak deterministic and nondeterministic polynomial time respectively are now denned as in [2]. Also, we define the class NP W D of weak digital nondeterministic polynomial time by requiring the guesses in NP W to be elements in I*. This kind of nondeterminism describes the complexity of discrete search as appears for instance in the Travelling Salesman or the Knapsack problems (see [6]). In the sequel, unless otherwise stated, all the complexity classes are in the weak model. The adjective full as opposed to weak will be applied to the notions as they were introduced in [2]. A first result concerning weak nondeterministic polynomial time is that it coincides with full nondeterministic polynomial time. Consequently, we derive the NP w -completeness of the full NP-complete problems of [2]. Let us recall that QS is the set of systems of quadratic equations having a real solution, and that 4FEAS is the set of degree 4 polynomials having a real root. Theorem 2. We have that NP R = NP W where NP R is the class of sets decided in full nondeterministic polynomial time. Proof. We first observe that the reductions given in [2] to reduce any problem in NP R to 4FEAS, work in weak polynomial time. This can be seen either checking the weakness at the proof given in [2] or realizing that the quoted reductions does not use noninteger constants and seen as a Boolean algorithm (dealing with the input x and the machine constants as symbols) it is performed in polynomial time and thus, according to Lemma 3 of [10], that it works in weak polynomial time. Now, since 4FEAS can be trivially solved in weak nondeterministic polynomial time, we have an NP W algorithm for solving all problems in NP R by composing for any SeNP R the reduction to 4FEAS with the algorithm for solving this last problem. D A side consequence of this last proof is the following result. Theorem 3. The sets QS and 4FEAS are NP^-complete for reductions in P w . Let us introduce now a parallel computational model. Definition 4. An algebraic circuit % over R is a directed acyclic graph where each node has indegree 0,1 or 2. Nodes with degree 0 are either labeled as input or with elements of R (we shall call the last ones constant nodes). Nodes with indegree 2 are labeled with the arithmetic operations of R, i.e. +, •, - and /. Finally, nodes with indegree 1 are
1499 6
F. Cucker et al.
of a unique kind and are called sign nodes. There is a set of m ^ 1 nodes with outdegree 0 called output nodes. In the sequel the nodes of a circuit will be called gates. To each gate we inductively associate a function of the input variables in the usual way (note that sign gates return 1 if their input is greater or equal to 0, and 0 otherwise). In particular, we shall refer to the function associated to the output gates as the function computed by the circuit. For an arithmetic circuit <, the size s(#) of <€, is the number of gates in <€. The depth d(^) of <£, is the length of the longest path from some input gate to some output gate. The cost of an arithmetic gate is defined as before and the cost of a path in the circuit is the sum of the costs of their gates. We define the weak running time of a circuit on an input x to be the maximum of the costs of their paths. The weak running time of the circuit is defined then as before. Given an algebraic circuit #, the canonical encoding of # is a sequence of 4-tuples of the form (g, 0p,,,gp)eR4 where g represents the gate label, op is the operation performed by the gate, g, and g, are 0 if gate g is an input gate, and gr is 0 if gate g is a sign gate (whose input is then given by gt) or a constant gate (the associated constant being then stored in gt). Also, we shall suppose that the first n gates are the input gates and the last m the output gates. Let {#n}n€i\i be a family of circuits. We shall say that the family is Pw-uniform if there exists a real Turing machine M that generates the ith coordinate of the encoding of #„ with input n-l
(i,l,.. .,1) in weak polynomial time in n. We shall say that the family is EXPw-uniform when there is a real Turing machine M as above but working in time weak exponential in n. We now define PARW to be the class of sets S such that there is a Pw-uniform family of circuits {#„} having size exponential in n and weak polynomial running time such that the circuit <€„ computes the characteristic function of the set of elements in S with size n. The class PEXPW of sets decided in weak exponential parallel time is defined in an analogous manner. The next proposition is a weak model version of the main theorem in [4]. Proposition 5. LetfneR [Xu..., Xn~] be a family of irreducible polynomials such that for all n the zero set %{fn) is a variety of dimension n— 1 and dcg( f„)^d(n). Then, any family of circuits deciding the set S= {xeRa&iyjX|(x) = 0} has a weak running time greater than d(n). Proof. Let us assume that there exists a family of circuits #„ having running time bounded by r(n) and deciding S.
1500 7
Separation of complexity classes in Koiran's weak model
For each n we consider the size N of <€n and we call "configuration" any point in R" and "initial configuration" the point H-n
(xu...,x,,,0,...,0) At each step of the computation we modify some of the coordinates of the current configuration replacing them by the result of operating (via one of ( + , —,*,/)) on two other coordinates. Those modifications can depend on Boolean conditions of the form G,(xi,...,x11)>Q,
where Qt{Xu... ,X„) is a rational function (whose coefficients depend on the output of previous sign gates, and therefore on the actual input xu... ,x„) and (2,-(xj,... ,x„) is the content of coordinate i in R". At the end of the computation the Nth coordinate of thefinalconfiguration will be 0 or 1 according to the truth of a large (butfinite)system of the form \ / f /\Qj.,(Xu...,X„H0/\ j-l \i-l
A
Qij(Xu...,X„)
i-jj+1
/
where the degrees of the numerator and denominator of the QitJ are bounded by r(n). By expressing the sign of a quotient in terms of the signs of numerator and denominator we can replace the rational functions by polynomials with the same bound for the degrees. Also, expressing an inequality like F(Xu...,Xn)>0, as the disjunction F(Xu...,Xm)-0V
F(Xu...,Xm)>0
and then distributing, we can describe S as a union of sets given by systems of polynomial inequalities of the form A Ff (*,,...,*„) = () A / \
Gj(Xlt...,X„)>0.
Now, since the zero set &{fn) has dimension n— 1, one of those sets must contain a subset of dimension n — 1. Since the set described by the G/s is open, it must be nonempty, and then it defines an open subset of R". But our zero set has dimension n— 1, and therefore we must have s>0. Finally, all the polynomials Fit i = 1,..., s, vanish on that (n - l)-dimensional subset of the variety. But, since the variety is irreducible, this implies that every Ft must vanish on the whole variety. Using the fact that the ideal (/„) is the definition ideal of &(f„) (see [3, Theoreme 4.5.1]) we conclude that all the F, are multiples of/„ and thus, that their degrees are greater than d(n). Since these degrees are a lower bound of the running time of the circuit <€n we deduce the proposition's statement. D
1501 8
F. Cucker et al.
Theorem 6. (a) PR £PAR W , (b) NP W £PAR W . Proof, (a) Let us consider the set S = {x6R to |xf" = x 2 where n = \x\). Note that for each n the subset of elements of S having size n is an irreducible variety of dimension n — 1. Thus, S cannot be decided in weak polynomial parallel time because of the preceding proposition. On the other hand, it clearly belongs to P„. (b) Trivial since P„ £ NP„ = NP W . Note that it can also be shown observing that the following sentence 3y\iy2---iy*-\xl=y\
A.y?=.y 2 A...
A^2_,=X2
is equivalent to Xi" = x 2 and can be checked in weak nondeterministic polynomial time. □ Corollary 7. The inclusion NP WD c NP W is strict. Proof. Trivial since NP WD c PARW (the parallel machine just tests the exponential number of possible guesses independently). D Corollary 8. The problems QS and 4FEAS cannot be solved in weak polynomial time even allowing parallelism or digital nondeterminism. Theorem 6 provides a result quite unusual in complexity theory since either in the Boolean setting or in the full real setting the class NP is included in its corresponding PAR. We can prove however that in the weak model nondeterministic polynomial time can be solved in deterministic exponential time. Lemma 9. If a set S c RK can be decided in (full) parallel time t(n) then it can be decided in weak deterministic time 2°"<")). Proof. The weak machine simply simulates the parallel one. This takes full time 20
1502 Separation of complexity classes in Koiran's weak model
The preceding results can be summarized in the following picture: NP W = NP R EXPV *PARV where an arrow -» means inclusion, an arrow -► means strict inclusion and a crossed arrow *♦ means that the inclusion between the corresponding complexity classes does not hold. Theorem 2 asserts that the classes NP W and NP„ coincide. This is not necessarily the case for their subclasses NPC W and NPC R of complete problems, since in the first case the reductions considered are in P w and in the second in P R . In fact, it is trivial that NPC W £ NPC R . The converse, however, seems less trivial to prove according to the consequences that it has. Theorem 12. If NPC W = NPC R then P R # N P „ . Proof. Let us suppose that PR = NP R . Then we have that NP R = NPC R and in particular, that NP WD c NPC R . By hypothesis this entails that NP W D £ NPC W and thus that NP WD = NP W , contradicting Corollary 7. □
3. Weak machines and Boolean complexity classes Definition 13. Given a class # of subsets of R°°, we shall call its Boolean part the class of subsets of I * obtained by considering for any SeC its subset of elements belonging tor*. One of the main results in [10] states that the Boolean part of P w is P/poly. This result was used then to show that if P W = NP W then the polynomial hierarchy of Meyer and Stockmeyer collapses, a consequence that now becomes meaningless since we know that P W # N P W . The main ideas of the paper (and the techniques used there) remain however very interesting since, as we shall see, they can still be fruitfully used. We begin by recalling the main technical tool obtained in [10]. Theorem 14. Let S a Uk be a semialgebraic set defined by a system P,(Xu...,Xk)>0.
«=l,...,N,
1503 10
F. Cucker et ai
with PieZ[Xu...,Xt'], and let D be the maximum degree of the P^s and H the maximum height of their coefficients. If S^0, there exists a rational point xeQ* belonging to S and having height bounded by aHDh with a and b depending only on k. Theorem 15. The Boolean part of PARW is PSPACE/poly. Proof. Let SePAR w and let us consider its subset S of elements in I*. We shall see that S belongs to PSPACE/poly. Since S belongs to PARW, there is a family of algebraic circuits {#„} having weak running time bounded by a polynomial q(n). Moreover there is a real Turing machine M that given (n, i) produces the ith gate of #„ within a time that we can suppose to be also bounded by q(n). Let the size of #„ be bounded by s(n) = 2"' and ai,... ,a.k be the constants of M. With the exception of the constant value for the constant gates, all the remaining values computed by M are positive integers with polynomial height. Without loss of generality we will suppose that M first produces a base two representation of these numbers and then - without using the real constants au...,at - computes the corresponding integers. This property allows us to suppose that the integer value returned at the end of a computational path only depend on the path itself and not on the constants xu...,ak. On the other hand, the constant gates depend on a,,...,a k and their associated constant y in can then be expressed as r„,,(<*!,... ,ak) where rnJ is a rational function having polynomial degree and coefficient heights since the weak running time of M is polynomially bounded. Also, let 'V..;(ai,---,ajt)SJ0,
r«.i.A*i,---,a-k)<0
v'=l,...A
the rational functions that determine the computation path followed by M on input (nj) and let )'„., be the system of inequations resulting by replacing the a,,...,a k by the indeterminates Xu... .Xk. For any n, and for any element uel*, the computation done by #„ on input u can be described by a set of equations
where the f„,uA and the g„,„,, are rational functions having polynomial degree of coefficient heights (because of the polynomial weak running time of <#„) and the third subindex runs over the sign gates of #„. Let us replace in E„.„ each oecurrence of a ynJ by its correspondent rational function rni(Xl Xk) and let {„.„ be the resulting system of inequations. If we now
1504 Separation of complexity classes in Koiraris weak model
11
define
s.-( U r,«W U C..A we obtain a system of inequations that has the real solution (o^,... ,a k ). In order to apply the preceding theorem to ensure the existence of a small rational solution we use Koiran's trick to get rid of the equalities (see [10, Section 5.2.]). We obtain then a new system E„ having only strict inequalities and such that if a point r=(rlt... ,rk) is a solution of the system the machine M * obtained by replacing a( by r< produces a circuit <#'„ whose outcome for any uel" is the same as that of <$„. We can now deduce the existence of a point r = (r 1( ... ,r k )eQ* satisfying the system E„ such that each component has height polynomial in n. Therefore the computations done by M r over binary inputs can be carried out by a Turing machine in polynomial time (see [10, Lemma 3]). On the other hand, the circuit tfj can be readily transformed into a Boolean circuit having polynomial depth. From the classical equivalence between parallel polynomial time and PSPACE in the Boolean setting, we deduce that S belongs to PSPACE/po/y. On the other hand, and using the same equivalence, one trivially shows that any set in PSPACE/po/y can be accepted by a Pw-uniform family of circuits in weak parallel time. D Theorem 16. (i) The Boolean part of NP W D is NP/po/y. (ii) The Boolean part of N P W D ^ C O - N P W D is (NPnco-NP)/po/y. Proof. They are done in a similar manner as the preceding one.
D
Some consequences follow from the preceding theorems. In order to state them, let us recall that we denote by PH the polynomial hierarchy of Meyer and Stockmeyer and by Z{ its fcth level for any fceN (see [13] and [1, Ch. 18]). Corollary 17. (i) If PW = NP W D then the polynomial hierarchy collapses at its second level. (ii) / / NP W D = PARW then PSPACE = If. Proof. If P w = NP W D then we have that P/po/y = NP/po/y, from where we deduce (see [9] Theorem 6.1) the first statement. For the second statement we use that if NP W D = PARW then, since PARW is closed under complements, we must indeed have that (NP WD nco-NP WD ) = PAR w . This entails on the one hand that PSPACE/po/y £= NP/po/y and thus, because of a slightly modified version of [9, Theorem 4.2] (that can be found in [15, Corollary 4.29]) that PSPACE = r 3 . But on the other hand our assumption implies that N P s ( N P n c o NP)/po/y and thus, because of [8] 4.9 that PH = Zf. From both equalities we conclude the desired result. □
1505 12
F. Cucker et al.
4. The power of alternation A common computational resource in Complexity Theory is alternation. It consti tutes a strengthening of nondeterminism in the sense that the machine can now alternate existential - i.e. nondeterministic - guesses with universal ones. It is then not surprising that the complete problems for polynomial alternating time generalize the NP-complete problems in a very precise way. Thus, while in the Boolean setting the classical NP-complete problem - SAT - can be seen as the decision of the existential theory of Boolean logic, the most well known complete problem for polynomial alternating time turns out to be the decision of the unrestricted Boolean logic. A similar situation holds in the real setting (see [5]). In this section we shall see that there is a doubly exponential lower bound in the parallel time needed to solve some problems solvable in polynomial alternating time. Definition 18. We shall say that a set S is accepted in weak Polynomial Alternating Time if there exists a polynomial p and a machine M such that for every yeR°°, yeS iff 3x,Vz,• • • 3xpmVzp(|,.„M
accepts
(y,xuzu...,xpiM),zpm))
in weak time p(\y\) and we shall denote this fact by SePAT w . A variation of an argument already used in [16] and in [7] together with proposi tion 5 allows us to prove the following result. Theorem 19. PATW <£PEXPW. Proof. Let us consider the set S = {xe R001 x?2" = x2 where n = |x|}. Because of Proposi tion 5 this set is not in PEXP W . On the other hand, for any neN we consider the formula
1506 Separation of complexity classes in Koiran's weak model
13
5. The unordered case and the Knapsack problem In [12] it is shown that nondeterministic polynomial time is strictly more powerful than deterministic polynomial time for real Turing machines that only perform scalar multiplications and only branch on equalities. In [11] this result is improved by showing that the Knapsack problem can be solved in nondeterministic polynomial time but not in deterministic polynomial time by these kind of machines. In this section we further extend this last result to weak machines with the same kind of branching. In the rest of this section all the real Turing machines branch according with tests of the kind x=0. Let us recall from [2] that the real Knapsack problem is defined to be the set n
KP = {xeR°°|3u 1 ,...,u B 6r s.t. £ "i*i=l where n = |x|}. i- 1
Theorem 20. The Knapsack problem cannot be solved in weak polynomial time. Proof. Let M be a machine solving KP in time t(n). For any n we consider the polynomial //(*„...,*„)=
[] «>,
(blXi + - +
baX„-l)
b.)eF
that has degree 2". Clearly, for any (xu...,x„) we have that (xj x,)eKP iff H(x 1 ,...,x.) = 0. Now, for any n we consider the algebraic computation tree TM_H and its canonical path, which is obtained by answering # at all the branching nodes v. Moreover, let us consider the rational function / v ( X „ ...,X„, a,,...,a*) in the variables XU...,XH associated to each one of them and let F be the product of their numerators. The set of points following the canonical path is a n dimensional subset of R". Thus, since the set of points in R" satisfying KP has dimension n—1, we deduce that its corresponding leaf must be labeled REJECT. Thus, if a point xeR" is in KP then we have that F(x) = 0 i.e. the rational function F vanishes on the zero set-of H. But this implies that the degree of F must be bigger than the degree of H, and from the weakness of M we deduce that t(n)2 >2". □ If we denote by P^, and NP WD the classes of weak deterministic and digital nondeterministic polynomial time for real Turing machines that branch on equalities, our last result separates Pw from NP WD since KP is certainly in NP WD . Corollary 21. P W * N P W D . It is an open problem whether this separation holds for machines with arbitrary branching.
1507 14
F. Cucker et al.
References [1] J.L. Balcazar, J. Dias and J. Gabarro, Structural Complexity I, EATCS Monographs of Theoretical Computer Science, Vol. 11 (Springer, Berlin, 1988). [2] L. Blum. M. Shub and S. Smale. On a theory of computation and complexity over the real numbers: AT-completeness, recursive functions and universal machines. Bull. Amer. Math. Soc. 21 (1) (1989) 1-46. [3] J. Bochnak, M. Coste and M.-F. Foy, Geometrie algebrique reelle, Ergebnisse der Math., 12 (Springer Berlin, 1987). [4] F. Cucker, P R * N C R . J. Complexity 8 (1992) 230-238. [5] F. Cucker, On the complexity of quantifier elimination: the structural approach, The Computer Journal 36 (1993) 400-408. [6] F. Cucker and M. Matamala. On digital nondeterminism, preprint, 1993. [7] J.H. Davenport and J. Heints, Real quantifier elimination is doubly exponential, J. Symbolic Comput. 5(1988)29-35. [8] J. Kamper, Nonuniform proof systems: a new framework to describe nonuniform and probabilistic classes, Theoret. Comput Sci. (1991) 85 305-311. [9] R.M. Karp and R.J. Upton. Turing machines that take advice, Enseign. Math. 28 191-209 1982, 1988. [10] P. Koiran, A weak version of the Blum. Shub & Smale model, in: Proc. 34th Found. Comput. Sci. (1993) 486-495. [11] P. Koiran, Computing over the reals with addition and order, Theoret. Comput. Sci. 133 (1994) 35-47, this volume. [12] K. Meer, A note on a P # N P result for a restricted class of real machines, J. Complexity 8 (1992) 451-453. [13] A. Meyer and L. Stockmeyer, The equivalence problem for regular expressions with squaring requires exponential time, in: Proc. 13ih Symp. on Switching and Automata Theory (1973) 125-129. [14] J. Renegar, On the computational complexity and geometry of the first order theory of the reals, parts I. II and III J. Symbolic Comput. 13 (1992) 255-352 [15] U. Schoning, Complexity and Structure, Lecture Notes in Computer Science, Vol. 211 (Springer Berlin, 1988). [16] L. Slockmeyer and A Meyer, Word problems requiring exponential time, in: Proc. 5th Symp. on Theory of Computing (1973) 1-9.
1508
ON THE INTRACTABILITY OF HILBERTS NULLSTELLENSATZ AND AN ALGEBRAIC VERSION OF "NP # p?" MICHAEL SHUB AND STEVE SMALE Section 1. Introduction. In this paper, we relate an elementary problem in number theory to the intractability of deciding whether an algebraic set defined over the complex numbers (or any algebraically closed field of characteristic zero) is empty. More precisely, we first conjecture: The Hilbert nullstellensatz is intractable. The Hilbert nullstellensatz is formulated as a decision problem as follows. Given flt ..., ft: C -» C, and complex polynomials of degree d„ i = 1, ..., t, decide if there is a z e C" such that f,(z) = 0 for all i. There is an algorithm for accomplishing this task. From Hilbert, the answer is no if and only if there are polynomials g(: C* -»C, i = 1 t with the property
I *«/« = !•
C)
i-i
Brownawell [2] has made the most decisive next step by finding a good bound on the degrees of these gt. With that, one may decide if (*) has a solution by linear algebra, since (*) is afinite-dimensionallinear system with the gt's as unknowns. This procedure is called the "effective nullstellensatz." To say what "intractable" means in our conjecture, it is necessary to have a formal definition of algorithm in this context. That is done in Blum-Shub-Smale [1]. In that paper, algebraic algorithms (called algorithms over C) are described in terms of "machines," which make arithmetical computations and branch according to whether a variable (the first state variable) is zero or not. In [1], furthermore, one has the concept of a polynomial time algorithm, in particular, a polynomial time decision algorithm over C. This may be expressed in the present setting as T(f)<s(fY
(the c power of s(/) all / ) .
Here / = (fu...,/,) is the input, T{f) is the number of operations (arithmetic, branching) used to accomplish the decision, and s(f) is the total number of coeffiReceived 26 May 1994. The authors were partially supported by National Science Foundation grants. The work was carried out at Instituto de Matematica Pure e Aplicada, Rio de Janeiro. 47
1509 SHUB AND SM ALE
48
cients of the /, (input size). Also c is a universal constant The input size is
<"-£(";*)• Our conjecture is now formally the mathematical statement: There is no poly nomial time algorithm over C which decides the Hilbert nullstellensatz. An algebraic version of the computer science problem "NP # PT is also intro duced in [1]. From that paper, it follows that the nullstellensatz is a universal decision problem in a certain sense. It is "NP complete over C." It follows that "NP # P over C" if and only if our main conjecture is true. In other words, we may assert that the algebraic version of NP # P is true if and only if the Hilbert nullstellensatz is intractable. Valiant [9] also has an algebraic theory of NP completeness that differs from ours in his focus on "formula size," which is not equivalent to a computational notion. Moreover, his model is not uniform and does not permit branching on a variable x ^ 0. A computation of length ( of the integer m is a sequence of integers x0, xu ..., X/ where x„ = 1, x( = m and given k, 1 ^ k < I, there are i, j , 0 < i, j < k such that xk = x, o xj where ° is addition, subtraction, or multiplication. We define T: Z -> N (the natural numbers) by saying x(m) is the minimum length of a com putation of m. The following is easy to check. PROPOSITION.
r(m) =$ 2 log m.
If m is of the form 2 2 \ then r(m) = log log m + 2. The same is essentially true even if m is any power of 2. (All logs are to the base 2) We raised the question as to whether x{m) < (log log nif, where c is indepen dent of m. Welington de Melo and Benar F. Svaiter [6] showed by a counting argument that the answer is no. H. Lenstra also tells us that Jeff Shallit answered our question as well. Carlos Gustavo Moreira [7] subsequently gave quite sharp estimates on this problem. Yet our second question remains unanswered. Problem. Is there a constant c such that z(kl) < (log kf
allfc?
Given a sequence of integers ak, we say that ak is easy to compute if there is a constant c such that r(ak) ^ (log kf>a\\k> 2, and is hard to compute otherwise. We say that the sequence ak is ultimately easy to compute if there are nonzero integers mk such that mkak is easy to compute and ultimately hard to compute otherwise.
1510 INTRACTABILITY OF {ALBERT'S NULLSTELLENSATZ
49
MAIN THEOREM. / / the sequence of integers k\ is ultimately hard to compute, then the Hilbert nullstellensatz is intractable, and consequently the algebraic ver sion of "NP # P" is true.
To prove the Main Theorem, we consider an intermediate decision problem which we call twenty questions: Input {k, ht(k), z) e N x N x C. Decide if z E {1,2,..., k}. Here ht(k) is defined to be the largest natural number less than or equal to log k. THEOREM 1. / / the Hilbert nullstellensatz is tractable, i.e., if NP = P over C, then there is a machine M over C (in the sense of [I]) and a constant c such that Jt decides twenty questions in time bounded by (log kf. THEOREM 2. If a machine over C (in the sense of [1]) decides twenty questions in time bounded by (log kf for some constant c, then the sequence k\ is ultimately easy to compute.
The Main Theorem follows immediately from Theorems 1 and 2. Theorem 1 is fairly simple in our computational setting; its proof is carried out in Section 2. Most of the substance of our paper is in the proof of Theorem 2. For this T must be extended to polynomial rings. The algebraic and transcendental con stants used by the machine must be circumvented. These arguments are carried out in Sections 3 and 4. The complexity of deciding twenty questions was considered in a slightly dif ferent context in Shub [8]. The paper by Heintz-Morgenstern [5] is related to our work here. Section 2. Proof of Theorem 1 (of Section 1). We prove Theorem 1 by em bedding "twenty questions" in a decision problem {Y, Yyt,) which is in NP over C. Then if NP = P over C, (Y, Yy„) is in P, and there is a machine Jt that decides twenty questions in time bounded by (log kf, c a constant. Here J( is the restric tion of the machine which decides (Y, Yr„) in polynomial time. The decision problem (Y, Yy„) is described as follows. Y = C°° and Yytt = UkeN^e..* where Yyck = {(*. nt(k), zu...,
zM4)( 0,.. .)|zi e {1,..., k}}.
The embedding of twenty questions in (Y, Yytt) is simply (k, ht(k\ z) -* (z, ht(k\ z, 1,.... 1,0,0,...), where the number of ones is ht(k) — 1.
1511 50
SHUB AND SMALE
The proof of Theorem 1 is finished by the next proposition. PROPOSITION.
(Y, Yyes) is in NP over C.
Proof. The NP machine operates on variables (uu u2, z,,..., z„, w 0 ,..., wn, vi0, ...,vJn for j = 1, 2, 3,4). It checks if u2 is an integer by addition of ones. It checks if the input size (given with the input by definition) is 6w2 + 5. If so, n = u2. It checks if w„ = 1, Wi(w, - 1) = 0, and vSi(v}i — 1) = 0 for i = 0, .... n and ;' = 1, 2, 3, 4. It checks if "I =Z"=o2Vj. It sets x, = £"=o 2'tyi for ; = 1, 2, 3, 4. Finally, it checks if u i = 2 i + X;=i xj- If so> it outputs yes. Note that if the tests are verified, the w's and v's are 0 or 1; «1? the x}, and hence zx are nonnegative integers, and u2 = htiut). The time required is a constant times u2. Finally, we show that every element of Yy„k has a positive test. Let (k, ht(k), z,,..., zm), 0,...) e yyM>». Then z, is a nonnegative integer so that k — z is sum of four integers squared: k - z, = x\ + x\ + xl + x\.
D
Section 3. Easy to compute sequences in rings. In this section, we prove the facts about easy to compute sequences that are needed for the proof of Theorem 2 of the introduction. These concepts are close to those of algebraic complexity theory (see, for example, [4], [5]). Given a ring (or field) R and generators g0,..., g„ of R, a computation of length £ of the element r e R is a sequence of elements /•_„, ..., r0, rlt ..., r(, where r_; = 0, for 0 ^ j < n, rt = r. Moreover, given k between 1 and £ (inclusive), there are p, q with —n^p, q < k, such that rk = rp o rq, where o is the operation of addition, subtraction, or multiplication (or division by a nonzero element if R is a field). Define T = xg g : R -»N by i(r) is the minimum length of a computation of r. Note that the T: Z -»N of the introduction is a special case. PROPOSITION 1. Let (g0, ...,gn) and (h0,..., hm) be two sets of generators of a ring R. Then there is a constant c > 0 such that
\
oSr^\
*„(r) + c>
allreR
-
The proof is straightforward. Proposition 1 allows one to define hard and easy sequences of elements of R, independently of the choice of generators, exactly as in the introduction for Z.
1512 INTRACTABILITY OF HILBERT'S NULLSTELLENSATZ
51
PROPOSITION 2. Let G and H be finitely generated rings (or fields). Let j:G-*H be a ring homomorphism of G onto H. If gkeG is an easy to compute sequence, then so is
Proof. L e t « ! , . . . , e, be a set of generators of G. Then 4(ex\ .... ^(e,) is a set of generators of H. Thus, *«.,)
«*J.*(9k)) < t„
n(gk),
for all k.
D
PROPOSITION 3. Let R be a finitely generated integral domain and K its quotient field. (i) Iffk eK is an easy to compute sequence in K, then there are easy to compute sequences pk, qk in R such that fk = pjqk for all k. (ii) Let / t e K [ t , t„, klt..., k^] be an easy to compute sequence where tlt ..., tm are variables and At A„, are elements of an extension of K. Then there are easy to compute sequences pk e R[tlt..., t„, At, ..., A,] qk e R such that fk = Pk/
Proof. We can assume that the generators are glt..., gt, tu ..., t„ Xlt..., A^ where the gt generate R. Now use the instructions for computing fk to compute (pk, qk). For example, (Pk> 9k) + (Pj> Qj) = (Pt9j + PjQi, Qtlj)THEOREM 1. Let f, e Q(t, ku..., A„) be an easy to compute sequence of nontrivial rational functions in the variable t and transcendentally independent complex numbers kt km. Then there is an easy to compute sequence of integral poly nomials Pi e Z[f] such that p( ^ 0 for all i and for z e Q, pt(z) = 0 whenever fi(z>K - 0 = 0.
For the proof we use two lemmas. LEMMA 1. Let a polynomial fe Z[t0, tu..., t„] have degree d. Iff is zero on every integer point in the cube in R*+1 centered at (0,..., 0) with side having length (d + l\ then f is identically zero.
The proof is a straightforward induction on m. LEMMA 2. Let fieZ[t,kl,..., Am] be an easy to compute sequence of nontrivial integral polynomials in the variable t and transcendentally independent complex numbers klt..., km. Then there is an easy to compute sequence of nontrivial integral polynomials pt 6 Z[f\ such that for z e Q, pt(z) = 0 whenever ft{z, ku..., k^) = 0.
For the proof, we may assume that 1, t, kt, ..., A. are the generators of Z[t, kly..., A,,] that we use for defining computational length. Let n, be the computational length of /,. Then the degree of /, is less than 2"< + 1. Using Lemma 1, considering the A, as variables, there is an (m + l)-tuple
D
1513
52
SHUBANDSMALE
of integers (fcto, ...,klm) = k such that /,(k to ,..., fc,J / 0 and \kUJ\ ^ 2"< + 1, for all i, j . By the proposition of Section 1, T(ktJ) < 2(n, + 1). Write /,(t, Xu.... AJ = Z / a i./( f )^ / where / is a multi-index (a finite sum, of course). Since XJtj= 1,..., m, are transcendentally independent, f,(z, A 1( ..., Am) = 0 for a rational number z if and only if a(/(z) = 0 for all /. Let k[ = (kn,.... fcfc,). Then if /,(z, Aj A.J = 0 for some rational z, we see that p{(t) = Enu./WW) 7 vanishes at t = z. Note also that p,(k,iP) / 0. Finally, by computing ky first and substituting ku, j = 1, ..., m in the in structions for computing /,, p, is computed with computational length at most nt + 2m(n, + 1), and so p, is an easy to compute sequence. Now we return to the proof of Theorem 1. By Proposition 3 (i), onefindseasy to compute sequences p„ q[ in Q[t, Xu..., Am] such that p'Jql = f{. By Proposi tion 3 (ii), we find easy to compute sequences p" e Z[t, kx A,,], q, e Z such that p\ = Pt/q'i'. Thus p" is not zero. Also p'{{z, k\,..., Xm) = 6 if z is rational and /,(z, A,,..., Xm) = 0. Now using p" in Lemma 2finishesthe proof. Section 4. Proof of Theorem 2 (of Section 1). The proof is preceded by two propositions. Let a machine M solve a decision problem (K*, Y) where K is a field of char acteristic 0 branching on x = 0 or x # 0. Suppose there is a sequence nk of posi tive integers with M halting at time T(k) on inputs of size nk. Let K"* c K°° be the nt-fold cartesian product of K and Yk= Yn K\ Under the hypothesis that Yk is a proper, nonempty subvariety of K\ we define the Jfcth canonical path as the computation path which at each branch node is taken by a Zariski-dense set of inputs in X"*. Thus, a canonical path may be described as a certain sequence yiy2'"}V> t < T(k) where each y, is a branch node, and yJ+l is the node encountered by almost all inputs subsequent to y}. We omit y, in the case that all inputs arriving at y} take the same branch (see Cucker-Shub-Smale [3]). Branching is determined by a condition xx = 0 or not. Then Xj is represented by a rational function G} defined almost everywhere on X"*. It is easy to see that the computational length of Gj is bounded by c1j + c2 where cuc2 are constants. Let Hk = TlGj. PROPOSITION 1. The rational function Hk defined almost everywhere on X"* vanishes on Yk, but is not identically zero. Its computational length is bounded by CjTW + Cz.
Proof. The machine must answer no on a Zariski-dense set of points of K**, so Yk must be contained in the union of the varieties V} = {x\Gj(x) = 0}. This proves the first assertion. The last assertion is a special case of the remark preceding Proposition 1. PROPOSITION 2. Let M be a machine over a field K which is afinitealgebraic extension of afield K. Then there is a machine M over K and a constant c > 0 such
1514 INTRACTABILITY OF HILBERT'S NULLSTELLENSATZ
that for any decision problem (y, y y J, 7 c K°°, decided byJ^YnK™, is decided by Jt; and, moreover, the stopping time satisfies: TAy)*cTj(y),
53 Yy„ n K°°)
yeYnK°°.
Proof. We may assume that at any computation node, the computation per formed is either addition, multiplication, subtraction, or division of two elements of K. We may regard K as a vector space over K of some fixed dimension q. Thus, K can be represented as K* where the embedding K c K i s the inclusion of K in IC as the first coordinate. Now addition and multiplication are represented byfixedsymmetric bilinear maps B + :K«x K « - K « Bx:K"xK''
-*K*.
Division of b by a is accomplished by solving the linear system Bx(a, y) = b for y by Gaussian elimination. This requires on the order of q3 steps. To define Jt, replace K°° by (X*)00. The input of Kx in K°° is replaced by the input K°° as the first coordinates in (X4)00. Multiplication nodes are replaced by Bx and addition nodes by B+. Subtraction nodes are replaced by — 1 followed by B+. Branching is done on the coordinates of K1. One can see using the isomorphism between Kq and K that on inputs ye Y n K"", Jt gives the same answers as Jt with the desired time bound where c is on the order q3. D Proof of Theorem 2. Assume that Jt decides twenty questions in time bounded by (log kf. Let nlt..., \i( be the nonrational constants of Jt so that we may view Jt as a machine over Q(^,,..., /*,). Now Q(/i,,..., /*,) is a finite alge braic extension of afinitelygenerated purely transcendental extension Q(A t ,..., AJ of Q. Thus, by Proposition 2, there is a machine M over Q(A,,..., XJ) which solves twenty questions restricted to Q(A U ..., Am). Thus, on the input of (k, ht(k), z), 2 e Q ( ^ , ...,Am), ^decides if z e {l, ...,&} in time bounded by c1(logfc)c. Then, by Proposition l, there are nontrivial rational functions fkeQ(t, A.lt ...,Xm) which vanish on {l,...,/c} and whose computational length in Q(t, Xlt .... Xm) is bounded by c2(log kf + c3. Here c3 is a bound for the computational length of rational constants introduced by the machine M, and c 3 depends only on M. Therefore, fk is an easy to compute sequence of nontrivial rational func tions. By Theorem l of Section 3, there is an easy to compute sequence of nontrivial polynomials pk e Z[t] vanishing on {l,..., k}. By Lemma 1 of Section 3, there is an integer m such that pk(m) # 0, \m\ «* 2e + 1 where t is the computa tional length of pk. So T(TO) < t + 1. We may assume |m| is minimal with these properties. Then pk is zero at each integer between zero and m. Evaluating pk at m gives a computational length of at most 2/ + 1 for pk(m). Since pk{m) is an integral multiple of k\, the sequence k\ is ultimately easy to compute. □
1515 54
SHUB AND SMALE REFERENCES
[1] L. BLUM, M. SHUB, AND S. SMALE, On a theory of computation and complexity over the real numbers: NP-completeness, recursive functions and universal machines, Bull. Amer. Math. Soc. 21 (1989), 1-46. [2] W. D. BROWNAWELL, Bounds for the degrees in the Nullstellensatz, Ann. of Math. (2) 126 (1987), 577-591. [3] F. CUCKER, M. SHUB, AND S. SMALE, Separation of complexity classes in Koiran's weak model, to appear in Theoret. Comput. Sci. [4] J. HEINTZ, "On the computational complexity of polynomials and bilinear mappings: A survey," in Applied Algebra, Algebraic Algorithms and Error-Correcting Codes, Lecture Notes in Comput. Sci., 356, Springer-Verlag, Berlin, 1989,269-300. [5] J. HEINTZ AND J. MOROENSTERN, On the intrinsic complexity of elimination theory, J. Complexity 9 (1993), 471-498. [6] W. DE MELO AND B. F. SVAITER, The cost of computing integers, to appear in Proc. Amer. Math. Soc. [7] C. G. MOREIRA, On asymptotic estimates for arithmetic cost functions, to appear. [8] M. SHUB, "Some remarks on Bezout's Theorem and complexity theory" in From Topology to Computation: Proceedings of the Smalefast, ed. by M. W. HntscH, J. E. MARSDEN, and M. SHUB, Springer-Verlag, New York, 1993,443-455. [9] L. G. VALIANT, "Completeness classes in algebra," Conference Record of the Eleventh Annual ACM Symposium on Theory of Computing, ACM, New York, 1979, 249-261. SHUB: T. J. WATSON RESEARCH CENTER, IBM RESEARCH, YORKTOWN HEIGHTS, NEW YORK 10598,
USA; [email protected] SMALE: DEPARTMENT OF MATHEMATICS, CITY UNIVERSITY OF HONG KONG, TAT CHEE AVE., KOWLOON, HONO KONG; [email protected]
1516
uteroational Journal of Bifurcation and Chaoi, Vol. 6, No. 1 (1996) 3-26 5 World Scientific Publishing Company
COMPLEXITY AND REAL COMPUTATION: A MANIFESTO** LENORE BLUM* International Computer Science Institute, 1947 Center St., Berkeley, CA 94704, USA E-mail: lblumQicsi.berkeley.edu FELIPE CUCKER* 1 Universitat Pompeu Fabra, Balmes 138, Barcelona 08008, Spain E-mail: cucker6upf.es MIKE SHUBl IBM T.J. Watson Research Center, Yorktovn Heights, NY 10598-0818, USA E-mail: [email protected] STEVE SMALE* Mathematics Department, City University of Hong Kong, Tat Chee Ave., Kowloon, Hong Kong E-mail: [email protected] Received August 10, 1995; Revised August 18, 1995 Finding a natural meeting ground between the highly developed complexity theory of computer science — with its historical roots in logic and the discrete mathematics of the integers — and the traditional domain of real computation, the more eclectic less foundational field of numerical analysis — with its rich history and longstanding traditions in the continuous mathematics of analysis — presents a compelling challenge. Here we illustrate the issues and pose our perspective toward resolution.
1. A i m The classical theory of computation had its origins in work of logicians — of Godel, Turing, Church, Kleene, Post, among others — in the 1930s. The model of computation that developed in the follow-
ing decades, the Turing machine, has been extraordinarily successful in giving the foundations and framework for theoretical computer science. Our point of view is that the Turing model (we will call it "classical") with its dependence on 0's
'Webster: A public declaration of intentions, motives, or views. 'This article is essentially the introduction of a book with the same title to appear shortly. 'Partially supported by the Letts-Villard Chair at Mills College. 'Partially supported by DGICyT PB 920498, the ESPRIT BRA Program of the EC under contract nos. 7141 and 8S56, projects ALCOM II and NeuroCOLT. 'Partially supported by NSF grants. 3
1517 4 L. Blum et oi.
and l's is fundamentally inadequate for giving such a foundation to the theory of modern scientific com putation, where most of the algorithms — with ori gins in Newton, Euler, Gauss, et al. — are real number algorithms. Our viewpoint is not new. Al ready in 1948, John von Neumann, in his Hixon Symposium lecture, articulated the need for "a de tailed, highly mathematical, and more specifically analytical theory of automata and of information". In this lecture, von Neumann was particularly crit ical of the limitations imposed on the "theory of automata" by its foundations in formal logic: There exists today a very elaborate sys tem of formal logic, and specifically, of logic as applied to mathematics. This is a disci pline with many good sides, but also serious weaknesses Everybody who has worked in formal logic will confirm that it is one of the technically most refractory parts of mathe matics. The reason for this is that it deals with rigid, all-or-none concepts, and has very little contact with the continuous concept of the real or of the complex number, that is, with mathematical analysis. Yet analysis is the technically most successful and bestelaborated part of mathematics. Thus for mal logic, by the nature of its approach, is cut off from the best cultivated portions of mathematics, and forced onto the most dif ficult part of the mathematical terrain into combinatorics. The theory of automata, of the digital, all-or-none type as discussed up to now, is certainly a chapter in formal logic. It would, therefore, seem that it will have to share this unattractive property of formal logic. It will have to be, from the mathematical point of view, combinatorial rather than analytical. We propose a formal theory of computation which integrates major themes of the classical the ory and builds on the classical foundations, yet at the same time is more mathematical, perhaps less dependent on logic, and more directly applicable to problems in mathematics, numerical analysis and scientific computing. We propose a theory of real computation. We propose to do this in a way which preserves the Turing theory as a special case of the new the ory, i.e. by an appropriate choice of the fields ad mitted. In this way, results from computer science
give insight to numerical analysis and the reverse holds as well.
2. Seven Examples 2.1. 2.2. 2.3. 2.4. 2.5.
The Mandelbrot Set A Julia Set Newton's Method The Knapsack Problem The HUbert Nullstellensatz as a Decision Problem 2.6. 4-Feasibility 2.7. Linear Programming, Integer Programming.
2.1.
Is the Mandelbrot decidable?
set
The British mathematical physicist, Roger Penrose, in the popular book, "The Emperor's New Mind", p. 124, writes: Now we witnessed... a certain extraor dinarily complicated looking set, namely the Mandelbrot set (Fig. 1). Although the rules which provide its definition are surprisingly simple, the set itself exhibits an endless vari ety of highly elaborate structures. Could this be an example of a non-recur sive (i.e., undecidable) set, truly exhibited before our mortal eyes?
Fig. 1. The Mandelbrot set. (Courtesy of H.-O. Peitgen and Springer-Verlag.)
1518 Complexity and Real Computation: A Manifesto 5
Fig. 2. On the boundary in Seahorse Valley (magnification of a tiny portion of the right boundary within the aquare inset in Fig. 1). Courtesy of H.-O. Peitgen and Springer-Verlag.
It is known that the boundary of the Mandel brot set has a rich and complex structure. (See for example Fig. 2 where a part of this boundary is shown.) Hence Penrose's query seems reasonable. Penrose is motivated to ask this question to make an argument against artificial intelligence. While we find this use of mathematics not com pelling, the question of the decidability of the Man delbrot set has another justification. It can partly answer and give insight to the question: Can one decide if a differential equation is chaotic? The Mandelbrot set M is defined as the set of complex numbers c such that the sequence c, c 2 -t-c, (c2 + c) 2 + c , . . . remains bounded. More formally, for c 6 C, the complex numbers, let p c (z) = z 2 + c and let p"{z) be the nth iterate of p c applied to z. That is, P?(*) = Pc(- • • Pcbc{Pc(z)))), 2
2
So p c (0) = c, p (0) = c + c,.... complement of the set
n times. Then M is the
M' = {c e C|p?(0) -♦ oo as n -> 00} . The set M may also be described as the set of all inputs c which do not halt for the flow chart in Fig. 3.
This is because if ever the sequence c, tp + c, (c2 + c) 2 + c escapes the disk of radius 2, it will go off to infinity.' To answer Penrose's query, one needs a "ma chine" or "algorithm" that, given input c, a com plex number, will decide in a finite number of steps whether or not c is in M. (See Fig. 4.) After asking his question, Penrose acknowledges being somewhat inexact (p. 125). The classical theory of computation presupposes that all the un derlying sets are countable and hence ipso facto cannot handle these questions about subsets which are uncountable. Next, Penrose seeks ways to bypass this prob lem. One way is to use computable real numbers (to describe the appropriate complex numbers). This would be the approach of recursive analysts, an area 'This fact is utilized in designing computer algorithms for drawing "pictures" of the Mandelbrot set: Let N be a large integer. For given point c, generate up to S elements of the sequence c, c2 + c, (c2 + c) 2 + c,... along with their magnitudes. If and when some magnitude is greater than 2, color c white, else color c black. Note that white points are definitely in M' while black points are possibly in M with our confidence level partly dependent on Af. (For more sophisticated algorithms, see the book by Peitgen and Saupe (1988]).
1519 6 L. Blum et at.
Input c
x,i/)«-(0,c ) (x,y) <-(l 2 + y,y) ■
Ms«2? - No Yes Halt and Output 1 Fig. 3.
A flow chart associated with the Mandelbrot set.
Input c
. c€M> Yes/ Output 1 Fig. 4.
\ ^No Output 0
Desired decision machine for M.
originating with the early work of Turing. Here one might imagine a Turing machine being input a real number bit by bit. Using its internal instructions, the machine op erates on what it sees, possibly every so often outputting a bit. The resulting sequence, if any, would be considered in the limit the (binary expansion of the) real output. Problems arise here when one wants to decide if two numbers are equal and so Penrose rejects this approach. As he points out on p. 126, "One impli cation of this is that even with such a simple set as the unit disc,..., there would be no algorithm for deciding for sure . . . whether the computable number i 2 + y1 is actually equal to 1 or not, this being the criterion for deciding whether or not the computable complex number x + iy lies on the unit circle. . . . Clearly that is not what we want." Another tack might be to consider the rational or algebraic skeleton of the problem. Thus, we could
rephrase Penrose's question: given a complex num ber c whose real and imaginary parts are rational, or algebraic, decide whether or not c is in M. Indeed, this has been a tack used by theoretical computer scientists to deal with problems whose natural un derlying spaces are the real numbers (such as the linear programming problem) or the complex num bers. However, this approach is also problematical. For example, the curve i 3 + y3 = 1 has no rational points with both x and y positive. So, the rational skeleton provides no useful information about the given curve. After exploring several such approaches, Penrose p. 129 concludes: " . . . one is left with the strong feeling that the correct viewpoint has not yet been arrived at." Thus Penrose's question, "Is the Mandelbrot set decidable?" makes no sense! Now note that the flow chart of Fig. 3 could be interpreted as a machine with "halting set" which is precisely the complement M' of M. (The set M' might be said to be "semi-decidable".) This machine has the power to accept complex num bers, perform basic arithmetic operations on com plex numbers and to compare magnitudes. It is an example of a machine (to be defined formally later) over the real numbers K not C since it uses the real comparison \z\ > 2. What is not clear but is true, is whether there is not a similar kind of machine with M as its halt ing set. The theory of real computation proposed makes precise and formal some of the suggestions here. In this theory, Penrose's question becomes well-defined and we answer it: the Mandelbrot set is not decidable over IR. 2.2.
Example
of the Julia
set
of T(z) = 3? + 4 We are looking at a polynomial map T : C —» C from the point of view of iteration or as a complex dynamical system. Thus we write T 2 (z) = T(T(z)) and Tk for the composition of T with itself k times. Let us specify T(z) = z2 + 4. Observe that if \z\ > 2, \Tk(z)\ -> oo as k -» oo. Consider the flow chart in Fig. 5. Call this machine M. In M there are 4 nodes (the boxes) which are called as we descend in the di agram: input node, computation node, branch node and output node. Again we have an example of a machine over the real numbers IR since its branch ing depends on real inequality comparisons.
1520 Complexity and Real Computation: A Manifesto 7
Input z
Input z
z +- T(z)
zeJ?
Ho
Yes
1*1 > 2 ?
Halt and output 1
No
Y«s
Fig. 6.
J is a halting set (supposing J is decidable).
Output z Fig. 5. A Julia set flow chut.
The halting set QM of M is the set of inputs z C C such that by following the flow of the flow chart, we eventually halt (or output). For example QM contains the set of all z with \z\ > 2. Moreover 0, ±1, ±2 are all in flu. However fixed points of T (so T(z) = z, or z2 + 4 = z) are not in the halting set. In fact any periodic point of T (so Tk(z) = z some k = 1, 2, 3,...) is not in fijf. A little thought will show that QM is an open set of complex numbers, so that J = C - HM must contain the closure of the set of periodic points of T. J is the Julia set of T in the terminology of complex dynamical systems and it can be proved that J is the closure of the set of periodic points of T and is homeomorphic to a Cantor set. A question again suggested by the classical the ory of computation is: Is J decidable, or equivalently, is there a real machine with the halting set J? To see the equivalence of these questions, note that a decision machine for J can be converted into a machine with halting set J as indicated in Fig. 6. On the other hand, suppose J were the halting set of some machine Mj. Recall, the flow chart machine M given in Fig. 6 has halting set C - J. We construct a decision machine for J by hooking together M and Mj (see Fig. 7). To decide if z € J, input z into both M and Mj and run the two machines in tandem. One and only one of these machines will halt, the one that does decide the membership of z. (Schematically, we have indicated a parallel process. This could be turned into a sequential process by alternating in turn between operations of M and Mj.)
Input z
My
M If M halts
Output Yas
Output Ho Fig. 7.
J is decidable (supposing J is a halting set).
With a formal development of machines over IR, one can answer the above questions (the answer is, J is not a halting set and hence not decidable). 2.3.
Newton's
method
The previous two examples raised questions con cerning the existence of machines that would decide Yes or No to queries of the form: Given z € C, is z € M (or J)? On the other hand, often we want algorithms that search for solutions to problems of the form: Given a polynomial /, find C such that /(C) = 0. Newton's method is the "search algorithm" sine qua non of numerical analysis and scientific compu tation. Here we briefly recall Newton's method for finding (approximate) zeros of polynomials in one variable. Given a one variable polynomial f(z) over the complex numbers C, define the Newton endomorphism Nf : C -» C by: Nf(z) -
/(«) '/'(*)'
This map is defined as long as f'(z) / 0.
(1)
1521 8 L. Bhm et ai.
Input z
zi-Nf(z)
l/WI < e ? No Y«s Output z Fig. 8. The Newton machine for / .
Now for Newton's method: Pick an initial point ZQ € C and generate the orbit ZQ, *i = Nf(zo), 22 =Nf(zi),...,
2*+1
plies there are local neighborhoods about the zeros of / that contract under (iterates of) Nf. Hence, any point z 6 C that, under the action of Nf, even tually enters one of these contracting neighborhoods will eventually approach a zero of / . This is the ba sis of Newton's method. If C is a simple zero of / , then N'.(Q = 0, i.e. C is a superattractingfixedpoint of Nf. It follows that there is an open set of points whose orbits un der the Newton endomorphism eventually converge quadraticaily to (. That is, starting at any of the points in this open set, Newton's method will even tually converge to £, and moreover, at some stage, the precision of the approximation will double with each successive iteration. Decidability questions also arise in this context. It is well known that Newton's method is not gen erally convergent. The main obstruction to general convergence is the existence of attracting periodic points of period at least 2. For example, consider the cubic polynomial f(z) = z3 - 2z + 2.
= yv/(.k) = yv*+1(2o),... Some stopping rule such as "stop if |/(2*)| < c and output Zk" is given. In practice, if the procedure has not stopped with an output after a certain number of iterates, or it becomes undefined at some stage, a new initial point is chosen. We may represent Newton's method schemati cally as in Fig. 8. In this simple machine, we assume that / , Nf and e are built in. An initial point z 6 C is "input" to the machine. Later we will consider machines which allow / and E to be input as well. Proposition 1. (a) /(C) = 0 if and only if Nf(Q = C and (b) Nf(0 = C implies |JVJ(C)| < 1. Part (a) of the proposition follows directly from the formula defining Nf (and noting that there are polynomials g and h such that jr = % and if /(C) = 0, then p(C) = 0 but ft(C) * 0). To show (b) we observe that, for C a zero of / of multiplicity m,
To see this, note that N', = ffm and evaluate it using the Taylor expansion of / about (, /(z) = am(z - O m + higher order terms a™ / 0. Thus the zeros of / are the fixed points of Nf and the fixed points of Nf are attracting. This im
Here Nf(z) = z - x\'J'-f ■ So Nf(0) = 1 and Nf(l) = 0, i.e. 0 is a point of period 2 under the Newton endomorphism. Also, by the chain rule we seCi (Nf)'(Q) = 0. So there is a neighborhood of points about 0 whose orbits under the Newton map fail to converge to a zero of / . These points are juperattracted to the periodic orbit: 0, 1, 0, 1, So, given / , it is natural to ask: Is the set of "good" starting points for Newton's method decidable? Here we will say that a point z € C is good if its orbit under Nf converges to a zero of / . It can be shown that the set of good starting points coincides with the halting set of some Newton machine M (with built in e that depends on / ) . So as before we can rephrase our question: Is the set of bad points a halting set? With a theory of real machines doing exact arithmetic it can be shown that the answer is no. We include here a picture indicating the good and bad points for Newton's method applied to the polynomial f(z) = (2 2 -l)(2 2 +0.16). In this picture (see Fig. 9), bad points are colored black and the set of them includes the Julia set for the Newton endomorphism. 2.4.
The Knapsack
problem
Now suppose R is a commutative ring with unit. For specificity, one may suppose R is the ring of
1522 Complexity and Real Computation: A Manifesto 9
Fig. 9. The dynamics of the Newton endomorpbiam for f(z) Springer- Verlag.
integers Z, the rationals Q, the reals R, or the com plex numbers C. Consider the following problem: Given x i , . . . , x n 6 R decide if there is a non-empty subset S C { 1 , . . . , n} such that
Zies « = 1We shall call this problem the Knapsack Problem (KP) over R. In the classical theory, it is also known as the Subset Sum Problem: Given positive integers i i , . . . , i n , c decide if there is a subset 5 C { 1 , . . . , n} such that E i e s x> = cHere we imagine c to be the capacity of a knapsack, x, the weights of given items, and the question is: Can one fill the knapsack to capacity with some subset of the items? Note that in our formulation of the Knapsack Problem, the ring R need not be ordered. Thus, the machines we consider here branch only on equality comparisons over R. In many ways, the Knapsack Problem is similar to our previous decidability problems. Let Kn be the Knapsack set which we write in the following form: Kn = {xeRn\3b€{0,
1}" such that £ 6 ^ = 1 } .
(2) Now we are seeking a machine that will decide, given x 6 ft" if x 6 K„.
(z2 - 1)(*2 + 0.16).
Courtesy of S. Sutherland and
But unlike our previous problems, KP is eas ily seen to be decidable: Given x € R", succes sively enumerate the non-zero elements of {0, 1}" and evaluate the corresponding Y. t>ixi- If and when an evaluation is 1, halt and output 1 (yes). Other wise, halt and output 0 (no). (See Fig. 10.) Given input x G R" this algorithm (or machine) will stop with a correct decision in at most 2 n - 1 enumerations. Fixing n, we can view the input space for our algorithm as the finite dimensional space R". How ever, since the algorithm is "uniform" in n, we can naturally view the input space as R°°, the infi nite direct sum of R (essentially, the space of fi nite, but unbounded sequences over R). Thus, let ting n vary, we have really sketched an algorithm to decide membership in K = \J K„. In the fi nite dimensional case, the polynomial tests at the branch nodes can be built into the machine. In the "infinite-dimensional" case, the tests at the branch nodes are computed by subroutines. The exhaustive search algorithm does not at all seem satisfactory. We may very well be unlucky and have to go through an exponential number of iterations before we halt with an answer. The big question is: can we do better? Is there an algorithm for deciding the Knapsack Problem in a polynomial number of "steps"? By this we mean, is there a uni form machine M and a positive integer c such that
1523 10 I Blum et at. Input (n, x) j«-l ENUMERATION SUBROUTINE with input (n,j), output b>, the jth non-jero sequence in {0,1}"
E2.i*J* = i ' Halt and output 1
j = 2" - 1 j «- j + 1
Halt and output 0
Fig 10. Exhaustive search machine for solving the Knapsack problem. for each n and x e fl", M will decide if i 6 K in less than n c steps? We will be more precise in the sequel as to how to deflne a uniform machine and how to count steps. For the moment, it is the number of nodes we traverse (counting multiple visits) in a flow chart machine from input to output, given input x. Notice that the exhaustive search algorithm does not use multiplication. A major unsolved prob lem is: Can multiplication speed up the process? Suppose R has no zero divisors. Let kn(x) € R[xi,..., x„] be the polynomial
*»(*)= II
(£6,*,-i)
(3)
6€{0.1}'
and V*. = { i € Rn\kn(x) = 0}. Then Vk„ = K„. So, for each n we could construct a machine with ife„ built in. (See Fig. 11.) The input space of this machine is fl". It de cides if x e K immediately (in one step) by evaluat ing the polynomial kn(x) and checking if the result is 0 or not. The degree of kn(x) is 2" - 1. So we have traded an exponential search for an evaluation of a polynomial whose degree is exponential in n. If R is ordered, a big question remains: Do order comparisons enable significant speed-ups?
Input (n, z)
n (E^- 1 ' 1 >€(0,1)" i - 1
Halt and output 1
Halt and output 0
Fig. 11. An algebraic machine for solving the Knapsack problem. While it may be hard to decide in general if i € K we can quickly verify it is, if we are given a good "witness" 6 € {0, 1}". Just sum up the corresponding li's. We express this property of the Knapsack Problem by saying KP is in the class "NP". Over the integers, KP is a universal problem with this property - it is "NP-complete over Z".
2.5.
The Hilbert as a decision
NulUtellensatz problem
Newton's method tackles the search problem: Given a polynomial / over C, find a zero of/.
1524 Complexity and Real Computation: A Manifesto 11
Here we consider the decision problem: Given a finite set of polynomials /; in n vari ables over C, decide if the /; have a common zero. By the Fundamental Theorem of Algebra, for one polynomial in one variable, the answer is always yes. This is not true for several polynomials in several variables. We call this decision problem the Hilbert Nullstellensatz over C and we denote it by HN/C. Thus, one seeks an algorithm, in fact an algebraic algo rithm over C, which on input / = {/i, • • •, /*} pro duces ye if and only if there is a C € Cn such that /i(C) = 0 for all t. By an algebraic algorithm we have in mind a machine whose computations, like Newton's meth od, involve the basic arithmetic operations, but whose branching now depends only on equality comparisons and not on order, i.e. a machine now over C. The input / to such a machine can be thought of as the vector of coefficients of the fc in C" where N is given by the formula
"-§("*)•
di = d e g / j ,
t = l,...,i
Thus, this N represents the size S(f) of the in put/. There are algorithms which accomplish this task. In the linear case, i.e. when deg/, = 1 for t = 1,..., k, this is a simple linear algebra prob lem. In the general case, Hilbert has shown that the answer is no if and only if there exist polynomifcls 9i i • • •, 9k in n variables with the property
£ 9ifi = 1
(4)
where the equality is equality as polynomials. Bounds on the degrees of the g, have been proved and in fact we may take degj, < D" where D = max{3, d i , . . . , in). Thus, Eq. (4) becomes a fi nite dimensional linear algebra problem; to find the coefficients of the ,. As one can readily see, the number of Gaussian Elimination steps required is exponential in the size S(f) of the input vector. This suggests a problem which we formulate as a main conjecture. Conjecture. intractable.
The Hilbert Nulhtellensatz over C is
By this we mean, there is no algorithm which solves HN/C with the number of arithmetic opera tions A(f) satisfying the bound
A(f) < S(f)c where c is a universal constant. But a precise for mulation of the conjecture awaits a formal definition of machine over C. With a formal theory of machines over C it can be shown that HN/C is in NP over C by noting that given / and a test point { € C , we may test f for a zero, i.e. whether ft(Q = 0 for » = 1,..., k, in a number of arithmetic operations that is polyno mial in S(/). Moreover HN/C can be shown to be universal with this property so that any problem in NP over C can be quickly reduced to HN/C. That is to say HN/C is NP complete over C. From these considerations it can be shown: The Hilbert Nullstellensatz over C is intrac table if and only if P ^ NP over C. The above sketches our ideas on formulating complexity issues in an algebraic framework. How does this relate to the P # NP problem of classical complexity theory? By abuse of notation write HN/Zj for the problem: Given a finite set of polynomials /; in n vari ables over Z3, decide if the /< have a common zero in (Zj)". By considerations similar to the above, HN/Zj can be seen to be intractable if and only if P ^ NP (in the classical sense). Since our algorithms in the above formulation are defined in terms of algebra (characteristic 2 algebra), it could be said that we are placing classical complexity into an algebraic setting in addition to extending it to new problems such as P / NP over C. These new problems are interesting in their own right. But also, by posing problems such as P ^ NP within a broader frame work, we may be able to employ new mathematical tools to study the classical case as well as gain new insights by analogy or by direct connections. 2.6.
Feasibility
of real
polynomials
Let us replace the field of the complex numbers by the real numbers in the problem above and con sider polynomials / 1 , . . . ,/* € R[Xi,..., X„]. The problem at hand now is to decide if there exists a
1525 12 I. Blum et al.
common root ( € R " . Since the reals have a nat ural order, it is natural here to consider algorithms that branch on order comparisons. A particular feature of the real case, not shared by the complex one, enables us to consider the same problem with only one polynomial at the cost of slightly increasing the degree. We associate with the polynomials / i , . . . , / * the single polynomial g = 53?=i Si- Now g has the property that for ev ery f € i t " , £ is a common root of all the U 'f a ° d only if g(£) = 0. Therefore, solving our problem for the fi turns out to be equivalent to solving it for g. Moreover, if not all the / , are linear then the de gree of g is at least 4. Let us restrict our attention to degree 4 polynomials and consider the following problem denoted by 4-FEAS: Given a degree 4 polynomial in n variables with real coefficients, decide whether it has a real zero. Again, the input g for this problem can be seen as a vector in RN where
»-(*:') is the size S(g) of this input. An algorithm for solving this problem was first given by Tarski in the context of exhibiting a deci sion procedure for the theory of real numbers. In the context of complexity theory, Tarski's algorithm is highly intractable. The number of arithmetic op erations performed by this algorithm grows in the worst case by an exponential tower of n 2's. Later on, Collins devised another algorithm that solved 4-FEAS within a number of arithmetic operations bounded by More recent algorithms achieve single exponential bounds. These algorithms are quite elaborate. Again, this suggests a problem which we for mulate as another conjecture. Conjecture.
The 4-FEAS problem is intractable.
That is, we conjecture that there is no algo rithm which solves 4-FEAS with the number of arithmetic operations A(f) satisfying the bound
AS) < SUY where c is a universal constant. But again, a pre cise formulation of the conjecture awaits a formal definition of a machine over R .
As with the Hubert Nullstellensatz, 4-FEAS is seen to be in NP over R . It is also universal with this property, that is, 4-FEAS is NP complete over1R. 2.7.
Linear integer
programming and programming
Over the reals consider the problem: Given a set of m linear inequalities in n variables AiX >b{
i = 1,..., m
where An = ^2 oijXj ,
dij € R and 6i e R
(5)
decide if there is a point x € R n satisfying (5). We call this problem the (real) linear programming feasibility (LPF) problem. We now write the system of inequalities (5) as Ax>b, where A is the m x n matrix whose ith row is Ai and b the m-vector whose tth entry is 6,. Then we may rewrite L P F / R as: Given the m x n real matrix A and 6 € R m , decide if there is an x € R n such that Ax > b. This problem is simpler than M N / R in that the functions are linear, but more complicated by the fact that we consider inequalities. The set of solu tions of Ax > b is called a polyhedron. The (real) linear programming optimization (LPO) problem is: With input (A, b, c) minimize c • x subject to Ax > b where A is an m x n real matrix, 6 € R m and c 6 R n , or decide no minimum exists. Replacing the reals by the integers or the rationals in LPF we have the integer programming (IPF) and the rational linear programming feasibil ity problems: Given an m x n matrix A with integer (re spectively rational) entries a,, and 6 € Z m
1526 Complexity and Real Computation: A Manifesto 13
(respectively Q m ), determine if there is an x e Z n (respectively Q n ) such that Ax > b. The corresponding integer programming (IPO) and rational linear programming optimization prob lems are: minimize c • x subject to Ax > b or determine that no minimum exists. Here A is an m x n integer (or rational) matrix, 6 6 Z m (or Q m ), c € Z" (or Q") and x € Z m (or Q m ). Algorithms are known which solve all six of these problems, but there is a great deal of differ ence in what is known about the efficiency of algo rithms which solve them. Integer programming is set apart from real or rational linear programming by the fact that the solution of linear equations with integer coefficients are not necessarily integers, e.g. 2 i = 1. But it is also set apart from the reals by the notion of the size of the input. For the reals, the input size S(A, b) or S(A, 6, c) of the problem is the number of real variables involved, mn + m for the feasibility prob lem and mn + m + n for the optimization problem. No algorithm is known for either the feasibility or the optimization problem with the total number of arithmetic operations A(A, b) or A(A, 6, c) satisfy ing the following complexity bounds:
entries of the matrix A and 6, the components of the vector b. For the optimization problem we take input size Sht(A, b, c) to be equal to S(A, 6, c) times the max imum height of all the integers
The situation for the rationals, rational linear programming feasibility and optimization, is dra matically different. The height ht(x) of a rational number x = | where p and q are relatively prime is defined as max(ht|p|, ht\q\). Otherwise the defini tions of the input sizes Snt(A, b), S/,t(A, 6, c), and costs Cnt(A, 6), Cht(A, b, c) are the same as in in teger programming. There are algorithms for both problems and a constant d > 0 such that
A(A, b) < S(A, b)d or A(A, b, c) < S{A, 6, c)d for a universal constant d. Thus an outstanding problem is: Is linear programming tractable over R ? Again, to make this precise one needs a formal definition of an algorithm, or machine, over H . Complexity estimates for algorithms for integer programming traditionally take the binary lengths of the integers used into account both in the input size of a problem "instance" and the "cost" of the algorithm in that instance. The height of an integer i , ht(x), is the first integer greater than or equal to log(|i| + 1). For the feasibility problem we take input size SM(J4, 6) to be equal to be S(A, b) times the max imum height of all the integers dy, i = 1 , . . . , m, j = 1 , . . . , n and ij, i = 1 , . . . , m where dy are the
6, c)d.
C„,M, 6) < Sht(A,
b)d
and Ch,(A, b, c)<Sht(A,
3.
6, c)d.
T h e Classical T h e o r y of Computation
As we have noted, the classical theory of compu tation had its origins in the work of logicians in the 1930s. Of course at that time, there were no computers as we know them. While this work, in particular Turing's (1937), clearly anticipated the development of the modern digital general purpose computer, a primary motivation for the logicians was to formulate and understand the concept of de cidability, or of a decidable set. In particular, the aim was to make sense of such questions: "Is the set of true sentences of arithmetic decidable?" or
1527 14 L. Blum et at.
FINITE STATE CONTROL
"^
Tape
iwmw
/
read-write head
ifltnioiiiiioiiioin
lllllBl
Fig. 12. A Turing machine.
"Is the set of diophantine equations with integer solutions decidable?" 2 Intuitively, a set S is decidable if there is an "effective procedure" that given any element u of U (some natural universe containing S) will decide in a finite number of steps whether or not u is in 5, i.e. if the characteristic function of S (with respect to U) is "computable". To put the first query in this format, U would be the set of arithmetic sentences, S the true ones. For the second, U would be the set of polynomials with integer coefficients and S the subset of those with integer solutions. The models of computation designed by these logicians were intended to capture the essence of this concept of effective procedure or computation. The idea was to design theoretical machines with operations, and finitely described rules for proceed ing step by step from one operation to the next, so simple and constructive that it would be self-evident that the resulting computations were effective. A number of distinct formal models of compu tation were proposed. A primary example is the Turing machine. (See Fig. 12). Here we have a finite state control device with a read-write head and a two-way infinite tape con sisting of an infinite number of cells. The control device is regulated by a program which is a finite set of instructions of the form (g, s, o, q'). Here q and 'The latter question is known as Hilbert's Tenth Problem, posed by David Hilbert (along with 22 other seminal prob lems) at the second International Congress of Mathemati cians in Paris on August 8, 1900 It was originally taken for granted, by mathematicians in general, and Hilbert in par ticular, that the answers to the above questions were both affirmative. The queries were actually posed as tasks: Pro duce decision procedures for the given sets. The incompleteness/undecidability results of Godel in 1931 in the first place, and of Matiyasevich in 1971 on the unsolvability of Hilbert's Tenth Problem in the second, show such tasks eannot b* car ried out.
q' belong to a finite set {qo, ■ ■ ■, g/v} called the set of states of the machine, s is a symbol 0, 1, or B (for blank), and o is one of the following operations: R (move right one cell), L (move left one cell), 0 (print 0), 1 (print 1), or B (print B). The instruction is interpreted as follows: If the device is in state g with read-write head scanning a cell containing symbol s, then the device performs operation o and goes into state q'. (If o is a print operation then it is implicit that the head erases the current symbol before printing.) We assume the program is con sistent, i.e. for each q and s there is at most one instruction starting with the pair (q, s). The machine operates as follows: Given an in put string x, a finite sequence of 0s and Is written on consecutive tape cells (with Bs everywhere else), the head is placed over the leftmost symbol of i . (If i is the empty string the head is placed on any cell.) The control device is started in initial state q0 and proceeds according to the program instruc tions until it can no longer proceed (i.e. it reaches a state q while scanning a symbol s for which there is no instruction starting with the pair q, s). If and when this occurs the output is the string of 0s and Is starting at the current scanned cell (and going right) until the first occurrence of a B (which may be the current cell, in which case the output is the empty string). One might consider input and out put strings as natural numbers written in binary (a convention here could be that the empty string is interpreted as 0 and a non-empty input string al ways has 1 in its left most place). Here, and in each formalism for computation, a function / from the natural numbers N to N is defined to be computable if it is the input-output map of some such machine. Thus we can now say formally: a set of natural numbers is decidable if its characteristic function is computable, in this case by a Turing machine.
1528 Complexity and Real Computation: A Manifesto 15
A fundamental object of study is the halting set of a machine. This is the set of all inputs for which the machine halts, i.e. produces an output. It is clear that the halting sets are exactly the semidecidabte sets: a set S of natural numbers is semidecidable if there is a machine which outputs 1 when input an element of 5, and otherwise outputs 0 or does not halt. A little "programming'' shows that S is decidable if and only if both it and its comple ment are semi-decidable. (Schematically, see Figs. 6 and 7.) This notion of computability can be naturally extended to the integers, Z, the rational numbers, Q, or any domain that can be "effectively encoded" in N. Thus, for example, by "godel" coding sen tences (of a first order language) as natural num bers, one can begin to formally ask (and answer) within the formalism questions about the decidabil ity of the set of true sentences of various mathemat ical theories. It is quite remarkable that even though the for malisms we just described and the others proposed were often markedly different, in each case, the re sulting class of computable functions — and hence decidable (as well as semi-decidable) sets — was exactly the same. Thus, the class of computable functions appears to be a natural class, indepen dent of any specific model of computation.3 And consequently, the answers to the basic questions of decidability will be independent of formalism. This gives one a great deal of confidence in the theoretical foundations of the theory of com putation. Indeed, what is known as Church's the sis is an assertion of belief that the classical for malisms completely capture our intuitive notion of computable function. Thus for example, in the light of Church'8 thesis, the negative solution to Hilbert's Tenth Problem can be gotten by showing there is no Turing machine to decide the solvability in integers of diophantine polynomials. Compelling motivation clearly would be required to justify yet a new model of computation. 4. Toward a Mathematical Foundation of Numerical Analysis Our perspective is to formulate the laws of compu tation. Thus we write not from the point of view of *In fUitiml terminology, theee functions are often called the recursive function!, decidable lets are the recursive tett and ■emi-decidable aeti are the recursively enumerate tett.
the engineer who looks for a good algorithm which solves his problem at hand, or wishes to design a faster computer. The perspective is more like that of a physicist, trying to understand the laws of sci entific computation. Idealizations are appropriate, but such idealizations should carry basic truths. Scientific computation is the domain of compu tation which is based mainly on the equations of physics. For example, from the equations of fluid mechanics, scientific computation helps understand better design for airplanes, or assists in weather pre diction. The theory underlying this side of compu tation is called numerical analysis. There is a substantial conflict between theoreti cal computer science and numerical analysis. These two subjects with common goals have grown apart. For example, computer scientists are uneasy with calculus, while numerical analysis thrives on it. On the other hand numerical analysts see no use for the Turing machine. The conflict has at its roots another age-old conflict, that between the continuous and the dis crete. Computer science is oriented by the digital nature of machines and by its discrete foundations given by Turing machines. For numerical analysis, systems of equations, and differential equations are central and this discipline depends heavily on the continuous nature of the real numbers. The developments described in the previous (and next) sections have given a firm foundation to computer science as a subject in its own right. Use of Turing machines yields a unifying concept of algo rithm, well-formalized. Thus this subject has been able to develop a complexity theory which permits discussion of lower bounds of all algorithms without ambiguity. The situation in numerical analysis is quite the opposite. Algorithms are primarily a means to solve practical problems. There is not even a formal defi nition of algorithm in the subject. One is reminded of how the development of the definition of differentiable manifold was so important in the history of differentiable topology. The history of algebraic geometry gives us a similar lesson. Thus we view numerical analysis as an eclec tic subject with weak foundations; this certainly in no way denies its great achievements through the centuries. A major obstacle to reconciling scientific com putation and computer science is the present view of the machine, i.e. the digital computer. As long as the computer is seen simply as a finite or discrete
1529 16 L. Blum tt at.
object, it will be difficult to systematize numerical analysis. We believe that the Turing machine as a foundation for real number algorithms can only obscure concepts. Toward resolving the problem we have posed, we are led to expanding the theoretical model of the machine to allow real numbers as inputs. There has been great hesitation to do this because of the dig ital nature of the computer. Here, we might learn a lesson from the history of science. In particular, Isaac Newton was faced with an analogous problem in writing his Principia. At the time of Newton, scientists assumed that the world was atomistic, as viewed by the ancient Greek, Democritus. Newton accepted that picture according to which all matter is composed of indivisible particles, a finite number in each bounded region. On the other hand, New ton's mathematics was continuous as was Euclid's. Moreover, the differential equations Newton needed for his theory involved calculus and the continuum, contrasting with the corpuscular view of the uni verse. It was a substantial problem for Newton to reconcile the discrete world with the continuous mathematics. The resolution was produced by an alyzing the effect of replacing an object (e.g. the earth) by a finite number of particles, then mak ing a better approximation with a larger number of particles. In the limit, the mathematics becomes contin uous. Thomas Kuhn in The Copernican Revolution writes: In 1685, Newton proved that, whatever the distance to the external corpuscle, all the earth corpuscles could be treated as though they were located at the earth's centre. That surprising discovery, which at last rooted gravity in the individual corpuscles, was the prelude and perhaps the prerequisite to the publication of Principia. And Kuhn adds: At last it could be shown that both Ke pler's Law and the motion of a projectile could be explained as the result of an in nate attraction between the fundamental corpuscles of which the world machine was constructed. Now our suggestion is that the modern digital computer could be idealized in the same way that
Newton idealized his discrete universe. The ma chine numbers are rational numbers, finite in num ber, but they fill up a bounded set of real num bers (e.g. between -1000 and 1000) sufficiently densely that viewing the computer as manipulating real numbers is a reasonable idealization, at least in a number of contexts. Moreover, if one regards computer graphical output such as our picture of the Mandelbrot or Julia sets with their apparently fractal boundaries and asks to describe the machine which made these pictures one is driven to the idealization of machines which work on real or complex numbers in order to give a coherent explanation of these pictures. For a wide variety of scientific computations the contin uous mathematics which the machine is simulating is the correct vehicle for analyzing the operation of the machine itself. These reasonings give some justification for tak ing as a model for scientific computation, a machine model which accepts real numbers as inputs. Of course a great many issues such as round-off error must be dealt with. Moreover the ultimate justifi cation is: does the model developed this way give new insights and understanding to the use of the big machines? 5. Classical Complexity Theory and Its Extension A cornerstone of classical complexity theory is the theory of NP-completeness and the fundamental P * NP? problem. A main goal of this book is to extend this theory to the real and complex numbers, and in particular, to pose and investigate the fundamental problem within a broader mathematical framework. The foundations for such a theory shall be de veloped in the next chapters. But here we give some background and briefly and informally intro duce some of the classical notions. Since the 1930s, much of the work of logicians focused on identifying and classifying decidable and undecidable problems. A prevailing view was that once a problem was known to be decidable (or solv able), then by and large, it was not terribly deep or interesting. In contrast, there was a great deal of interest and activity designed to untangle and understand the rich hierarchy amongst the unde cidable problems (the "degrees of unsolvability"). To relate the notion of solvable problem to our earlier discussion of decidability (in Sec. 3), we can
1530 Complexity and Real Computation: A Manifesto 17
view a decision problem as a pair (X, Xy„). Here X is the set of problem instances, and Xy^ the sub set of yes-instances. Thus X plays the role of the universe U and Aye, of the subset 5. The problem is decidable (or solvable) if XytM is decidable (by a machine that on input z € X will output 1 if x is in Xye. and 0 if not). So, for example, for Hubert's Tenth Problem, X would be the set of diophantine equations (polynomial equations with integer coef ficients) and Xy„ the subset of those with integer solutions. With the advent of the digital computer, and its promise of solving hitherto intractable problems, interest perked in the realm of the solvable with the quest for efficient algorithms. Although there were many successes, it soon became apparent that a number of problems (such as the famous Travel ing Salesman Problem) while solvable in principle, defied efficient solution. These problems seemed in essence intractable. Thus, amongst the solv able, there appeared to be yet another rich and natural hierarchy, with the dichotomy of tractability/intractability mirroring the earlier dichotomy of decidability/undecidability. And so, the theory and field of computational complexity was born.4 The foundation of this theory was developed in the 1960s, primarily by researchers originally trained in mathematics and logic but who found
'Again in his Hixon Symposium lecture, von Neumann voiced the need for such a theory: Throughout all modern logic, the only thing that is important is whether a result can be achieved in a finite number of elementary steps or not. The si2e of the number of steps which are required, on the other hand, is hardly ever a concern of formal logic. Any finite sequence of correct steps is, as a matter of principle, as good as any other. It is a matter of no consequence whether the number is small or large, or even so large that it could not possibly be carried out in a lifetime, or in the presumptive lifetime of the stellar universe as we know i t . . . . [On the other hand] in the case of an automaton the thing which matters is not only whether it can reach a certain result in a finite number of steps at all but also how many such steps are needed. A primary concern here for von Neumann was his conviction that the cumulative effect of the small but non-zero prob ability of component failure "may (if unchecked) reach the order of magnitude of unity — at which point it produces, in effect, complete unreliability." In fact, to the contrary, the phenomenon of error build-up due to computer failure haB not posed difficulties anywhere near the magnitude posed by the (apparent) intractability phenomenon.
more hospitable environments for these interests in the newly emerging computer science departments. The theory began in an abstract setting with the formulation of axiomatic complexity measures yield ing surprising speed-up theorems, and then became more concrete with the NP-completeness results in the early 1970s. It is primarily this latter work, showing the equivalence of literally thousands of of ten seemingly unrelated difficult problems, that has captured the attention of researchers from many fields. These problems have the property that an efficient solution to any one can be easily converted to an efficient solution to any other. The formalisms of classical complexity theory are founded on the models and formalisms of clas sical computation theory. Formal measures of com plexity are intended to indicate various degrees of difficulty inherent in problems. These difficulties could be measured by the amount of information necessary to describe a problem (descriptional or informational complexity), the power of the lan guage needed (descriptive complexity), or the amount of resources, such as time or space, required to solve the problem. In this book we will pri marily follow the tradition of computational com plexity which studies the cost of computation with regard to time, or number of steps, for solution. The complexity of a problem is then measured in terms of the complexity of machines for solving it. Paramount here is that complexity is given as a function of input word size L, classically measured in bits. A machine M is said to be in class P if there are positive integers c and q such that for all in puts i , costM(x) < c(size(i))'. Here costjvf(x) denotes the number of 6o»»c oper ations performed by machine M from input x to output. A decision problem (X, Ayes) is in class P, or solvable in polynomial time, if it is decidable by a machine in class P. Polynomial-time is an attempt to capture a no tion of tractability and is what is meant in this discussion when we use qualifiers such as "quick", "efficient", "short" and "fast". Note that to give upper bounds on complex ity or to show a problem is tractable it is sufficient to demonstrate one appropriate machine. On the other hand, to claim a lower bound g for complex ity or that a problem is not in class P (and hence intractable) is more problematic. For now we must
1531 18 L. Blum ct al.
demonstrate that every machine for solving it has a complexity function that grows faster than g or, in the latter case, faster than any polynomial. The more subtle concept of class NP is meant to capture the notion that some problems have the property that for each yes-instance there exists a quick verification, or short proof, of this fact. Since a quick decision also serves as a quick veri fication, that class P is contained in class NP. It is natural to ask the converse: If a yes-instance has a short proof, can we find some such proof quickly? This is the essence of the fundamental P = NP? problem. Again, as in the case of decidability and coraputability, for all this to be reasonable and natural, we must have some degree of assurance that these notions and classes are independent of most "rea sonable" formalisms. To illustrate some of these ideas, we consider probably the most well known problem of classical complexity theory, the Traveling Salesman Problem (TSP): Given n cities, the distances (a^) between them and a positive number k, does there ex ist a tour through all the cities with total dis tance less than or equal to k? and the related Shortest Path Problem (SPP): Given n cities, the distances (aij) between them, two specified cities I and m, and a pos itive number k, does there exist a path from I to m with total distance less than or equal to Jt? The SPP is solvable in order n 2 operations.5 On the other hand, the TSP appears not at all to be tractable. All known solutions essentially re quire us to enumerate the (n — 1)! possible tours. By Sterling's formula, n! is asymptotically equal to (n/e)n\/2itn which is exponential in n. Let's look at these problems a bit more for mally. First, we can easily pose them as decision *We indicate a solution to the special case when the distances between distinct cities are either 1 or k + 1 : At stage 0, label city I with the number 0. At stage s + 1, label all unlabeled cities that are distance 1 from the cities labeled s by the number a + 1. If no such cities exist, terminate process and answer uyesn if city m is labeled by a number < k, otherwise answer ano."
problems in the above form. For example, for the TSP let X = {(A, k)\A — (aij) is an n x n matrix of distances, k > 0} and Xye, ={(A, it) € X|there is a tour T with Dist(i4, T) < Jt} Here T is a cyclic permutation of {1, 2 , . . . , n} and Dist(yt, T) = Efei 1 a T , n+1 + aTnTl. Notice that X is the set of all problem instances, for all n. This reflects the fact that we are interested in solving problems uniformly. Notice also that, in stating these particular problems, we have made no assumption that the distances are integers; it makes perfectly good sense to talk about these particular problems over the re als or any ring with order. Over R, a natural measure of the size of a TSP or SPP instance would be n 2 (the number of entries in the matrix A of distances). Over Z, a more nat ural measure would be n2b where b is the maximum of the heights (or binary lengths) of the distances (aij) and k. This size roughly reflects the number of symbols needed to describe the instance (or its bit length) and is essentially the classical measure. The classical measure of cost, the bit cost, is the number of Turing machine operations for solution. Thus the bit costs of the above solutions for SPP and TSP are of order n26 and (n - 1)!, respectively. Over the reals, a natural measure of cost could be the number of arithmetic computations and com parisons, and so for the above solutions to SPP and TSP, of order n 2 and (n - 1)!, respectively. Thus, over R or over Z, the cost of these solutions (as a function of size of instance) is linear in the case of SPP and exponential in the case of the TSP. Now the TSP, while not known to be in class P over any ordered ring, is seen to be in class NP in any reasonable sense. Although we may not be able to easily tell if a TSP instance has a "good" tour (i.e. one of total distance bounded by it), if we are handed a good one we can quickly check it out: First check that it is indeed a tour and then sum up the n distances along the tour and compare with k. The TSP is HP-complete over Z, i.e. it is uni versal for NP problems over Z: If (X, Ay«.) is a problem in NP over Z, then problem instances
1532 Complexity and Real Computation: A Manifesto 19
x e X can be efficiently encoded as Travelling Sales man instances Tz such that x € Xy^ if and only if Tx has a good tour. Thus an efficient solution to the TSP will yield an efficient solution to any other NP problem. Hence the importance of the TSP in clas sical complexity theory — not only because it is one of the ubiquitous problems of discrete optimization, but also because of its NP-completeness! The Knapsack Problem (KP) introduced in Sec. 2.4 can also be posed as a decision problem in the above form. Let n>0
Xym=[x€X\3b€{0,l}n
such that £ f c x i = l } (6)
Over any ring R, KP is in class NP in any reasonable sense. Moreover, the Knapsack Problem is also NPcomplete over Z and hence equivalent to the TSP with respect to complexity. We propose a theory of NP and NP-complete ness over an arbitrary ring. In such a framework one can obtain both old and new NP-completeness results. The classical satisfiability problem is NPcomplete over the field Z2. The integer program ming problem is naturally NP-complete over Z. In non-classical domains the Hilbert Nullstellensatz is NP-complete over C or I t Hierarchies of complexity classes can be developed over the real numbers. 6. Complexity Theory in Numerical Analysis It is natural to be skeptical about machines using exact arithmetic in numerical analysis. Most nu merical problems can only be solved to within an accuracy of t. Round-off error is an important fact in the use of actual machines for solving scientific problems. Does it make sense to try to extend the complexity theory of computer science to numerical analysis? We recall that computer scientists say that an algorithm defined by a machine M is tractable (or polynomial time, or in P) if the computation time T(x) associated to input x satisfies the bound T(x) < c(8ize(x))* all inputs x
(7)
where the constants c and q depend only on M. Here time is the number of Turing machine opera
tions and size is the number of bits. A problem is tractable if there is a tractable algorithm solving it. Some of the algorithms of numerical analysis are quite immediately tractable in a natural ex tension of this definition. Consider the problem of solving a linear system of equations Ax = b. The input of the problem is a non-singular n x n ma trix A and a vector 6 € R n . Gaussian elimination produces an output x, solving this problem in less than en3 arithmetic operations. Therefore one can apeak of the "tractability" of Gaussian elimination where Turing operators are replaced by arithmetic operations (and comparisons). Also the size of the input now becomes naturally the number of input variables. Thus T(A, b) < c(size(A, B)) 3 ' 2 for Gaussian elimination. More generally in numerical analysis it is im portant to take into account the desired accuracy c of an approximate solution. This is because most problems cannot be solved exactly, even using exact arithmetic. Thus one must modify the concept of tractable and one way to do this is consider e < 1 as an additional (special) input to the problem. Then one demands that the time T of computation satisfy T(e,x)<(|log£|+size(x))«,
e
(8)
Much recent work on solving non-linear equa tionsfitsinto this framework, where even sometimes I log e| is replaced by log | log e\ in (8). Frequently among the set of inputs to a prob lem, there is a subset of "ill-posed" problems where the main algorithms fail and may even fail in prin ciple. In general as an input gets closer to this illposed set the time of computation becomes larger. A "condition number", a function on the in put, has been defined traditionally to deal with this phenomenon. If the condition number of a certain input is large, then the time of computation can be expected to become large and the effects of round off error to become substantial. It has often been shown that there is a relation between the condi tion number of an input and the reciprocal of the distance to the ill-posed set. Then the desired com plexity results have the form T(e, x) < (| log e| + log n(x) + size(z))« where /1 is the condition number of x.
1533 20 L. Bhtm et al.
7.
Summary
As we have said, our proposal is to develop a theo ry of machines which will take real numbers as inputs. Generally speaking, mathematical theories are built on plausible abstractions and simplifications intended to capture the essence of, rather than precisely describe, phenomena they are attempt ing to model. We hope that the basic assumptions reflect fundamental underlying principles, and that the results inferred from these assumptions reveal new truths. Justification for our proposal will ultimately depend on how well this last task is accomplished. The basic arithmetic operations (+, —, x, /) are to be taken as primary in the structures of com putation. This point of view bestows an algebraic emphasis so that it becomes natural to suppose that the inputs and states of the machines are numbers (or finite sequences of numbers) in a field (mathe matical sense of the word). In the main case this is the field of real numbers. But certainly the field of complex numbers is also important. There are natural situations where division can not be done as within the integers Z. So to cover those cases we propose a model of computation of machines over a ring. Now formulating a theory of computation in this manner, i.e., over a field K, one can include and extend the classical theory by taking K = Zi (the field of 2 elements). In this way, the classical theory takes on an algebraic setting. By choosing A" to be the real numbers R, we are able to obtain a setting which provides a foundation of numerical analysis. The notion of an algorithm over R, be comes well-defined as a mathematical object in its own right. So we will have developed an extension of the classical theory to a new theory which can be specialized to the study of real number algorithms. This theory by the nature of our development is pri marily algebraic. More precisely, when the field K is an ordered field as is the case of R, the compar isons include <, and the geometry becomes what is called semi-algebraic. The classical algorithms of mathematics and of computer science naturally fit into this framework. It is important to remark that a fundamental property of classical computation is that the ma chines are finite objects, even though they operate on inputs which have no a priori bound on size. This property is satisfied by the machines suggested here.
8. Brief History The ideas we have presented are at the confluence of different traditions in mathematics and computer science. On the one hand, there is the work of classical computability and complexity. The initial motivat ing force here was — as we have already remarked — the question of the decidability of the arithmetic, and also the tenth problem posed by Hilbert at the second International Congress of Mathematicians in 1900. A common characteristic of these prob lems is the possibility of expressing their underly ing objects (arithmetic sentences and diophantine equations) in a language over a finite alphabet. Not surprisingly, the host of theoretical computational models that were subsequently proposed to formal ize the notion of decidability were designed to act on finite strings over a finite alphabet.' This was the case with the general recursive functions of Kleene [1936), the A-computable func tions of Church [1936], the computable functions of Turing [1936] and the canonical systems of Post [1943], to mention just the most influential models. Perhaps less expectedly, all models were equivalent in the sense that they defined the same class of com putable functions. This gave rise to Church's thesis discussed in Sec. 3. On the other hand, there is a long standing tra dition of decidability results in algebra and analysis that we refer to as the numerical tradition. This theory leads to algorithms — like Newton's method discussed in Sec. 2 and Gaussian elimination for solving linear systems of equations — as well as to several undecidability results. A paradigm here is Galois' result on the nonsolvability by radicals of polynomial equations of degree 5 or more. It is im portant to notice that these algorithms manipulate real numbers in much the same way proposed here. With the arrival of the digital computer, atten tion shifted from decidability to complexity issues, and the first of the traditions described above pro duced a sophisticated theory of complexity of which the P versus NP question described in Sec. 5 became central. We owe to this research concepts and tools that enable us to classify computational problems 'Furthermore, at the turn of the century, with recent discov eries of paradoxes in the foundation of mathematics well in mind, there was an understandable preoccupation with ques tions of consistency. This surely was a factor in stipulating, in the early computational models, that the simplest operations were to be performed on the simplest objects.
1534 Complexity and Real Computation: A Manifesto 21
into complexity classes reflecting different resource requirements, and then to discover structural rela tions among these classes. During the 1960s, Rabin [1960b], Hartmanis and Stearn [1965] and Blum [1967] developed the notion of measuring the complexity of a problem in terms of the number of steps required to solve it with an algorithm. This led in a natural way to the association by Cobham [1964], Edmonds [1965] and Rabin [1966] of the concept of "feasible" or "tractable" problems to the class P. Simultaneously, it was observed that a large class of search problems seemed to defy the existence of algorithms signifi cantly better than brute force. Independently Cook [1971] and Levin [1973] characterized this class (Cook named it NP) and proved the existence of complete problems for it. Cook exhibited the first NP-complete problem, the Satisfiability Problem of propositional logic. Shortly afterwards, Karp [1972] showed that a series of familiar problems from dif ferent areas of discrete mathematics were also NPcomplete. This gave strong impetus to the subject that was reflected, on the one hand, in work ex hibiting hundreds of NP-complete problems and, on the other hand, in attempts to prove the inequality P ^ NP leading to results on the structure of the class NP. A lively exposition on the P versus NP question (containing a large list of NP-complete problems) can be found in the already classic book by Garey and Johnson [1979]. A survey of the state-of-theart of this question is given in Sipser [1992]. In this latter article, a recently discovered letter of Godel to von Neumann dated 1956 is reproduced in which Godel stated the P versus NP question in the form of the time required by a Turing machine to test whether a formula of the predicate calculus has a proof of a given length. The rise of complexity issues in the numerical tradition is less attached to the advent of the dig ital computer. Early in 1937, in a short note of Scholz [1937], complexity questions arose under the form of the number of additions needed to produce a given integer starting from 1. Seventeen years later Ostrowski [1954] conjectured the optimality of Homer's rule for evaluating univariate polynomials. In order to do so, he defined a formal model of com putation and associated to it an idea of cost. This was followed by a flow of results concerning lower bounds (including the proof of Ostrowski's conjec ture by Pan [1966] for computational models with the following two characteristics:
(i) they take their inputs from i f where R is a ring, and (ii) their basic operations are arithmetic and com plexity is measured by how many such opera tions are performed. In most cases, the ring R was chosen to be the field of real numbers IR and this choice, together with the second characteristic above, reflected the kind of computations done in numerical analysis. However, these models were essentially nonuniform. This fact, useful for the search of lower bounds, becomes an obstruction to developing a theory of complexity for general purpose algorithms. Two very influential papers at the end of the 1960s were those of Winograd [1967] and Strafien [1969]. They helped to make this search for lower bounds in al gebraic problems an independent subject of study, now known as algebraic complexity. Some central examples of algebraic computational models along with lower bounds for them are given by Steele and Yao [1982], Ben-Or [1983] and Smale [1987]. Two early books on algebraic complexity are the ones by Borodin and Munro (1975] and by Winograd [1980]. A recent survey of the subject can be found in Strafien [1990]. Complexity issues are at the forefront of current research related to designing algorithms for finding zeros of polynomials and determining the solvability of polynomial systems. Amongst the major refer ences here are: Collins [1975], Shub k Smale [1993a, 1993b, 1993c, 1993d, 1994], Schonhage [1982], BenOr, Kozen k Reif [1986], Renegar [1987, 1992], Pan [1987, 1995], Grigor'ev k Vorobjov [1988], Canny [1988], and Heintz, Roy k Solerno [1990]. This work may be considered as the modern counterpart to algorithmic investigations begun earlier in the cen tury by Hermann [1926], Van der Waerden [1949], and Tarski [1951], in particular related to elimina tion theory for real closed fields. While the history of numerical analysis provides us with a great mo tivating force towards our efforts here, we will only give the reference [Goldstine 1977]. The computational model proposed in this ar ticle has a candidate in Blum, Shub k Smale [1989]. This candidate lies on the traditions of both com puter science and numerical analysis since it in corporates the universality of universal machines and of NP-complete problems, while keeping the as sumptions of the numerical one (real numbers given as an entity and unit cost of arithmetical opera tions) that make it suitable for modelling continu ous algorithms. There is a growing body of work —
1535 22 L. Blum tt al.
by Cucker [1992a, 1992b, 1993], Koiran [1993], Meer (1990, 1992, 1993, 1994], Michaux [1989, 1991], and Poizat [1995] among others — giving a broad devel opment to this point of view. In addition to the work already described, there are many more contributions by mathematicians and computer scientists which predate the afore mentioned model. We proceed now to review some of them. Close to the classical approach, Rabin [1960a] developed a theory of computable algebra and fields in which the underlying domains can be effectively coded by natural numbers and are thus, necessarily countable. On the other hand, the theories of computa tion over abstract structures, are quite general. See e.g. Engeler [1967] (contained also in Engeler [1993]), Friedman [1971] (or as discussed by Shepherdson in Harrington et al. [1985]), Tiuryn [1979], and Moschovakis [1986]. These general approaches both exploit and explore the logical properties of pro cedures. But, when applied to specific structures such as the reals, they do not yield the concrete mathematical results (such as the undecidability of the Mandelbrot set, NP-completeness of the Hilbert Nullstellensatz) that will quite naturally follow from the model we propose here. There is yet another possible approach to the complexity of real valued problems known as re cursive analysis, originating with Turing's seminal paper [Turing, 1936]. Indeed, in this paper Turing introduced the notion of computable real numbers before, and as a means to, defining computation over the integers.7 Here the machine model is the classiThis fact seems not well known, so it is of considerable his torical interest to examine the very first paragraph of Turing's paper: The "computable" numbers may be described briefly as the real numbers whose expressions as a decimal axe calculable by finite means. Although the subject of this paper is ostensibly the computable numbers, it is almost equally easy to define and investigate com putable functions of an integral variable or a real or computable variable, computable predicates, and so forth. The fundamental problems involved are, how ever, the same in each case, and I have chosen the computable numbers for explicit treatment as involv ing the least cumbrous technique. I hope shortly to give an account of the relations of the computable numbers, functions, and so forth to one another. This will include a development of the theory of functions of a real variable expressed in terms of computable numbers. According to my definition, a number is computable if its decimal can be written down by a machine.
cal Turing machine and one deals with real numbers that, roughly speaking, are fed to the machine bit by bit. This contrasts with the numerical tradition where real numbers are viewed not as their decimal (or binary) expansion, but rather as mathematical entities. Text references for recursive analysis are Ko (1991] for complexity matters and Weihrauch [1987] for computability issues. Other references are Friedman k Ko [1982], Pour-El & Richards [1983], Hoover [1987] and Kreitz k Weihrauch [1982]. More closely related to our perspective are the register machines of Shepherdson k Sturgis [1963] and the RAM's or random access machines. These were originally defined over the integers (see Aho, Hopcroft k Ullman [1974] for a definition) but with an algebraic character in their ground operations. An extension of the RAM to the real numbers is sug gested in the book of Preparata k Shamos [1985]. The goal of the model is primarily to describe algo rithms in computational geometry and the formal development of a theory of computability or com plexity is not pursued. Also, in the book mentioned above by Borodin k Munro [1975], the authors state that their underlying model of computation will be the RAM. However, immediately afterwards they say that this "code can be 'unwound' and separate programs can be written for each 'degree' of the desired class of functions", justifying therefore the subsequent use of models of fixed dimension. Again, the formal development of a theory of computabil ity and complexity over fields is not pursued. A different algebraic/logical approach to computabil ity, called the combinatory programme, has been developed by Engeler and his students [1995]. Perhaps closest to our approach is the work of Herman k Isard [1970] on computability over arbi trary fields. Here, some finite dimensional problems over the reals are shown to be undecidable in a man ner similar to our proof of the undecidability of the Mandelbrot set. Also close is the work of Tucker [1980] and Tucker k Zucker [1992] who employ the theory of computing over abstract structures to ob tain computability and non-computability results in line with ours. Friedman k Mansfield [1992] have also specialized the abstract theory to specific struc tures to good avail. Another model, again close in spirit, is a the ory of real Turing machines outlined by Abramson [1971]. The machine model developed here can op erate on arbitrarily long vectors of real numbers. The main thrust of the article is to develop a
1536 Complexity and Real Computation: A Manifesto 23 hierarchy of non-computable functions according to their use of a greater-lower-bound operation. Yet another approach is information-based complexity, developed in Traub et al. [1988]. A paradigm problem whose complexity is analyzed here is: given a function / of class C in [0, 1]", com pute L ji. / . As one notes, inputs for this problem cannot be given in general by a finite vector of real numbers. So one must assume the existence of a routine that given x € [0, l ] n returns f(x). The complexity is evaluated in terms of the operations done as well as in terms of the number of times this routine is used. Again we are in the realm of the nu merical tradition since the arithmetic is performed on real numbers at a constant cost and the main issue is the search for lower and upper bounds. We close this section with some general refer ences to the topics introduced in this chapter. A good reference for the Mandelbrot and Julia sets is Devaney [1989]. For the undecidability of the Mandelbrot set see Blum & Smale [1993]. A major reference for Hubert's Tenth Problem is Matiyasevich [1993]. For the Nullstellensatz, see Lang [1993] or Kendig [1977]. More advanced books in algebraic geometry are those of Hartshorne [1977] or Shafarevich [1977]. These references however, only deal with the qualitative aspect of the Nullstellensatz. Exponential bounds for the degrees in the Nullstel lensatz were proved by Bronawell [1987] and refined by Kollar [1988] and Caniglia, Galligo k Heintz [1988]. Real polynomials, real algebraic sets and semi-algebraic sets are exposed in the monographs of Benedetti k Risler [1990] and by Bochnak, Coste k Roy [1987]. For Newton's method see Smale's survey article (Smale [1985]). A classical reference for linear and integer programming is the book of Schrijver [1986]. For the classical theory of computability and Turing machines see the books by Davis [1965], Rogers [1967] and Cutland [1980]. Classical com plexity theory is a younger subject. The books by Balcazar, Diaz k Gabarrd [1988, 1990] or by Papadimitriou [1994] offer very good introductions to its achievements. Other references for this chap ter are Blum [1990, 1991], Smale [1988, 1990] and Hirsch, Marsden k Shub [1993]. The quotations from von Neumann, Penrose k Kuhn are taken from von Neumann [1963], Penrose [1991] and Kuhn [1957]. The reference list, while extensive, is not meant to be exhaustive.
References Abramson, F. [1971] "Effective computation over the real numbers," in 12th Annual IEEE Symp. on Switching and Automata Theory, pp. 33-37. Aho, A., Hopcroft, J. k UUman, J. [1974] The Design and Analysis of Computer Algorithms (Addison-Wesley). Balcazar, J., Diaz, J. k Gabarro, J. [1988] Structural Complexity I, EATCS Monographs on Theoretical Computer Science, Vol. 11 (Springer-Verlag). Balcazar, J., Diaz, J. k Gabarro, J. [1990] Structural Complexity II, EATCS Monographs on Theoretical Computer Science, Vol. 22 (Springer-Verlag). Ben-Or, M. [1983] "Lower bounds for algebraic computa tion trees,'' in 15th Annual ACMSymp. on the Theory of Computing, pp. 80-86. Ben-Or, M., Kozen, D. k Reif, J. [1986] "The complexity of elementary algebra and geometry," J. of Computer and Systems Sciences 18, 251-264. Benedetti, R. k Risler, J.-J. [1990] Real Algebraic and Semi-Algebraic Sets (Hermann). Blum, L. [1990] "Lectures on a theory of computation and complexity over the reals (or an arbitrary ring)," in Lectures in the Sciences of Complexity II, ed., E. Jen (Addison-Wesley), pp. 1-47. Blum, L. [1991] "A theory of computation and com plexity over the real numbers," in Proc. of the Int. Congress of Mathematicians (Springer-Verlag), pp. 1491-1507. Blum, L., Shub, M. & Smale, S. [1989] "On a theory of computation and complexity over the real num bers: NP-completeness, recursive functions and uni versal machines," Bulletin of the Amer. Math. Soc. 21, 1-46. Blum, L. k Smale, S. [1993] "The Gddel incomplete ness theorem and decidability over a ring," in From Topology to Computation: Proc. of the Smalefest, eds. Hirsch, M., Marsden, J. k Shub, M. (Springer-Verlag), pp. 321-339. Blum, M. [1967] "A machine-independent theory of the complexity of recursive functions," J. ACM 14, 322-336. Bochnak, J., Coste, M. k Roy, M.-F. [1987] Giomitrie algibrique rielle (Springer-Verlag). Borodin, A. k I. Munro [1975] The Computational Com plexity of Algebraic and Numeric Problems (Elsevier). Brownawell, W. [1987] "Bounds for the degrees in the Nullstellensatz," Annals Math. 126, 577-591. Caniglia, L., Galligo, A. k Heintz, J. [1988] "Borne sim ple exponentielle pour les degres dans les theoremes de zeros sur un corps de caracteristique quelconque," C. R. Acad. Sci. Paris 307, 255-258. Canny, J. [1988] "Some algebraic and geometric compu tations in PSPACE," in 20th Annual ACM Symp. on the Theory of Computing, pp. 460-467. Church, A. [1936] "An unsolvable problem of elementary number theory," Amer. J. of Math. 58, 354-363.
1537 24 L. Blum tt al. Hartshorne, R. [1977] Algebraic Geometry (SpringerCobham, A. [1964] T h e intrinsic computational diffi Verlag). culty of problems,'' in Int. Congress for Logic, Method ology, and the Philosophy of Science, ed. Bar-Hillel, Y. Heintz, J., Roy, M.-F. k Solerno, P. [1990] Sur la com (North-Holland), pp. 24-30. plexity du principe de Tarski-Seidenberg," Bulletin de la Socitti Mathimatiquc de Prance 118, 101-126. Collins, G. [1975] Quantifier Elimination for Real Closed Fields by Cylindrical Algebraic Deccomposition,Herman, G. k Isard, S. [1970] "Computability over ar bitrary fields," J. London Math. Soc. 2, 73-79. Vol. 33 of Lecture Notes in Computer Science Hermann, G. [1926] "Die Frage der endlich vielen Schritte (Springer-Verlag), pp. 134-183. in der Theorie der Polynomideale," Math. Ann. 95, Cook, S. [1971] T h e complexity of theorem proving 736-788. procedures," in 3rd Annual ACU Symp. on the The Hirsch, M., Marsden, J. k Shub, M. (eds.) [1993] From ory of Computing, pp. 151-158. Topology to Computation: Proceedings of the Smaledicker, F. [1992a] T h e arithmetical hierarchy over the fest (Springer-Verlag). reals,'' J. of Logic and Computation 2, 375-395. Hoover, H. [1987] "Feasibly constructive analysis," Ph.D. Cucker, F. [1992b] " P R ^ NC R ," J. of Complexity 8, Thesis, Department of Computer Science, University 230-238. of Toronto. Cucker, F. [1993] "On the complexity of quantifier elim Karp, R. [1972] "Reducibility among combinatorial prob ination: The structural approach,'' The Computer lems," in Complexity of Computer Computations, eds. Journal 36, 400-408. Miller, R. k Thatcher, J. (Plenum Press), pp. 85-103. Cutland, N. [1980] Computability (Cambridge University Kendig, K. [1977] Elementary Algebraic Geometry Press). (Springer-Verlag). Davis, M. [1965] The Vndecidable (Raven Press). Devaney, R. [1989] Chaotic Dynamical Systems (Addison- Kleene, S. [1936] "General recursive functions of natural numbers," Math. Annalen 112, 727-742. Wesley). Edmonds, J. [1965] "Paths, trees, andflowers,''Canadian Ko, K. [1991] Complexity Theory of Real Functions (Birkhauser). J. of Math. 17, 449-467. Koiran, P. [1993] "A weak version of the Blum, Shub k Engeler, E. [1967] "Algorithmic properties of structures," Smale model," in S4th Annual IEEE Symp. on Foun Maths. Systems Theory 1, 183-195. dations of Computer Science, pp. 486-495. Engeler, E. [1993] Algorithmic Properties of Structures Kollar, J. [1988] "Sharp effective Nullstellensatz," J. of (World Scientific). Amer. Math. Soc. 1, 963-975. Engeler, E. in collaboration with Aberer, K., Amrhein, Kreitz, C. k Weihrauch, K. [1982] "Complexity theory B., Gloor, O., von Mohrenschildt, M., Otth, D , of real numbers and functions," in Theoretical Com Schwatrzler, G. k Weibel, T. [1995] The Combinatory puter Science, eds. Cremers, A. k Kreigel, H. vol. 145, Programme (Birkhauser). Lecture Notes in Computer Science (Springer-Verlag), Friedman, H. [1971] "Algorithmic procedures, general pp. 165-174. ized turing algorithms, and elementary recursion the Kuhn, T. [1957] The Copernican Revolution: Planetary ory," in Logic Colloquium 1969, eds. Gandy, R. k Astronomy in the Development of the Western Yates, C. M. E. (North-Holland), pp. 361-390. Thought (Harvard University Press). Friedman, H. k Ko, K. [1982] "Computational complex Lang, S. [1993] Algebra 3rd edn. (Addison-Wesley). ity of real functions," Theoretical Computer Sci. 20, Levin, L. [1973] "Universal sequential search problems," 323-352. Probl. Pered. Inform. 1X3, 265-266. (in Russian, En Friedman, H. k Mansfield, R. [1992] "Algorithmic proce glish translation in Problems of Information Trans. dures," Trans. of the Amer. Math. Soc. 332, 297-312. 9(3); corrected translation in Trakhtenbrot [1984]). Garey, M. k Johnson, O. [1979] Computers and Intract Matiyasevich, Y. [1993] Hubert's Tenth Problem (MIT ability: A Guide to the Theory of NP-Completeness Press). (Freeman). Meer, K. [1990] "Computations over Z and R: A com Goldstine, H. [1977] A History of Numerical Analysis parison," J. of Complexity 6, 256-263. from the 16th through the 19th Century (Springer- Meer, K. [1992] "A note on a P ^ NP result for a re Verlag). stricted class of real machines," J. of Complexity 8, Grigoriev, D. k Vorobjov, N. [1988] "Solving systems of 451-153. polynomial inequalities in subexponential time," J. of Meer, K. [1993] "Real number models under various sets Symbolic Computation 5, 37-64. of operations," J. of Complexity 9, 366-372. Harrington, L., Morley, M., Seedrov, A. & Simpson, S. Meer, K. [1994] "On the complexity of quadratic pro (eds.) [1985] Harvey Friedman's Research on the Foun gramming in real number models of computation," dations of Mathematics (North-Holland). Theoretical Computer Sci. 133, 85-94. Michaux, C. [1989] "Une remarque a propoe des ma Hartmanis, J. k Steams, R. [1965] On the computational chines sur 1R introduces par Blum, Shub et Smale," complexity of algorithms," Trans, of the Amer. Math. C. R. Acad. Sci. Paris 809, Serie I, 435-437. Soc. 117, 285-306.
1538 Complexity and Real Computation: A Manifesto 25
Michaux, C. [1991] "Ordered rings over which output sets are recursively enumerable," Proc. Amer. Math. Soc. 112, 569-575. Moschovakis, Y. [1986] Foundations of the theory of al gorithms, Draft. Ostrowski, A. [1954] "On two problems in abstract al gebra connected with Horner's rule,'' in Studies in Mathematics and Mechanics presented to Richard von Mises (Academic Press), pp. 40-48. Pan, V. [1966] "Methods of computing values of polyno mials," Russian Math. Surveys 21, 105-136. Pan, V. [1987] "Sequential and parallel complexity of approximate evaluation of polynomial zeros," Cornput. Math. Appl 14, 591-622. Pan, V. [1995] "Optimal (up to polylog factors) sequen tial and parallel algorithms for approximating com plex polynomial zeros," in £7th Annual ACM Symp. on the Theory of Computing, pp. 741-750. Papadimitriou, C. [1994] Computational Complexity (Addison-Wesley). Peitgen, H.-O. k Saupe, D. (eds.) [1988] The Science of Fractal Images (Springer-Verlag). Penrose, R. [1991] The Emperor's New Mind (Penguin Books). Poizat, B. [1995] Us Petits Cailloux (Alea). Post, E. [1943] "Formal reductions of the general com binatorial decision problem," Amer. J. Math. 65, 197-268. Pour-El, M. k Richards, I. [1983] "Computability and noncomputability in classical analysis," Trans, of the Amer. Math. Soc. 275, 539-560. Preparata, F. k Shamoe, M. [1985] Computational Ge ometry: An Introduction, Texts and Monographs in Computer Science (Springer-Verlag). Rabin, M. [1960a] "Computable algebra, general theory and theory of computable fields," Trans, of the Amer. Math. Soc. 95, 341-360. Rabin, M. [1960b] "Degree of difficulty of computing a function and a partial ordering of recursive sets," Tech. Rep. 2, Hebrew University of Jerusalem. Rabin, M. [1966] "Mathematical theory of automata," in 19th ACM Symp. in Applied Mathematics, pp. 153-175. Renegar, J. [1987] "On the efficiency of Newton's method in approximating all zeros of systems of complex poly nomials," Math, of Oper. Research 12, 121-148. Renegar, J. [1992] "On the computational complexity and geometry of the first-order theory of the reals. Part I," J. of Symbolic Computation 13, 255-299. Rogers, H. [1967] Theory of Recursive functions and Effective Computability (McGraw-Hill). Scholz, A. [1937] "Aufgabe 253," Jahresber. Deutsch. Math.-Verein. 47, il-\2. Schonhage, A. [1982] "The fundamental theorem of algebra in terms of computational complexity," Tech. Rep., Math. Institut der. Univ. Tubingen.
Schrijver, A. [1986] Theory of Linear and Integer Pro gramming (John Wiley k Sons). Shafarevich, I. [1977] Basic Algebraic Geometry (Springer-Verlag). Shepherdson, J. k Sturgis, H. [1963] "Computability of recursive functions," J. of the ACM 10, 217-255. Shub, M. k Smale, S. [1993a] "Complexity of Bezout's theorem I: Geometric aspects," J. of the Amer. Math. Soc. 6, 459-501. Shub, M. k Smale, S. [1993b] "Complexity of Bezout's theorem II: Volumes and probabilities," in Computa tional Algebraic Geometry, eds. Eyssette, F. k Galligo, A., vol. 109 Progress in Mathematics (Birkhauser), pp. 267-285. Shub, M. k Smale, S. [1993c] "Complexity of Bezout's theorem III: Condition number and packing," J. of Complexity 9, 4-14. Shub, M. k Smale, S. [1993d] "Complexity of Bezout's theorem IV: Probability of success, extensions," to appear in SI AM J. of Numer. Anal. Shub, M. k Smale, S. [1994] "Complexity of Bezout's theorem V: Polynomial time," Theoretical Computer Sci. 133, 141-164. Sipser, M. [1992] "The History and Status of the P versus NP Question," in 24th Annual ACM Symp. on the Theory of Computing, pp. 603-618. Smale, S. [1985] "On the efficiency of algorithms of anal ysis," Bulletin of the Amer. Math. Soc. 13, 87-121. Smale, S. [1987] "On the topology of algorithms I," J. of Complexity 3, 81-89. Smale, S. [1988] "The Newtonian contribution to our un derstanding of the computer," in Newton's Dream, ed. Stayer, M. (McGill-Queens University Press). Smale, S. [1990] "Some remarks on the foundations of numerical analysis," SIAM Rev. 32, 211-220. Steele, J. k Yao, A. [1982] "Lower bounds for algebraic decision trees," J. of Algorithms 3, 1-8. StraBen, V. [1969] "Gaussian elimination is not optimal," Numer. Math. 13, 354-356. Strafien, V. [1990] "Algebraic complexity theory," in Handbook of Theoretical Computer Science, ed., van Leeuwen, J., vol. A (MIT Press/Elsevier), pp. 633-672. Tarski, A. [1951] A Decision Method for Elementary Algebra and Geometry (University of California Press). Trakhtenbrot, B. [1984] "A survey of Russian approaches to perebor (brute-force search) algorithms," Annals of the History of Computing 6, 384-400. Traub, J., Wasilkowski, G. k Wozniakowski, H. [1988] Information-Based Complexity (Academic Press). Tucker, J. [1980] "Computing in algebraic systems," in Recursion Theory, its Generalizations and Applica tions, eds. Drake, F. k Wainer, S. London Math. Soc. (Cambridge University Press).
1539 26 L. Blum et at Tucker, J. & Zucker, J. [1992] "Examples of semicomVan der Waerden, B. [1949] Modern Algebra (F. Ungar putable sets of real and complex numbers," in ConPublishing Co.). structivity in Computer Science, eds. O'Donnell, M. von Neumann, J. [1963] Collected Works V, ed. Taub, A. & Myers Jr, J., vol. 613, Lecture Notes in Computer (MacMillan). Science (Springer-Verlag), pp. 179-198. Weihrauch, K. [1987] Computability, EATCS Mono Turing, A. [1936] "On computable numbers, with an ap graphs on Theoretical Computer Science, vol. 9 plication to the Entscheidungsproblem," Proc. London (Springer-Verlag). Math. Soc. Ser. £, 42, 230-265. Winograd, S. [1967] "On the number of multiplications Tyurin, J. [1979] "A survey of the logic of effective defini required to compute certain functions," Proc. National tions," in Logic of Programs, ed. Engeler, E., vol. 125, Acad. Sci. 58, 1840-1842. Lecture Notes in Computer Science (Springer-Verlag), Winograd, S. [1980] "Arithmetic complexity of compu pp. 198-245. tations," SIAM Reginal Conf. Ser. Appl. Math. 33.
1540 Lectures in Applied Mathematics Volume 3 2 , 1996
Algebraic Settings for the Problem "P ^ NP?" Lenore Blum, Felipe Cucker, Mike Shub, and Steve Smale ABSTRACT. When complexity theory is studied over an arbitrary unordered field K, the classical theory is recaptured with K = Zj. The fundamental result that the Hilbert Nullstellensatz as a decision problem is NP-complete over K allows us to reformulate and investigate complexity questions within an algebraic framework and to develop transfer principles for complexity theory. Here we show that over algebraically closed fields K of characteristic 0 the fundamental problem "P / NP?" has a single answer that depends on the tractability of the Hilbert Nullstellensatz over the complex numbers C. A key component of the proof is the Witness Theorem enabling the elimination of transcendental constants in polynomial time.
1. Statement of Main Theorems We consider the Hilbert Nullstellensatz in the form HN/Jf: given a finite set of polynomials in n variables over a field K, decide if there is a common zero over K. At first the field is taken as the complex number field C. Relationships with other fields and with problems in number theory will be developed here. This article is essentially Chapter 6 of our book Complexity and Real Compu tation (to be published by Springer). Background material can be found in [Blum, Shub, and Smale 1989]. Only machines and algorithms which branch on uh{x) = 0?" are considered here. The symbol < is not used. Thus the development is quite algebraic, eventually using properties of the height function of algebraic number theory. A main theme is eliminating constants. The moral is roughly: using transcendental and algebraic numbers doesn't help much in speeding up integer decision problems. Let Q be the algebraic closure of the rational number field Q. The following will be proved. THEOREM
1. 7/P = NP overC, then P = NP over Q, and the converse is also
true. 1991 Mathematics Subject Classification. 68Q (Computer Science, Theory of Computation), 11G (Number Theory, Arithmetic Algebraic Geometry). Blum was partially supported by the Letts-Villard Chair at Mills College. Cucker was par tially supported by DGICyT PB 920498, the ESPRIT BRA Program of the EC under contracts no. 7141 and 8556, projects ALCOM II and NeuroCOLT. Cucker, Smale, and Shub were partially supported by NSP grants. (£) 1996 American Mathematical Society
125
1541 126
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE REMARK
1. Here C may be replaced by any algebraically closed field contain
ing Q. Now we are going to define an invariant T of integers (and polynomials over Z) which describes how many arithmetic operations are necessary to build up an integer starting from 1. More precisely, a computation of length I of the integer m is a sequence of integers, X O , X J , . . . , X J where io = 1, xj = m and given k, 1 < k < I, there are i,j, 0 < i,j < k such that x/t = x< o Xj where o is addition, subtraction or multiplication. We define r : Z —♦ N by r(m) is the minimum length of a computation of m. The following is easy to check, where here and in the sequel log denotes log2. PROPOSITION 1. For allm&N
one has T(m) < 21ogm.
2
If m is of the form 2 , then r (m) = log log m +1. The same is essentially true even if m is any power of 2. Open Problem. Is there a constant c such that r(A;!)<(logfc) c allfceN? We remark that if "factoring is hard" using inequalities then the open problem has a negative answer.1 DEFINITION 1. Given a sequence of integers a* we say that a* is easy to com pute if there is a constant c such that r(ofc) < (logfc)c, all k > 2, and hard to compute otherwise. We say that the sequence a* is ultimately easy to compute if there are non-zero integers m* such that m^aic is easy to compute and ultimately hard to compute otherwise.
In Sections 5 and 6 we will prove: THEOREM 2. If the sequence of integers k\ is ultimately hard to compute, then HN/C, the Hilbert Nullstellensatz over C, is intractable and hence P ^ NP over C. Thus in that case, P ^ NP over Q.
Next consider the analogous situation for polynomials with integer coefficients / € Z[i\. A computation of length I of / is a sequence of i*j e Z[t] where i*o = 1, «i = t, Ui = f and given k, 1 < k < I there are i,j, 0 < i,j < k such that uic = Uj o Uj where o is addition, subtraction or multiplication. Define r : Z[t] —» N by r(f) is the minimum length of a computation of / . Let Zer(/) be the number of distinct integral zeros of / . The following has a certain plausibility: HYPOTHESIS. Zer(/) < r ( / ) c for all non-zero / 6 Z[t]. Here c is a universal constant. We don't know if the hypothesis is true or false, even, for example, with the constant c = 1. 1 Here is a sketch of the proof. Suppose to the contrary that k\ is easy to compute and n is the product of primes p and q where p < fc < q. We will show how to easily factor n. Let x o . i i , . . . ,zj = fc! be a short computation of fc!, I < (logfc)c. Then we induce a short computation of r = fc! mod n using
(io mod n, i i mod n
r = i j mod n).
By the Euclidean Algorithm, y = gcd(r, n) may be easily computed. By our hypothesis it follows that y = p, and thus our assertion is proved.
1542 ALGEBRAIC SETTINGS FOR THE PROBLEM "P ^ NP?" THEOREM
127
3. // the above Hypothesis is true then NP / P over C, and NP ^ P
over Q. 2. Eliminating constants: Easy cases In this section we begin a study of the problem of eliminating the constants of a computation without an exponential increase in the time. Our first result asserts that this can always be done if the constants lie in an algebraic extension of the given field K. The result holds for fields of any characteristic. To start we briefly and informally recall some notation and definitions. Suppose M is a machine over a commutative ring (or field) R with unit. Then both the input and output spaces of M are R°°, and the state space is Roo- Here R°° is the disjoint union R
= Un>oi?
and Roc is the b\-infinite direct sum space over R. Elements of Roo have the form i = (... , x_ 2 ,x_i,x 0 . Xi,x 2 ,.. •) where x* 6 R for all integers i, x* = 0 for |fc| sufficiently large, and . is a dis tinguished marker between XQ and x\. We call Xi the "i-th coordinate" of x and x i , . . . , x n the "first n coordinates" of x. For brevity, and when the intent is clear from context, we sometimes omit the negative coordinates and write elements of the state space as (n, x\,X2,... , x n , 0 , . . . ) or (xi, X2,... , x n , 0 , . . . ) , or even (xi,x 2 ,... ,x„). The machine's input mop, associated with its input node, takes a point x € Rn C R°° and maps it to (...,0,0
0,0,...)GRoo.
Here n is the size of x. The output map, associated with the machines's output node, takes a point x = (... , m . X\, X2,...) 6 Roo and maps it to ( x i , . . . , x m ) if m is a positive integer, to the unique point of R° if m = 0, and is undefined otherwise. These maps make sense if the characteristic of R is 0. If the characteristic is positive, we replace n and m here by the appropriate number of 1 's to the left of the distinguished marker. Machines also have computation, shift and branch nodes with associated oper ations (polynomial or rational maps, right/left shifts and the identity) and associ ated next state and next node maps. Without loss of generality, we assume that at branch nodes, machines branch right or left depending on whether or not the first coordinate xi of the current state x is 0. A decision problem over R is a pair (Y,Yo) where Y0 C Y C R°°. Here Y is the set of problem instances and Vo is the set of yes instances. For example, KN/K is a decision problem over K where Y = {finite polynomial systems over K} and YQ — {finite polynomial systems over K that are solvable over K}. A finite polynomial system over K is represented as an element of K°° assuming some standard listing of its coefficients. A machine M over R decides the problem (Y, Vo) if, for all inputs y 6 Y, M outputs 1 if y e Vo and 0 if not. The halting time, T\f(y), is the length of the computation path, or sequence of nodes, traversed from input y to output. The problem (Y, Vo) is in class P or in polynomial time over R if it can be decided by a
1543 128
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
machine with halting time bounded by a fixed polynomial in the size of y, for all y e Y. Polynomial time is our notion of tractability. The problem (Y, Yo) is in class NP over R if there is a machine M' such that for all y € Y, y e Yo if and only if there is a witness w € R°° such that, given input (y,w), M' outputs 1 in time bounded by a fixed polynomial in the size of y. We say (Y, Vo) is l^Y-complete over R if it is in class NP/R and every problem in class NP/i? can be encoded in (V, Y0) in polynomial time. It follows from [Blum, Shub, and Smale 1989] that, for any field K, HN//C is NP-complete over K. DEFINITION 2. Suppose K c L arefieldsand (Y, Vo) is a decision problem over L. The restriction of (Y, Y0) to K is (Y n K°°, Y0 n K°°). The same applies to the case where K is a ring. PROPOSITION 2. Let M be a machine over a field L which is an algebraic extension of a field K. Then there is a machine M' over K and a constant c > 0 (depending on M) with the following property. For any decision problem (Y, Yo) over L decided by M, the restriction of (Y,Y0) to K is decided by M', and the halting time satisfies
TM'(y) < cTM(y), for all y e Y n K°°. PROOF. Since M has only a finite number of constants, then by restriction, M is also a machine over a subfield of L that is a finite algebraic extension of K. Thus, our proposition will follow if we assume that L is a finite algebraic extension of K, and show it for this case. So we make this assumption. Consider L as a vector space over K of dimension q. Thus L may be represented as Kq where the inclusion K C L is represented as the inclusion of K in Kq as the first coordinate. We now construct a machine M' over K that on inputs from K°° simulates M on these inputs with halting time increased by no more than a multiplicative constant. The state space of M' is considered as (Kq)oc so that it also represents Loo. An initial subroutine of M' in effect takes an input from K°° and writes it as the first coordinates in {Kq)oc. Without loss of generality, we may assume that at any computation node of M, the computation performed is either addition, multiplication, subtraction or division of two elements of L. (Any machine can be so converted with at most a multiplicative constant increase in halting time.) Since addition and multiplication in L are represented by fixed symmetric bilinear maps over K B+ : Kq x Kq -» Kq Bx : Kq xKq ^ Kq, M' can simulate the addition and multiplication nodes of M by incorporating these polynomial maps in computation nodes. Subtraction nodes of M are simulated in M' by multiplication by (-1) followed by B+. Division of b by a is accomplished by solving the linear system Bx (a, y) = b for y by Gaussian elimination. This requires on the order of g3 steps. In each of these simulations, constants from L that occur in M are replaced by their corresponding g-tuples over K. Since x = (x\,... ,xq) represents the zero element in L if and only if X{ = 0 for i = 1,... ,q, branching in M is simulated by checking if the first q coordinates of an element in the state space of M' are zero. Shifting right or left in M is simulated by shifting right or left q times in M'. Care is taken to keep track of the
1544 ALGEBRAIC SETTINGS FOR THE PROBLEM "P # NP?'
129
intended lengths of sequences in the computation. A final subroutine ensures that the appropriate finite sequence (xi,x 9 + i,X2,+i,... ,x m ,+i) of the coordinates of the "final state" x in a computation is output. Using the isomorphism between Kq and L, one can see that on inputs y € Y n K°°, M' gives the same answers as M with the desired time bound where c is on the order of q3. O The proof doesn't require M to be a decision machine, merely that the outputs of M are defined over K, i.e. are in K°°, for inputs over K. PROPOSITION 3. Let R be an integral domain and K its quotient field. Let (Y, YQ) be a decision problem solved by a machine M over K in time T. Then M can be replaced by an equivalent machine without division, with constants only from R, and with halting time cT for some c € N. Thus, there is a machine over R solving the restriction of (Y, YQ) to R in time cT for some c 6 N. PROOF. A machine M' without division that simulates M is obtained by "dou bling" the space used in the computation. An initial subroutine of M' takes an input (xi,X2,... ,xs) to (2s,xi,l,X2,1,. •• , x , , l , 0 , . . . ) in the state space of M'. Note that elements x = (xi, X2,... , x,, 0,...) in the state space of M may be represented (non-uniqueley) by elements in the state space of M' of the form
(x^,x^,x2n\x^,...,x^,x^0,--) where x< = -jjj- for all i < s. Since K is the quotient field of R, elements of AQQ x
i
have representation in Rx,. Computation nodes of M' perform the natural modification of the operations associated with the computation nodes of M which, as above, are assumed to be the basic arithmetic operations over K. In particular, a computation node in M that performs a division /(xi,X2) = (xi/x^) is replaced in M' by a computation node with associated map 9\x\
! x l ix2
>X2 I
=
\x\
x
2
>xl
x
2
)
Constants from K in M are replaced in M' by pairs of constants from R. So for example, a computation node of M with associated map /(xi) = fcxi with constant A: € if, is replaced in M' by a computation node with associated map g(Xi,xy) = (pxy, gxp) for some p,q€ R with k = p/q. Thus, if an initial input to M' comes from R°°, all states in the subsequent computation will be in Roo. A branch node in M that tests if xi = 0 is replaced in M' by one that tests if x[n) = 0 and x[f jt 0. Finally, M accepts an input x, i.e. M outputs the value 1 given input x, if the first coordinate of the final state in the computation is 1 (and the 0-th coordinate is 1). Consequently, M' is designed to accept an input if the first and second coordinates of the final state in the computation are equal but not 0 (and the 0-th coordinate is 2). Similar considerations apply for rejecting an input. The overall slowdown of M' with respect to M is linear. Thus, M' has the requisite properties for both conclusions in the statement of the proposition. □
1545 130
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
3. Witness Theorem We need an algebraic theorem, which we call the Witness Theorem, for the proof of our main results. This section is devoted to that theorem. The first step is to extend the definition of r to polynomials in several variables over Z. Let G S Z[(i, ...,tn]. Quite similarly to the one variable case, consider finite sequences (ti,...,tn,l,ui,...,us
= G)
where for 1 < k < s, u* = v o w for some v, w € {
In situations we encounter, / is presented so that it is not obvious if it is zero. THEOREM 4 (Witness Theorem). Let F(x,t) = F(z\,...,xr,ti,...,ti) be a polynomial in r + I = n variables with coefficients in Z and let Fx € Q[ti, • • • ,t{\ be defined by Fx(t) = F(x, t) for each x S Q . Suppose that N is a positive integer satisfying:
log AT > 4nr 2 + 4, r = T(F). Then for x € Q , there exists an algebraic number tui tn {2 N ,xf,... ,x?} such that the point w = (w\,... ,u>i) where Wi = tujlj, i = 2 , . . . , / is a witness for FxeQ{tu...,ti}. Our proof of the Witness Theorem depends heavily on the use of heights of algebraic numbers. The height H : Q —► Q is a function whose properties are summarized in the following proposition. PROPOSITION 4.
(a):
H(l)
= H(0)
= 1, H(2)
= 2, H{w)
> I, H(-w)
=
H(w), H ( i ) = H(w) (b): H(v + w)< 2H(v)H(w) (c): H(wk) = H(w)k, H(vw) < H(v)H(w) (e):
H(vw)>$$ifw?0.
A definition of H and proofs of (a) and (c) are given in [Lang 1991]. Moreover (b) is proved in the appendix to this section. Note that (d) follows from (b) by H(v) = H((v + w) - w) < 2H(v + w)H(w). Now divide by 2H(w). Similarly we obtain (e) from (c) by H(v) = H ((vw)-j
< H(vw)H{w).
Note that, in general, \i=0
/
i=0
1546 ALGEBRAIC SETTINGS FOR THE PROBLEM "P # NP?"
131
All that is used in this section is the existence of a function H : Q —► Z+ with properties (a), (b) and (c) (and hence also (d) and (e)). It is a good exercise to prove Proposition 4 for Q with H(r) = max(|p|, |g|) where r = jj and gcd(p,q) = 1. If g S Q[t] is a one variable polynomial, and g(t) = 53t=oa»*'> define H(g) = PROPOSITION
5. For all g e Q[t] and allw 6 Q 2dH(w)dH(g).
H(g(w)) < PROOF. Use Homer's argument H l^aiw'
I
=
H(ao + w(a\ + w(a2 H
h w(ad-i + wad)) •• •))
2dH(ao)H(w)H(ai)H(w)H(a2)...
< d
= 2 d J | H(ai)H(w)d.
D
i=0 a
If G(x) = 5^ aax
is a polynomial in n variables over Q, let
H(G) = l[H(aa). a PROPOSITION
6. For G € Z[tu...,
tn], let r = T(G).
Then
H(G)<22^\ Toward the proof we have the following lemma whose proof is simple and straightforward. LEMMA 1. The degree o/Gis less than or equal to 2 T . The number of mono mials in G, indexed by a, is less than Dn, where D = 2T. □ We prove now Proposition 6. PROOF OF PROPOSITION 6. It goes by induction on r. One checks it by in spection for T = 1. Now let G = FF' where T(F), T(F') < r (the case G = F + F' or G = F - F', is even simpler). Write F(x) = ^,aaxa, F'(x) = J2bpx13 and G(x) = Y^c-y1^Then
0
Note that by Lemma 1, the degrees of F, F' and G are less than or equal to D and the number of terms in F, F' and G is even less than Dn. Then
H^)
<
l[2H(a^0)H(b0) P
<
2D"
H(F)H(F').
Thus H(G)
<(2D"'H(F)H(F'))D".
1547
132
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
By the induction hypothesis H{G) < 2°2n • 2 D " - 2 - < ^ " ^ . so < D2n +
log H(G)
<-■
n2nr _i_ tynrty2n(T-l)2
< for
T
Dn-22n{T-1)2+1 +l
22nT*
> 2.
□
The next proposition while simple, with a short proof, is crucial for the Witness Theorem. PROPOSITION 7. Let g e Q[t] be a non-constant polynomial in one variable of degree d. Then for every x € Q, H(x)
PROOF.
Write d
g(t) = J2aiti,
ad^0,
d>0.
i=0
Then
H(g(x)) = H (a^ + >
1
Y,^)
H(adxd)
H^l
> 1
2* H(ad)H(x)d-iH(ao)...
^(od_0
> l^£l 2*H(g)Here we have used Propositions 4 and 5. □ d COROLLARY 1. For g € Q[t], if H(x) > 2 H(g), then g(x) / 0 unless g is zero. □ For x € Q let #(z) =
max
H(x{).
l
For G € Q [ < i , . . . , t n ] and i = ( i i , . . . , x r ) € Q r , r < n, let G' X l i ... i I r .(t r +i,..., tn) = G(xi,... PROPOSITION
8. For any G G Q[t\,...,t„]
,xr,tr+i,...
and x = ( i i , . . . , x r ) € Q r with
r
,tn).
H(G)(2H(x))D"+1 ,xr).
1548 ALGEBRAIC SETTINGS FOR THE PROBLEM "P / NP?"
133
Let G{t) = £ Q a Q t Q . N o t e t h a t G *i *r e Q[*r+i>...,t n ] is a polynomial whose coefficients may be indexed by ( a r + i , . . . , a n ) and, for each ( a r + i , . . . , a n ), have the form PROOF.
where the sum is over a = ( a i , . . . , a n ) such that the last n- r entries of a are ( a r + i , . . . ,a„). We must estimate the product of the heights of these coefficients to obtain the proposition. The estimate is similar to that used in Proposition 5. The estimate for the height of a coefficient of GXl Xr is
< 2°r n
%)%r.-%)'"
o=(oi,...,a r )
< 2° r
H(aa)H(x)D.
J] <*=(<»l.."i<»r)
Take the product over all the coefficients to get H(GXl
Ir)
2D"H(G)H(x)Dn+1
<
yielding the necessary estimate.
D
For the proof of the Witness Theorem we may assume that wi is one of 2N, x f , . . . , X? with largest height so H{w1)>mBx(2N,H(xi)N). Then H(wi) > 1 and H{wi) > H(wi-i). Now with these x,w as in the Witness Theorem, for each j = 1 , . . . , / and 0 = (Pj+i,...,/3i), we will define a one variable polynomial C# so that we will be able to apply the one variable lower bound of Proposition 7. Write
F(x,t)=
<*a.0Xat0
Y, ar=(oi
ar)
then define a=(ai,...,a r )
0=(0u...,0j,0j+l
A)
LEMMA 2. For each j = 1 , . . . , I and J3 as above H(Wj) > PROOF.
2DH(Gp.
Fix j . It is sufficient to prove that H(w])>2DH(FXiWl
„,_,),
or yet by Proposition 8 that H(wj)>2DH(F)(2H{xu...,xr,w1,...,wj-l))D'x+\ Now use Proposition 6. The needed estimate is:
"
\2D-
2 2 ' nT 2(max(2, F ( x ) ) ) ° " + '
if j = 1.
1549 134
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
Now use D = 2 r , and verify that log N > T + 2nr2 + 2(n + 1)T.
□
The Witness Theorem is almost proved. Use Proposition 7 with j = I in Lemma 2. We obtain Gli(t) = Fx,Wl
„,,_,(«), 0 = 0
and #(«;,) > 2 D tf (G^). Therefore Fx>Wl
„,,_, is zero. So for each 0i, a=(a 1 ,...,o„) 0=(0i ft-1.A)
Continuing the same process for / — 1,1 — 2 , . . . , 1, we obtain eventually for any /3 = (/3,,...,/3,)that a
This yields our theorem.
□ Appendix to Section 3
In this appendix we prove part (b) of Proposition 4. We could not find it in the literature. To do so, we need to use some notions from algebraic number theory. References for these notions can be found in Section 9. DEFINITION
4. HK{U) =
J J max(l,\u\^),
where MK is the set of valua-
tions of K. If v restricts to no then define Nv = \KV : FVo], the degree relative to the completions. DEFINITION 5. H{u) = HK(U)^^, independent of K.
LEMMA 3.
^
for u e K. It can be shown that this is
Nu = [K : Q], where MK = Mjf UM£, M£> the archimedean
1/6 M£?
valuations and MK the non-archimedean valuations.
□
Now for the proof of part (b) of Proposition 4. PROOF OF PROPOSITION 4 ( B ) . We may write:
HK(x + y)
=
J]
maxOUs + yl?') ] } max(l,|* + y|?")
<
J]
2*" J ] (max(l,|x|„|y,|)) N - J ] (max(l, \x\„,
from properties of | I,/, archimedean and non-archimedean respectively.
\yl))N"
1550 ALGEBRAIC SETTINGS FOR THE PROBLEM "P / NP?"
135
Since max(l, |x|„, \y\„) < max(l, |x|„) max(l, \y\„), HK{x
+ y)<2^M*N"
J ] {max(l,\x\„))N'
J ] (max(l,|y| w ))^.
Using the lemma it follows that HK(x
+ y) <2l*
^HK(x)HK(y).
Taking roots we thus obtain H(x + y)<2H(x)H(y).
D
4. Elimination of Constants: General Case The main focus of this section is the following proposition. PROPOSITION 9 (Elimination of Constants). Let K c L be fields where K c Q. Let (Y, YQ) be a decision problem solved by a machine M over L. Then there is a machine M' over K solving the restriction of (Y, YQ) to K and a constant c € N such that TM,(y) < TM(y)c for ally £Y r\K°°. LEMMA 4. For the proof of the preceding proposition, it is sufficient to consider the case L = K(si,... ,si) where s\,..., si is a transcendence base for L over K (i. e., the Si are algebraically independent over K). PROOF. The machine M over L uses a finite number of constants n\,..., T]r g L. Therefore M can be considered as a machine over K(T)I,. ..,nr) C L by restric tion. By a standard theorem of algebra (in field theory) one may rewrite the sequence 771,..., r)T as S\,..., si, n\,..., fiq where the Si,..., s; form a transcendence basis for K(si,... ,si) over K and the /xi,... ,/x, are algebraic over K(s\,... ,si). Now apply Proposition 2 to obtain a machine over K(si,.. .,si) with the same values on inputs from K°° as M and only a constant multiple increase in time. □ PROOF OF PROPOSITION 9. We give the proof of the Elimination of Constants Proposition where L has the form given in Lemma 4. By Proposition 3 we may suppose that each computation node of the machine M is an arithmetic node (+, —, x) and that each constant in M is a polynomial in s = (s\,..., sj) over K. Let a = (QI , . . . , am) be a sequence of all the coefficients occuring in these polynomials. We may then suppose each constant in M is of the form p(a, s) where p is a polynomial over Z. Let C be the sum of the r(p) over all constants in M. We construct a machine M' over K that given input y = (j/i,-..,y n ) 6 Y n AT00 generates a computation path that simulates the computation path 7,, = (770, • • • , T]t,...) generated by M on input y. The critical construction is to simulate the branching structure of 7 y , and to do this with at most a polynomial increase in time. So suppose Tjt is a branch node and gt the associated step t branching polyno mial. That is, gt is the composition of the successive computations occuring along the computation path 7 y through step t. We may consider ft as a polynomial in y, a, and s over Z with r(gt)
1551 136
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
We construct M' so that given input y = (j/i,..., yn) and "time" t, M' gener ates elements W\,..., wi in K to replace « i , . . . , s; and thus obtain a machine over K. To produce ui\, let N = 4(n + m + l)(t + C) 2 + 4 and repeat squaring of each 2,2/i, • • •, yni <*i i • • • , c*m -N times (not giving a unique Wi, but a set of them). Let wi,...,wi be as in the Witness Theorem. Now test if gt{y, a,w) = 0 successively for each one of the n + m + 1 choices. If any one of these gt(y,a,w) / 0 then gt(y,a,s) ^ 0 and we branch ac cordingly. On the other hand by the Witness Theorem, if y G K°° and all the gt(y,a,w) = 0, then gt{y,ct,s) = 0. It is easy to check that the total increase in time is polynomial so that we have proved our proposition. □ Denote the decision problem HN over a field K by (YK, YQK). LEMMA
5. If K c L are algebraically closed fields then (YL n K°°, Y0L n K°°) = (YK, Y0K).
That is, the restriction o/HN/L to K is HN/A". Write YL = (J*£,n,M w n e r e yz,,n,/t,d is the space offc-tuplesof poly nomials, each of degree < d, in n variables over L. Let Fo,L,n,fc,d be the subset with a common zero. Since clearly YL n K°° = YK it is sufficient to show that yo,L,n,k,d H K°° = Y0^K,n,k,dThis latter amounts to showing that if {/t}f=1 € V L ^ M 1 " 1 K°° have a common zero £ € Ln, then the fi must have a common zero in Kn. But this follows directly from the model completeness of the theory of algebraically closed fields. Alterna tively, if the fi have no common zero in A"n, then by Hilbert's Nullstellensatz, there exist gi,i = l,...,k, polynomials in n variables, such that Yl&fi = 1- Evaluation at £ gives a contradiction. This proves Lemma 5. □ PROOF.
Now we can prove the first statement of Theorem 1. PROOF OF THEOREM 1 ( " I F " DIRECTION). Suppose P = NP over C. Then HN/C € P by a machine M over C. Then M "solves" HN/Q, (inputs from Q°°) by Lemma 5, but M is still a machine over C. Now apply Proposition 9 to obtain a machine over Q solving HN/Q in polynomial time. By the NP-completeness of this problem, P = NP over Q. D
5. Twenty Questions Toward the proofs of Theorems 2 and 3 we introduce a decision problem we call "Twenty Questions" which is of independent interest. Let R be a ring (integral domain) or field of characteristic 0 which we consider without order and let N be the positive integers. Then Twenty Questions over R is the problem: Given input (Jfc, ht(fc), z) G N x N x R, decide if z € {1,2,..., k}. Here ht(fc) is defined to be the largest natural number less than or equal to logfc. Even if R happens to be an ordered ring as Z, we continue to branch only on equality tests. Twenty Questions over any ring R can be decided in time 3A; by the machine in Figure 1. Can one do better? We don't know. But if R = Z, and branching on order is permitted, then the decision time is approximately log k, with the algorithm used in the parlor game called Twenty Questions.
1552
ALGEBRAIC SETTINGS FOR THE PROBLEM "P £ NP?"
137
j - 1
X=]
No No
Yes,
Output Yes
J+-J + 1
j=k+l? Yes Output No
FIGURE
1. A machine for Twenty Questions.
We say that Twenty Questions over R is tractable if it can be decided in time (logfc)c over R where c is some constant (depending only on R). The next theorem shows that if Twenty Questions over Z is tractable, then so is the order relationship itself. THEOREM 5. / / Twenty Questions over Z is tractable, then on input (x,y) € Z x Z, one can decide if x < y in time polynomial in max (log |x|, log \y\). P R O O F . Figure 2 shows a machine that solves the problem. This machine halts after visiting at most 3A; + 2 nodes where k is the first integer greater than max(log |x|, log \y\). Of these nodes, 2fc are Twenty Questions for 2 , 2 2 , . . . , 2fc twice each, hence the total time is
2 ^ j c +fc+ 2 which is less than or equal to
( * ^ ) ' +fc+ 2.
□
THEOREM 6. IfP = NP over C, then Twenty Questions over C is tractable. PROOF. The method is to embed Twenty Questions in a decision problem (K, yye,) which is in NP over C. Then if NP = P over C, (Y, Yyes) is in P over C and there is a machine M which decides Twenty Questions in time bounded by (logfc)c, c a constant. Here M is the restriction of the machine which decides (Y, Yye8) in polynomial time.
1553 138
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
FIGURE 2. A machine computing < in Z. The decision problem (Y, Yyea) is described as follows: Y = C°° and Yye8 = \J Yyea,k where fceN
Yyes,k = {(fc,ht(fc),zi,...,z ht(fc )) | zi e {l,...,fc}}. The embedding of Twenty Questions in (Y, Yyea) is simply: (fc,ht(*),2)-»(*,ht(*),z,l,... I l) where the number of ones is ht(fc) — 1. The proof is finished by the next lemma. LEMMA
□
6. (V, YyeB) is in NP over C.
PROOF. The NPc machine operates on variables (ui,U2,zi,...,zn,wo,...,wn,vj0,...,Vjn)
for j = 1,2,3,4.
It checks if w2 is an integer by addition of l's. It checks if the input size (given with the input by definition) is 6U2 + 5 . If so n = u 2 . It checks if wn = 1, Wi(wi — 1) = 0 and Vji(vji — 1) = 0 for i = 0,... ,n and j = 1,2,3,4. It checks if U l = £"=02*iUi. It sets Xj = E,n=02*t>ji for j = 1,2,3,4. Finally it checks if «i = z\ + $^1 =1 £?. If so it outputs Yes. Note that if the tests are verified,
1554
ALGEBRAIC SETTINGS FOR THE PROBLEM "P ? NP?"
139
the w's and v's are 0 or 1; ui, the Xj and hence z\ are non-negative integers and «2 = ht(ui). The time required is a constant times «2• Finally we show that every element of YyeStk has a positive test. Let (fc,ht(fc),zi,...,zht(fc)) S
*yes,A:*
Then z\ is a non-negative integer so that k — z, is sum of four integers squared, k - z\ = x\ + x% + X3 + 14.
□
REMARK 2. The result and proof of Theorem 6 are valid if C is replaced by Q everywhere in the statement. THEOREM 7. // Twenty Questions over C is tractable, then Twenty Questions over Z is tractable.
PROOF. It follows immediately from the elimination of constants in Sections 2 and 4. □ 6. Proof of Theorems 2 and 3 PROOF OF THEOREM 3. Suppose that P = NP over C. Then by Theorems 6 and 7, Twenty Questions over Z is tractable. Thus there is a machine over Z deciding
Given (k, ht(fc), z ) e Z x N x N x Z , does z e [1, A]? in time (logfc)c. By the Canonical Path Theorem for each k there is a one variable non-trivial polynomial gk 6 2[t] vanishing on the set {1,2, ...,&} with r(gk) < (logfc)c. Observe that the hypothesis preceeding Theorem 3 is now violated. That is Zer(flfc) > k > (logfc)c > TfoO for k > k0.
a We now prove Theorem 2. PROOF OF THEOREM 2. We know that for each k, the degree of gk is less than or equal to 2T<-9kK So there is an integer /, |{| < 2r(9fc> with gr(l) ^ 0. We may assume |/| is minimal satisfying gk(l) ^ 0. By Proposition 1, T(/) < 2r(gk) so that T(1) < 2(logfc)c. Then gk is zero at each integer between 0 and I. Observe that gk(l) has A! as a factor by checking the 2 cases I < 0 and I > k. Moreover by evaluating gk at /,
T(9k(i))<3(\ogky. Let m/t = gk(l)/k\ (l depends on k also) in the definition of ultimately hard to compute. This finishes the proof of Theorem 2. □ 7. Main Theorem, An Algebraic Proof of the Converse Let A" be an algebraically closed field and L a field, K c L. A set S C K[t\,..., t„] determines an algebraic set VK C Kn by x € VK if and only if f(x) = 0 for all / 6 S. Moreover S also determines an algebraic set Vi, C Ln by x e Vi, if and only if f(x) = 0 all / € 5.
1555 140
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE LEMMA
7. With S and notation as above, let V[ be the algebraic set defined by
V'L = {xe (L)n | f(x) = 0allf€
K^,...,
tn] B f = 0 on VK).
ThenVi = VL. PROOF. Since clearly V'L C Vi, it is sufficient to show that any f e K[ti,... ,tn] vanishing on VK must also vanish on Vi. But by the Hilbert Nullstellensatz such an / satisfies, for some / > 0, / ' € IK{S), the ideal generated by S over K. Therefore / ' also vanishes on VL and hence / does. □
Therefore VL is determined by VKPROPOSITION 10. Let S c K[x\,...,xn] and gi,---,gi € K[xlt... ,xn}. Let VK and Vr, be defined by S. If there is a point z € VL such that g,(z) ^ 0,/or all i = I,...,I then there is a point z' € VK such that gt(z') ^ 0, Vi = 1,..., /. PROOF. We first prove the proposition in case VK is irreducible. Now proceed by induction on /. The case / = 0 is already done in the proof of Lemma 4. By induction we suppose the assertion proven for I — 1 and establish it for I. Assume that z € VL and gi{z) =£ 0 for all i = 1,... ,1. By induction the set U of z' e VK such that gt(z') ^ 0, for all i = 1,... ,1 — 1 is non-empty and Zariski open. If there is no z' € U such that gi(z') ^ 0 then o/ is zero on U and hence zero on VK by the irreducibility of VK- Hence by the Nullstellensatz there is an m such that 5,m is in the ideal IK(S) generated by S in K\x\,... ,xn). Hence g™ is also in the ideal IL{S) generated by S in L[xi,... ,xn] and ; vanishes on VL which is a contradiction. The general case is finished by the next lemma. □
LEMMA 8. Let VK C Kn be an algebraic set xuiih VK the union of algebraic sets V\ and V^- Then VL = VU U V2.L. PROOF. For i - 1,2, the ideals satisfy I(Vi) D I(VK)Thus if x £ Li%L, i = 1 and 2, then x 6 VL- On the other hand if x $ V\,L U V2>L, then there exist fi € I{VUK), i = 1,2 such that f{{x) ? 0. Thus f\h{x) / 0 and fxf2 i I(V1)UI(V2)
= I(VK)SOX^VL-
□
A basic quasi-algebraic formula over a ring R is: /i(x)=0,...,//(i)=0
* = l,...,/,
ft(i)/0,
j = l,...,*}.
A basic quasi-algebraic formula over Z defines a basic quasi-algebraic set over Z in K1" for any field K. A subset of Km is quasi-algebraic over R if it is the union of a finite number of basic quasi-algebraic sets over R. Quasi-algebraic sets over R in Km are closed under finite union, finite intersection and the operation of taking complements.
1556 ALGEBRAIC SETTINGS FOR THE PROBLEM "P ^ NP?"
141
PROPOSITION 11. Given n,m there is a finite set of basic quasi-algebraic for mulas over Z such that: given any field K, n x m matrix A over K, and vector b e Kn then the linear equation A(x) — b has a solution in Km if and only if (A, b) is in the quasi-algebraic set in Knxm+n defined by these formulas. PROOF. The system A(X) = b has a solution if and only if there are k columns of A such that the (n x k) matrix B determined by them has rank k while the nx (k + 1) matrix obtained by adjoining the column 6 also has rank k, 0 < k < m. This condition is expressed in terms of the determinants of the minors of A which are polynomial over Z in the coefficients of A. □
COROLLARY 2. Given m, n, and a vector of degrees d= (d\,... ,dm), there is a finite set of basic quasi-algebraic formulas over Z such that for any algebraically closed field K, the system of equations /i(z) = 0 , . . . , /„(*) = 0, deg fi = dt has a solution in Kn if and only if the coefficients of the fi lie in the quasi-algebraic set determined by these formulas. PROOF. By the effective Nullstellensatz, the system f\{x) = 0 , . . . , / m ( x ) = 0 has no common zero if and only if there exist
algebraically closed fields. / / P = NP over K, then
PROOF. It suffices to show that the machine M which decides Hubert's Null stellensatz over K in polynomial time decides it over L with the same polynomial time bounds. Fix n, m and d. Let Kn,m>d be the set of corresponding inputs of HN/TsT, and L„,m,d for HN/L. Thus / € Knimj consists of m polynomials f \ , . . . , fm of K[t\,... ,t„] with degree fi = a\. The yes subset of ifn,m,d will be denoted by Kn,m,d,o, and the yes subset of Ln<m,d by L„,m,d,0. Assume M has two output nodes, yes and no and that the time bound for inputs of An,m,d is T. Consider a yes instance y of HN/L and let NV>T be the node of M in the orbit of y at time T. Since Ani7n,di0 and Ln,m,d,o are defined by the same sets of basic quasi-algebraic formulas over Z and the node is determined by the basic quasi-algebraic formulas over K determined by the branch nodes in the orbit of y up to time T, Proposition 10 implies that there is a yes instance of .Kn.m.d at node Ny^ at time T. Thus Nyj- is the yes node. The same argument applies to a no instance, interchanging yes and no. □
8. Main Theorem, A Model Theoretic Proof of the Converse In this section we give an alternate proof of Theorem 8 using model theoretic results and techniques. Assuming K C L are algebraically closed fields, it suffices to prove the following two lemmas.
1557 142
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
LEMMA 9. If M is a polynomial time machine over K that outputs the value 0 or 1 when input an element of K°°, then the same is true when K is replaced by L (and hence by any field extension of K). LEMMA 10. If M is a time-bounded machine over K that decides HN/K, then the set of inputs to M from L°° that output the value 1 is exactly the set of yes instances o/HN/L.
Lemmas 9 and 10 follow easily from the Model Completeness (Strong Transfer Principle) of the theory of algebraically closed fields: Suppose K C L are algebraically closed fields and $ is a first order sentence in the language of fields with constants from K. Then $ is true when interpreted in K if and only if $ is true when interpreted in L. To prove Lemma 9, let p be the polynomial time bound for M over K and let H be the computing endomorphism of M over K. We apply the Strong Transfer Principle to each sentence $„, n > 0 (seen easily to be writable as a first order sentence over K): Vy3zo • • • 3zv(n)3w[zo = (1, y) fcj^ zk = H(zk-\)
& Zp(n) = (N, w) & (0(w) = 0 or 0(w) = 1)]
where y = (yi,... , yn) and w = (u>i,... , wp{n)). The sentence $ n asserts that for each input to M of size n, the computation halts in time bounded by p(n) with output value 0 or 1. Each sentence $ n is true in A", so each is true in L. We use the same technique to prove Lemma 10. For each m, d, n let /i(» 1 ,*) = 0 , . . . , / r a ( y m , * ) = 0 be the general system of m polynomial equations of degree d in n variables x = ( i i , . . . , xn) and variable coefficients y' = ( y ' j , . . . , y*;), t = 1,.. - , m (here / de pends on d and n). Let p(n) be a (not necessarily polynomial) time bound for M. We apply the Strong Transfer Principle to each sentence $m,d,n. m,d,n> 0: Vy 1 .. .Vy m {3i(&- ,/<(»', x) = 0) < ^ 3zo... 3zp{ml)3w[zo = (1, ( y \ . . . ,y m )) &&I 0 zk = #(«*_,) & zp(ml) = (N,w) & 0{w) = 1]} The sentence $ m ,d, n asserts that for each sequence of coefficients y 1 , . . . ,y m (from the given field), the system fi(yl,x) = 0,... ,/ m (j/ m .x) = 0 has a solution (in the given field) if and only if M with input (y 1 ,... , ym) halts with output 1. Each such sentence is true in K, therefore each is true in L. 9. Additional comments and bibliographical remarks The part of Theorem 1 asserting P = NP over C implies P = NP over Q, is proved here for the first time. The same is true for the Witness Theorem of Section 3 and Proposition 9 as well. The converse in Theorem 1 is due to Michaux [1994] who gave a model theoretic proof similar to ours. Much of the rest is from [Snub and Smale TA]. In particular Theorems 2 and 6 are proved in that paper. A version of Theorem 5 is used in [Shub 1993].
1558 ALGEBRAIC SETTINGS FOR THE PROBLEM "P £ NP?"
143
The function T is a version of standard concepts in algebraic complexity theory as for example in Heintz and Morgenstern [1993] There is also a simpler function without multiplication in the old subject of additive chains (see Scholz [1937] and Knuth [1981]). Some results on r are in [de Melo and Svaiter TA] and in [Moreira 1995]. The relationship of the open problem in Section 1 to factoring was first pointed out to us by Don Coppersmith. For related results on factoring see [Strassen 1976]. For the necessary material on heights needed in Section 3 and its appendix see [Lang 1991]. Lang [1993] is a good background in general for the algebra and in particular for the field theory (e.g. Lemma 4 of Section 4). REMARK 3. Michaux [1994] also proves that if C C K c L where K is alge braically closed, then P = NP over L implies P = NP over K. REMARK
4. Bruno Poizat has pointed out the following result.
THEOREM
9. IfP = NP over an infinite field K, then K is algebraicaly closed.
The proof is based on a result of Angus Mcintyre [1971] stating that if an infinite field admits elimination of quantifiers then it is algebraicaly closed. Then the idea is that if P = NP over K, HN/if is solved by a time bounded machine over K. Then it can be shown that K admits elimination of quantifiers. An analogue of Mcintire's result to ordered and valued fields can be found in [Mcintyre, McKenna, and van den Dries 1983]. REMARK 5. It follows from Theorem 1 and the previous remark that the prob lem P = NP over K reduces to the single problem P = NP over Q in characteristic zero.
Open Problem real fields?
Does a similar result prevail in characteristic p ^ 0? And for References
BLUM, L., M. SHUB, and S. SMALE (1989). On a theory of computation and
complexity over the real numbers: NP-completeness, recursive functions and universal machines. Bulletin of the Amer. Math. Soc. 21, 1-46. DE MELO, W. and B. SVAITER (TA). The cost of computing integers. To appear in Proceedings of the Amer. Math. Soc. HEINTZ, J. and J. MORGENSTERN (1993). On the intrinsic complexity of elimi nation theory. Journal of Complexity 9, 471-498. KNUTH, D. (1981). The Art of Computer Programming, Volume 2. AddisonWesley. LANG, S. (1991). Diophantine Geometry. Springer-Verlag. LANG, S. (1993). Algebra, 3rd edition. Addison-Wesley. MCINTYRE, A. (1971). On u>i-categorical theories offields,fund. Math. 71,1-25. MCINTYRE, A., K. MCKENNA, and L. VAN DEN DRIES (1983). Elimination of quantifiers in algebraic structures. Adv. in Math. 47, 74-87. MICHAUX, C. (1994). P ^ NP over the nonstandard reals implies P / NP over R. Theoretical Computer Science 133, 95-104. MOREIRA, C. (1995). On asymptotical estimates for arithmetical cost functions. Preprint. SCHOLZ, A. (1937). Aufgabe 253. Jahresber. Deutsch. Math.-Verein. 47, 41-42.
1559
144
L. BLUM, F. CUCKER, M. SHUB, AND S. SMALE
M. (1993). Some remarks on Bezout's theorem and complexity theory. In M. Hirsch, J. Marsden, and M. Shub (Eds.), From Topology to Computation: Proceedings of the Smalefest, pp. 443-455. Springer-Verlag.
SHUB,
SHUB, M. and S. SMALE (TA). On the intractability of Hubert's Nullstellensatz
and an algebraic verion of "P = NP". To appear in Duke J. of Math. V. (1976). Einige Resutate iiber Berechungskomplexitat. Jber. Deutsch. Math.-Verein. 78, 1-8.
STRASSEN,
LENORE BLUM, INTERNATIONAL COMPUTER SCIENCE INSTITUTE, 1947 CENTER S T . , BERKELEY, CA 94704, U.S.A., AND MATHEMATICAL SCIENCES RESEARCH INSTITUTE, 1000 CENTENNIAL DRIVE,
BERKELEY, CA 94720
E-mail address: lblumCicsi.berkeley.edu or lblumQmsri.org FELIPE CUCKER, UNIVERSITAT POMPEU FABRA, BALMES 132, BARCELONA 08008, SPAIN
E-mail address: cuckerOupf. es MIKE SHUB, IBM T. J. WATSON RESEARCH CENTER, YORKTOWN HEIGHTS. NY
10598-0218,
U.S.A. E-mail address: shubCwatson.ibm.com STEVE SMALE, DEPARTMENT OF MATHEMATICS, CITY UNIVERSITY OF HONG KONG, TAT CHEE AVE, KOWLOON, HONG KONG
E-mail address: masmaleCsobolev.cityu.edu
1560
Ada Numerica (1997), pp. 523-551
© Cambridge University Press, 1997
Complexity theory and numerical analysis Steve Smale Department of Mathematics City University of Hong Kong Hong Kong E-mail: [email protected]
CONTENTS 1 Introduction 2 Fundamental theorem of algebra 3 Condition numbers 4 Newton's method and point estimates 5 Linear algebra 6 Complexity in many variables 7 Probabilistic estimates 8 Real machines 9 Some other directions References
523 525 528 531 533 538 540 542 546 546
Preface Section 5 is written in collaboration with Ya Yan Lu of the Department of Mathematics, City University of Hong Kong. 1. I n t r o d u c t i o n Complexity theory of numerical analysis is the study of the number of arith metic operations required to pass from the input to the output of a numerical problem. To a large extent this requires the (global) analysis of the basic algorithms of numerical analysis. This analysis is complicated by the existence of illposed problems, conditioning and round-off error. A complementary aspect ('lower bounds') is the examination of efficiency for all algorithms solving a given problem. This study is difficult and needs a formal definition of algorithm.
1561 524
S. SMALE
Highly developed complexity theory of computer science provides some inspiration to the subject at hand. Yet the nature of theoretical computer science, with its foundations in discrete Turing machines, prevents a simple transfer to a subject where real number algorithms such as Newton's method dominate. One can indeed be sceptical about a formal development of complexity into the domain of numerical analysis, where problems are solved only to a certain precision and round-off error is central. Recall that, according to computer science, an algorithm defined by a Turing machine is polynomial time if the computing time (measured by the number of Turing machine operations) T(y) on input y satisfies: T(y) < K(size(y))c.
(1.1)
Here, size(y) is the number of bits of y. A problem is said to be in P (or tractable) if there is a polynomial time algorithm (i.e. machine) solving it. The most natural replacement for a Turing machine operation in a nu merical analysis context is an arithmetical operation, since that is the basic measure of cost in numerical analysis. Thus, one can say with little objec tion that the problem of solving a linear system Ax = b is tractable because the number of required Gaussian pivots is bounded by en and the input size of the matrix A and vector b is about n 2 . (There remain some crucial ques tions of conditioning to be discussed later.) In this way complexity theory is part of the tradition of numerical analysis. But this situation is no doubt exceptional in numerical analysis in that one obtains an exact answer, and most algorithms in numerical analysis solve problems only approximately with, say, accuracy e > 0, or precision l o g e - 1 . Moreover, the time required depends more typically on the condition of the problem. Therefore it is reasonable for 'polynomial time' to be recast in the form: T(y, e)
( M (y) + size(y) - log ej.
(1.2)
Here, y = (j/i,- • • ,yn), with y^ e R is the input of a numerical problem, with size(y) = n. The accuracy required is e > 0 and /x(y) is a number representing the condition of the particular problem represented by y (/x(y) could be a condition number). There are situations where one might replace /x by log/i or l o g e - 1 by logloge - 1 , for example. Moreover, using the notion of approximate zero, described below, the e might be eliminated. I see much of the complexity theory ('upper bound' aspect) of numerical analysis conveniently represented by a two-part scheme. Part 1 is the es timate (1.2). Part 2 is an estimate of the probability distribution of n, and
1562 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
525
takes the form
prob{y.fi(y)>K}<^y,
(1.3)
where a probability measure has been put on the space of inputs. Then Parts 1 and 2 combine, eliminating the /z, to give a probability bound of the complexity of the algorithm. The following sections illustrate this theme. One needs to understand the condition number fi with great clarity for the procedure to succeed. I hope this gives some immediate motivation for a complexity theory of numerical analysis and even to indicate that, all along, numerical analysts have often been thinking in complexity terms. Now, complexity theory of computer science has also studied extensively the problem of finding lower bounds for certain basic problems. For this one needs a formal definition of algorithm, and the Turing machine begins to play a serious role. That makes little sense when the real numbers of numerical analysis dominate the mathematics. However without too much fuss we can extend the concept of a machine to deal with real numbers, and one can also start dealing with lower bounds of real number algorithms. This last is not so traditional for numerical analysis, yet the real number machine leads to exciting new perspectives and problems. In computer science, consideration of polynomial time bounds led to the fundamentally important and notoriously difficult problem 'P = NP?'. There is a corresponding problem for real number machines, namely 'P = NP over R?'. The above is a highly simplified, idealized snapshot of a complexity theory of numerical analysis. Some details follow in the sections below. Also see Blum, Cucker, Shub and Smale (1996), referred to hereafter as the Mani festo, and its references for more background, history and examples.
2. Fundamental theorem of algebra The fundamental theorem of algebra (FTA) deserves special attention. Its study in the past has been a decisive factor in the discovery of algebraic num bers, complex numbers, group theory and more recently in the development of the foundations of algorithms. Gauss gave four proofs of this result. The first was in his thesis which, in spite of a gap (see Ostrowski in Gaxiss), anticipates some modern algorithms (see Smale 1981). Constructive proofs of the FTA were given in 1924 by Brouwer and Weyl. Further, Peter Henrici and his co-workers have given a substantial de velopment for analysing algorithms and a complexity theory for the FTA. See Dejon and Henrici (1969) and Henrici (1977). Also, Collins (1975) gave
1563 526
S. SMALE
a contribution to the complexity of FTA. See especially Pan (1996) and McNamee (1993) for historical background and references. In 1981-82, two articles appeared with approximately the same title, Schonhage (1982) and Smale (1981), which systematically pursued the issue of complexity for the FTA. Coincidentally, both authors gave main talks at the International Congress of Mathematicians, Berkeley 1986, on this subject; see Schonhage (1987) and Smale (1987a). These articles fully illustrate two contrasting approaches. Schonhage's algorithm is in the tradition of Weyl, with a number of ad ded features which give very good polynomial time complexity bounds. The Schonhage analysis includes the worst case and the implicit model is the Tur ing machine. On the other hand, the methods have never extended to more than one variable, and the algorithm is complicated. Some subsequent de velopments in a similar spirit include Renegar (19876), Bini and Pan (1987), Neff (1994), and Neff and Reif (1996). See Pan (1997) for an account of this approach to the FTA. In contrast, in Smale (1981), the algorithm is based on continuation meth ods such as Kellog, Li, and Yorke (1976), Smale (1976), Keller (1978), and Hirsch and Smale (1979). See Allgower and Georg (1990, 1993) for a survey. The complexity analysis of the 1981 paper was a probabilistic polynomial time bound on the number of arithmetic operations, but much cruder than Schonhage's. The algorithm, based on Newton's method, was simple, ro bust, easy to program, and extended eventually to many variables. The implicit machine model was that of Blum, Shub and Smale (1989), here after referred to as BSS (1989). Subsequent developments along these lines include Shub and Smale (1985, 1986), Kim (1988), Renegar (19876), Shub and Smale (1993a, 19936, 1993c, 1996 and 1994), hereafter referred to as Bez I-V, respectively, and Blum, Cucker, Shub and Smale (1997), hereafter referred to as BCSS (1997). Here is a brief account of some of the ideas of Smale (1981). A point z is called an approximate zero if Newton's method starting at z converges well in a certain precise sense; see Section 4 below. The main theorem of this paper asserts the following. Theorem 2.1 A sufficient number of steps of a modified Newton's method to obtain an approximate zero of a polynomial / (starting at 0) is polynomially bounded by the degree of the polynomial and l/<7, where a is the probability of failure. For the proof, an invariant fi = //(/) of / is defined akin to a condition number of / . Then the proof is broken into two parts. Part 1: A sufficient number of modified Newton steps to obtain an approx imate zero of / is polynomially bounded by /i(/).
1564 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
527
The proof of Part 1 relies on a Loewner estimate related to the Bieberbach conjecture. P a r t 2: The probability that fx(f) is larger than k is less than k~c, some constant c. The proof of Part 2 uses elimination theory of algebraic geometry and geometric probability theory, Crofton's formula, as in Santal6 (1976). The crude bounds given in Smale (1981), and the mathematics too, were substantially developed in Shub and Smale (1985, 1986). Here is a more detailed, more developed, complexity theoretic version of the FTA in the spirit of numerical analysis. See BCSS (1997) for the details. Assume given (or input): a complex polynomial f(z) = ^ diZx in one complex variable, a complex number zo> and an e > 0. Here is the algorithm to produce a solution (output) z* satisfying
l/(OI < e-
(2-1)
Let to = 0, U = ti-i + At, where At = 1/fc, for some positive integer k; thus tfe = 1, and we have a partition of [0,1]. For any polynomial g, we define Newton's method by Ng(z) = z- ^pr, g'{z)
for all z € C, such that g'(z) ^ 0.
Let ft(z) = f(z) — (1 — t)f(zo). Then, generally, there is a unique path Q such that ft(Ct) = 0 all t € [0,1] and Co = ZQ- Define inductively Zi = Nft.(zi-1),
i = l,...,k,
z* = zk.
(2.2)
It is easily shown that for almost all (/,20), Zi will be defined, i = l,...,k, provided At is small enough. We may say that k = 1/At is the 'complexity'. It is the main measure of complexity in any case: the problem at hand is, 'how big may we choose At and still have z* satisfying (2.1) and (2.2)?' (i.e. so that the complexity is the lowest). Next a 'condition number' /x(/, 20) is defined which measures how close £t is to being ill-defined. (More precisely /z(/, zo) = cosec 0 where 6 is the supremum of the angles of sectors about /(20) for which the inverse / - 1 mapping f(zo) to ZQ is defined.) T h e o r e m 2.2 A sufficient number k of Newton steps defined in (2.2) to achieve (2.1) is given by fc<26/x(/,^o)(log^^
+ l)
1565 528
S. SMALE
Remark 2.1 (a) We are assuming 0 < e < 1/2. (b) Note that the degree d of / plays no role, and the result holds for any (f,zo,e). (c) The proof is based on 'point estimates' (a-theory) (see Section 4 below) and an estimate of Loewner from Schlicht function theory. Thus it doesn't quite extend to n variables. It remains a good problem to find the connection between Theorem 2.2 and Theorem 6.1. For the next result suppose that / has the form d t=0
Theorem 2.3 The set of points ZQ € C, \ZQ\ = R > 2, such that /i(/, ZQ) > b, is contained in the union of 2(d — 1) arcs of total angle 2/1
+Sm
. _,
d{b
1
\
R=i)-
This result is an estimate on how infrequently poorly conditioned pairs (/, ZQ) occur. It is straightforward to combine Theorems 2.2 and 2.3 to eliminate the H and obtain both probabilistic and deterministic complexity bounds for approximating a zero of a polynomial. The probabilistic estimate improves the deterministic one by a factor of d. Theorem 2.3 and these results are in Shub and Smale (1985, 1986), but see also BCSS (1997), and Smale (1985). Remark 2.2 The above-mentioned development might be improved in sharpness in two ways. (A) Replace the hypothesis on the polynomial / by assuming as in Renegar (19876) and Pan (1996) that all the roots of / are in the unit disk. (B) Suppose that the input polynomial / is described not by its coefficients, but by a 'program' for / . 3. C o n d i t i o n n u m b e r s The condition number as studied by Wilkinson (1963), important in its own right in numerical analysis, also plays a key role in complexity theory. We review it now, especially some recent developments. For linear systems, Ax = b, the condition number is defined in most basic numerical analysis texts. The Eckart and Young (1936) theorem is central, and may be stated as \\A-l\\-*
=
df(A,Zn),
1566 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
529
where A is a non-singular n x n matrix, with the operator norm on the left and the Frobenius distance on the right. Moreover, E n is the subspace of singular matrices. The case of 1-variable polynomials was studied by Wilkinson (1963) and Demmel (1987), among others. Demmel gave estimates on the condition number and the reciprocal of the distance to the set of polynomials with multiple roots. We now give a more general context for condition numbers and give exact formulae for the condition number as the reciprocal of a distance to the set of ill-posed problems following Bez I, II, IV, Dedieu (1997a, 19976, 1997c) and BCSS (1997). Consider first the context of the implicit function theorem:
F:RfcxRm^Rm, OF m
-^-(ao, 2/0) : K -»• R m oy
C\
F(a0,y0) = 0, non-singular.
Then there exists an open neighbourhood U of ao in Rk and a C 1 map G : U -► R m such that G(a0) = t/o and F(a, G(a)) = 0, for aeU. Regard Fa : R m —► R m , Fa(y) = F(a,y), as a system of equations para meterized by a £ R*. Then a might be the input of a problem Fa(y) = 0 with output y; G is the 'implicit function'. Let us call the derivative DG(ao) : Rfc —► R m the condition matrix at (ao,2/o)- Then the condition number n(ao,yo) = /i, as in Wilkinson (1963), Rice (1966), Wozniakowski (1977), Demmel (1987), Bez IV, and Dedieu (1997a), is defined by Mao,yo) = ||i?G(ao)||, the operator norm. Thus n(ao, yo) is the bound on the infinitesimal output error of the system Fa(y) = 0 in terms of the infinitesimal input error. It is important to note that, while the map G is given only implicitly, the condition matrix , dF ,dF. n_, 1 DG(a0) = —{a0,yo) — (oo,y0) is given explicitly, as is its norm, the condition number /i(ao,2/o)An example, given by Wilkinson, is the case where Rfc is the space of real polynomials / in one variable of degree < A; — 1, and R m = R the space of C, F(f, C) = /(C)- O n e m a Y compute that in this case
For the discussion of several variable polynomial systems, it is convenient to use complex numbers and homogeneous polynomials.
1567
530
S. SMALE
If / : C n — ► C is a polynomial of degree d, we may introduce a new vari able, say 2 0 , and define / : C n + 1 -> C by f(l,zi,...,zn) = f{z\,...,zn) and /(Azo, \z\,..., Xzn) = Xdf(zo, ■ ■ ■, zn). Thus / is a homogeneous poly nomial. If / : C" -> C", / = ( / i , . . . , /„), deg /j = dj, i = 1 , . . . , n, is a polynomial system, then by letting / equal ( / i , . . . , / n ) , we obtain a homogeneous sys tem / : C" + 1 — C n . Any zero of / will also be a zero of / and justification can be made for the study of such systems in their own right. Thus now we will consider such systems, say / : C n + 1 —► C n and denote the space of all such / by Hd, d = (du ..., d„), degree ft = di. Recall that an Hermitian inner product on C n + 1 is defined by n
z,w€Cn+l.
(z,w) = ^2ziWi,
Now, define for degree d homogeneous polynomials / , g : C n + 1 —► C,
= 5 Z ( Q )
£*&*'
where f(z) = £
faza,
gaza.
g(z) = £
\a\=d
\a\=d
Here a — (a.\,..., a n +i) is a multi-index and fd\
d\
\aj
ai!---Q„+i!'
^ "
^ '
The weighting by the multinomial coefficient is important, and yields unitary invariance of the inner product, as below. Proposition 3.1 (Reznick 1992) Let f,Nx : C n + 1 -► C be degree d homogeneous polynomials, where Nx(z) = (x,z)d. Then f(x) = {f,Nx). Corollary 3.1
|/(x)| < H/ll ||iV,|| < H/ll ||x||d. For f,g£Hd,
define
= E ^ f * .
ii/n = (/'/)1/2-
Dedieu has suggested weighting by 1/dj to make the Condition Number Theorem below more natural. The unitary group U(n + 1) is the group of all linear automorphisms of C n + 1 which preserve the Hermitian inner product.
1568 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
531
There is an induced action of U(n + 1) on Hd defined by
{crf){z) = f(a-1z),
z€Cn+1,
feHd.
Then it can be proved (see, for instance, BCSS 1997) that
(o-f,o-g) = (f,g),
f,g€Hd,
This is unitary invariance. There is a history of this inner product going back at least to Weyl (1932), with contributions or uses in Kostlan (1993), Brockett (1973), Reznick (1992), Bez I-V, Degot and Beauzamy (1997), Stein and Weiss (1971), Dedieu (1997a). Now we may define the condition number n(f, C) for / G Hd, C £ C n + 1 , /(C) = 0 using the previously defined implicit function context. To be technically correct, one must extend this context to Riemannian manifolds to deal with the implicitly defined projective spaces. See Bez IV for details. The following is proved in Bez I (but see also Bez III, Bez IV). Condition Number Theorem 1 M(/,0 =
Let / G Hd, C € C n + 1 , /(C) = 0. Then
«*((/, O.Ec)"
Here the distance d is the projective distance in the space {g G Hd ■ s(C) = 0} t o the subset where C is a multiple root of g. The proof uses unitary invariance of all the objects. Thus one can reduce to the point C, = (1,0, • • •, 0), and then to the linear terms, and then to the Eckart-Young theorem. Dedieu (1997 a) has generalized this result quite substantially, and has considered sparse polynomial systems (Dedieu 19976). Thus a formula for the eigenvalue problem becomes a special case.
4. Newton's method and point estimates Say that z G C n is an approximate zero of / : C" —♦ C n (or R n —► R n , or even for Banach spaces) if there is an actual zero C of / (the 'associated zero') and
IN-CII
(4.1)
where Z{ is given by Newton's method
zi = NJ{zi-l),
zo = z,
Nf(z) =
z-Df(z)-1f(z).
Here Df(z) : C n -> C n is the (Frechet) derivative of / at z. An approximate zero z gives an effective termination for an algorithm provided one can determine whether z has the property (4.1).
1569
532
S. SMALE
Towards that end, the following invariant is decisive 7 = 7 ( / . z) = sup k>2
Df{z)-lD^f{z)
1±
k\
Here D^f(z) is the fcth derivative of / considered as a /c-linear map and we have taken the operator norm of its composition with Df(z)~1; if the expression is not defined, then use 7 = 00. See Smale (1986), Smale (1987a) and Bez I for details of this development. The invariant 7 turns out to be a key element in the complexity theory of non-linear systems. Although it is defined in terms of all the higher deriv atives, in many contexts it can be estimated in terms of the first derivative, or even the condition number. Theorem 4.1 (Smale 1986; see also Traub and Wozniakowski 1979) Let / : C n ^ C n , ( G C n with /(C) = 0. If
7(/,0ll*-Cll<^^, then z is an approximate zero of / with associated zero £. Now let a = a(f,z)
= (3(f,z)j(f,z),
0(f,z)
= \\D f {z)~l f {z)\\.
Theorem 4.2 (Smale 1986) There exists a universal constant ao > 0 such that: if a(f, z) < ao for / : C n —> C", z € C n , then z is an approximate zero of / (for some associated actual zero C of / ) . Remark 4.1 This is the result that motivates 'point estimates'. One uses it to conclude that z is an approximate zero / by checking an estimate at the point z only. Nothing is assumed about / in a region or / at £. Remark 4.2 For this definition of approximate zero, the best value of ao is probably no smaller than 1/10. See developments, details and discussions in Smale (1987a), Wang (1993), Bez I, and BCSS (1997). Now how might one estimate 7? In Smale (1986, 1987a), there is an estimate in terms of the first derivative of / , but an estimate in Bez I seems much more useful. In the context of Section 3, let / € Hd, C G C n + 1 , /(C) — 0) snd 7o(/, C) = HCll7(/)0- The last is to make 7 projectively invariant. Recall that D = max(dj), d = ( d i , . . . , d„), a\ = deg/*. Theorem 4.3 (Bez I)
70(/,C)<^U/,0Recall that /x(/, £) is the condition number.
1570
COMPLEXITY THEORY AND NUMERICAL ANALYSIS
Remark 4.3
533
One has a similar estimate without assuming f(Q = 0.
As a corollary of Theorem 4.3 and a projective version of Theorem 4.1, one obtains the following. Theorem 4.4 (Separation of zeros, Malajovich-Munoz 1993, BCSS 1997, Dedieu 19976, 1997d) Let / € Hd, and C, C' be two distinct zeros of/. Then
D u,(f)
= max(deg/i), / = (/i,... ,/„), = max u(f, C) is the condition number of / , C,/(C)=o and d is the distance in projective space.
Remark 4.4
One has also the stronger result
Remark 4.5 The strength of Theorem 4.4 lies in its global aspect. It is not asymptotic even though n is defined just by a derivative. We end this section by stating a global perturbation theorem (Dedieu (19976)). Theorem 4.5
Let / , g : Cn -» C n , C € C n with /(C) = 0. Then, if a{9,0
<
^
and
WI-DfiO-'DgiOWK9-^^-, there is a zero £' of g such that
IIC-C'll<2/i(/,C)ll/-5llHere everything is affine including /x(/, C). This uses Theorem 4.2. 5. Linear algebra Complexity theory is quite implicit in the numerical linear algebra literature. Indeed, numerical analysts have studied the execution time and memory requirements for many linear algebra algorithms. This is particularly true for direct algorithms that solve a problem (such as a linear system of equations) in a finite number of steps. On the other hand, for more difficult linear algebra problems (such as the matrix eigenvalue problem) where iterative methods are needed, the complexity theory is not fully developed. It is our
1571
534
S. SMALE
belief that a more detailed complexity analysis is desirable and such a study could help lead to better algorithms in the future. 5.1. Linear systems Consider the classical problem of a system of linear equations Ax = b, where A is a. n x n non-singular matrix, b is a column vector of length n. The standard method for solving this problem is Gaussian elimination (say, with partial pivoting). The number of arithmetic operations required for this method can be found in most numerical analysis textbooks: it is 2ra3/3 + 0(n2). Most of these operations come from the LU factorization of the matrix A, with suitable row exchanges. Namely, PA = LU, where L is a unit lower triangular matrix (whose entries satisfy \lij\ < 1), U is an upper triangular matrix, and P is the permutation matrix representing the row exchanges. When this factorization is completed, the solution of Ax = b can be found in 0(n2) operations. Similar operation counts are also avail able for other direct methods for linear systems, for example, the Cholesky decomposition for symmetric positive definite matrices. Another method for solving Ax = b, and, more importantly, for least squares problems, is to use the QR factorization of A. The number of required operations is 4n 3 /3 + 0(n2). All these direct methods for linear systems involve only a finite number of steps to find the solution. The complexity of these meth ods can be found by counting the total number of arithmetic operations involved. A related problem is to investigate the average loss of precision for solving linear systems. It is well known that the condition number n of the matrix A bounds the relative errors introduced in the solution by small perturbations in b and A. Therefore, log K is a measure of the loss of numerical precision. To find its average, a statistical analysis is needed. The following result for the expected value of log K is obtained by Edelman. Theorem 5.1 (Edelman 1988) Let A be a random nxn matrix whose entries (real and imaginary parts of the entries, for the complex case) are independent random variables with the standard normal distribution, and let K = ||J4|| ||^4 _1 || be its condition number in the 2-norm; then E(logn) = logn + c +o(l),
for
n —► oo,
where c « 1.537 for real random matrices and c « 0.982 for complex random matrices. The above result on the average loss of precision is a general result valid for any method, as a lower bound. If one uses the singular value decomposition to solve Ax = 6, the average loss of precision should be close to ^(log K) above. For a more practical method like Gaussian elimination with partial
1572 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
535
pivoting, the same average could be larger. In fact, Wilkinson's backward error analysis reveals that the numerical solution x obtained from a finite precision calculation is the exact solution of a perturbed system (A + E)x — b. The magnitude of E could be larger than the round-off of A by an extra growth factor p(A). This gives rise to the extra loss of precision caused by the particular method used, namely, Gaussian elimination with partial pivoting. Well-known examples indicate that the growth factor can be as large as 2 n _ 1 . But the following result suggests that large growth factors only rarely appear exponentially. Conjecture 5.1 (Trefethen) For any fixed constant p > 0, let A be a random n x n matrix, whose entries (real and complex parts of the entries for the complex case, scaled by y/2) are independent samples of the standard normal distribution. Then, for all sufficiently large n, Prob ( p ( A ) > n a ) < r r p , where a > 1/2. For iterative methods, we mention that a complexity result is available for the conjugate gradient method (Hestenes and Stiefel 1952). Let A be a real symmetric positive definite matrix, #o be an initial guess for the exact solution x» of Ax = b, and Xj be the jih iterate of the conjugate gradient method. Then the following result is well known (Axelsson (1994), Appendix B): ||zj-s,|U<2f^|^J
||so-x.|U,
where the A-nona of a vector v is defined as ||U||.A = (vTAv)ll2. one easily concludes that if
log?
n(y/«*
From this,
A
then ||xj - x,|U < e||x0 - **IU5.2. Eigenvalue problems In this subsection, we consider a number of basic algorithms for eigenvalue problems. Complexity results for these methods are more difficult to obtain. For a matrix A, the power method approximates the eigenvector corres ponding to the dominant eigenvalue (largest in absolute value). If there is one dominant eigenvalue, for almost all initial guesses xo, the sequence generated by the power method Xj = J 4 J X O / | | A J X 0 | | converges to the dom inant eigenvector. A statistical complexity analysis for the power method
1573
536
S. SMALE
tries to determine the average number of iterations required to produce an approximation to the exact eigenvector, such that the angle between the approximate and exact eigenvectors is less than a given small number e (edominant eigenvector). These questions have been studied by Kostlan. The average is first taken for all initial guesses XQ and a fixed matrix A, then extended to all matrices for some distribution. Theorem 5.2 (Kostlan 1988) For any real symmetric n x n matrix A with eigenvalues |Ai| > |A2I > ••• > |An|, the number of iterations rt(A) re quired for the power method to produce an e-dominant eigenvector, averaged over all initial vectors, satisfies log cote ^ _ , AS ^ \[ip(n/2) - V(l/2)] + log cote < rt(A) < ' , ,. , 1 ix 1 +!» log |Ai| — log |A2| log|Ai| — log |A2| where ip(x) =
r'(x)/T(x).
When an average is taken for the set of n x n random real symmetric matrices (the entries are independent random variables with Gaussian dis tributions of zero mean, the variance of any diagonal entry is twice the variance of any off-diagonal entry), the required number of iterations is in finite. However, a finite bound can be obtained if a set of 'bad' initial guesses and 'bad' matrices of normalized measure 77 are excluded. Theorem 5.3 (Kostlan 1988) For the above n x n random real sym metric matrix, with the probability 1 — 77, the average required number of iterations to produce an e-dominant eigenvector satisfies
Te < M
" ±Ji ^ Wn/2) ~ V'(1/2) + 21ogcote) •
Similar results hold for complex Hermitian matrices. Furthermore, a fi nite bound on random symmetric positive definite matrices is also available. Statistical complexity analysis for a different method of dominant eigen vector calculation can be found in Kostlan (1991). In practice, the Rayleigh quotient iteration method is much more efficient. Starting from an initial guess XQ, a sequence of vectors {XJ} is generated from or
^\oc' 1
?/
For symmetric matrices, the following global convergence result has been established. Theorem 5.4 (Ostrowski 1958, Parlett and Kahan 1969, Batterson and Smillie 1989) Let A be a symmetric n x n matrix. For almost any choice of XQ, the Rayleigh quotient iteration sequence {XJ} converges to an
1574 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
537
eigenvector and Umj-^ooQj+i/Q] < 1> where 9j is the angle between Xj and the closest eigenvector. A statistical complexity analysis for this method is still not available. In fact, even for a fixed symmetric matrix A, there is no upper bound on the number of iterations required to produce a small angle, say, 0j < e for a small constant e. In general, for a given initial vector xo, one can not predict which eigenvector it converges to (if the sequence does converge). On the other hand, for nonsymmetric matrices, we have the following result on non-convergence. Theorem 5.5 (Batterson and Smillie 1990) For each n > 3, there is a nonempty open set of matrices, each of which possesses an open set of initial vectors for which the Rayleigh quotient iteration sequence does not converge to an invariant subspace. Practical numerical methods for matrix eigenvalue problems are often based on reductions to the condensed forms by orthogonal similarity trans formations. For a n n x n symmetric matrix A, one typically uses Householder reflections to obtain a symmetric tridiagonal matrix T. The reduction step is a finite calculation that requires 0(n3) arithmetic operations. While many numerical methods are available for calculating the eigenvalues and eigen vectors of symmetric tridiagonal matrices, we see the lack of a complexity analysis for these methods. The QR method with Wilkinson's shift always converges; see Wilkinson (1968). In this method, the tridiagonal matrix T is replaced by si + RQ (still symmetric tridiagonal), where s is the eigenvalue of the last 2 x 2 block of T that is closer to the (n, n) entry of T, and QR = T — si is the QR factorization of T — si. Wilkinson proved that the (n, n — 1) entries of this sequence of T always converge to zero. Hoffman and Parlett (1978) gave a simpler proof for the global linear convergence. The following is an easy corollary of their result. Theorem 5.6 Let T be a real symmetric n x n tridiagonal matrix. For any e > 0, let m be a positive integer satisfying m > 61og2 \ + log 2 (^ n _ 1 T n 2 _ l i n _ 2 ) + 1. Then, after m QR iterations with Wilkinson's shift, the last subdiagonal entry of T satisfies \Tn,n-l\
< ۥ
It would be interesting to develop better complexity results based on the higher asymptotic convergence rate. Alternative definitions for the last subdiagonal entry to be sufficiently small are desirable, because the usual de-
1575
538
S. SMALE
coupling criterion is based on a comparison with the two adjacent diagonal entries. The divide and conquer method suggested by Cuppen (1981) calculates the eigensystem of an unreduced symmetric tridiagonal matrix based on the eigensystems of two tridiagonal matrices of half size and a rank-one updating scheme. The computation of the eigenvalues is reduced to solving the following nonlinear equation n
c2
where {dj} are the eigenvalues of the two smaller matrices and {CJ} are re lated to their eigenvectors. This method is complicated by the possibilities that the elements in {dj} may be not distinct and the set {CJ} may con tain zeros. Dongarra and Sorensen (1987) developed an iterative method for solving the nonlinear equation based on simple rational function approx imations. See Bini and Pan (1994) for a complexity analysis of a related algorithm. A related method for computing just the eigenvalues uses the set {dj} to separate the eigenvalues and a nonlinear equation solver for the charac teristic polynomial. In Du, Jin, Li and Zeng (19976), the quasi-Laguerre method is used. An asymptotic convergence result has been established in Du, Jin, Li and Zeng (1997a), but a complexity analysis is still not available. The method is complicated by the switch to other methods (the bisection or Newton's method) to obtain good starting points for the quasi-Laguerre iterations. For a general real nonsymmetric matrix A, the QR iteration with Fran cis's double shift is widely used to triangularize the Hessenberg matrix H obtained from the reduction by orthogonal similarity transformations from A. In this case, there are simple examples for which the QR iteration does not lead to a decoupling. In Batterson and Day (1992), matrices where the asymptotic rate of decoupling is only linear are identified. For normal Hessenberg matrices, Batterson discovered the precise conditions for decoup ling under the QR iteration. See Batterson (1994) for details. To develop a statistical complexity analysis for this method is a great challenge.
6. Complexity in many variables Consider the problem of following a path, implicitly defined, by a computa tionally effective algorithm. Let Hd be as in Section 3. Let F : [0,1] - Hd x C" + 1 , F(t) = (ft,(t), satisfy ft(Q) = 0, 0 < t < 1, with the derivative Dft(Ct) having maximum rank. For example, Ct could be given by the implicit function theorem from ft and the initial Co with /o(Co) = 0.
1576 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
539
Next, suppose [0,1] is partitioned into k parts by to = 0, U = tj_i + At, At = l/jfc; thus tk = 1. Define via Newton's method Njt Zi = Nfu(zi-i),
i = l,...,k,
20 = Co-
(6.1)
For sufficiently small At, the Z{ are well defined and are good approximations of Ci- But k = 1 /At represents the complexity, so the problem is to avoid taking At much smaller than necessary. What is a sufficient number of Newton steps? Theorem 6.1 (The main theorem of Bez I)
The biggest integer k in
2 2
cLD n
is sufficient to yield Z{ by (6.1) which is an approximate zero of ft{ with associated actual zero C^, each i = 1 , . . . , k. In this estimate c is a rather small universal constant, L is the length of the curve ft in the projective space, P(7id), 0 < t < 1, D is the max of the di, i = 1 , . . . ,n and fi = maxo
z+
{yeCn+1:(y,z)=0}.
As a consequence of the Condition Number Theorem and Theorem 6.1, the complexity depends mainly on how close the path (ft, Ct) comes to the set of ill-conditioned problems. An improved proof of Theorem 6.1 may be found in BCSS (1997). For earlier work on complexity theory for Newton's method in several variables, see Renegar (1987a). Malajovich (1994) has implemented the algorithm and developed some of the ideas of Bez I. The main theorem of the final paper of the series Bez I-Bez V is as follows. Theorem 6.2 The average number of arithmetic operations sufficient to find an approximate zero of a system / : C" —► C n of polynomials is poly nomial^ bounded in the input size (the number of coefficients of / ) . On one hand, this result is surprising, because it gives a polynomial time bound for a problem that is almost intractable. On the other hand, the 'algorithm' is not uniform: it depends on the degrees of the (fi) and even the desired probability of success. Moreover, the algorithm isn't known! It is only proved to exist. Thus Theorem 6.2 cries out for understanding and development. In fact, Mike Shub and I were unable to find a sufficiently good exposition to include in BCSS (1997).
1577
540
S. SMALE
Since deciding if there is a solution to / : C n —* C n is unrflcely to be accomplished in polynomial time, even using exact arithmetic (see Section 8), an astute analysis of Theorem 6.2 can give insight into the basic problem 'What are the limits of computation?' For example, is it 'on the average' that gives the possibility of polynomial time? A real (rather than complex) analogue of Theorem 6.2 also remains to be found. Let us give some mathematical detail about the statement of Theorem 6.2. An 'approximate zero' has been defined in Section 4, as, of course, exact zeros cannot be found (Abel, Galois, et al.). Averaging is performed relative to a measure induced by the unitarily invariant inner product on homogenized polynomials of degree d = (di,... ,dn), where di = deg/i, / = (/l, • • • > /n) (see Section 3). If N = N(d) is the number of coefficients of such a system / , then unless n < 4 or some di = 1, the number of arithmetic operations is bounded by cN4. If n < 4 or some di = 1, then we get cN5. An important special case is that of quadratic systems, when dj = 2 and so N < n 3 . Then the average arithmetic complexity is bounded by a polynomial function of n. 'On the average' in the main result is needed because certain polynomial systems, even affine ones of the type / : C2 —> C 2 , have one-dimensional sets of zeros, extremely sensitive to any (practical) real number algorithm; one would say such / are ill posed. The algorithm (non-uniform) of the theorem is similar to those of Section 2. It is a continuation method where each step is given by Newton's method (the step size At is no longer a constant). The continuation starts from a given 'known' pair g : C n + 1 —► C n and C € C n + 1 , g(Q = 0 . It is conjectured in Bez V that one could take for g, the system defined by gi(z) = ZQ Zi, i = 1 , . . . , n and £ = (1> 0 , . . . , 0). A proof of this conjecture would yield a uniform algorithm. Finally, we remark that in Bez V, Theorem 6.2 is generalized to the prob lem of finding £ zeros, when t is any number between one and the Bezout number n?=i di and the number of arithmetic operations is augmented by the factor i2. The proof of Theorem 6.2 uses Theorem 6.1 and the geometric probability methods of the next section.
7. Probabilistic estimates As described in the Introduction, our complexity perspective has two parts, and the second deals with probability estimates of the condition number. We have already seen some aspects of this in Sections 2 and 5. Here are some further results.
1578 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
541
Section 3 describes a condition number for studying zeros of polynomial systems of equations. We have dealt especially with the homogeneous setting and defined projective condition number /z(/, C) for / 6 Hd, d = (d\,..., dn), degree ft = du and C € C n + 1 with /(C) = 0. Then
The unitarily invariant inner product (Section 3) on Hd induces a probab ility measure on Hd (or equivalently on the projective space P(Hd))- With this measure the following is proved in Bez II. T h e o r e m 7.1 Probability {/ e Hd : /*(/) > - } < Cde4 n
Cd = n3{n + l){N-l)(N-2)T>,
N = dimHd,
V =
]\di i=l
In the background of this and a number of related results is a geometric picture (from geometric probability theory), briefly described as follows. It is convenient to use the projective spaces P(Hd), P(Cn+l) and their product for the environment of this analysis. Define V to be the subset of ordered pairs (system, solution): V = {(/,C) € P(Hd) x P(C" + 1 ) : /(C) = 0}. Let 7Ti : V -> P(Hd), n2 : V -+ P(C n + 1 ) be the restrictions of the corres ponding projections, as shown below. V C P(Hd) x P ( C n + 1 )
P(Cn+1)
P(Hd)
T h e o r e m 7.2 (Bez II) / JxeP(Hd)
Let U be an open set in V, then
#(7rf 1 (x)n[/) = / V
'
detfDG(a)DG(a)*Y 1/2
/ 1
1
V2eP(C"+ )-/(a,2)€7r2- (i)nt/
V
'
Here DG(a) is the condition matrix, DG(a)* its adjoint and # means car dinality. This result and the underlying theory is valid in great generality (see Bez II, IV, V, BCSS (1997)).
1579
542
S. SMALE
There is one aspect of these results and arguments that is quite unsettling and pervades Bez II-V: the implicit existence theory is not very constructive. Consider the simplest case (Bez III). For the moment, let d > 1 be an integer and Hd the space of homogeneous polynomials in two variables of degree d. It follows from the above geometric probability arguments that there is a subset Sd of P(Hd) of probability measure larger than one-half such that, for / G Sd, /x(/) < d. Problem 7.1 (Bez III) that
Construct a family {fd e Hd ■ d = 2,3,...} so
H(fd) < d,
or even
/z(/ d ) < dc,
for c any constant. By 'construct', we mean to provide a polynomial time algorithm (e.g. in the sense of the machine of Section 8) which, given input d, outputs fd satisfying the above condition. (This amounts to constructing elliptic Fekete polynomials.) See also Rakhmanov, Saff and Zhou (1994, 1995). Another example of an application of the above setting of geometric prob ability is the following result. For d = (d\,..., dn), let H* denote the space of real homogeneous systems (f\,..., /„) in n+1 variables with degree fc = d{. One can average just as before and obtain the following. Theorem 7.3 (Bez II) The average number of real zeros of a real homo geneous polynomial system is exactly the square root of the Bezout number T> = riiLi di (D being the number of complex solutions). See Kostlan (1993) for earlier special cases. See also Edelman and Kostlan (1995). For the complexity results of Bez IV, V, Theorem 7.1 is inadequate. There one has similar theorems where the maximum of the condition number along an interval is estimated.
8. Real machines Up to now, our discussion might be called the complexity analysis of al gorithms, or upper bounds for the time required to solve problems. To complement this theory one needs lower bound estimates for problem solv ing. For this endeavour, one must consider all possible algorithms that solve a given problem. In turn this needs a formal definition and the development of algorithms and machines. The traditional Turing machine is ill-suited for this purpose, as is argued in the Manifesto. A 'real number machine' is the most natural vehicle to deal with problem-solving schemes based on Newton's method, for example.
1580
COMPLEXITY THEORY AND NUMERICAL ANALYSIS
543
There is a recent development of such a machine in BSS (1989) and BCSS (1997), which we will review very briefly. Each input is a string y of real numbers of the form •••0003/i-yn000--; the size S(y) of y is n. These inputs may be restricted to code an instance of a problem. An 'input node' transforms an input into a state string.
Input node
y •
y *- f(y) No
■
Computation node
■
2/i > 0 ?
Branch node
Yes y
Output node
Fig. 1. Example of a real number machine
The computation node replaces the state string by a shifted one, right or left shifted, or does an arithmetic operation on the first elements of the string. The branch nodes and output nodes are self-explanatory. The definition of a real machine (or a 'machine over R') is suggested by the example and consists of an input node and a finite number of computation, branch, and output nodes organized into a directed graph. It is the flow chart of a computer program seen as a mathematical object. One might say that this real number machine is a 'real Turing machine' or an idealized Fortran program. The halting set of a real machine is the set of all inputs such that, acting on the nodal instructions, we eventually land on an output node. An inputoutput map
1581
544
S. SMALE
A machine has polynomial time complexity (sometimes with a restricted class of inputs) if it enjoys the property T{y) < S(y)c,
for all inputs y,
(8.1)
where c is independent of y. In this estimate, T(y) is the time to the output for the input y measured by the number of nodes encountered in the computation of (j>{y). Recall that the size S(y) of y is the length of the input string y. If the size of the inputs is bounded, and there are no loops, i.e., the machine is a tree of nodes, then one has a tame machine, or an algebraic computation tree. These objects have been used to obtain lower bounds for real number problems. One such development is that of Steele and Yau (1982) and Ben-Or (1983), based on a real algebraic geometry estimate of Oleinik and Petrovski (1949), Oleinik (1951), Milnor (1964) and Thorn (1965). Another is that of Smale (19876) and Vassiliev (1992), and based on the cohomology of the braid group. Lower bounds tend to be modest and difficult to obtain, but are necessary for the understanding of the fundamental problem: 'What are the Umits of computation?' Note that the definition of a real machine is valid with strings of numbers lying in any field if one replaces the branch node with the question, ly\ = 0?' If this field is the field of two elements, one has a Turing machine, and the size becomes the number of bits." If one uses complex numbers, then one has a 'complex machine'. Side remarks 8.1 The study of zeros of polynomial systems plays a cent ral role in both mathematics and computation theory. Deciding whether a set of polynomial equations has a zero over R is even universal in a formal sense in the theory of real computation. This problem is called 'NP-complete over R' and hence its solution in polynomial time is equival ent to 'P = NP over R.' For machines over C, this problem is that of the Hilbert Nullstellensatz, and Brownawell's (1987) work was critical in get ting the fastest-known algorithm (but not polynomial time!) The relation to NP-complete over C and 'P = NP over C is as in the real case. The same applies to the field Z2 of two elements and 'P = NP over Z2?' is the same as the classical Cook-Karp problem 'P = NP?' of computer science. See BCSS (1997). My own belief is that this problem is one of the three great unsolved prob lems of mathematics (together with the Riemann hypothesis and Poincare's conjecture in three dimensions). The rest of Section 8 is more tentative, as we present suggestions in the direction of a 'second generation' real machine. For an input y of a problem, an extended notion of size still denoted by
1582
COMPLEXITY THEORY AND NUMERICAL ANALYSIS
545
S(y) could be convenient. The extended notion would be the maximum of the length of the string (i.e. the previously defined size) and other ingredi ents, as follows: (i) the condition number /i(y), or its log, or similar invariants of y (ii) the precision l o g e - 1 , where e is the required accuracy (or perhaps, depending on the problem, e, or even log loge - 1 ) of the output (iii) for integer machines, the number of bits. It is convenient to consider the traditional size of the input as part of the input (BSS 1989, BCSS 1997). Should the same hold for the extended size? We won't try to give a definitive answer here. Part of this answer is a question of convenience, part interpretation. Should the algorithm assume that the condition number is known explicitly? Probably not, at least very generally. On the other hand, if one has a good theoretical result on the distribution, one can make some guess about the condition number. This can to some extent justify taking the condition number of the particular problem as input. It is analogous, for example, to running a path-following program inputing an initial step size as a guess. Let me give an example of an open problem that fits into this framework. Let d = (di,...,dm) and Vn,d be the space of m-tuples of real polynomials / = (/ii---./m) in n variables with deg/i < d{. Put some distance D on ~Pn,d- Say that / is feasible if the system of inequalities /j(x) > 0, all i = 1 , . . . , m has a solution x € R n . Let the 'condition number' of / be defined by: fx(f) = (
inf
£>(/, g))
yp not feasible
/i(/) = (
inf
\g feasible
if / is feasible,
J
D(f,g))
if / is not feasible. J
Let the extended size S(f) of / G Vn,d be the maximum (perhaps oo) of dm\.Vn>d and //(/). P r o b l e m 8.1 Is there a polynomial time algorithm deciding the above feasibility problem using the extended size? The problem is formalized in terms of the real machines described above, using exact arithmetic in particular. We now propose an extension of the earlier notion of real machine to allow round-off error in the computation. A round-off machine over R is a real machine, together with a function of inputs that, at each input, computation and output node, adds a state vector of magnitude less than some positive constant 6. One has no a priori knowledge of the added state vector (it's an adversary). This idealization
1583
546
S. SMALE
has the virtue of simplicity; we hope this compensates for its ignorance of important detail. A problem will be called robustly solvable if it can be solved for inputs of finite extended size by a round-off machine, no matter what the round-off error. More important is the concept of robustly solvable in polynomial time. In addition to the estimate (8.1) with extended size, S(y), one adds a require ment such as
W)fi S{yr
(a2)
One can now sharpen Problem 8.1 to ask for a decision which is robustly solvable in polynomial time. The above gives some sense of the notion of a robust or numerically stable algorithm, perhaps improving on the attempts in Isaacson and Keller (1966), Wozniakowski (1977), Smale (1990) and Shub (1993).
9. Some other directions Many aspects of complexity theory in numerical analysis have not been dealt with in this brief report. We now refer to some of these omissions. A general reference is Renegar, Shub and Smale (1997), which expands on the previous topics and those below. There is the important, well-developed field of algebraic complexity the ory, which relates very much to some of our account. I have the greatest admiration for this work, but will only mention here Bini and Pan (1994), Grigoriev (1987), and Giusti et al. (1997). Also well-developed is the area of information-based complexity. In spite of its relevance and importance to our review, I will only mention Traub, Wasilkowski and Wozniakowski (1988), where one will find a good introduc tion and survey. Another area in which the mathematical foundation and development are strong is the science of mathematical programming, or optimization. I believe that numerical analysts interested in complexity considerations can learn much from what has happened and is happening in that field. I especially like the perspective and work of Renegar (1996).
REFERENCES E. Allgower and K. Georg (1990), Numerical Continuous Methods, Springer. E. Allgower and K. Georg (1993), Continuation and path following, in Acta Numerica, Vol. 2, Cambridge University Press, pp. 1-64. O. Axelsson (1994), Iterative Solution Methods, Cambridge University Press. S. Batterson (1994), 'Convergence of the Francis shifted QR algorithm on normal matrices', Linear Algebra Appl. 207, 181-195.
1584 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
547
S. Batterson and D. Day (1992), 'Linear convergence in the shifted QR algorithm', Math. Comp. 59, 141-151. S. Batterson and J. Smillie (1989), 'The dynamics of Rayleigh quotient iteration', SIAM J. Numer. Anal. 26, 624-636. S. Batterson and J. Smillie (1990), 'Rayleigh quotient iteration for nonsymmetric matrices', Math. Comp. 55, 169-178. M. Ben-Or (1983), Lower bounds for algebraic computation trees, in 15th Annual ACM Symposium on the Theory of Computing, pp. 80-86. D. Bini and V. Pan (1987), 'Sequential and parallel complexity of approximating polynomial zeros', Computers and Mathematics (with applications) 14, 591622. D. Bini and V. Pan (1994), Polynomial and Matrix Computations, Birkhauser, Basel. L. Blum, F. Cucker, M. Shub and S. Smale (1996), 'Complexity and real computa tion: a manifesto', Int. J. Bifurcation and Chaos 6, 3-26. Referred t o as the Manifesto. L. Blum, F. Cucker, M. Shub and S. Smale (1997), Complexity and Real Computa tion, Springer. To appear. Referred t o as BCSS (1997). L. Blum, M. Shub and S. Smale (1989), 'On a theory of computation and complexity over the real numbers: JVP-completeness, recursive functions and universal machines', Bull. Amer. Math. Soc. 21, 1-46. Referred t o as BSS (1989). R. Brockett (1973), in Geometric Methods in Systems Theory, Proceedings of the NATO Advanced Study Institute (D. Mayne and R. Brockett, eds), D. Reidel, Dordrecht. W. Brownawell (1987), 'Bounds for the degrees in the Nullstellensatz', Annals of Math. 126, 577-591. G. Collins (1975), Quantifier Elimination for Real Closed Fields by Cylindrical Algeb raic Decomposition, Vol. 33 of Lect. Notes in Comp. Sci., Springer, pp. 134183. J. J. M. Cuppen (1981), 'A divide and conquer method for the symmetric tridiagonal eigenproblem', Numer. Math. 36, 177-195. J.-P. Dedieu (1997a), Approximate solutions of numerical problems, condition num ber analysis and condition number theorems, in Proceedings of the Summer Seminar on 'Mathematics of Numerical Analysis: Real Number Algorithms', AMS Lectures in Applied Mathematics (J. Renegar, M. Shub and S. Smale, eds), AMS, Providence, RI. To appear. J.-P. Dedieu (19976), Condition number analysis for sparse polynomial systems. Pre print. J.-P. Dedieu (1997c), 'Condition operators, condition numbers and condition number theorem for the generalized eigenvalue problem', Linear Algebra Appl. To appear. J.-P. Dedieu (1997d), 'Estimations for separation number of a polynomial system', J. Symbolic Computation. To appear. J. Degot and B. Beauzamy (1997), 'Differential identities', Trans. Amer. Math. Soc. To appear. B. Dejon and P. Henrici (1969), Constructive Aspects of the Fundamental Theorem of Algebra, Wiley.
1585 548
S. SMALE
J. Demmel (1987), 'On condition numbers and the distance to the nearest ill-posed problem', Numer. Math. 5 1 , 251-289. J. J. Dongarra and D. C. Sorensen (1987), 'A fully parallel algorithm for the sym metric eigenvalue problem', SIAM J. Sci. Statist. Comput. 8, 139-154. Q. Du, M. Jin, T. Y. Li and Z. Zeng (1997o), 'The quasi-Laguerre iteration', Math. Comp. To appear. Q. Du, M. Jin, T. Y. Li and Z. Zeng (19976), 'Quasi-Laguerre iteration in solving symmetric tridiagonal eigenvalue problems', SIAM J. Sci. Comput. To appear. C. Eckart and G. Young (1936), 'The approximation of one matrix by another of lower rank', Psychometrika 1, 211-218. A. Edelman (1988), 'Eigenvalues and condition numbers of random matrices', SIAM J. Matrix Anal. Appl. 9, 543-556. A. Edelman and E. Kostlan (1995), 'How many zeros of a random polynomial are real?', Bull. Amer. Math. Soc. 32, 1-38. C. F. Gauss (1973), Werke, Band X, Georg Olms, New York. M. Giusti, J. Heintz, J. E. Morais, J. Morgenstern and L. M. Pardo (1997), 'Straightline program in geometric elimination theory', Journal of Pure and Applied Algebra. To appear. G. Golub and C. van Loan (1989), Matrix Computations, Johns Hopkins University Press. D. Grigoriev (1987), in Computational complexity in polynomial algebra, Proceed ings of the International Congress Math. (Berkeley, 1986), Vol. 1, 2, AMS, Providence, RI, pp. 1452-1460. P. Henrici (1977), Applied and Computational Complex Analysis, Wiley. M. R. Hestenes and E. Stiefel (1952), 'Method of conjugate gradients for solving linear systems', J. Res. Nat. Bur. Standards 49, 409-436. M. Hirsch and S. Smale (1979), 'On algorithms for solving f(x) = 0', Comm. Pure Appl. Math. 32, 281-312. W. Hoffman and B. N. Parlett (1978), 'A new proof of global convergence for the tridiagonal QL algorithm', SIAM J. Numer. Anal. 15, 929-937. E. Isaacson and H. Keller (1966), Analysis of Numerical Methods, Wiley, New York. H. Keller (1978), Global homotopic and Newton methods, in Recent Advances in Numerical Analysis, Academic Press, pp. 73-94. R. Kellog, T. Li and J. Yorke (1976), 'A constructive proof of Brouwer fixed-point theorem and computational results', SIAM J. Numer. Anal. 13, 473-483. M. Kim (1988), 'On approximate zeros and rootfinding algorithms for a complex polynomial', Math. Comp. 5 1 , 707-719. E. Kostlan (1988), 'Complexity theory of numerical linear algebra', J. Comput. Appl. Math. 22, 219-230. E. Kostlan (1991), 'Statistical complexity of dominant eigenvector calculation', J. Complexity 7, 371-379. E. Kostlan (1993), On the distribution of the roots of random polynomials, in Prom Topology to Computation: Proceedings of the Smalefest (M. Hirsch, J. Marsden and M. Shub, eds), Springer, pp. 419-431. G. Malajovich (1994), 'On generalized Newton algorithms: quadratic convergence, path-following and error analysis', Theoret. Comput. Sci. 133, 65-84.
1586 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
549
G. Malajovich-Munoz (1993), On the complexity of path-following Newton al gorithms for solving polynomial equations with integer coefficients, PhD thesis, University of California at Berkeley. J. M. McNamee (1993), 'A bibliography on roots of polynomials', J. Comput. Appl. Math. 47(3), 391-394. J. Milnor (1964), On the Betti numbers of real varieties, in Proceedings of the Amer. Math. Soc., Vol. 15, pp. 275-280. C. Neff (1994), 'Specified precision root isolation is in N C \ J. Comput. System Sci. 48, 429-463. C. Neff and J. Reif (1996), 'An efficient algorithm for the complex roots problem', J. Complexity 12, 81-115. O. Oleinik (1951), 'Estimates of the Betti numbers of real algebraic hypersurfaces', Mat. Sbornik (N.S.) 28, 635-640. In Russian. O. Oleinik and I. Petrovski (1949), 'On the topology of real algebraic surfaces', Izv. Akad. Nauk SSSR 13, 389-402. In Russian; English translation in Transl. Amer. Math. Soc. 1, 399-417 (1962). A. Ostrowski (1958), 'On the convergence of Rayleigh quotient iteration for the computation of the characteristic roots and vectors, I', Arch. Rational Mech. Anal. 1, 233-241. V. Pan (1997), 'Solving a polynomial equation: some history and recent progress', SIAM Review. To appear. B. N. Parlett and W. Kahan (1969), 'On the convergence of a practical QR al gorithm', Inform. Process. Lett. 68, 114-118. E. A. Rakhmanov, E. B. Saff and Y. M. Zhou (1994), 'Minimal discrete energy on the sphere', Mathematical Research Letters 1, 647-662. E. A. Rakhmanov, E. B. Saff and Y. M. Zhou (1995), Electrons on the sphere, in Computational Methods and Function Theory (R. M. Ali, S. Ruscheweyh and E. B. Saff, eds), World Scientific, pp. 111-127. J. Renegar (1987a), 'On the efficiency of Newton's method in approximating all zeros of systems of complex polynomials', Math, of Oper. Research 12, 121-148. J. Renegar (19876), 'On the worst case arithmetic complexity of approximating zeros of polynomials', J. Complexity 3, 90-113. J. Renegar (1996), 'Condition numbers, the Barrier method, and the conjugate gradi ent method', SIAM J. Optim. To appear. J. Renegar, M. Shub and S. Smale, eds (1997), Proceedings of the Summer Sem inar on 'Mathematics of Numerical Analysis: Real Number Algorithm', AMS Lectures in Applied Mathematics, AMS, Providence, RI. B. Reznick (1992), Sums of Even Powers of Real Linear Forms, Vol. 463 of Memoirs of the American Mathematical Society, AMS, Providence, RI. J. R. Rice (1966), 'A theory of condition', SIAM J. Numer. Anal. 3, 287-310. L. Santalo (1976), Integral Geometry and Geometric Probability, Addison-Wesley, Reading, MA. A. Schonhage (1982), The fundamental theorem of algebra in terms of computational complexity, Technical report, Math. Institut der Universitat Tubingen. A. Schonhage (1987), Equation solving in terms of computational complexity, in Pro ceedings of the International Congress of Mathematicans, AMS, Providence, RI.
1587 550
S. SMALE
M. Shub (1993), On the work of Steve Smale on the theory of computation, in Prom Topology to Computation: Proceedings of the Smalefest(M. Hirsch, J. Marsden and M. Shub, eds), Springer, pp. 443-455. M. Shub and S. Smale (1985), 'Computational complexity: on the geometry of poly nomials and a theory of cost I', Ann. Sci. Ecole Norm. Sup. 18, 107-142. M. Shub and S. Smale (1986), 'Computational complexity: on the geometry of poly nomials and a theory of cost II', SIAM J. Comput. 15, 145-161. M. Shub and S. Smale (1993a), 'Complexity of Bezout's theorem I: geometric aspect', J. Amer. Math. Soc. 6, 459-501. Referred to as Bez I. M. Shub and S. Smale (19936), Complexity of Bezout's theorem II: volumes and probabilities, in Computational Algebraic Geometry (F. Eyssette and A. Galligo, eds), Vol. 109 of Progress in Mathematics, pp. 267-285. Referred to as Bez II. M. Shub and S. Smale (1993c), 'Complexity of Bezout's theorem III: condition num ber and packing', J. Complexity 9, 4-14. Referred to as Bez III. M. Shub and S. Smale (1994), 'Complexity of Bezout's theorem V: polynomial time', Theoret. Comput. Sci. 133, 141-164. Referred to as Bez V. M. Shub and S. Smale (1996), 'Complexity of Bezout's theorem IV: probability of success; extensions', SIAM J. Numer. Anal. 33, 128-148. Referred to as Bez IV. S. Smale (1976), 'A convergent process of price adjustment and global Newton method', J. Math. Economy 3, 107-120. S. Smale (1981), 'The fundamental theorem of algebra and complexity theory', Bull. Amer. Math. Soc. 4, 1-36. S. Smale (1985), 'On the efficiency of algorithms of analysis', Bull. Amer. Math. Soc. 13, 87-121. S. Smale (1986), Newton's method estimates from data at one point, in The Merging of Disciplines: New Directions in Pure, Applied, and Computational Math ematics (R. Ewing, K. Gross and C. Martin, eds), Springer, pp. 185-196. S. Smale (1987a), Algorithms for solving equations, in Proceedings of the Interna tional Congress of Mathematicians, AMS, Providence, RI, pp. 172-195. S. Smale (19876), 'On the topology of algorithms I', J. Complexity 3, 81-89. S. Smale (1990), 'Some remarks on the foundations of numerical analysis', SIAM Review 32, 211-220. J. Steele and A. Yao (1982), 'Lower bounds for algebraic decision trees', Journal of Algorithms 3, 1-8. E. Stein and G. Weiss (1971), Introduction to Fourier Analysis on Euclidean Spaces, Princeton University Press. R. Thorn (1965), Sur l'homologie des varietes algebriques reelles, in Differential and Combinatorial Topology (S. Cairns, ed.), Princeton University Press. J. Traub and H. Wozniakowski (1979), 'Convergence and complexity of Newton iteration for operator equations', J. Assoc. Comput. Mach. 29, 250-258. J. Traub, G. Wasilkowski and H. Wozniakowski (1988), Information-Based Complex ity, Academic Press. L. N. Trefethen (preprint), Why Gaussian elimination is stable for almost all matrices.
1588 COMPLEXITY THEORY AND NUMERICAL ANALYSIS
551
V. A. Vassiliev (1992), Complements of Discriminants of Smooth Maps: Topology and Applications, Vol. 98 of Transl. of Math. Monographs, AMS, Providence, RI. Revised 1994. X. Wang (1993), Some results relevant to Smale's reports, in Fmm Topology to Com putation: Proceedings of the Smalefest (M. Hirsch, J. Marsden and M. Shub, eds), Springer, pp. 456-465. H. Weyl (1932), The Theory of Groups and Quantum Mechanics, Dover. J. Wilkinson (1963), Rounding Errors in Algebraic Processes, Prentice-Hall. J. Wilkinson (1968), 'Global convergence of tridiagonal QR algorithm with origin shifts', Linear Algebra Appl. I, 409-420. H. Wozniakowski (1977), 'Numerical stability for solving non-linear equations', Numer. Math. 27, 373-390.
1589 JOURNAL OF COMPLEXITY 14, 454-465 (1998) ARTICLE NO. CM980486
Some Lower Bounds for the Complexity of Continuation Methods Jean-Pierre Dedieu* LAO, Universiti Paul Sabatier, 31062 Toulouse Cedex 04, France E-mail: [email protected]
and Steve Smalef Department of Mathematics, City University of Hong Kong, 83 Tat Chee Avenue, Hong Kong E-mail: [email protected] Received June 4, 1998
1. INTRODUCTION AND MAIN RESULTS In this note we consider the zero-finding problem for a homogeneous polynomial system, /:C" + , -»C"»,
with m^n,f = (fi,..., fm), ftS3fd: the space of homogeneous polynomials f,: C + 1 -»C with degree(/,) = d,. The well-determined (m = n) and underdetermined (m
Nf(x) =
x-(Df(x)\x,)-lf(x)
* This paper was completed when J.-P. Dedieu was visiting at the City University of Hong Kong in Spring 1998. 1 Corresponding author. 454 0885-064X/98 $25.00 Copyright O 1WS by Acadamic Pttm AH rifbtl of nproductioa in tny fonn nMrved.
1590
SOME LOWER BOUNDS
455
when m = n, with Df{x)\xi the restriction at x^ of the derivative of/at x. Here xL denotes the space orthogonal to x in C + 1. When m^n, we take
Nf(x)=x-(DAx)\x,yf(x), where, for any linear operator A: E-» F between two Hermitian spaces, A* denotes its Moore-Penrose inverse. A* = A*(AA*)~1 when A is onto, which is the case considered here. A Newton continuation method sequence (NCM sequence) is a sequence of pairs (/(,Qe(^)*x(C"+1)*,
0<*
(given a vector space E, E* denotes the set of nonzero vectors) satisfying the conditions
and oc(/i+i, Q < a 0 ,
with associated zero £<+i.
0
This last condition, we will make it precise later, implies that the projective Newton's sequence •*o = C/,
X
P+I = M/UI(XP),
p^O,
converges quadratically towards C<+iThe complexity of an NCM sequence (/„ {,), 0 < / < £ , is measured by k. Upper bounds for the complexity of NCM sequences have been given by Shub and Smale in their papers [7-9], about the complexity of Bezout's theorem. They give an upper bound, depending mainly on the degree D of the considered system and on the condition number of the homotopy. The case of sparse polynomial systems is studied in [2] by Dedieu; the case of homogeneous polynomial systems by Malajovich [5] and Blum, Cucker, Shub, and Smale in their book [1, Chap. 14], the case of multihomogeneous underdetermined polynomial systems by Dedieu and Shub [3] and the case of overdetermined polynomial systems by Dedieu and Shub [4]. Our main results here are two lower bounds for the complexity of a NCM sequence. In the first one we relate this complexity to the degree:
1591
DEDIEU AND SMALE
456 THEOREM
1. For any NCM sequence (f„ £,), O^i^k, k^cmax
one has
( 1, —— ) dR(C0, C*),
where c>0 is a universal constant given below and dR(C0,Ck) the Riemannian distance in P ( C + 1) between (0 and (k. The Riemannian distance in P ( C + I ) is defined by dR(u, v) = arc cos ll«ll II"II +1
for any u, ue(C" )*. Remark 1. A direct computation shows that this bound is sharp and this complexity is obtained for the family of systems defined by
where C, = (1 - 0 0 , 0,.... 0) +1(\, a „ ..., aH), aeC given. For any feJf?d\et us define Ef= {jce(C" +I )* : rank Df(x) <m). In our second theorem, we give a lower bound for the complexity of an NCM sequence in terms of the arithmetic mean of the distances of £, from Z/f. THEOREM
2. For any NCM sequence (f„ £,), O^i^k, k>c-
one has
d
^o,Ck)
where c>0 is (another) universal constant. Remark 2. This lower bound shows that the complexity of an NCM sequence increases with the proximity of singular points. This proximity is measured here by the arithmetic mean of the distances dR(Ct, £/), 1 <»<&. COROLLARY 1. Let e > 0 be given. For any NCM sequence {f„C), 0^iK:k, such that
<**(&, £>,Ke,
KU*,
1592
SOME LOWER BOUNDS
457
we have k^ce~ldR{C0,Ck)y with c as in Theorem 2. The proofs of these theorems are based on Smale's alpha-theory intro duced by Smale in [10]. We use here its homogeneous version as described in Dedieu and Shub [3]. When Df(x) is onto, we define /
II y(f x) = max 1, ||x|| max \\(Df(x)\xJ P(f,x)=\\x\\-1 <x(f,x) =
UDfix)^
Dkf(x) Wk-V\ - ^
f(x)\\,
0(f,x)y(f,x).
In the definition of y{f, x), || || is the operator norm with respect to the canonical Hermitian structure over C" +1 . These three quantities are invariant under scaling and under unitary transformations, *(/, *) = *(/, **) = MV, x) = Mfou, u~l(x)), with *e{a,/?, y}, for any xe(C+l)*, AeC*, and any unitary transfor mation u in C + 1 . When Df(x) is not onto, we take *{f,x)=p(f,x)
= y(f,x) = ao.
The following theorem [ 3, Theorem 1 ] justifies our definition of a NPC sequence. THEOREM 3. There is a universal constant a 0 > 0 with the following property, for any homogeneous system fe 3Vd and xe(C+1)*, ifa(f, x)< a0, then the projective Newton sequence,
x0 = x,
xk + j = Nf(xk),
is defined and satisfies \\xk+i-xk\\/\\xk\\^(lf-lfi{f,x) for any k^O. This sequence converges to a zero Ce(C+l)*
dM,xk)*a{\)*~1 fHJ,x)
of fond
1593
DEDIEU AND SMALE
458 with
We can take &0= 1/137.
2. PROOFS OF THEOREMS 1 AND 2 These proofs are the consequences of the three following propositions. In the first one we compute the minimum value of y(f, £)• PROPOSITION
1. We have max f 1, —^— J = min y(f, C),
where the minimum is taken over all pairs ( £ , / ) e ( C + 1 ) * x(jfd)* with AC) = 0. Proof. We first prove that (D — l)/2 is a lower bound for y(f, (). We can suppose that rank Df(Q = m, since, otherwise, y(f, £) = °o- Using the invariance properties of y(f, () under scaling and unitary transformations, we also can suppose that ( = (1, 0,..., 0). Since /(£) = 0 we have f,{z) = z$-1{aulzl
+ ■■■ +aitHzH) + gt(z),
l^i^m,
with degree(gt,z0)Kid, — 2. Let us denote by A the mxn matrix with entries aUJ. Thus, Df(Q = (0 | A). The second derivative of/ = {fu ..., fm) is given by D2M) = (d,-l)^T
^
+ D*gl(C),
where A, = (aul,...,aUH) and On is the nxn degree(g/( z0) <
(di-l)£aUJvJ y-i
zero matrix. Since
1594 SOME LOWER BOUNDS
459
so that
D2f(C)(C,v) = Diag(dt-\)Av,
0=
j \v,
Here Diag(rf, — 1) denotes the mxm diagonal matrix with diagonal enties dt—\. This gives
and, consequently, \\A* Diag(rf,- 1) A||/2 = ||(Z)/(C)|cx)t D2f(Q/2\\ ^y(f, Q. Let us consider the matrix B = A* Diag(d,— 1) A. Let A = ULV be a singular value decomposition of A: U and K are unitary mxm and « x n matrices and 27 = (zJ 10) with J = Diag(«r/), <x, > ••• ^am>0 the singular values of A. We have B=V*(*Q
\{J* Diag(rf,- 1) U{A\0) V
A
= v*(
~Xu* D i a 8(4-V U A
°\v<
so that the eigenvalues of B are rf, — 1,..., dm — \ and 0. Since \B\ ^p(B) (its spectral radius), we get ||B|| = M t D i a g ( 4 - l M | | ^ . D - l and this proves the inequality K/.O^max^l,^1 In order to prove the converse inequality, we study the example //(z) = 2o ,_,z <.
Ki'^m,
z = (z0,.... z j ,
1595 DEDIEU AND SMALE
460
and we consider ( = (1,0,..., 0), so that /(£) = 0. The derivative of f is given by DAC) =
(o,im,oH_j,
with Im the mxm identity matrix and On_m the (n-m)x(n-m) matrix. The other derivatives are given by
I Sml
Z>*/,(C)(K\ ...,«*) =
zero
»!.-<,
0
KJKk +I
for any «',..., u* 6 C" . These partial derivatives are equal to 0, except when, for some j = 1 • • • k, we have i"i= ••• = « / _ , = /,+ , = ••• = »* = 0,
ij = i.
In this case, this partial derivative is equal to {d,— 1) • •• (dt — k + 1). Thus,
D'mw
«*)=E
(rf,-i)-(rf,-*+i)«i-«r,«/«i+i-«j
and HD^CKu1,..., ii*)
I < - l
I(rf,-l)-(rf 4 -*+l) J-l 2\l/2
■«s ^(D-l)---(D-k+l) k
/ m
l«'---«>- 1 «/^ + , ---«Sl 2 )
1/2
< ( Z ) - l ) - - ( D - J t + l)Jt/2 when Hi/11|= ••• = ||t/*|| = 1, as can be proved by induction over k^2. Consequently, k
max *>2
mow D f(C) k\ f
I/(*-D
/i/fl-l\\w-i)
^.j
1596 SOME LOWER BOUNDS
461
using the fact that the sequence (j{t_i)) I/( * !)» k^2, is decreasing. This yields y{f, f)<max(l, (D —1)/2) and achieves the proof of Proposition 1.
I 2. There is a universal constant c>0 with the following property: For any fe(JtTd)*, C and x e ( C + , ) * ) if a.(f,x)^a0 with associated zero C, then PROPOSITION
dAC,x)y(f,C)
O ^ M ^ I - ^ .
This function is decreasing from 1 at u = 0 to 0 at u = 1 — -^2/2. We first start with a linear algebra lemma. Its proof may be found in [3, Lemma 2a]. LEMMA 1. Let X and Y be Hermitian spaces and A,B:X-* Y linear operators with B onto. If
\\BHB-A)\\^X<1 then A is onto and \\A
\\(Df(y)\^ m*)\*
K^rrrijf(u)
Proof Df(y) = Df(x) + Zk>2k(Dkf(x)/k\)(y-x)k-i
so that
(m*)\*)HDny)\*L-Wx)\*)I. k(Df(x)\^^T^(y-x)k-1 U *>2
K
-
1597
462
DEDIEU AND SMALE
If we take the operator norm of both sides we get k7(f,x)k-l\\y-xf
||(D/(*)Li)t(£>/(>OLi-z>/(*)l*i)ll< £ *>2
_1_
= I
^ - ' = 7 1 - ^2 - 1
£ 2 "~
d-«)
and this number is < 1 since u < 1 — v /7/2. By Lemma 1, £>/(.y)|Jci is onto and
IWj')U),iW»)UI<,_((w!.tf)_„-^. I 3. Let x and Ce(C+x)* te g/ue/i JMCA fAaf ||x|| = ||f|| = 1 and «(/, x) ^ <x0 with associated zero £. Then LEMMA
X/.CX;! ( 1 - 0 0 0 ) ^ ^( 0 0 0 ) ' w/iere <7 anrf a0 are fAe constants appearing in Theorem 3. Proof. Since/(C) = 0 we have C£<= ker £>/(£) so that {Df(Q\c±)r = Df(tf is the minimum norm right inverse of Df(Q. Thus, / II Dkf(C)\\1/lk~X)\ y(f, C) = max ^1, max |(Df(C)\ c ^-j£p| J D*/"(nil1/(*_1,\ <max l.max ( ^ ( O l ^ - f P \ *>2 II *• II / Since a(/, x)
Dkf( HII
II
II
Dkf(C)
j!
9\, •
(DAC)\^y-^\U\\(Df{o\x,yDf(x)\\\\(Df(x)\^-{^ AC I
II
Let us denote w= ||x-CII y(f, x). Since oif/^J^a,,, by Theorem 3, we have dK(C,z)^opXf,x). Moreover, since ||x|| = IICII = 1, we also have ||x-CKrfj,(C,«)»othat u < dR(C, x) y(f, x) < <7<x(/, x) <
1598 SOME LOWER BOUNDS
463
Thus, by Lemma 2
\\(Df(0\x^Df(x)\\<
d-«)2
We also have Dkf(H)\\ AC.
||
Dk+'f(x)
II ; > 0
||
\\x-cw
K. / .
(* + ')!.,,„,*+,_,„.. /J>0
fcI/I
,,,/
r(/.*)*-'
J^^'-Mlx-Cl^j-j-,^,.
Thus 2
I
A
V *" y(/,C)^max l.max (l-«) y(/,x)**>2V *(«) ( l - « ) t + 1 and, using the inequality u ^ <7<x0, we are done.
I)
|
Proof of Proposition 2. Since <x(/, x) ^ a 0 , by Theorem 3, we have dn(C, x) < <x/?(/, x). Using Lemma 3 we obtain <xoc
>,*) 0 <**(£*) rt/.CXfffl/,*) (1 - <xa (1 -
and we are done.
|
Our last ingredient is a corollary of the following theorem (gammatheorem for homogeneous polynomial systems); see [3, Theorem 2] and [1, Chap. 14, Theorem 1] for the case m = n. THEOREM 4. There is a universal constant y0 with the following property: letCe{Cn+1)* be a zero offe(jPd)* and *e(C" + 1 )*. If
ll*-CM/,0/IICKyo then the projective Newton sequence, x0 = x,
xk+i = Nf(xk),
is defined and converges to a zero C e ( C + 1 ) * off and
d*(C.xk)*o(i)*-lfi(f,x).
1599
464
DEDIEU AND SMALE
3. For any Ce(C" /(£) = 0 we have PROPOSITION
+I
)* and any fe(Jt?d)* such that
d«(C,Z/)y(f,C)>yoProof If xe£f then the projective Newton sequence x0 = x, k + i = N/(Xk) is not defined. By Theorem 4 we have necessarily
x
ll*-Clly(/C)/IICII>yo. When <*, C>=0 then dR(C, x) = n/2 and the conclusion holds. When <x, C> #0, scaling x such that <x — C, *> = 0, we obtain *(*. C)> II*-CII/IICII
and we are done.
|
Proofs of Theorems 1 and 2. Given an NCM sequence (/„ (,), 0 < i < k, since a(/ / + 1 , C/X«o ^vith associated zero C/+i> we get by Proposition 2 <**(t/-n. C/) ?(//+1, C/+iX c. By Proposition 1 we obtain max ( l . ^ ) « k ( C H . l , C , X c , so that / D-\\ / £>-l\*_1 maxf 1,——jdR(Co, C*)<max( 1,—— 1 £ *(£,+1, £/)<<* and this proves Theorem 1. As previously we have dACt+u tt) y(ft+u Ct+i) ^c and by Proposition 3 we have
1600 SOME LOWER BOUNDS
465
so that
/-o this proves Theorem 2.
\* <-o |
REFERENCES 1. Blum, L., Cucker, F., Shub, M., and Smale, S. (1997), "Complexity and Real Computa tion," Springer-Verlag, New York/Berlin. 2. Dedieu, J. P. (1997), Condition number analysis for sparse polynomial systems, in "Foun dations of Computational Mathematics" (F. Cucker and M. Shub, Eds.), Springer-Verlag, New York. 3. Dedieu, J. P., and Shub, M. (1997), Multihomogeneous Newton's method, Math, of Comp., to appear. 4. Dedieu, J. P., and Shub, M. (1998), Newton and predictor-corrector methods for overdetermined systems of equations, preprint 5. Malajovich, G. (1994), On generalized Newton algorithms, Theor. Comput. Sci. 133, 65-84. 6. Shub, M. (1993), Some remarks on Bezout's theorem and complexity, in "Proceedings of the Smalefest" (M. V. Hirsch, J. E. Marsden, and M. Shub, Eds.), pp. 443-455, SpringerVerlag, New York. 7. Shub, M., and Smale, S. (1993), Complexity of Beout's theorem. I. Geometric aspects, J. Am. Math. Soc. 6, 459-501. 8. Shub, M., and Smale, S. (1996), Complexity of Bezout's theorem. IV. Probability of success, extensions, SIAM J. Numer. Anal. 33, 128-148. 9. Shub, M., and Smale, S. (1994), Complexity of Bezout's theorem. V. Polynomial time, Theor. Comput. Sci 133, 141-164. 10. Smale, S. (1986), Newton's method estimates from data at one point, in "The Merging of Disciplines: New Directions in Pure, Applied and Computational Mathematics" (R. Ewing, K. Gross, and C. Martin, Eds.), Springer-Verlag, New York/Berlin.
1601
J. Symbolic Computation (1999) 27, 21-29 Article No. jsco. 1998.0242 Available online at http://www.idealibrary.com on IIEj^l
@
A Polynomial Time Algorithm for Diophantine Equations in One Variable FELIPE CUCKERtl, PASCAL KOIRAN*U AND STEVE SMALEt§ ^Department of Mathematics, City University of Hong Kong, 83 Tat Chee Avenue, Kowbon, Hong Kong *LIP, Ecole Normale Supirieure de Lyon, 46, allee d'ltalie, 69364 Lyon Cedex 07, France
We exhibit an algorithm computing, for a polynomial / € Z [t], the set of its integer roots. The running time of the algorithm is polynomial in the size of the sparse encoding of / . © 1999 Academic Press
1. Introduction The goal of this paper is to prove the following. THEOREM 1. There is a polynomial time algorithm which given input / £ Z[(] decides whether f has an integer root and, moreover, the algorithm outputs the set of integer roots of f. Here we are using sparse representation of polynomials and the classical (i.e. Turing) model of computation and complexity. That is, for / £ Z[t], f = adtd H
h ait + a 0 ,
we encode / by the list of pairs {(i, at) | 0 < i < d and Oi ^ 0}. The size of the sparse representation of / is denned by size(/) = ^
( ht (°*) + **(»))
where ht(a) = log(l + \a\) is the (logarithmic) height of an integer a € Z. Thus, size(/) is roughly the number of bits needed to write down the list representing / . Polynomial time means that the number of bit operations to output the answer is bounded by c(size(/))rf for positive constants c, d. Note that the degree of / is at most 2?ize(f) and this exponential dependence is sharp in the sense that there is no q G N such that the degree of / is bounded by (size(/)) 9 for 1
E-mail: macuckertmath.cityu.edu.hk "E-mail: Pascal.KoiranCens-lyon.fr ^E-mail: masmaleCmath.cityu.edu.hk
0747-7171/99/010021 + 09
$30.00/0
© 1999 Academic Press
1602 22
F. Cucker et al.
all / . In particular, evaluating / at a given integer x may be an expensive task since the size of f(x) may be exponentially large as a function of size(/) and size(x). Algorithms for sparsely encoded polynomials (or just sparse polynomials as they are usually called) are usually much less efficient than for the standard (dense) representation in which / is represented by the list {ao,ai,...,a<j}. This is due to the fact that some polynomials of high degree can be represented in a very compact way. For dense polynomials, the existence of a real root can be decided efficiently (by Sturm's algorithm). It seems to be an open problem whether this can also be done in polynomial time with the sparse representation. Theorem 1 states that the existence of an integer root for sparse polynomials can be decided in polynomial time. In fact, all integer roots can be computed within that time bound. Our algorithm relies in particular on an efficient procedure for evaluating the sign of / at a given integer x. The (efficient) sign evaluation problem seems to be open for rational values of x. We note here that a version of Theorem 1 is well-known for dense polynomials. For a general overview on computer algebra for one-variable polynomials see Akritas (1989) and Mignotte (1992). 2. Computing signs of sparse polynomials The main result of this section is the proof that one can evaluate the sign of a polyno mial / at x G Z in polynomial time. That is, given / G Z[t] and x G Z, we can compute the quantity
if f(x) < 0
f -1 sign(/(x)) = { 0 [ 1 in time polynomial in size(x) and size(/).
if /(*) = 0 if f(x) > 0
2. There exists an algorithm which given input x G Z and f € Z,[t] computes the sign of f(x). The halting time of this algorithm is bounded by a polynomial in size(x) and size(/).
THEOREM
Recall that a straight-line program with one variable is a sequence V = {ci,... ,Ck, t, u\,... ,ue} where c\,...,Ck € Z, and for i < £, Uj = a*b with * G {+, —, x} and a,b two elements in the sequence preceding rij. Clearly, ue may be considered as a polynomial f(t); we say that V computes f(t). For every polynomial f(t) there exist straight-line programs computing f(t). Thus, straightline programs are regarded as yet another way to encode polynomials which turns out to be even more compact than the sparse encoding. We define the size of V to be k
size(P) = I + ^2 size(c<). t=i
1. Let V be a straight-line program in one variable of size s computing f(t) and x G Z such that \f{x)\ < T for some T > 0. Then f(x) can be computed in time polynomially bounded in s and size(T).
LEMMA
PROOF. One performs the arithmetic operations (there are at most s of them) in the
1603 Polynomial Time Algorithm
23
ring Z2T of integers modulo 2T. Each operation in this ring is done with a number of bit steps polynomial in size(T). The result, /(x), is the value of f(x) modulo 2T and therefore, by hypothesis, the value of f(x) if f(x) > 0 and the value 2T + f(x) if f{x) < 0. Subtracting 2T from /(x) if f(x) > T we get f(x). D LEMMA 2. There is an algorithm which given input (x,a) € Z 2 , x > 0, a > 0 outputs £ 6 Z, £ > 0, suc/i tfiar. 2 / _ 1 < xQ < 2 ' + 1 . 77ie /laittns time is bounded by a polynomial in size(x) andsize(a). PROOF. We want to compute £ € Z, ^ > 0 satisfying £ - 1 < a logx < £ + 1. To do so, it is enough to compute an approximation y of logx such that \y — logx| < l/(2a) since in this case - « < cty - alogx < and we may take £ to be the closest integer to ay. Working in base 2, \y - logx| < l/(2a) is satisfied if y is computed with [log a + log log x] + 1 bits of precision. Here, for a real number z, \z~\ denotes the smallest integer greater than or equal to z. Define n = 2[loglogx + loga]. By Theorem 6.1 of Brent (1976) we can compute the first n bits of logx in time 0(M(n)logn) where M(n) is the time required to multiply two positive integers of height at most n. This finishes the proof. D P R O O F OF THEOREM 2. We can assume that x > 0 since if x < 0 then f(x) is g(~x) where g is obtained from / by changing the sign of the coefficients of the monomials with odd degree. Also, if x = 0 the problem can be solved by looking at the constant term of / . Thus, suppose x > 0. Let Jfc be the number of monomials of / so that
/ = ait01 +■■■+ aktPk
with fa > fc > • • • > & > 0.
Then, f(x) can be evaluated using Homer's rule as follows. Let a* = /3k and Qj = Pj - /3J+i for j = 1,...., k - 1. Then, /3j = ct, + aj+i + • • • + ak for j = 1 , . . . , fc. Now we inductively define po = 0 and Si = Pi-i + a<
and
pi = s<xai
for i = 1 , . . . , k. We then have pk = /(x). The precise evaluation of /(x) using the sequence of operations given by Homer's rule is not achieved in polynomial time since the intermediate results can be too large. Instead, we will inductively compute a sequence of rough approximations of s^ and pi, with the right sign and of small (i.e. polynomially bounded) size. More precisely, we will produce a sequence of pairs (m,i,Mi) € N2 and (t>j, Vj) € N 2 and a sequence of integers Oi, with i = l,...,k with the following properties. For i = l,...,fc, <7j 6 {-1,0,1} and m
<,2M<] M m Pi € [ - 2 - , - 2 « ] Pi = 0 Pie[2
i{(n = l if tTi = - 1 if ai = 0.
(1)
1604 24
F. Cucker et at.
Moreover, 0 < Mi - m, < 3i.
(2)
Note that, since mi < log |p<|, we can write m< with a number of bits which is polyno mial in S = max{size(x),size(/)}. The same holds for M< since Af< < m< -I- 3i. The same properties hold for Si and (v<, VJ). That is, for i = 1 , . . . , k,
{
Sie[2v',2v'}
ifa< = l
^€[-2^,-2"-] Si = 0
if
(3)
and 0 < Vi - Vj < 3i - 2.
(4)
The general appearance of the algorithm is the following. For input (x,f), compute a i , . . . , a * as above and let OQ = 0. Then, inductively, for i = 1 , . . . , k (a) compute vi7 Vi andCTJfrom nij_i,M<_i and o^-i (b) compute rrii and Mi from Vj, Vi and a*. Output
It is immediate to check that, if mi_i,Mi_x and CTJ_I satisfy conditions (1) and (2), then Vi,Vi and a, satisfy conditions (3) and (4). All lines in the above algorithm are executed in polynomial time. This is immediate except for the computation of the exact value of p^ But the algorithm in Lemma 1 has a halting time bounded by a polynomial in size(P) and size(T) for any V computing p<(i). In our case one can take any straightline program computing pi of size polynomial in the size of / (Homer's rule as described above provides one with 2% — 1 operations) and we note that the size of T, is about Afj_i,
1605 Polynomial Time Algorithm
25
and Af4_i < m^ + 3(» - 1) < Iog(2|a{|) + 3(t - 1) which is also polynomial in size(/). For (b), we proceed as follows. Compute £ such that 2 < _ 1 < xa< < 2* +1 as in Lemma 2. If ai jt 0 then let mi = Vi + £ - 1 and Mi = Vi + E + I. Notice that in (a) we do not use the values of m<_i and Mj_! if ot = 0. Consequently, we do not compute them in (b) if this is the case. □ REMARK 1. It is an open problem whether one can compute the sign of f(x) in poly nomial time if / is given as a straight-line program. This is so even allowing the use of randomization, in which case the state of the art is an algorithm for deciding whether fix) = 0 in randomized (one-side error) polynomial time (see Schwartz (1980)). This algorithm, however, does not tell, in case f(x) ^ 0, whether f(x) > 0 or f(x) < 0.
3. Proof of Theorem 1 First we give a preliminary lemma. It is a well known result (cf. Mignotte (1992, Ch. 5.3)) but we prove it here for sake of completeness. In the following we count roots without multiplicity, that is, the expression "fc roots" means k different roots. LEMMA
3. Let f e R[t] have k monomials. Then f has at most 2k real roots.
PROOF. If k = 1 the statement is true. If k > 1 write / = xap with p(0) ^ 0. Then p', the derivative of p, has k - 1 monomials and, by induction hypothesis, at most 2(fc - 1) roots. From this we deduce, by Rolle's theorem, that p has at most 2k - 1 real roots and hence / has at most 2k. □ DEFINITION 1. Let p G Z[t] and M G Z, M > 0. Let C = {[ui,i>t]}t=i,...,/v be a list of closed intervals with integer endpoints satisfying tij < 14+1 and Vi = Ui or v< = Uj + 1 for all i. We say that C locates the roots ofp in [-M, M\ if for each root C of p in [-Af, M] there is i < N such that C € [tti,t>i]. Note that in this case p has no roots in (vi,Ui+i) for all i.
Let g G Z[<] and M G Z, M > 0. Write g = tap with p(0) ^ 0 and suppose that C = {[ut,Ui]}i=i,...,Ar locates the roots of p1 in [-M,M\. Then, for each i < TV, p has at most one root in the interval (vi,Ui+i) since, by Rolle's theorem, if p has two roots in (vi, tij + i) the p1 must have a root in this interval as well. Moreover, p has a root in this interval iff p(v<)p(u<+i)<0. This is so since if p(vi)p(ui+i) > 0 and p has some root in (ui,u <+ i) then either p has (at least) two roots in [^,14+1] or it has a double root in (u<, Uj+i). In both cases p' has a root in (vi, Uj+i) contradicting the choice of C. PROPOSITION 1. There is an algorithm which, given input g,p G Z[t], M,N and C as above computes a list C locating the roots ofp in [-M, M]. The list C has at most N + 2k
1606 26
F. Cucker et al.
intervals where k is the number of monomials of g. The halting time of the algorithm is polynomially bounded in size(M), size(p) and N. P R O O F . Using ...,UN,VN,M.
the algorithm of Theorem 2 compute the sign of p at the points —M, u\, vi,
Let [x,y] be any of the N + 1 intervals [—M,u\], [vi,ti2],.. •, [VN-I,UN], [ V ^ , M ] . If p(x)p(y) > 0 we know that there are no real roots of p in [x,y]. Otherwise, there is only one root which can be located in an interval of the form [u, u+1] by applying the classical bisection algorithm with integer mid-points (the interval has the form [u, u] if we find a mid-point u such that p(u) = 0). We form C by adding to C these intervals. Since the total number of roots of p is bounded by 2fc it follows that the number of intervals in C is at most N + 2k. The bound for the halting time is proved as follows. Bisection is applied to N + 1 intervals at most. Each of these intervals has length at most 1M and therefore, the number of sign evaluations is of the order of logM, that is, it is linear in size(M). Finally, all the sign evaluations (the 2(N +1) first ones and the ones performed during the bisection process) are done in polynomial time in size(M) and s\ze(g) by Theorem 2. □
P R O O F OF THEOREM 1.
Let
/ = a j ^ 1 + • • • + aktPk with (i\ > fa > ■ ■ ■ > /3fc > 0. Then, we can define polynomials p< inductively by / = Vkp\ p'j = tr",-1p2
p\ (0) ^ 0 and p\ has k monomials P2(0) ^ 0 and pi has k — 1 monomials
j/ fc _i = t^Pk
pk e Z, pk ± 0
where j k = Pk and 7 1 , . . . , 7k_i only depend on /?i,..., fa it L is a. bound for the absolute value of the coefficients of / , the coefficients of pj are bounded by L(5{~v for j = 1 , . . . , k. Therefore, since pj has exactly k — j + 1 coefficients, we deduce that size(pj) < (k - j + l)(j - 1) size(/?i) -I- size(/) which is bounded by 2(size(/)) 3 for all j = 1 , . . . , k. Now we note that if £ is an integer root of / , then either £ = 0 or £ divides a/t. To prove this, suppose that /(C) = 0 and £ ^ 0. Then we have 0l £/3i-A.
+ ...+ ak^0k-i-0k
= -a f c .
Since £ divides the left-hand side, it must divide at. Thus, all integer roots of / are in the interval [—\ak\, |a/t|] and we can restrict our search to this interval. Consider the algorithm
1607 Polynomial Time Algorithm
27
input / Compute p i , . . . ,Pfc. LetCfc = [0,0]. For i = k — 1 , . . . , 1, inductively compute d locating the roots of p< in [— |afc|, |ajfc|] using Proposition 1 with input C<+i. Let S = 0. For each endpoint x of an interval in C\, if f(x) = 0 then let S = S U {x}. Output S The list Ck isolates the roots of pk- Then, by k - 1 applications of Proposition 1, the list C\ isolates the roots of pi and since it contains the interval [0,0], the roots of / . This ensures the correctness of the algorithm. The polynomial bound for the halting time follows from Proposition 1 plus the fact that size(pj) < 2(size(/)) 3 for all j = 1 , . . . , k. Notice that pi + i is computed from pt by first computing the derivative pj — which is done with 2(fc — i) arithmetic operations — and then dividing by a power of t — which is done with k - i arithmetic operations. Thus, the sequence p i , . . . ,Pk can be computed with 0{k2) arithmetic operations. Since all the operand have polynomial size in size(/), the sequence is computed in polynomial time. □ 4. A Refinement a
De a n m t e
er
Let / = ^r=o »*°'' g polynomial with ao < ot\ < ■ ■ ■ < a„ and all a^'s nonzero. Given k € { 1 , . . . ,n - 1}, one can write uniquely / as / = r* + xak+lqk where »> and qk are integer polynomials, and deg(7>) = a* (of course, r* = £)t=o a«*a" axi^ qk = Xir=fc+i aitai~ak). With this notation, we have the following simple known fact. 2. Let Mk = s u p o ^ ^ |OJ|. Ifx is an integer root of f and \x\ > 2, x must also be a root of qk and rk provided that ajt+i — a* > 1 + log Mk-
PROPOSITION
PROOF.
Since x is a root of / , |rfc(x)| = \qk(x)\ ■ \x\ak+1. Moreover, |r fc (x)| < Mfc(l + M + • • ■ + | x n = Mfc'
.
From these two relations we obtain \qk(x)\ ■ \x\a^-a"
< M fc |x|/(|x| - 1) < 2Mk
since |x| > 2. Finally, qk(x) ^ 0 implies (ak+i -a f e )log|x| < l + logMfc since \qk(x)\ > 1 in this case. This is in contradiction with the hypothesis ak+i - a* > 1 + logMj.. We conclude that qk{x) = 0, and rjt(x) = 0 follows immediately. □ This proposition applies in particular to polynomials that have a small number of terms compared to their degree (of course these are precisely the polynomials for which the sparse representation is interesting). Specifically, if / is a polynomial of degree d = an with a nonzero constant coefficient (i.e. ao / 0) and M = Mn = s u p ^ ^ , , \ai\, there must
1608 28
F. Cucker et al.
exist a gap of at least d/n between two consecutive powers of / . Therefore one can always apply this proposition when & > 1 + log M. In any case, if the proposition applies we can first compute the integer roots of r* (or qic) and then check whether any of these roots is also a root of q^ (or fk). This can sometimes speed up the algorithm described in the previous sections, in particular when either qk or r* is of small size compared to / . For instance, if / is of the form f(x) = x2 — 3 + x5q(x), only —1 and 1 can possibly be integer roots of / . And if / is of the form f(x) = x2 — 9 + x7q(x), all integer roots are in {—3, —1,1,3}.
5. Final Remarks Natural extensions of Theorem 1 would consider the existence of rational or real roots of / . For rational roots, the arguments in Section 3 can be extended. If a rational p/q is a root of / then p divides the constant term and q divides the leading coefficient. Thus, the number of possible roots is again exponential in size(/) and the bisection method applies. However, it is an open question whether one can compute the sign of f(p/q) in polynomial time. For real roots the situation seems even more difficult since bisection only may not detect multiple roots. In another direction, one could consider diophantine equations in several variables. For sparse polynomials in several variables, sign determination seems to be a difficult question, and it is not clear whether Theorem 2 can be generalized. Actually, right now it is not known whether any algorithm exists to decide diophantine equations in two variables. Recall that the (logarithmic) height of an integer x is defined by ht(x) = log(l + |x|). Let / G Z [ t i , . . . , tn], f = ^2aeAaQta with A a finite subset of N n , aa ^ 0 for a G A, and ta = t"1 ■ • -t£" if a = ( a i , . . . , a n ) . The sparse representation of / is the sequence of pairs (a,aa), and the size of / for this representation is defined by size(/) = ^ ( h t ( a ) + ht(a a )) where ht(a) = ht(ai) H h ht(a n ). It is well known that / can be evaluated at a point x G Z n in time polynomial in size(/) and size(i) if / is considered with the dense representation. 1. Given / G Z [ t i , . . . , t„\ and x G Z n , is it possible to compute sign(/(x)) in polynomial time in size(x) and size(/) for the sparse representation of / ?
PROBLEM
Theorem 2 solves this problem for the case n = 1. For any fixed n, Shub (1993) solves it using Baker's (1975) theorem in case / has only two monomials (but the halting time depends exponentially in n). Moreover he poses a question akin to Problem 1. Worse, the problem of deciding feasibility of diophantine equations in many variables is well-known to be undecidable (cf. Matiyasevich (1993)). Thus we consider the twovariable case. Since this problem looks much harder than in one variable, we would be happy with a single exponential algorithm for dense polynomials. If / e Z[t\,..., t„] has degree d G N, the dense representation of / is the sequence of coefficients {aa} for all a G N n with |a| = a i H \-an
1609 Polynomial Time Algorithm
29
in N n . Then, the size of the dense representation of / is size(/) = ^^ size(o a ). \a\
Here size(a) = ht(a) if a ^ 0 and size(O) = 1. We propose the following conjecture. CONJECTURE 1. The feasibility of any diophantine equation P(x,y) = 0 can be decided in time 2C" where C is a universal constant and s is the size of P for the dense repre sentation. This would follow from certain height estimates. Height bounds are a topic of current interest in number theory, but there are more conjectures than theorems. For instance, the Lang-Stark conjecture (Lang, 1991) proposes the upper bound |x| < Cmax(|a| 3 ,|6| 2 )* (C and A; are universal constants) on the height of all solutions of equations of the form y2 = x 3 + ax + b with 4a 3 + 2762 ^ 0. Here we only need a bound on the smallest height of a solution, though. References Akritas, A. (1989). Elements of Computer Algebra with Applications. New York, John Wiley & Sons. Baker, A. (1975). Transcendental Number Theory. Cambridge, Cambridge University Press. Brent, R. (1976). Fast multiple-precision evaluation of elementary functions. J. ACM, 23, 242-251. Lang, S. (1991). Number Theory III, Volume 60 of Encyclopaedia of Mathematical Sciences. Berlin, Springer. Matiyasevich, Y. (1993). Hilbert's Tenth Problem. Cambridge, MA, The MIT Press. Mignotte, M. (1992). Mathematics for Computer Algebra. Berlin, Springer. Schwartz, J. (1980). Fast probabilistic algorithms for verification of polynomial identities. J. ACM, 27, 701-717. Shub, M. (1993). Some remarks on Bezout's theorem and complexity theory. In Hirsch, M., Marsden, J. and Shub, M., eds. Prom Topology to Computation: Proceedings of the Smalefest, pp. 443-455. Berlin, Springer. Originally Received 30 October 1997 Accepted 30 July 1998
1610
Complexity Estimates Depending on Condition and Round-off Error Felipe Cucker and Steve Smale Department of Mathematics City University of Hong Kong 83 Tat Chee Avenue. Kowloon HONG KONG e-mail: {macucker ,masmale}Omath. cityu. edu. hk
This paper has two agendas. One is to develop the foundations of round-off in computation. The other is to describe an algorithm for deciding feasibility for polynomial systems of equations and inequalities together with its complexity analysis and its round-off properties. Each role reinforces the other. Categories and Subject Descriptors: F.2.1 (Analysis of Algorithms and Problem Complex ity]: Numerical Algorithms and Problems; G 5 [Numerical Analysis]: Roots of Nonlinear Equations General Terms: Algorithms, Theory Additional Key Words and Phrases: Feasibility of Polynomial Systems, Error Analysis, Condi tioning, Iterative Methods
1. INTRODUCTION Solving systems of polynomial equations by computer has become a principal task in many areas of science. In most of this usage of the machine, there is a loss of precision at many steps of the process. For example a round-off error occurs when two numbers with 10 digit precision are multiplied to obtain a new number with 10 digit precision. The main goal of this paper is to give a theoretical foundation for solving systems of equations (with inequality constraints permitted as well) which takes into account this loss of precision. We give an algorithm with reasonable complexity estimates for solving general real polynomial systems of equalities and inequalities and which succeeds in the presence of round-off error, provided the This work was supported by CERG grants 9040393 and 9040189. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or direct commercial advantage and that copies show this notice on the first page or initial screen of a display along with the full citation. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, to republish, to post on servers, to redistribute to lists, or to use any component of this work in other works, requires prior specific permission and/or a fee. Permissions may be requested from Publications Dept, ACM Inc., 1515 Broadway, New York, NY 10036 USA, fax +1 (212) 869-0481, or p»raiision««acm.org.
1611
2
•
F. Cucker and S. Smale
precision of the computation satisfies a bound polynomial in some parameters (to be described). Moreover we strive to give some foundations for the round-off error in the setting of real number machines as in [Blum et al. 1998]. Modifying these "real Turing machines" to include round-off error helps to make a more compelling case for their use in the foundations of numerical analysis, developing a suggestion in [Smale 1997]. Our Main Theorem states that the feasibility of any system of equalities and inequalities in n variables can be decided in time polynomial in the condition num ber and the number of input variables and singly exponential in n. Moreover, a polynomial bound is given for the required precision. We prove the following. Main Theorem (Theorem 9 below) follows
(
Let
fi{a,x) = 0 t = l , . . . , m 9j{a,x) > 0 j = l , . . . , r hk(a,x)>0 k = l,...,q
where fc, gj, hk are polynomials in a i , . . . , a/, x\,..., xn with integer coefficients. There is a machine M over R which decides, on input (tp,a), a £ R ' , if there is a point x € R " such that (a,x) satisfies
with c a universal constant. The number of processors is bounded by the expression for the sequential time above. The polynomial bound for the precision remains the same. The concept of condition number has played a large role in numerical analysis. For the problem of solving linear systems of equations one can easily prove that the relative error in the solution can be bounded by the relative error in the input times a quantity K(A) defined by K(A) = \\A\\\\A->\\.
1612 Complexity Estimates depending on Condition and Round-off Error
•
3
Here A is the matrix of the system and the norm is the operator norm. In other words, if Ax = y and A(x + 5) — (y + S) then
iivii-
(
Vir
This K(A) is called the condition number of A. The condition number (i*(<pa) of the Main Theorem is essentially a far reaching generalization of this K(A). Its development starts in Section 3 below. The analysis of round-off and conditioning goes back to the work of von Neumann and Goldstine [1947, 1951], and Turing [1948], where the expression "condition number" was coined.1 A further important development of condition number was done by Wilkinson [1963]. Concerning a priori bounds for the condition number, we take the view (not uncommon) that computing the condition number may be as difficult as computing the solution. So the situation suggests a study of probability distributions on the condition number. Developing this would carry us too far afield, but results of this kind are e.g. in [Blum et al. 1998]. Algorithms for deciding the feasibility of systems of polynomial equations and inequalities have existed since the work of Tarski [1951]. In more recent times new algorithms were devised to improve Tarski's since the latter has hyperexponential complexity. In the seventies, Collins [1975] and Wuthrich [1976] independently de vised algorithms whose complexity has a doubly exponential dependence on n. A breakthrough was made later by Grigoriev and Vorobjov [1988]«with the introduc tion of an algorithm working in single exponential time. This algorithm is sequential and only works for systems of polynomials with rational coefficients. The cost mea sure considered is the bit cost. Subsequent developments appeared in [Heintz et al. 1990; Renegar 1992; Basu et al. 1994]. These articles provide algorithms which work also with arbitrary real numbers at unit cost and can be efficiently parallelized. All these algorithms assume that the computations are exact, i.e. no round off is produced during the computation. The ones working in single exponential time have a fairly complicated description and correctness proofs. In contrast, the algorithm we describe here is quite simple to describe and works under round-off computations for all well-posed inputs. A key point is that round-off error accumulates mainly in the linear algebra stage for both traditional algorithms and ours. But in the former case, the dimensions of the matrices are exponentially larger than in ours. For example, for the feasibility of / : R n -► R n with deg(/j) = d, we have a n n x n matrix to deal with while Renegar's matrices have order
1613
4
F. dicker and S. Smale
machines process real numbers and this is their main mode. The general feasibility problem for a sparse set of polynomial equalities and inequalities in Euclidean space taking round-off into account is quite subtle. Thus the basic algorithm is first studied in the simpler case of systems of homogeneous polynomial equations, postponing its round-off analysis until Part II. Even this case generalizes greatly the real projective Hilbert Nullstellensatz. Our algorithm is based on Newton's method and presented for the first time in this form. After the completion in Section 8 of the complexity analysis of this case, the results are extended to allow inequalities and deal with affine systems (both introduce new difficulties) in Sections 10 to 12. The round-off environment and the round-off analysis are established in Part II. TABLE OF CONTENTS
1. Introduction PART I. 2. Feasibility of homogeneous sparse polynomial systems 3. Condition numbers 4. Deciding homogeneous sparse systems of equalities 5. Bounds on the invariants 6. Newton's method and ^proximate zeros 7. Point estimates for sp "se systems 8. Proof of Theorem 1 9. Proof of Theorem 2 10. Feasibility of affine semi-algebraic systems 11. Deciding basic semi-0 jebraic systems 12. Proof of Theorem 4 PART II 13. Round-off machines 14. Sets of sparse systems and uniform complexity 15. Round-off algorithms and Linear Algebra 16. Sparse systems and round-off algorithms 17. Proof of the Main Theorem 18. Round-off algorithms for problems in N P R
PARTI 2. FEASIBILITY OF HOMOGENEOUS SPARSE POLYNOMIAL SYSTEMS
We discuss the homogeneous algebraic case first and subsequently show how this can be used in the affine semi-algebraic case in the Main Theorem above. Let N > 0 and HN be the Euclidean space of dimension N. The norm || Ife in N R , which is the norm associated to the dot product in R " , is defined by N \
«=1
1614 Complexity Estimates depending on Condition and Round-off Error
'
5
Other norms on R N include N
IMIi = 5Z1**1
and
INloo = max |x<|.
In this paper we shall consider the above three norms. For points x in Euclidean spaces, if no confusion is possible, we shall denote ||x||2 simply by ||x||. Denote by B(RN) the unit box in R w , that is B(R") = { i e R
w
|
max |x<| = 1}.
Recall that a function / : R ' + 1 x R n + 1 -> R is bihomogeneous of degrees c and d when f(Xa,fix) = Xcfidf(a,x) for any (o,x) G R* + 1 x R n + 1 and A,/i 6 R . DEFINITION 1. [H(c,d)] Let / i , . . . , / m € 7Z[ao,...,ai,xo,---,x„] be bihomo geneous polynomials of degrees C\,..., Cm in the a's and of degrees d\,..., dm in the x's respectively. Let / = ( / i , . . , / m ) - For a given a G R* + 1 con sider these polynomials as polynomials in the variables x 0 , . . •, xn with coeffi cients in 7L[ao,... , a<] C R and denote this system by fa : R n + 1 -> R m so that fa(x) = f(a,x). We say that a pair (/, a) is feasible (or that /„ is feasible) if there is a point x G R " + 1 , x ^ 0, such that fa(x) = 0. We call / = ( / i , . . . , fm) a sparse polynomial system. Given c = ( c i , . . . , Cm) € N m and d = ( d i . . . . , d m ) G l^"1 we shall denote the set of all sparse systems by H(c.d)REMARK 1. Let D = max{di dm}. In the rest of this paper, for simplicity of statements, we will assume that D > 2. EXAMPLE 1. The dense case may be thought of as a special case. Consider / = (/i, • • • i fm) with each ft of the form
/< = 5Z a°xa|a|=d.
Here a denotes a multiindex ( a n , . . . , a n ) G W + 1 with |a| = ao + . . . + a n = d, and xa = XQ° ■ • ■ x£ n . In this case the number of a variables is
Moreover, for each i = l , . . . , m , Cj = l, that is the polynomials /< are linear in the a's. To decide whether such a system / has a non-zero real root can be seen as a decision version over the reals of the projective Hilbert Nullstellensatz. EXAMPLE
2. Consider m = l, I = 3,n = l, and polynomials / of degree d with
the form / ( a , x ) = aoXi + o i x f - 1 x 0 + o 2 xiXo - 1 + o 3 roThis is a simple example of sparsity in which some monomials are not allowed. Again, / is linear in 00,01,02,03.
1615
6
*
F. Cucker and S. Smale
In the two examples above, / is linear in the variables Oj. But of course, this needs not be so. For some time to come we will fix / € 'H(cd) (&&d thus, we fix the parameters c, d, £, m and n) and consider the problem: Decide on input a € B ( R / + 1 ) if /„ is feasible. Eventually / will be allowed to vary. Related with this feasibility problem is another problem. Say /„ is e-quasifeasible if there is an a' € B(Re+1) such that \\a - a'\\ < t and /„< is feasible. This notion is related to "backward error analysis" of numerical analysis. Thus, /„ is quasifeasible if there exists a nearby feasible problem. We shall also study the problem of deciding, given a and e > 0, whether /„ is e-quasifeasible. 3. CONDITION NUMBERS Recall we fixed / G 'H(cd) ■ The goal of this section is to introduce three invariants and Kf. They are important ingredients for describing our feasibility algorithms. Since we want to consider underdetermined systems besides the well-determined systems, generalized inverses to surjective linear maps will appear naturally. We now recall the definition and main properties of one such generalized inverse. M(/O), LJ
DEFINITION 2. Let A : V -> W be a surjective linear map of finite dimensional real vector spaces with inner products. The Moore-Penrose inverse A^ : W —► V of A is defined to be the inverse of A restricted to (ker >4)-L, i.e. the right inverse of A with smallest operator norm. It is also characterized by A* = A*{AA*)~X where A* : W —> V is the adjoint (i.e. transpose in a matrix representation) of A. For more details see [Campbell and Meyer 1979; Allgower and Georg 1990; Shub and Smale 1996]. The first invariant, n(fa), x e B ( R n + 1 ) define
is a kind of condition number. For a 6 B(Rl+i)
and
K(/ 0 ,x) = max{l,||/||||D/ 0 (z)t||} where Dfa(x) : fftn+1 -+ R m is the derivative of /„ at x and we have taken its Moore-Penrose inverse. The norm in the expression ||.D/ 0 (x)t|| is the operator norm
where we are using the Euclidean norm on R m as well as on R n + 1 . If Dfa(x)^ is not defined (i.e. Dfa{x) is not surjective) we take ||I?/a(x)*|| = oo. The quantity ||/|| in the above expression for n(fa,x) is the norm of / we now describe. Let Q = (QO, • •. ,Q/) G N / + 1 and denote by \a\ the sum ao + • ■ • + at- Define similarly |/3| for Q € N n + 1 . For / a bihomogeneous polynomial,
f(a,x)= £ / 0 / 3 a V , \0\=d
1616 Complexity Estimates depending on Condition and Round-off Error define a norm
ll/lll = £
\faffl
83 Our norm on H(c,d) is defined by
-En/* »=i for / = ( / i , . . . , / m ) e
U(C4).
For / in Example 1, ||/|| <
) where
D = maxd,. i<m
For / in Example 2, ||/|| = 4. For fa feasible define «(/o) =
min z€B(R"+1) /.(x)=0
n(fa,x).
For /o infeasible define A
2
(/«) = 77771 i™> 1 ll/.(*)lloo. ll/H zestR""-
The condition number we consider is If \ - j *(-k) M ~ 1 STCJ
if if
/» iS /• i s infeasible
feaSible
/, v
[i)
If «;(/„) = oo we say that /x(/0) = oo. Note that, because of the compactness of B(R n + 1 ), the minimum in the definition of A(/ 0 ) is attained and we have A(/ 0 ) ^ 0. Therefore, /i(/ 0 ) < oo if / 0 is infeasible. We will see in Corollary 1 in Section 5 that n(fa) > 1 for all / € n{c
U
=
m ax
+1
x,»6B(R" )
\Ma>*)-Ma>v)\ ll/tllllk-J/IU
o€B(R'+>)
and let L/ = max{l,L/l,...,L/m}. We will show soon that Lf < D.
(2)
1617
F. Cucker and S. Smale The third quantity to play a role in the study of our problem is the following, Kf = max < 1,
Dkfa(X)
max
11/11*!
(o,z)eB(R'■^■')xB(Il"+ , )
&
(3)
where Dkfa(x) is the kth derivative of /„ at x. Recall that Dkfa(x) is a fc-linear map and that its operator norm is defined by iin*/ M\\ -
» H J '/«( g )("i.---»"*)ll
l | J ) / a ( l )ml l Y m , x
"
"w~S-
Here the maximum is over all points (y\,... ,«*) 6 (R n + 1 - {0})*. The definition of K/ does not communicate an immediate geometric meaning as do fi(fa) and Lf. We will show in Section 5 that Kf can be bounded by \D2 where D is as above. REMARK 2.
(1). For simplicity above and below, we use extensively the normalizations Halloo = 1 and II^HQQ = 1. To a certain extent we could have defined scale in variant quantities on R / + 1 x R n + 1 in their place. For example in the definition of «(/<>) we could have considered the linear map Diag(||a||-«|Ml£'W.(*) (4) in the place of Dfa(x). Here, for Ai,..., Ar € R, Diag (A*) denotes the diagonal matrix with Aj in the tth position in the diagonal, t = l,...,r. Note that the expression (4) is homogeneous of degree 0 both in a and x. A similar though less simple modification can be done to the kth derivative Dkfa(x). A more satisfying development could perhaps follow Dedieu-Shub [1997] "multihomogeneous Newton." (2). We also remark that /x(/0), •£/, and Kf are all homogeneous expressions of degree 0 in / . (3). The choice of the box £ ( R n + 1 ) instead of the spheres S
If we take representatives of a and x on 5(R 2 ) we get
a=
"nd
(7!'7i)
x=
(^'^)'
Then, Dfa(x) is given by the matrix
and therefore, 2*1
Dfa(x)
•-S(A)
1618 Complexity Estimates depending on Condition and Round-off Error
•
9
and
\\Dfa(xn = -^-. Had we defined /x(/ a ) using the points (a, x) in unit spheres we would have /i(/ 0 ) =
\\Dfa(xV\\\\f\\ = ^ .
On the other hand, if we take representatives of a and i on B(fL2) a= (1,1) and i = (1,1). Then, Dfa{x) is given by the matrix (d, —d) and therefore,
we get
and
Thus M.) = max(l, | | D / 0 ( i ) ' | | | | / | | } = maxfl, =§} = 1. Note the exponential dependence on d for /*(/<,) in the first case and the lack of such dependence in the second. We now consider the corresponding invariants for the problem of quasifeasibility. So consider again / € 'H(c.d) as m Definition 1 of Section 2. For (a,x) 6 B ( R ' + 1 ) x fl(R"+1) define K(/,(a,z))=max{l,||/||||D/(a,x)t||} +1
where Df(a.x) : R ' to both a and x). For /„ feasible let
x R n + 1 -> R m is the derivative of / at (a,x) (with respect
«(/o) =
min
z€B(Rn+1)
/c(/,(a,x)).
/.(*)=o
Since quasi-feasibility is a weaker assertion than feasibility, we are able to deal with it more generally. This is reflected by the fact that «(/„) is finite in many cases when /c(/o) = oo, and that Jc(/Q) < «(/„)• For / € H{c,d) and a e B ( R / + 1 ) define -it
\ - l ^(/ Q )
~ I sfe
if
/»
is f e a s i b l e
if ia is mfeasible
,s>
[
where A(/ G ) is as above. Again, if Jc(/0) = oo we say that fi(fa) = oo. Note that £(/„) < (*{fa) for all a 6 B ( R n + 1 ) and itsfiniteness*is a less restrictive condition than that of n(fa). For many important situations /x(/o) can be estimated while /x(/ a ) can not. Thus Theorem 2 dealing with quasifeasibility, below, covers a wide range. In Proposi tion 4 we will see that for dense systems, i.e. systems / as in Example 1 above, as well as for the / in Example 2, we have «(/ Q ) <
1619
10
*
F. Cucker and S. Smale
We also replace Kj by a new value, namely K/ = max < 1,
max (o,i)6B(F'+')xB(R-+1)
Dkf(a,x)
sir
ll/ll*!
(6)
where now Dkf(a,x) : ( R ' + 1 x R n + 1 ) * -¥ R m is the kth derivative of / at (a,x). One has that Kf > Kf but we shall see that Kj < \{C + D)2 where D is as above and C = max d. l
4. DECIDING HOMOGENEOUS SPARSE SYSTEMS OF EQUALITIES We state here our first two theorems. THEOREM 1. Let f e H{C,d)- There is an algorithm, to be described, such that with input a e B ( R ' + 1 ) , halts correctly with output:
■ (&) "fa is infeasible", or ■ (b) "fa *5 feasible" and gives an actual zero of fa represented by an explicit approximate zero, or else . (c) doesn't halt. The last case occurs if and only if /x(/o) = oo. The number of basic steps is bounded by
{bn(fa)2L,KfVn-)n where b is a universal constant. REMARK 3. The notion of approximate zero is made precise in Section 6. The notion of basic step is the evaluation of certain functions at points x € i?(R" + 1 ) and is made precise in the proof of Theorem 1 in Section 8. The number of arithmetic operations involved in a basic step depend only on c, d, I, n and m of Definition 1 and is bounded by a polynomial in these parameters. THEOREM 2. Let f 6 ^(C,d)- There is an algorithm, to be described, such that with input (e.o), 0 < e < 1, a € B ( R / + 1 ) , halts correctly with output:
. (a) "fa is infeasible", or ■ (b) "fa is e-quasifeasible", or else . (c) doesn't halt. The last case occurs only t//i(/ 0 ) = oo. The number of basic steps is bounded by
fbti(fa)2LfKfVIn~y where b is the universal constant of Theorem 1. Moreover in case (b) the output includes a' such that \\a' - a\\ < e and fa> is feasible. Explicitly, the output includes (a,x) which is an approximate zero representing (a',x') with fa'(x') = 0.
1620
11
Complexity Estimates depending on Condition and Round-off Error REMARK 4.
(a). For several important cases, such as the Nullstellensatz of Example 1, case (c) above does not occur. As we already remarked, we will prove in the next section that «(/„) < ll/H for all / € V-(Ctd) feasible and all a € £ ( R ' + 1 ) . Consequently, £(/„) < oo for all / € H{c,d) and all o 6 B(Rt+1). (b). In Theorem 2 halting may occur even if £(/<») = oo. (c). Note that infeasibility and quasifeasibility are not mutually exclusive. (d). We observe that the bound above depends on the input (a,e). Since / is fixed, the time bound has the form (cp(/ 0 )/e) n , with c independent of a and e, which is polynomial in /2(/a) and in 1/e. Eventually, we will want / to vary. A similar remark holds regarding Theorem 1 (without the e). In Section 10 we will extend Theorems 1 and 2 above to the affine case and subsequently also to systems containing inequalities. First we give bounds for L/, Kf and Kj as well as for «(/„) for dense systems as in Example 1. 5. BOUNDS ON THE INVARIANTS The following proposition will be useful. PROPOSITION 1. Let f € R[ao,...,a*,xo,... ,x n ] be bihomogeneous of degree c and d in the a's and the x's respectively. Let a € 1R/+1 and x € IR n+1 . Then
N- \f{a,x)\ < H/HillalUMI-, (ii). For u € R / + 1 and v 6 R n + 1 \Df(a,x)(u,v)\
< H/lliNl^llxll^CcllxlUlulloo + d||a|UMIoo). 1
(Hi). \\Df(a,x)\\<\\fh\\a\\ So-'Nfc'(cllxHoo+dlMloo). ' (iv). ||Z>fc/(a,x)ll< ll/llil|a||£*IMl£ PROOF.
l)||o| l *llio"S-O) c • • • (c - j + 1) d ■ ■ ■ (d - k + j + lSoj
Part (i) is straightforward. For part (ii) we have Df(a,x)(u,v)
= Daf(a,x)(u) +
Dzf(a,x)(v).
For a multiindex a = (QQ, • • •,at) let 2j = (QQ, . . . , <*J — 1,..., at). Then df |U./(a,x)(u)| = £ ^ ( a , x ) t t i ox <=o *
Y,Y,*<*a*a°ixfiui «=0 a,0
IHI^Mlxll^ a,/3
t=0
= lloll^MNISolNlooll/llic. A similar argument proves that \Dxf{a,x)(v)\ < ||a||| o Hx||5J 1 ||v|| 0O ||/||idand there fore (ii) follows.
1621 12
*
F. Cucker and S. Smale
Part (iii) follows from the bound ||i||oo < IMI for every i € R " and part (iv) is proved similarly. □ REMARK 5. The proposition above is for total derivatives, but a similar result holds if we differentiate only with respect to the x's \Dfa(T)(v)\
<
H/IWIollSodllxll^lMloo
and \\Dkfa(x)\\ COROLLARY
<
H/IU||o||So||x||i-*d-..(d-k+l).
1. For all f e ft(c,d) and allae
£ ( R ' + 1 ) , /*(/„) > 1 and £(/„) >
1. PROOF. By definition, «(/„) and Jc(/0) are both greater or equal than 1. Now, for any x € B ( R n + 1 )
||/.(*)IU = max \fi(a,x)\ < max \\fi\U < ||/|| 15*Sm
l
from which we deduce 277~y ^ *•
CD
We now give a bound for Lf. PROPOSITION
2. Let f e U(c,d)-
For each i = l , . . . , m , all a € B ( R / + 1 ) and
i,yefi(R^) |/i(a,x)-/,(a,y)|<||/,||1di||x-y||00. 77ius, Lj < D where D = max dj. «<m
PROOF. Let i < m and consider a € B ( R / + 1 ) and x,y € B ( R n + 1 ) . Denote by fi.a the function / i>a : R n + 1 -+ R x -> fi(a,x). By the mean value theorem, there is a point £ in the line segment xy joining x and y such that /i(a,x)-/i(a,y) = £»/l.o(0(x-y). Thus. |/<(a,x)-/i(o,y)| = |£>/,,o(0(x-y)|. By Remark 5 we have \Dfi
- ylU- Since C € xy
The following is a "higher derivative estimate" (compare [Shub and Smale 1993a]). That both the a variables and the x variables take values in a compact space makes this estimate possible. PROPOSITION
3. Let f e Hied), C = m a x d and D = maxaV Then i<m
i<m
1622 Complexity Estimates depending on Condition and Round-off Error
13
and *f<
n
•
We will use the following. LEMMA 1. Let s eTi.
Then, for any
k>2
I±T PROOF.
Raising both sides toD the the (k (k — - l)th power our statement becomes
or yet s*-i(s-l)*-i 2
which is immediate.
*-i
^ a(8-l)---(a-k 2-3- k
+ l)
,
D
3. We first prove the bound for Kf. Consider any (a, x) G B ( R ' + 1 ) x £ ( R n + 1 ) and assume for the moment that m = 1. By Proposition 1 (iii) together with the equalities Halloo = ||a;||oo = 1 one has that || — S?'x' || is at most
P R O O F OF PROPOSITION
■frZ(J) =
|/!lif
c • ■ (c - j + 1) d ■ ■ ■ (d - k + j + 1)
k\
c!
y
j^j -{j-k)\{c-j)\{d-k
d\
+ j)\
The last line uses a well-known combinatorial identity (see, for instance, [Graham et al. 1989] (5.23)). Now if m > 1, we have
Dkf(a,x) ll/ll*!
1
m
J=I
ll/ll2
Dkfi(a,x)\ k\
and, by the* case m = 1,
1623 14
*
F. Cucker and S. Smale
Applying Lemma 1 we deduce the bound for Kj. The bound for Kj is obtained in a similar way using Remark 5 in the place of Proposition 1. □ REMARK
6. The proof above yields actually a slightly better bound for Kj
namely, ~ ^ (max{c + 4 } ) 2 K,< ^ • We end this section by proving that K ( / 0 ) < ||/|| for / as in Examples 1 or 2. We actually prove a stronger result namely, that for all (a,x) e £ ( R / + 1 ) x 2?(R n + 1 ), \\Df(a,xy\\
By Lemma 2, and since A1 =
PROOF.
A'(AA*)~l,
||^|| 2 = | | ^ M t | | = || {A'{AAmylY =
(A'(AA')-1)
||
and, since {AA")-1 is symmetric
UAA'r'AA'iAAT'W
= IK^*)'1!!This proves the first equality. The second one then follows since the eigenvalues of (AA')-1 are the inverses of the eigenvalues of AA*. □ 4. Let N >m and A : R N -+ R m be linear of the form (J, M) where J is a diagonal m x m matrix whose diagonal entries are either 1 or —I, and M is any m x (N - m) matrix. Then \\A^\\ < 1. LEMMA
By Lemma 3
PROOF.
\\Ai\\2 = \\(AATl\\
= \\{I +
MM')-l\\
where I is the m x m identity matrix. The conclusion follows since J + MM* is a symmetric matrix all whose eigenvalues are greater than or equal to 1. □ 4. Let f be as in Example 1 or as in Example 2. Then for all a 6 fl(R' ) andx € B ( R n + 1 ) , K ( / , ( O , X ) ) < ||/||. In particular, if fa is feasible, PROPOSITION +1
£(/«) = «(/.) < 11/11Consider points a € B ( R ' + 1 ) and x € B ( R n + 1 ) with |xj| = 1. For i — 1 , . . . , m, denote by a<j the coordinate of a corresponding to the monomial Xj{.
PROOF.
1624 Complexity Estimates depending on Condition and Round-off Error
*
15
Then, ■$£- (a, x) = 1 . Thus, up to reordering its columns, the matrix of Df(a, x) has the form of Lemma 4 and we have ||Z?/(a,:c)t|| < 1. □ REMARK 7. Proposition 4 has as a consequence that, in this case, we don't need to know the zeros to estimate the condition number.
6. NEWTON'S METHOD AND APPROXIMATE ZEROS
In this whole section we use the norm || H2, denoted by || ||, and associated operator norms because of the use of Moore-Penrose inverses. Let / : R N -¥ R m be analytic and x 6 HN such that the derivative Df(x)
:RN^Rm
is surjective. Then, the Moore-Penrose inverse of this derivative £»/(x)f : R m -4 TLN is well defined and we define the Newton operator at x to be Nf(x) = x-
Df{x)*f(x).
N
DEFINITION 3. Let x € R and consider the sequence ., = x, and x*+i = N/{xi) for i > 0. The point x is an approximate zero of / this sequence is well defined and there exists a point x' 6 HN such that f(x') = 0 and
l|s'-*ill<(0
l|x'-xo||.
In this case we say that x' is the associated zero of x and that x represents x'. A problem related to Newton's method (for a fixed function / : R w -> R m ) is the following. Given a point x € R w , does the sequence of Newton iterates converge to a zero of / ? Within the theory of computation over the reals (with exact arithmetic) introduced in [Blum et al. 1989] this problem is shown to be undecidable. A simple example of function / for which the problem above is undecidable is the polynomial f(x) = x3 - 2x + 2 (cf. Chapter 2 of [Blum et al. 1998] for more details). Another example, mentioned by Mike Shub [1994], is the complex polynomial f(z) = (z2 l)(z 2 + 0.16). Shub's paper provides, in addition, a computer*generated graphic distinguishing "good" from "bad" starting points in the complex plane. The fractal character of this distinction becomes apparent in this graphic. Sufficient conditions exist, however, for x to be an approximate zero of / in what is called point estimates or a-theory [Smale 1986]. An exposition of this theory in a simple context can be found in [Blum et al. 1998]. We now briefly sketch its main concepts and results. For / : R " -> R m and x 6 TLN define /?(/,x) = ||D/(x)V(x)||, 7(/,x) = rnax
D / W t ^Jfc!W
1625 F. dicker and S. Smale
16 and
a ( / , z ) = /?(/, x ) 7 ( / , x ) . t
If £>/(i) is undefined, we write a,/?,7 = 00. The following is the main theorem in the theory of point estimates (see [Smale 1986; Shub and Smale 1996]). THEOREM 3. There exists a universal constant ao, around g, with the following property. Let f : R w -¥ R m andx € I t " . Ifa(f,x) < Qo thenx is an approximate zero of f. In addition, ifx' denotes its associated zero then ||a;-x'|| < 20(f,x). □
Theorem 3 provides a sufficient condition for / to have a zero. The rest of the section is devoted to a criterion to bound ||X?/(x)*|| by 2||Z?/(i') t || when x' is close enough to x. This criterion, applied to a sparse system / , will enable us to guarantee that, if /„ is feasible and /*(/„) < 00, the algorithm of Theorem 1 produces a point x satisfying the hypothesis of Theorem 3 after a finite number of steps. We remark here that the remaining results in this section are a significant step toward the proof of Theorem 3. LEMMA 5. Let A, B : TRN -> JR.m be linear maps with A surjective and \\A^A — Al t ||. Let V C JR.N be the orthogonal complement to ker A Then, A* A is the projection of IR." onto V and its restriction to V is the identity map IvLet By be the restriction of B to V and r : V -> V be the linear map defined by v = Iv - A^By- Then, from our hypothesis, ||t>|| < \ and therefore the limit PROOF.
JZTLn vl exists and has norm less than For all n € N, Iv - v n + 1 = ( £ " = 0
Iv
-(£')
T—■ < 2. 1 - || v || v'K/v - v). By taking limits
(Iv-v)=
£t,'Ut£?v k«=0
and thus, C£j^ 0 v'M* is an inverse of By- This proves that By is bijective and therefore that B is surjective. Moreover, since ( 5 3 " 0 v')A^ is a right inverse of B, its norm is at least that of B^ and thus 00
11**11 <
5>' K
Pt||<2||^||.
1=0
□ For the proof of Proposition 5 we use the following elementary lemma. LEMMA 6. For 0 < r < 1, ^
*r fc_1 = j -
*=2
1 - « / | this sum is bounded by 5.
* ~
r j - 1. In particular, for 0 < r < r
'
D
The following extends Lemma 2 of [Smale 1986] to the Moore-Penrose case.
1626 17
Complexity Estimates depending on Condition and Round-off Error
7. Let x,x' e HN be such that u = \\x - x'||7(/,i') < 1. \\Df(x')lDf(x') - Df(z')iDf(x)\\ < ^ - 1. LEMMA
PROOF.
Then,
By expanding Df around x' one has
Dkf(x')(x - a:')*-1 (*-!)!
Df(x) = Df(x') + J2 Jk=2
and thus, i\k-i *=2
Taking norms \\Df{x')^Df{x)
- Df(x')lDf(x')\\
<
Y,
Df(x')Wkf{x'){x
k=2 oo
Df(x'VDkf(x')
fc=2
But
DI(X
'^\f
/ 11 Ar — 1
\x — x
k\
^ 7(/>*') f c _ 1 and thus, the last line is bounded by
"^''^ oc
oo
*=2
k=2
using Lemma 6.
- i')*_1
-
(1-u)2
D
The following result has not been expressed before. Because of its simplicity and generality it can play a natural role in the theory of point estimates. 5. Let / : R " -» K m and x' e R N . For all x € R " satisfying u = \\x- x'\\j(f,x') < i one has \\Df(x)i\\ < 2||D/(x') t llPROPOSITION
PROOF.
Since u = | | i — i ' | | 7 ( / a , x ' ) < £ we have, by Lemma 7, \\Df(x'VDf(x)-Df(xyDf(x')\\<
1 (1-u)
-1
but i < 1 - w | and therefore, by Lemma 6, the last is smaller than | . Now apply Lemma 5 with A = | | D / 0 ( i ' ) | | and B = | | D / 0 ( i ) | | to deduce the statement. D Moore-Penrose versions of Newton's method were discussed in [Allgower and Georg 1990]. 7. POINT ESTIMATES FOR SPARSE SYSTEMS
In this section we apply the general results of the preceding section to the case that / is a sparse system. We will use (the Moore-Penrose extension of) Newton's method applied to fa : R n + 1 -* R m with starting point x G B ( R n + 1 ) , x to be chosen in Section 8. But the iterates will not be in the box and will not be scaled in contrast to the use of projective Newton's method in [Shub and Smale 1993a]. Instead, in the design of our algorithms, 0(fa,x) will be used to ensure that the
1627
18
F. Cucker and S. Smale
associated zero is not the trivial one. The goal is to obtain a suflicient condition for the feasibility of fa which is easy to compute. We first give a bound for y{fa,x) for / 6 H{e,d) and (a,x) G S ( R ' + 1 ) x B ( R n + 1 ) depending only on the first derivative Dfa(x) and the invariant Kf (which, recall, is bounded by ^ - ) . PROPOSITION
6. Let f e n{e
For all (a,x) G B ( R / + 1 ) x £ ( R n + 1 )
-r(fa,x)
dry 7(/o,*) = max
Dfa(xV^l Dkfa(x)
= max Dfa(x)*\
ll/P
< max || I ? / . ^ < max||£>/ Q (2;) <
max rir
t
Dkfa(x)
11/11*!
*/
K{fa,x)Kf.
D
We use the estimate in Proposition 6 to provide a . ound for a and a sufficient condition for feasibility. Let J5(fa,x)
= 0(fa,x)K{fa,x)Kf
=
\\Dfa(x^fa(x)\\K(fa,x)Kf.
Then, by Proposition 6, <*{fa,x)
and the following is a consequence of Theorem 3. 7. Let f e H{Cid) and (a,x) G B(Ht+1) x B ( R n + 1 ) . 7/6T(/ a ,x) < Qo then there exists a zero x' G R n + 1 of fa such that x is an approximate zero of fa. In addition, if x' is its associated zero then ||x — x'|| < 20(fa,x). D PROPOSITION
We now provide a criterion, easy to compute, to bound ||D/ 0 (x)^|| by 2||P/ 0 (x')*|| when x' is a zero of fa close enough to x. 8. Let f G H{c,d), a G B ( R ' + 1 ) and x' G S ( R n + 1 ) . For all x G B ( R n + 1 ) such that ||x - x'|| < 6K(K]x,)kf one has ||2?/tt(x)+|| < 2||£>/ a (x') t ||. PROPOSITION
By Proposition 6 and the inequality ||x - x'|| < 6n(/.?»,)A' w e deduce u - ||x - x'||7(/ 0 ,x') < \. Then the statement follows from Proposition 5. □ PROOF.
Suppose that /„ is feasible. Taking x' above to be the zero of /„ minimizing \\Dfa(x'y\\ and noting that now /*(/„) = /c(/ Q ,x') > ||/||||r>/a(x') t || we obtain the following.
1628
Complexity Estimates depending on Condition and Round-off Error
*
19
2. Let x' 6 B(R n + 1 ) such that fa{x') =0 and n(fa) = /c(/ 0 ,x') > ||/||||£>/ 0 (x')t|| < oo. For allxe B(Rn+1) such that \\x - x'\\ < t^m)ii, one has COROLLARY
HO/a(x)t|| < 2 ^ f i .
D
The above discussion extends with the appropriate changes to the application of Newton's method to / : R / + 1 x R n + 1 -► R m . Now the starting point is (a,x) 6 B(Rt+1) x B(Rn+1). The next proposition is analogous to Proposition 6 and is proved similarly. PROPOSITION
9. Let f € "W(c,d). For all (a,x) e B(R / + 1 ) x £ ( R n + 1 ) 7(/, (a, * ) ) < « ( / , (a, * ) ) ^ / .
D Define now a(/,(a,x)) = 0(f,(a,x))K(f(a,x))K,
=
\\Df(a,x^f(a,x)\\K(f(a,x))Kf.
Then, by Proposition 9, a ( / , (a,x)) < a ( / , (a, x)) and, again, the following is a consequence of Theorem 3. 10. Let f e H(c,d) and (o,x) € fl(R/+1) x J3(R n+1 ). // a(/, (a,x)) < Qo then there exist a' € R' + 1 and x' € R n + 1 such that (a, x) is an approximate zero of f with associated zero (a',x'). Moreover ||(o,x) — (a',x')|| < 2/J(/,(a,x)). D PROPOSITION
8. PROOF OF THEOREM 1
The basic idea of the algorithm is the following. Consider a grid of points in 2?(R""tl) sufficiently dense and evaluate /„ at all points x of the grid. Let a be the minimum (in norm) of these evaluations. If a is small enough, we will deduce that fa is feasible. If a is large enough, we will deduce instead that /„ is infeasible. The key propositions which permit these deductions are Propositions 7 and 12 respectively. 4. Let Q be a finite set of points in B(R n + 1 ). We say that Q is a grid of mesh s if for each y e B(R n + 1 ) there is an x 6 Q such that ||x - j/||oo < a. DEFINITION
It is easy to produce grids of a given mesh. Call a face of B(R n + 1 ) a set of the form F / = {x € B(R n + 1 ) | Xj = v) where j e {0,...,n} and v is 1 or - 1 . A face in B(R n + 1 ) is an n-dimensional cube. For each face F of B(R n + 1 ) consider a uniform grid GF of mesh s and let Q be the union !
>
F
We will call the QF the faces of Q.
■
1629 20
•
F. Cucker and S. Smale
11. The set Q is a grid of mesh s in B(Rn+1). It contains at most 2(n + 1)([-J + l ) n points where, for a real number x, [z\ denotes the largest integer smaller or equal than x. PROPOSITION
For each face F of B ( R n + 1 ) , the grid QF has ( [ j j + l ) n points and there are 2(n + 1) such faces. □ PROOF.
The next proposition gives a condition for infeasibility in terms oia, Lf and ||/||. 12. Let Q be a grid in 2?(R n + 1 ) of mesh s and f e H^.d)ll/o(z)lloc > s | | / | | i / for all x € Q then fa is infeasible. PROPOSITION
Suppose fa(x') = 0, x' € B ( R " + 1 ) and let x € G with ||x Then, for each i = l , . . . , m PROOF.
\Ma,x)\
= \fi(a,x) - fi{a,x')\
< L A ||/i||i||x - x ' | U <
by (2), contradicting ||/ 0 (x)||oc > s\\f\\Lf. REMARK
I'IU
If < s.
Lf\\f\\s
D
8. Note that in the context of the proof of Proposition 12, we also have
(
m
\
^(i/JI/illilli-x'Hoo)2)
P R O O F OF THEOREM 1.
1
/
2
/
m
< sLf lj2Wfihj
\
J/2
(7)
Consider the algorithm
input a (i) compute a grid Q in S ( R " + ) of mesh s (ii) if ||/ o (a 5 )|| O0 > all/Hi, for all are (? halt and return i n f e a s i b l e (iii) if a(fa,x) < Qo and 0(fa,x) < | for some x e Q, then halt and return f e a s i b l e and x
(iv) - := f go to (i) Propositions 12 and 7 ensure that this algorithm answers correctly. For, if it halts in step (ii). fa is infeasible by Proposition 12 and if it halts in step (iii) the feasibility of fa is ensured by Proposition 7. Note that the condition 0{fa,x) < ^ in (iii) ensures that the associated zero of fa is not the trivial one. Assume that n(fa) is finite. We will then show that the algorithm terminates within the claimed time bounds. Our claim is that the algorithm halts for any value of s in (i) satisfying s
-snUayL}K,jh-
(8)
First suppose that fa is feasible. Then n(fa) = «(/<»)• Let x' € B ( R n + 1 ) be a zero of fa such that /x(/Q) = K ( / 0 , X ' ) > ||/||||Afa(z , ) t ll and i € 5 b e a point in the same face of Q and such that ||x - x'Hoo < s. Then, ||x - x'|| < s-Jn. Since Lf,(i(fa)
> 1, one has s <
—■= and by Corollary 2, it follows that uH\fa)Kj\/n \\Dfa(x)m\\f\\ < ||I>/„(x')t||||/|| < 2/i(/.) and also that /c(/ tt ,x) < 2/x(/ 0 ). Thus,
1630 Complexity Estimates depending on Condition and Round-off Error
•
21
by (7) and (8), 5 ( / f l , * ) < K(fa,x)\\Dfa{x)*\\\\fa(x)\\Kf
< Mfaf'LfKf
<
^ .
On the other hand, /?(/„,*) = ||U/.(*)t/ a (x)|| < ||U/.(x)t||||/ a (*)|| < 2fi(fa)sLf < i the last inequality also by (8), the next to last by (7). We conclude that if /„ is feasible then the algorithm will halt at (iii). If fa is not feasible, then /i(/Q) = ^h\ ■ In this case, by the definition of A(/ a ), for all x e B(Rn+1) ll/.(*)lloo > ^
> WfWLfS
the second inequality again by (8), and the algorithm will halt at step (ii). We next prove the bound on the number of basic steps. Define as basic step either the computation of ||/ 0 (x)|| 00 in step (ii) or the computation of 5(/ a ,x) and /?(/„, x) in step (iii). Let r be the smallest integer such that 2-< ~
22 SniUYLjKjy/H
l Thus, 2r < ' ■ The main loop of the algorithm is executed at most r a times, with s = 2 _ 1 , 2 - 2 , . . . , 2~T. At the tth execution of the loop, s = 2~l and, by Proposition 11, Q has at most 2(n + 1)(2 ,+1 + 1)" points. Thus, the total number of basic steps in the algorithm
is r
r
4(n + 1) £ ( 2 * + 1)" = 4(n -I-1) £ ( 2 * + *)n t=l
t=2
< 4(n+l)(^2r+1r < 4 ( n , l ) (
3
W ° ^ ^ )
n
from which the claimed bound follows. We finally prove that if the algorithm halts, then /i(/ a ) < oo. If the algorithm halts in (ii) then fa is infeasible and /i(/„) = ^A \ which is finite. If the algorithm halts in step (iii) we have seen that there exist x € G, x' £ R n + 1 such that x is an approximate zero of /„ with associated zero x'. Moreover, 5(/ a ,x) < an and x' ? 0. Thus,
||, - x'|| < 20(fa,x) = ^ 4 - < K(fa,x)Kf
^
.
6K{fa,x)Kf
Since 8QO < 1 the pair (x, x') satisfies the hypothesis of Proposition 8 (with the roles of x and x' exchanged) and we deduce that Dfa(x') is surjective. Now consider x" = jr^jr-. Then fa(x") = 0 and x" € B(R n + 1 ). But Dfa(x") differs from
1631 22
•
F. Cucker and S. Smale
Dfa(x') by multiplying by a nonsingular diagonal matrix (see Remark 2) so is surjective as well and (i(fa) < co. □
Dfa(x")
9. The algorithm above uses the values of Lf and Kf. If such values are not available, we may use D and ^ - in their place thanks to Propositions 2 and 3 respectively. In this case, all the arguments above hold replacing Lf by D and Kf by — and the bound for the number of basic steps becomes REMARK
(M/a)^3^". 9. PROOF OF THEOREM 2 To prove Theorem 2 we will apply Newton's method to the function / : R ' + 1 x R n + 1 -4 R and consider starting points ( a , i ) € B(R* + 1 ) x B ( R n + 1 ) with a fixed. A small enough 0 will then ensure that the associated zero satisfies the desired equasifeasibility condition. 8. Let a € J3(R' + 1 ) and a' e R ' + 1 such that \\a-a'\\ < e withO < e < 1. Let a" = | j ^ - . Then, \\a - a"\\ < eVTTT. □ LEMMA
P R O O F OF THEOREM 2.
Consider the algorithm
input (a, e) s s ■= i • 2
(i) compute a grid Q in B ( R n + 1 ) of mesh s (ii) if ||/(o, i ) | | 0 0 > all/Hi, for all x € C halt and return inf e a s i b l e (iii) if Q ( / , (a, i ) ) < a 0 and /?(/, (a, a;)) < 2J)+1 for some i e 5 , then halt and return c-quasif e a s i b l e and (a, x) (iv) s := § go to (i)
Proposition 12 guarantees that if the algorithm halts in step (ii) then /„ is infeasible. On the other hand, if the algorithm halts in step (iii), Proposition 10 ensures the existence of a point (a',x') € R / + 1 x R n + 1 such that f(a',x') = 0. Moreover, ||(a',x') - (a,x)|| < 20(f,(a,x)) < -jj— and therefore \\a' - a\\ < -fi—. Let a" = | ^ - . Then, by Lemma 8, \\a" - a\\ < e and the equality f(a",x') = 0 ensures that / is e-quasifeasible. Assume that /2(/a) is finite. We will then show that the algorithm terminates within the claimed time bounds. As in Theorem 1, we claim that the algorithm halts for any value of s in (i) satisfying
s<
y ,
.
o)
If fa is feasible, one shows as in Theorem 1 the existence of a point x 6 Q such that a{f, (a,x)) < a 0 and 0(f, {a,x)) < 2J}+1 ■ Thus, the algorithm halts in (iii).
1632 Complexity Estimates depending on Condition and Round-off Error
•
23
If fa is not feasible, one shows also as in Theorem 1 that the algorithm halts at step (ii). (Note that the algorithm may have stopped, for a previous value of s, in step (iii)). The bound on the number of basic steps is verified as in the proof of Theorem 1.
□
REMARK
10. Using the estimates for L/ and Kj the bound on the halting time
becomes nn(fa)2D(C
+
D)*y/l+\y/n\n
with 6 a universal constant. REMARK 11. As we noted in Remark 4, and since / is fixed (and thus, so are n,C,D and £), the complexity of the algorithm above on input (e,a) is bounded by hi(fi(fa)/e)h2 where h\ and /i 2 are positive constants depending only on / (e.g. /12 = 2n). In later sections we will be interested however in algorithms which solve the feasibility problem for varying / ' s , i.e. which take / as a part of the input. Then, the more careful description of /ij and /12 will be used.
10 FEASIBILITY OF AFFINE SEMI-ALGEBRAIC SYSTEMS
Let f\<9]
for i = l , . . . , m , forj = l , . . . , r , for k = l,...,q.
(10)
We may write these conditions as fa(x) = 0, ga(x) > 0, and ha(x) > 0, and we call the triple <pa = {fa,9a,ha) a basic semi-algebraic system. DEFINITION 5. We say that tpa is feasible if a point x exists as in (10) and we say it is infeasible otherwise. Such an x is said to satisfy <pa. REMARK 12. Note that we are denoting by a a point in R*, in contrast to the previous sections, where a denoted a point in R / + 1 . Later in this section we will also denote by a points in R* +1 . Since the ambient space to which such points belong will be made precise as needed we expect that this will not create any confusion. A similar remark hold for points x in R n .
We reduce this problem to the context of Theorem 1. First "bihomogenize". If p G 7L[a, x] is one of the polynomials above, a = ( a i , . . . , a*) and 1 = ( x i , . . . ,xn), denote by p the polynomial in 2Z[aQ,a,xo,x] obtained by bihomogenizing p with respect to the a's and the x's respectively. That is, if
a.0
1633 24
•
F. Cucker and S. Smale
then
~
V^
a 0 c-\a\
d-\0\
a,0
Here c and d are the degrees of p with respect to the a's and the I ' S respectively, a and 0 are multiindexes, and \a\ denotes the sum of the components of a. JThus, Jx is the bihomogenization^ of /< for t = l , . . . , m and we write / = (/i, • • •. /m) and similarly for g and h. Define, for t = 1 , . . . , m, Lf. =
|/i(o,«)-/<(o,»')| — .
max l
«.«'€fl
||/i||l||x-*'||oo
Similarly, define Lg. and L/,k for j = 1 , . . . , r and k = 1 , . . . , q. Now define L„ =
max { l . L / ^ L ^ . L ^ } . i=l,...,m J=l r *=1 «
If D is the maximum of the degrees with respect to x of the /j's, the
= f,(a,x) = 9j(a,x)-\\gj\\iLvy?
Hk(a,x,y,z)
= hk(a,x)
- HMi-^2*
i = l,...,m j = l,...,r * = l,...,g
where y i , . . . , yr, z\,..., zq are new variables. Again, F — [F\,..., F m ) , G — (G, Gr) and H = {Hu...,Hq). Let $ = (F,G,H). By abuse of notation, we shall consider $ both as a system and as a map * : R ' + 1 x R n + 1 x R r + « -> R m +'+«
(11)
(a,a:,j/,z) •-> (F a (a:,y,z),G a (a;,y,2),^o(a:,y,-z))REMARK
13. Note that * is determined by ip. Moreover, if ' -
(i,q)
ll(l,a)lloo then a € B ( R ' + 1 ) , and $„ is determined by <pa. The relationship between feasibility of <pa and feasibility of $„ is given in the next proposition. Let P = { ( z , j / , z ) € B ( R " + 1 ) x R r + « | i 0 > 0 a n d z t > 0 for k =
l,...,q.}.
Points in P will be said to be proper. PROPOSITION 13. For any a € R ' , y>a w feasible if and only if there exists (x, y,z) € P such that #5^2:, V, 2) = 0. □
1634 Complexity Estimates depending on Condition and Round-pff Error
*
25
If a zero (x,j/,z) of $5- as in Proposition 13 exists we say that $ 5 is feasible on P and that (x,y, z) is a proper zero of $5. Otherwise, we say that $a is infeasible on P. We now introduce appropriate versions of some of the invariants of Section 3 for W The condition number for <pa will be denned, up to some minor modifications, as the condition number of the associated system $„. To lighten notation, for the rest of this section and in the next two sections fi means fi, and the same holds for gj, hk, etc. Also, we shall write z > 0 to denote z* > 0 for all k = 1 , . . . ,q. Define ||$|| by m
r
«=1
For a € B ( R '
+1
9
j=\
k=\
n+1
) , x € fl(R
r+
) and (y,z) e R « define
K(*0,(Z,V,*))
= max{l,||*||||U* 0 (a:,y 1 z) t ||}.
Now, let a € R ' . If
./
x
{<pa) =
•
mm
(*,y.*)€P ♦ T (x.y,z)=0
—
K,(j>g,(x,y,z))
d(X,J/,z)
—
(12)
where d(xyy,z) — min{a,u,2i, zq}. Note that the minimum in (12) is achieved even though the set is not compact. If ipa is infeasible (i.e. $ 5 is infeasible on P) define ||*||
ix,y.z)€P
Again, the condition number we consider is .,
,
/ K*{ipa) if Vo is feasible is infeasible.
If it'(
,, K^, = maxil, 1
max Jt>i (o,i,v,j|68(R'+1)xB(R»+')xE'+f
Dk*a(x,y, z) ||*||fc!
^
Note that, for k > 1, the fcth derivative Dk$a(x,y,z) does not depend on y or 2. Thus, the maximum considered actually ranges over (a,x) € S ( R ' + 1 ) x B ( R n + 1 ) . This guarantees the existence of the maximum as well as the upper bound D2 since the proof of Proposition 3 applies to A-*.
1635 26
*
F. dicker and S. Smale
11. DECIDING BASIC SEMI-ALGEBRAIC SYSTEMS
In the sequel we shall concentrate on the feasibility of the map # ? : R n + 1 x R r + « -> R m +'+« for a € R ' , and 5 as in Remark 13. It is to this map that Newton's method will be applied. The extension of Theorem 1 to the semi-algebraic setting is the following. THEOREM 4. Fix if as in (10). There is an algorithm which, with input a 6 R*, halts correctly with output
. (a) Vo «s infeasible", or . (b) "ipa is feasible" and gives a point £ € R n satisfying <pa. This point is represented by an approximate zero (x,y,z) 6 R n + 1 + r +« of is! if(x',y',z') is the associated actual proper zero of $a then, £ is the image of x' under the projection * : R n + 1 -> R n (x0,...,x„) *-> ( i i , . . . , x n ) . or else . (c) doesn't halt. The last case occurs if and only if n*{y>a) = oo. The number of basic steps is bounded by {bS(Va)2Lv(K;fjn
+ r + q)n
where b is a universal constant. By Proposition 13, to decide the feasibility of (pa it is enough to do so for the feasibility on P of the associated $o From here and till the end of this section, o will denote an arbitrary element in B(R/+1). The ideas to prove Theorem 4 are therefore variations of the ones we used for Theorem 1. However, we do not want to search on grids in 2?(R n+1+p " , "«) since this would introduce an exponential dependence in r and q for the time bound. To avoid this, the idea is to search on grids in B ( R n + 1 ) and to malce y and z depend explicitly on x. So, for each j = 1 , . . . , r consider the function yj : B(Rn+1) -> R
(0 and for each k = l,...,q
if5>(o,x)<0
consider the function z* : B ( R n + 1 ) -¥ R ! if = )
/pS fi Ma x)>0
[0
if hk(a,x) < 0 .
Let C : R n + 1 -4 R n + 1 + r + « x i-4
(x,y(x),z(x))
1636 Complexity Estimates depending on Condition and Round-off Error
•
27
where j/(x) = (yi(x),... ,y r (x)) and z(x) = (21 (x),.. .,z,(x)). For any grid Q in B ( R n + 1 ) denote by Q' the set of points x in Q with x 0 > 0. The analog of Proposition 12, has the following form. 14. Let Q be a grid in £ ( R n + 1 ) of mesh a > 0 andae / / ||*.(C(*))||oo > 2a||*||I„ for all x SG*, then *„ is infeasible on P. PROPOSITION
B(R/+1).
The proof of Proposition 14 relies on three lemmas. 9. Letp beany ofgi,...,gr,hi,...,h9, a e B ( R / + 1 ) andx' e B ( R n + 1 ) such that p(a,x') > 0. Let s > 0. For any x € B ( R n + 1 ) such that ||x - x'Hoo < * andp(a,x) < 0 one has |p(a,x)| < sZ,,p||p||i. LEMMA
PROOF. Since | | x - I ' I I ^ < s we deduce that \p(a, x) -p(a,x')\ < sLp\\p\\i by the definition of Lp. But since p(a,x') > 0 and p(a,x) < 0 we must have that |p(a,x)| is bounded by sZ,p||p||i which in turn is at most sL v ||p||i. D
10. Let s > 0, a € B(Hl+1) and (x'.y',*') € B(Rn+1) x R r + « be such that * a ( x ' , y ' , z ' ) = 0. For all x 6 £ ( R n + 1 ) such that ||x - x'Hoo < s one has ll*.(C(*))ll < 4*\\K and, a fortiori, ||*.(C(*))||oo < » | | * H V PROOF. For t = l , . . . , m , Fi(a,((x)) = fi{a,x) and Fi{a,x',y',z') = fi(a,x') = 0. Then, \Ft(a,a*))\ < s\\fihL}i < s\\UhL^ For j = 1, . . . , r , if gj(a,x) > 0 then G,(a, ((x)) 0 and if >(a,x) < 0 then |Gj(a,£(x))| = | ^ ( Q , X ) | and one uses Lemma 9 to snow that this is bounded by s-MlfljIliThe same reasoning proves that |/i*(a, £(z))| < s||/i*||i.Ly,. Finally, LEMMA
m
ll*o«(x))ll
2
r i
2
= 53/ l(o,C(*)) + 5 ; G > ( a , a - : ) ) »=1 m
j=\ r
q 2
+ 5;ft(a.C(*))a k=\ 9
< 5>ll/IIS;lli^) 2 + J > U M i ^ ) 2 «=i <
^ | | * | |
>=i
*=i
2
D
The following is easily checked. 11. Let Q be a grid of mesh s in B ( R n + 1 ) . For every x' € B ( R n + 1 ) with xo > 0 there exists x € Q* such that ||x — x'||oo < 2s. D LEMMA
Suppose * a is feasible on P. Then, there exist (x',y',z') € P such that * 0 ( x ' , y ' , z ' ) = 0. Since xj, > 0, Lemma 11 ensures that there exists x € G' such that ||x - x'||oo < 2s. By Lemma 10, ||*o(C(*))IU < 2311*111,,, which contradicts the hypothesis of Proposition^. □ P R O O F OF P R O P O S I T I O N ^ .
Proposition 14 provides a sufficient condition for infeasibility on P. The next proposition provides a sufficient condition for feasibility on P. Implicit in it is Newton's method applied to * a : R " + 1 x R p + « -► R m + r + « starting at a point C(*) where x € £ ( R n + 1 ) .
1637
28
*
F. Cucker and S. Smale
PROPOSITION 15. Let * be as above and a e B(tLt+1). Let x e B ( R n + 1 ) . / / 5(*o, (C(z)) < " o , 0(*«, (C(*)) < ^ ^ arid C(x) € P, then * a is feasible on P. Proposition 7 guarantees that £(z) is an approximate zero of $„. In addition, if C' = (x',y',z') is the zero of ♦„ associated to C(x), we have PROOF.
IIC(*) - C'lloo < ||C(«) - Clla < 2/3(*.,C(x)) < a(C(x)). Therefore, x'0 > 0 and z' > 0, i.e. C' 6 P-
□
The next result will be used in the proof of Theorem 4. 12. Letx,x' € B ( R n + 1 ) such that Wx-x'W^ < s < 1 ond* 0 (C(x')) = 0. Then, ||C(*)-C(x')lloc
ze(v,a),log n(v,a))
15. ROUND-OFF ALGORITHMS AND LINEAR ALGEBRA
Our main result in this section, Proposition 17, estimates the error for a particular algorithm in linear algebra. It provides a detailed example of round-off analysis and will be essential in the next two sections. Let 5 be an m x m non-singular real matrix. Recall that a QR factorization of 5 is an expression of 5 as a product S = QR with Q orthogonal and R upper triangular. In this case, | | 5 | | = ||i?|| and | | 5 _ 1 | | = ||i? _ 1 || where, we recall, ||S|| denotes the operator norm of S for the Euclidean norm in R m . This is particularly useful to estimate | | 5 _ 1 | | since R is now easy to invert. The main result of this section is the following. PROPOSITION 17. There is a round-off machine M which, with input {m,S), computes a QR factorization of S, S = QR, and then inverts R. Moreover M satisfies the following property. For all u> > 0, let u = max{l,w}, and
(utmH)cim*2c"nS where C2,C3 are universal constants sufficiently large and H > 1 a bound for the absolute value of the entries of S. If the round-off functions of M satisfy 6{m,S),6i(m,S) < A then (i). If US"11| < LJ then Ill^MI < 2w, and («)■ If\\S-1\\>uthen\\R=^\\>