/p/ /p:/ /t/ /t:/ /k/
) , x, k>
/k:/ /b/ /b:/ /d/ /d:/ /g/ /g:/
) >
toro torre male palla gli [ńń]; figlio, maglia, aglio lama mamma nano nanna gnomo; stagno afa baffi lava davvero casa, extra cassa rosa, sbaglio scemo; pesce, sciame, lasciare zio zanzara cozza razzo; mazurca cena; pace; bacio lacci (pl.), laccio gerla, regione; jazz raggio, laggiù quello, guaio chiave, piano; yogurt*; Juventus* juventino papa pappa seta setta cane; chino; questo; extra; karate*; kit* becco; becchi; acqua bibita babbo guado freddo lago; laghi (pl.); hegeliano* leggo (I read); tegghia
16
Gerald Bernhard and Gabriel Altmann
with nx denoting the number of representing graphemes, and fx the number of phonemes with uncertainty Ux . Table 2: Orthographic uncertainty of Italian phonemes Phoneme
nx
/i:/, /e/, /ao/, /au/, /ai/, /Ei/, /Eu/, /;R/, /;R:/, /l/, /l:/, /m/, /m:/, /n/, /n:/, /˜n/, /f/, /f:/, /v/, /v:/, /s:/, /z/, /ts/, /dz/, /t:s/, /w/, /p/, /p:/, /t/, /t:/, /b/, /b:/, /d/, /d:/ /E/, /E:/, /a:/, /O:/, /o:/, /u:/, /ń/, /s/, /S/, /d:z/, /tS/, /t:S/, /d:ˇz/, /g:/ /a/, /O/, /o/, /u/, /dˇz/, /j/, /k:/, /g/ /e/ /i/, /k/
Ux
fx
1
0
34
2 3 4 5
1 1.58 2 2.30
14 8 1 2
The mean uncertainty can be computed as the average by means of 1 U¯ = ∑ fx Ux N x∈I
(2)
where N is the number of all representations. In our case U¯ = [34(0) + 14(1) + 8(1.58) + 1(2) + 2(2.32)]/59 = 0.5641. Comparing this number with the result from Swedish, U¯ = 0.797, and with ¯ 0.965 (cf. Best & Altmann 2005), one could conclude that the German U= Italian orthography is not so vague as the German or Swedish. In order to get a more objective image of these differences we set up an asymptotic test for the difference of two mean uncertainties, i.e. U¯ 1 − U¯ 2 z= V (U¯ 1 ) + V (U¯ 1 )
(3)
¯ is its variance and z is the Here U¯ is the empirical mean uncertainty, V (U) quantile of the normal distribution. The variance of U¯ can be derived using the Taylor expansion as below: 2 N 1 1 1 1 ¯ V (U) = V V (nx ) ∑ log2 nx = N 2 ∑ V (log2 nx ) = N 2 ln2 2 ∑ nx N x=1 E(nx )
The phoneme-grapheme relationship in Italian
σ2 ,
17
Since nx is the original variable whose expectation is E(nx ) = μ and V (nx ) = we obtain, after substituting in the above formula, ¯ = V (U)
σ2 Nμ2 ln2 2
(4)
which can be estimated by means of empirical values as ¯ = V (U)
s2 0.48N x¯2
(5)
In (5) one can easily see that the variance of the uncertainty is a function of only the well known variation coefficient. For Italian we obtain 1 x fx N∑ = [1(34) + 2(14) + 3(8) + 4(1) + 5(2))/59 = 1.694915 1 1 ¯ 2 fx = x2 fx − x¯2 s2 = ∑ (x − x) N N∑ = [12 (34) + 22 (14) + 32 (8) + 42 (1) + 52 (2)]/59 − 1.6949152 = 0.991669 . x¯ =
Finally, ¯ Ital = 0.991669/[0.48(59)1.6949152 ] = 0.012189. V (U) ¯ German = In the same way we obtain the variance for German as V (U) ¯ 0.012602 and for Swedish V (U)Swed = 0.022763. If we perform the above test on Italian and German, we obtain 0.9650 − 0.5641 = 2.55 z= √ 0.012189 + 0.012602 and this value is significant, i.e. Italian has a significantly smaller orthographic uncertainty than German. For the difference between Italian and Swedish we obtain 0.7970 − 0.5641 z= √ = 1.24 0.012189 + 0.022763 which is not significant i.e. Swedish and Italian have roughly the same orthographic uncertainty.
18
3
Gerald Bernhard and Gabriel Altmann
The distribution of graphemic representations
If a language using the Latin alphabet has fewer phonemes than there are letters in Latin, it can represent each phoneme by one letter. In such a case all frequencies are concentrated in point x = 1. We speak then of a deterministic distribution. But if a language has more phonemes than the Latin letter inventory, it must reach for different means in order to build corresponding graphemes. One method is introducing marks placed over or under the letters, as in Slavic languages; another is using some letters to signalize a special quality, such as in German for prolonging the vowel; still another is combining or redoubling some letters, e.g. , , or even in several languages. These new forms can, however, be chosen in such a way that each phoneme can be represented by one unique grapheme. This ideal state is usually considerably disturbed by the interference with morphology or by disregarding the phonological development of language. In this way phonemes acquire multiple representations. From the statistical point of view, a distribution of representation sizes of phonemes arises and it can be captured formally. Since up to now only a small number of languages has been processed in this way, we can start from simple assumptions. At first, we assume that the representation size decreases geometrically, i.e. it follows a distribution of the form Px = pqx−1 , x = 1, 2, 3, . . . (6) This is the 1-displaced geometric distribution. For Italian and Swedish this hypothesis would be adequate. However, in German we see (see Table 3) that the distribution does not decrease monotonically but has its mode at x = 2, i.e. more phonemes are represented by means of two graphemes than by one grapheme. Table 3: Distribution of representation size of phonemes in three languages x
Italian
German
1 2 3 4 5 6
34 14 8 1 2 −
10 18 7 3 0 1
Swedish 16 10 6 1 2 1
The phoneme-grapheme relationship in Italian
19
This circumstance can have different causes which must be analysed individually. In order to keep the original hypothesis, we modify (6) by means of the Gram-Charlier expansion (see Shenton & Skees 1970; Maˇcutek, this volume, pp. 75ff.) and obtain 1 x−1 1+a x− , x = 1, 2, 3, . . . (7) Px = pq p where q = 1 − p, 0 < p ≤ 1, 0 ≤ a ≤ 1/q − 1, q = 1 − p. This distribution is called either Gram-Charlier-geometric or Shenton-Skees-geometric distribution (cf. Wimmer & Altmann 1999). If in (7) p = 1, we obtain the deterministic distribution representing the ideal case, and if a = 0, we obtain the original geometric distribution. The fitting of (7) to the data in these languages can be seen in Table 4. Evidently the fit is in each case very satisfactory, but the hypothesis cannot be corroborated better until more languages have been examined. The fit is shown graphically in Figures 1a–1c. Table 4: Fitting the distribution (7) to data in Table 3 x
Italian
German
Swedish
1 2 3 4 5 6
33.31 14.92 6.37 2.64 1.76 −
9.99 18.00 7.54 2.47 0.73 0.27
15.79 9.99 5.35 2.64 1.24 1.00
p a FG χ2 P
4
0.6488 0.2398 2 1.55 0.46
0.7768 2.3323 1 0.12 0.73
0.6152 0.4588 2 1.36 0.51
Grapheme size
A grapheme can consist of one or more Latin letters. Because of unequal size of Latin letter and phoneme inventories of target languages, new letters (e.g. the German <ß>) and several additional marks (tilde, accents, etc.) were
20
Gerald Bernhard and Gabriel Altmann
40
20
35 30
15
25 f(x) NP(x)
20
f(x) NP(x)
10
15 10
5
5 0
1
2
3
4
0
5
1
2
3
(a) Italian
4
5
6
(b) German
20
15 f(x) NP(x)
10
5
0
1
2
3
4
5
6
(c) Swedish
Figure 1: Fitting (7) to Italian, German and Swedish data
introduced. Thus grapheme inventory can be measured in two ways: (i) as the number of Latin letters without considering additional marks, (ii) as the number of Latin letters plus additional marks. Consequently the German grapheme <ä> can consist of one symbol according to method (i) and of two symbols according to method (ii). For Italian we obtain the results on the basis of Table 1 as shown in Tables 5a and 5b. Table 5a: Size of Italian graphemes: method (i) Size
Grapheme
1
2
3
Number 30 36
5
In Table 5b six graphemes with accent passed from size 1 to size 2. The variable “size” has too small a support, which does not allow us to set up a testable model. For the time being it is enough to characterize the graphemics by its average and to compare it with other languages. Using method (i) we obtain the mean size of 1.65 lying between German (1.68) and Swedish (1.61)
21
The phoneme-grapheme relationship in Italian Table 5b: Size of Italian graphemes: method (ii) Size
Grapheme
1 2
< ì, é, è, à, ò, ù, hi, he, eh, ha, ah, ho, oh, hu, uh, ao, au, ai, ei, eu, rr, ll, gl, mm, nn, gn, ff, vv, ss, sc, zz, ci, cc, gi, gg, pp, tt, ch, cq, bb, dd, gh>
3
Number 24 42
5
while method (ii) yields the mean size of 1.70 lying also between German (1.78) and Swedish (1.67).
5
The graphemic load of letters
Latin letters are used with different frequencies in graphemes of target languages. The exploitation of letters for building graphemes can be designated as graphemic load. One can ask whether the letters present in graphemes have something to do with the phonemic relevance or whether they are merely historical relicts. In German, the letter occurs in 16 graphemes and its function is both segmental (there is a phoneme /h/), purely combinatorial (e.g. in the grapheme <sch>) or suprasegmental, e.g. to prolong the preceding vowel. In Italian occurs in 13 graphemes, but it plays only a secondary role: either it occurs in historically petrified forms or it helps to maintain the phonetic value of the preceding consonant (). In Table 6 the letters are ordered according to their graphemic load. It is not yet possible to set up hypotheses about this distribution because the empirical background is still very restricted and the class occupation very small. For the time being we must content ourselves with the computation of the mean load which results from the numbers in Table 6 as 98/25 = 3.92. For German we get 3.96, for Swedish 3.36. Italian lies between them.
6
Letter usefulness
The participation of letters in building graphemes can be weighted. We assume that the role of the letter is the more peripheral the later it appears in the grapheme. The historical and morphological roles of letters is neglected
22
Gerald Bernhard and Gabriel Altmann
Table 6: Graphemic load of Italian letters (Participation in grapheme forming) Component in x graphemes 1 2 3 4 5 6 7 8 9 13
Number of letters
Letter y, x, j, k r, m, f, v, z, p, t, q, b, d n l, s o u e, g a, c i h
4 10 1 2 1 1 2 2 1 1
in this case. The weighting has a purely positional character. The smaller the weight, the more useful the letter graphemically. Let us consider as an example the letter occurring in the following graphemes (see Table 5a/5b): . Let px gi be the product of the position (px ) of the letter <x> and the number of graphemes gi in which it occurs in this way. Then the positional participation of a letter can be defined as PP<x> = ∑ px gi . (8) gi ∈G
If it is weighted in each position by the position itself, then we find position 1 eight times and position 2 twice, i.e. PP = 1(8) + 2(2) = 12 . If this operation is performed for each letter, one obtains the results for Italian in Table 7, PP<x> denoting the weight, and fx the number of letters. There is a possible correlation between the relative frequency of individual letters and their graphemic usefulness. For the time being we can merely compute the mean positional weight of letters in the form PW (Language) =
1 fx PP<x> L∑ x
(9)
where L is the size of the letter inventory. For Italian we obtain PW (Italian) = [1(4) + 3(1) + 4(9) + . . . + 22(1)]/25 = 6.48 .
(10)
The phoneme-grapheme relationship in Italian
23
Table 7: Positional participation of letters in graphemes PP<x> 1 3 4 6 7 8 9 12 16 18 22
Letter
fx
y, x, j, k q r, m, f, v, z, p, t, b, d n, s l, o e, a u g c i h
4 1 9 2 2 2 1 1 1 1 1
Comparing with Swedish (5.41) and German (6.12) we see that Italian has a strong letter usefulness (great positional weight). However, it will not be possible to examine historical and morphological dependencies before many languages have been analysed. The same holds for the comparison of individual letters in languages using Latin script and the relationship with the letter/grapheme frequency of occurrence.
References Best, Karl-Heinz; Altmann, Gabriel 2005 “Some properties of graphemic systems.” In: Glottometrics, 9; 29–39. Maˇcutek, Ján 2006 “On the distribution of graphemic representations”. This volume, pp. 75– 78. Shenton, Leanne R.; Skees, P. 1970 “Some statistical aspects of amounts and duration of rainfall”. In: Patil, Ganapati P. (Ed.), Random Counts in Scientific Work. University Park: The Pennsylvania State University, 73–94. Wimmer, Gejza; Altmann, Gabriel 1999 Thesaurus of univariate discrete probability distributions. Essen: Stamm.
Graphemic representation of English phonemes Fan Fengxiang and Gabriel Altmann
1
Introduction
The grapheme-phoneme analysis of English is radically different from cases analyzed hitherto (German, Swedish, Italian, Slovak). This is caused by (i) the historical origins of English, (ii) its many national and regional varieties and (iii) the borrowing of many foreign words. The grapheme-phoneme mapping has been examined from different aspects (cf. Adams 1990; Berndt, Reggia, Mitchum 1987; Cunningham & Cunningham 1992; Fry 2004; Hanna et al. 1966; Patterson & Morton 1985; Seidenberg et al. 1984), and some probabilities have been computed. Such analysis is relevant not only for linguistics but also for cognition studies and pedagogy. We are interested here only in some measurable properties of the English phoneme-grapheme correspondence, in order to be able to study later on the divergence or convergence of this representation. Our analysis focuses on American English and the results cannot hold for other varieties, i.e., British English, though the methods used here can be applied directly. In order to work with controllable data, we adhere to the phonemic/graphemic analysis based on the American Carnegie Mellon Pronouncing Dictionary hereafter referred to as cmudict1 , which has 129 425 entries with phonological transcriptions. The phonological symbols used in the dictionary are listed below. On typographical grounds we adhere to this way of symbolizing phonemes. The dictionary uses 39 consonant and vowel phonemes. In his Essential Introductory Linguistics, Hudson (2000: 24ff.) uses 38 phonemes without the vowel er, which is used in the cmudict, as well as in the World Book Dictionary (Barnhart & Barnhart 1979). The words analyzed are from the 1 000 000-word Brown Corpus, which has 42 436 word types minus the Arabic numerals. Of these words, the cmudict covers 31 591; the uncovered part mostly consists of personal and place names, and non-word strings. This phonemic/graphemic analysis is the anal1. ftp://ftp.cs.cmu.edu/afs/cs.cmu.edu/data/anonftp/project/fgdata/ dict/
26
Fan Fengxiang and Gabriel Altmann
Table 1: List of phonemes Phoneme
Example
Transcription
/AA/ /AE/ /AH/ /AO/ /AW/ /AY/ /B/ /CH/ /D/ /DH/ /EH/ /ER/ /EY/ /F/ /G/ /HH/ /IH/ /IY/ /JH/ /K/ /L/ /M/ /N/ /NG/ /OW/ /OY/ /P/ /R/ /S/ /SH/ /T/ /TH/ /UH/ /UW/ /V/ /W/ /Y/ /Z/
odd at hut ought cow hide be cheese dee thee Ed hurt ate fee green he it eat gee key lee me knee ping oat toy pee read sea she tea theta hood two vee we yield zee
AA D AE T HH AH T AO T K AW HH AY D B IY CH IY Z D IY DH IY EH D HH ER T EY T F IY G R IY N HH IY IH T IY T JH IY K IY L IY M IY N IY P IH NG OW T T OY P IY R IY D S IY SH IY T IY TH EY T AH HH UH D T UW V IY W IY Y IY L D Z IY
Graphemic representation of English phonemes
27
ysis of these 31 591 word types using the pronunciation given by the cmudict. The pronunciation of each word in the cmudict is in the following form (0, 1 and 2 represent word stresses): LABORATORY
2
L AE1 B R AH0 T AO2 R IY0.
Data
The 31 591 word types from the Brown Corpus were automatically separated into graphemes and then paired with their corresponding phonemes with the computer in the following form: i|n|au|g|u||r|a|t|io|n|, /ih n ao g y ah r ey sh ah n/ , i:ih|n:n|au:ao|g:g|u:y ah|r:r|a:ey|t:sh|io:ah|n:n| Computerized analysis is error prone, even with the best commercialized state of the art software. There is no exception in this analysis. Although the result was manually checked, there still may be errors. In addition, there are indeterminable cases. For example, the first in the word LABORATORY is not pronounced in American English; should it be paired with the letter and the phoneme /b/ or the letter with the phoneme /r/? The was finally put together with to become → , or BO:B meaning the grapheme in this word is pronounced as /b/. Another possibility would be to consider as representing nothing, but in that case the analysis would be quite different. The third possibility would be to consider the grapheme as representing the cluster /br/, but in that case the analysis would produce an enormous number of phonemes, clusters and graphemic representations. We chose the first alternative, which yielded a reasonable image of this kind of English. On the other hand, we could not avoid the fact that some single graphemes represent a group of phonemes, for example in the word COMPUTER represents the group of phonemes /y uw/, or <x> in BOX represents /k s/. In such cases the given phonemes are (implicit) parts of the graphemic representation and these cases are marked with ∈, e.g. /k/ → ∈ <x>. Another problem was the representation of a phoneme by zero grapheme. For example ABLER has the pronunciation of /ey b ah l er/, in which /ah/ is present phonemically but not graphemically. These cases are interpreted as /ah/ being part of if stays in front of (also ) and
28
Fan Fengxiang and Gabriel Altmann
marked as /ah/ → ∈ . There are several cases of this sort as can be seen in Table 9 (see p. 43ff.). The following are the first ten cases of the automatic graphemic separation and grapheme-phoneme mapping by the computer. The graphemes are separated with “|”, and “||” means there is an ungraphemically represented phoneme; the word pronunciation is enclosed between “/”; and “:” pairs the phoneme with its corresponding grapheme: a|, /ah/, a:ah| a|b||l|er|, /ey b ah l er/, a:ey|b:b ah|l:l|er: er| a|b||le|, /ey b ah l/, a:ey|b:b ah|le:l| a|b|a|ck|, /ah b ae k/, a:ah|b:b|a:ae|ck:k| a|b|a|n|d|o|n|, /ah b ae n d ah n/, a:ah|b:b|a:ae|n:n|d:d|o:ah|n:n| a|b|a|n|d|o|n|ed|, /ah b ae n d ah n d/, a:ah|b:b|a:ae|n:n|d:d|o:ah|n:n|ed:d| a|b|a|n|d|o|n|i|ng|, /ah b ae n d ah n ih ng/, a:ah|b:b|a:ae|n:n|d:d|o:ah|n:n|i:ih|ng:ng| a|b|a|n|d|o|n|m|e|n|t|, /ah b ae n d ah n m ah n t/, a:ah|b:b|a:ae|n:n|d:d|o:ah|n:n|m:m|e:ah|n:n|t:t| a|b|a|t|e|d|, /ah b ey t ih d/, a:ah|b:b|a:ey|t:t|e:ih|d:d| a|b|d|a|ll|ah|, /ae b d ae l ah/, a:ae|b:b|d:d|a:ae|ll:l|ah:ah| All representations of phonemes by graphemes are shown in Table 9 (see p. 43ff.). Here “/. . . /“ symbolizes a phoneme,“<>” a grapheme, while “∈” means that the given phoneme is part of the grapheme cluster. The condition under which the given phoneme – usually a vowel – is placed (uttered) behind a grapheme is symbolized by a superscript “<>”, e.g. /ah/ * means that in some occasion /ah/ can be pronounced within , e.g. → /ey b ah l/. Though /ah/ is not overtly represented, is considered its representation. There is the possibility of simply ignoring the zero grapheme representation, but we decided for the above alternative. There are 289 different graphemic representations; many of them are used for different phonemes. The graph connecting the phonemes with graphemes is a bipartite graph which, because of its extent, cannot be presented here. 3
Uncertainty
The first impression of Table 9 (p. 43) is that each phoneme has multiple representations. One would tend to say that there is a very weak connection between the phonemes and graphemes, and that phonemes are represented
Graphemic representation of English phonemes
29
by combinations of Latin letters which contain mere phonetic orientations but nothing more. The situation can unreservedly be matched with Accadian writing, in which cuneiform symbols of different directions and sizes are combined, or, still better, with Chinese script containing always a phonetic guide. Hence, English writing resembles and is developing into a kind of linear hieroglyphic or logographic script. The extent of this development can be numerically expressed in different ways. Here we shall show only some very elementary methods.
3.1
Unweighted uncertainty
The variation or diversification of the way of representing graphically a phoneme can be called in general uncertainty. In the case that the individual representations are not weighted, uncertainty in information theory is defined as the dyadic logarithm of the number of representations, i.e. H0 = log2 K
(1)
where K is the number of representations. Consider e.g. the phoneme /AA/ having 19 different representations. Its uncertainty can be characterized as H0 (/AA/) = log2 19 = 4.25. There is no maximum of H0 because K is potentially infinite but its minimum is 0. The greater H0 , the more diversified the phoneme representation. The results for all phonemes are presented in the second and third column of Table 2. It can be seen easily that vowels have more diversified representations than consonants, though some of them are weakly diversified (/AE/, /AW/, /OY/). Under other conditions of sampling and interpretation, one would get another picture, but any version is merely an approximation because of the diversity of English. There is no trend, e.g. for normality of distribution of K or H0 because the graphemic representations did not arise by chance but by historical development and adaptation of foreign words. Seeing the phoneme-grapheme relations as a bipartite graph, the number of graphemic representations is, as a matter of fact, the degree of a vertex (phoneme). Thus uncertainty, diversification and vertex degree are in this case synonymous.
30
3.2
Fan Fengxiang and Gabriel Altmann
Weighted uncertainty
Even if a phoneme has a great number of representations (K), not all of them are of the same importance. Their relevance is weighted by their frequency of occurrence. This can be of two sorts: one based on the dictionary and the other based on texts. If one of the representations occurs 1 000 times, it is surely more relevant than one occurring only once. Hence, another measure of uncertainty is the entropy of first order taking into account the relative frequencies of individual representations. Usually one uses the Shannon entropy defined as K
H1 = − ∑ pi log2 pi
(2)
i=1
where K is the number of representations and we estimate pi by fi /N, where fi is the absolute frequency, N being the number of occurrences of representations of the given phoneme (N = Σ fi ). The more concentrated the frequencies, the smaller the uncertainty H1 . In order to illustrate the computation, we use the representations of /TH/, where there are N = 741 cases distributed to graphemic variants in proportions: 1, 736, 3, 1. Since formula (2) can be rewritten as H1 = log2 N −
1 K ∑ fi log2 fi N i=1
(2a)
we obtain H1 = log2 741 − (1/741)[1 log2 1 + 736 log2 736 + 3 log2 3 + 1 log2 1] = 0.0676 . Though there are four graphemic variants, the uncertainty is very low, because one of them, , occurs in the great majority of cases. Consequently, even if K or H0 are great, H1 yields a more adequate picture of uncertainty/diversity. The results are presented in the fourth column of Table 2. H0 and H1 are characteristics of uncertainty/diversity. The first shows the raw diversity, the second the exploitation of this diversity. Theoretically they could be independent but it can be shown that there is a tendential dependence of H1 on H0 . The t and the F-tests show that the curves H1 = aH b0 and H1 = aebx are adequate though the determination coefficient is not high enough.
Graphemic representation of English phonemes Table 2: Uncertainties of individual phonemes Phoneme /AA/ /AE/ /AH/ /AO/ /AW/ /AY/ /EH/ /ER/ /EY/ /IH/ /IY/ /OW/ /OY/ /UH/ /UW/ /B/ /P/ /M/ /F/ /V/ /W/ /D/ /T/ /N/ /TH/ /DH/ /S/ /Z/ /R/ /L/ /JH/ /CH/ /SH/ /ZH/ /Y/ /K/ /G/ /NG/ /HH/
K
H0
H1
R
19 7 60 19 7 16 18 29 20 22 21 19 6 13 31 5 8 9 12 5 8 5 16 12 4 2 18 16 9 7 7 13 10 8 17 21 8 6 3
4.2479 2.8074 5.9069 4.2479 2.8074 3.9999 4.1699 4.8580 4.3219 4.4594 4.3923 4.2479 2.5850 3.7004 4.9542 2.3219 3.0000 3.1699 3.5850 2.3219 3.0000 2.3219 4.0000 3.5850 2.0000 1.0000 4.1699 4.0000 3.1699 2.8074 2.8074 3.7004 3.3219 3.0000 4.0874 4.3923 3.0000 2.5850 1.5850
1.2742 0.0560 3.1340 1.8927 1.2608 1.0467 0.9189 2.0685 1.3999 1.0632 2.4710 1.0851 1.1616 2.0327 2.7422 0.3049 0.4945 0.5761 1.0432 0.7787 1.1800 0.7994 0.9753 0.4592 0.0676 0.3868 1.5637 1.2667 0.5337 0.9450 1.5289 1.5872 1.8879 1.2482 1.5770 1.9685 0.9152 0.5884 0.0966
0.4956 0.9898 0.1544 0.4166 0.5057 0.6783 0.7185 0.4144 0.5644 0.6196 0.2265 0.6735 0.4963 0.3115 0.2278 0.9158 0.8418 0.8319 0.6553 0.6592 0.5485 0.7292 0.7255 0.8764 0.9866 0.8601 0.5308 0.6300 0.8521 0.6609 0.4649 0.4130 0.3572 0.6181 0.5105 0.4185 0.7339 0.7692 0.9777
N 3989 4884 19983 2411 751 2617 5884 6384 3922 13664 7095 2812 306 535 2091 3790 5999 6379 3370 2711 1738 9082 13740 14163 741 185 12662 6193 10798 11279 3946 1116 2539 184 1060 9163 2380 3370 1510
31
32
3.3
Fan Fengxiang and Gabriel Altmann
Concentration
Another way of characterizing the diversity is the Herfindahl measure of concentration called repeat rate in linguistics. It is defined as the sum of squares of the probability of graphemic representations, K
R = ∑ p2i .
(3)
i=1
The probability is estimated by relative frequency, i.e. pi = fi /N, which gives R=
1 K 2 ∑ fi . N 2 i=1
(3a)
Here Kand N are different for each phoneme. This index shows the concentration of graphemic representatives. If all frequencies are concentrated in one grapheme, then R = 1. If all frequencies are equal, i.e. the diversity is maximal, it attains the value R = 1/K. From the geometrical point of view, (3) represents the Euclidean distance in a K-dimensional space, i.e. R is the coordinate of the phoneme. For example with /AE/ there are 7 graphemes but the frequencies are concentrated on , thus R = 0.9898. The smallest concentration (the greatest dispersion) is with /IY/ having R = 0.2265. It is possible to norm R in order to restrict it to the interval <0, 1>, but we leave in its original form. All results are presented in the fifth column of Table 2.
4
Grapheme length distribution
In a situation similar to English (i.e. where the number of phonemes is greater than the number of Latin characters), we expect graphemes of different length or modified graphemes like in Slavic languages. Since there are 26 Latin letters and the number of English phonemes is greater, there must be at least some graphemes consisting of two letters. However, borrowings from other languages automatically amplify the number of graphemes, and since English has other phonemes than the languages of origin, the exploitation of existing graphemes diversifies, i.e. some graphemes are polyphonic (used for different phonemes) and others synophonic (different graphemes used for the same phoneme). In Table 9 we see that the phoneme /AA/ can be represented by 19
Graphemic representation of English phonemes
33
synophonemic graphemes, and the polyphonemic grapheme can represent four phonemes. However, there is no absolute arbitrariness in assigning a grapheme to a phoneme, or lengthening a grapheme by adding more letters to it. If there were no restrictions to length, it would develop randomly according to a Poisson process. Now, the Poisson distribution Px = e−a ax /x! (x = 0, 1, . . .) can be represented by the recurrence formula a Px = Px−1 (4) x and its shape is determined by the parameter a. For a < 1 it is monotonically decreasing, for a = 1 it has two modes, and for a > 1 it is bell-shaped. The greater a, the longer the tail of the distribution. Evidently, a must be greater than 1 because the extent of phonetic changes, borrowing and graphemic conservatism in English produces a great number of graphemes. The first step in coping with this proliferation would lead to the exploitation of two-letter graphemes, but not each grapheme can be used to represent each phoneme. Some two-letter graphemes are not allowed. Hence some three-letter graphemes must be applied, etc. However, phonetic reasons are not the only causes restricting the proliferation of grapheme length. It is above all the requirement of economy (or optimality) which cares for balance in all domains of language (cf. Zipf 1935, 1949; Köhler 1986). Thus in formula (4) a more rapid convergence must be built in. Tentatively we replace the proportionality function a/x by the Zipfian function a/xb and obtain a Px−1 . xb
(5)
ax P0 , x = 0, 1, 2, . . . , (x!)b
(6)
Px = Solving (5) we obtain Px =
representing the Conway-Maxwell-Poisson distribution (cf. Wimmer & Altmann 1999) already used in linguistics, being a special case of the Wimmer-Altmann (2005) approach. Since there is no “zero-length” grapheme, we either solve (5) for x = 2, 3, . . . or we displace (6) one step to the right, in order to obtain ax−1 , x = 1, 2, 3, . . . , (7) Px = [(x − 1)!]b T ∞
where T = ∑
j=0
aj , ( j!)b
i.e. T is identical with (P0 )−1 .
34
Fan Fengxiang and Gabriel Altmann
Table 3: Grapheme length distribution and fitting the Conway-Maxwell-Poisson distribution (7) x
fx
1 2 3 4 5
26 159 94 8 2
NPx 25.18 162.86 88.98 11.45 0.54
a = 6.4689, b = 3.5656 χ2 = 0.73, DF = 1, P = 0.39
Applying (7) to our data we obtain the observed and computed values as given in Table 3. The normalizing constant is T = 11.4796. The result is presented graphically in Figure 1. Parameter a can be interpreted as the element of randomness (speaker creativity), conservatism of orthography, borrowing etc., while b means the braking mechanism, the balancing force of economy. 200
150 f(x) NP(x)
100
50
0
1
2
3
4
5
Figure 1: Grapheme length distribution: Conway-Maxwell-Poisson distribution (7)
5
Polyphonemics and synophonemics of graphemes
In Table 9 we see that a grapheme can represent several phonemes. Let us call this property graphemic polyphonemics. The great majority of graphemes is,
Graphemic representation of English phonemes
35
of course, monophonemic: it can be ascribed only to one phoneme. We distinguish direct ascription of a grapheme to a phoneme from the fact that a phoneme is part of a grapheme. E.g. /k/ is directly represented by <x> (as in “excel”) which differs from representing /k/ as part of ∈<x> (as in “affix”). Thus <x> as a framing grapheme can be ascribed to 6 phonemes and as a direct grapheme to 2 phonemes, i.e. it can be representative in 8 cases. The results of counting can be found in Table 4 (x = number of phonemes represented by a grapheme; fx = number of graphemes with x representations; NPx = computed number of graphemes with x representations). The numbers in the table are to be read as follows: There are 191 graphemes, each of which represents exactly 1 phoneme; there are 43 graphemes, each of which represents exactly 2 phonemes, etc. In order to set up a model, we simply start from the Zipfian assumption of setting (relative) frequency proportional to the frequency class using directly the function from (5), namely K , x = 1, 2, 3, . . . , (8) xb where K is the proportionality constant having the function of the normalizing constant, since we use (8) as a probability distribution. The parameter b is, again, a control parameter braking over-strong polyphonemy. Formula (8) is usually called Zipf’s law or zeta distribution. Now, theoretically (8) has an infinite support, which is a nonrealistic situation. For our purposes it will be truncated after x = 10 because no grapheme represents more than 10 phonemes (up to now or in this variant of English). Hence we obtain Px =
K , x = 1, 2, . . . , R, (9) xb R being the truncation parameter (here 10). Applying (9) to the data in Table 4 we obtain the result in its third column. The graphic display is in Figure 2. The fit is excellent but it can be made still simpler. In the last row of Table 4 we see that the value of parameter a is approximately 2. Replacing a = 2 in (8) we obtain the so-called Lotka distribution Px =
Px =
6 x2 π2
, x = 1, 2, 3, . . . ,
(10)
called also the ergodic distribution of population size (cf. Wimmer & Altmann 1999: 394). However, truncating it at the right side we obtain another
36
Fan Fengxiang and Gabriel Altmann 200
150 f(x) NP(x): (9) NP(x): (11)
100
50
0
1
2
3
4
5
6
7
8
9
10
Figure 2: Polyphonemics of English graphemes: right truncated zeta (9)
normalizing constant Px =
1 x2 [π2 /6 − Ψ(R + 1)]
, x = 1, 2, 3, . . . , R,
(11)
where Ψ (.) is the trigamma function (the normalizing constant is simply the sum of 1/x2 in the given definition domain). Using the right truncated Lotka distribution (11) we obtain the results in the last column of Table 4. The result of fitting is slightly better because we have one degree of freedom more. But more important is the fixed parameter value. Table 4: Graphemic polysemics in English x
fx
NPx (9)
NPx (11)
1 2 3 4 5 6 7 8 9 10
191 43 21 9 9 6 5 1 2 2
187.84 46.37 20.45 11.44 7.29 5.05 3.70 2.82 2.23 1.80
186.48 46.62 20.72 11.65 7.46 5.18 3.81 2.91 2.30 1.86
a = 6.4689, R = 10 X 2 = 3.09, DF = 7, P = 0.88
R = 10 X 2 = 3.12, DF = 8, P = 0.93
Graphemic representation of English phonemes
37
The graphemic synophonemics considers simply the numbers of graphemes representing an individual phoneme. Using Table 9 we get the following numbers in decreasing order: SS*60, 31, 29, 22, 21, 21, 20, 19, 19, 19, 18, 17, 17, 16, 16, 16, 13, 13, 12, 12, 10, 9, 9, 8, 8, 8, 8, 8, 7, 7, 7, 6, 6, 5, 5, 5, 4, 3, 2.* As can easily be seen, the frequencies of individual representations are rather uniformly distributed; they do not display the same pattern as graphemically “simpler” languages. The only possibility of searching for order is to consider their ranks as the independent variable. In that case, we obtain the data given in the first two columns of Table 5. Table 5: Rank-frequency distribution of English graphemic synophones Rank x
fx
NPx
Rank x
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
60 31 29 22 21 21 20 19 19 19 18 17 17 16 16 16 13 13 12 12
54.46 36.11 29.80 26.21 23.75 21.89 20.42 19.19 18.14 17.22 16.40 15.66 14.98 14.35 13.77 13.22 12.70 12.21 11.73 11.28
21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39
fx 10 9 9 8 8 8 8 8 7 7 7 6 6 5 5 5 4 3 2
NPx 10.84 10.41 9.99 9.58 9.18 8.77 8.38 7.98 7.57 7.17 6.75 6.33 5.89 5.42 4.93 4.39 3.79 3.07 2.10
K = 2.1178, M = 0.6708, n = 38 DF = 35, X 2 = 5.43, P ≈ 1.00
For ranking of linguistic units one usually uses the negative hypergeometric distribution or a distribution from the Lerch family containing also the Zipf, Zipf-Mandelbrot, zeta and other distributions (cf. Zörnig & Alt-
38
Fan Fengxiang and Gabriel Altmann 70 60 50 40
f(x) NP(x)
30 20 10 0
5
10
15
20
25
30
35
Figure 3: Synophonemics of English graphemes, negative hypergeometric
mann 1995; Köhler & Martináková-Rendeková 1998, Grzybek & Kelih 2003; Grzybek, Kelih, & Altmann 2004; Best 2005a,b,c). Here we adhere to the negative hypergeometric because its fitting turned out to be the best as can be seen in Table 5 and Figure 3, though zeta and Zipf-Mandelbrot both yield very satisfactory results. We use it in 1-displaced form Px =
6
M+x K −M+n−x x−1 n−x+1 , x = 1, 2, . . . , n + 1. K +n−1 n
(12)
Letter participation
The 26 letters of the Latin alphabet used in English are not exploited equally to build graphemes. Letters which in Latin had a vocalic value are used more often than those having consonant value. Again, for individual letters we get exact numbers but the only possibility to capture formally the set of nominal categories (letters) is to rank them according to their participation in graphemes. Since the set of letters is not too large, the best model is again the 1displaced negative hypergeometric distribution (12). For orientation and further research we present the letters not in alphabetic but in ranked order. The result of computing can be seen in Table 6 and Figure 4. The fit is excellent and corroborates once more the adequacy of this model.
Graphemic representation of English phonemes
39
Table 6: Letter participation in English graphemes: Fitting the 1-displaced negative hypergeometric distribution (12) Letter
Rank x
e u o h t a r s i l c g n
NPx (12)
fx
1 2 3 4 5 6 7 8 9 10 11 12 13
94 54 52 46 40 36 35 34 32 32 25 25 19
Letter
91.64 62.92 51.97 45.30 40.47 36.66 33.48 30.74 28.31 26.11 24.1 22.22 20.46
Rank x
p w d m y b f z k x q j v
14 15 16 17 18 19 20 21 22 23 24 25 26
fx
NPx (12)
17 14 13 12 11 11 10 9 8 6 5 3 3
18.80 17.21 15.68 14.21 12.78 11.39 10.02 8.66 7.32 5.98 4.61 3.21 1.73
K = 2.5453, M = 0.7095, n = 25, DF = 22, X 2 = 6.95, P = 0.9990 100 90 80 70 60
f(x) NP(x)
50 40 30 20 10 0
2
4
6
8
10
12
14
16
18
20
22
24
26
Figure 4: Letter participation in different graphemes: negative hypergeometric (12)
7
Weighted participation
A sightly different aspect is the evaluation of weighted participation of letters. In Section 6 we examined the participation of letters in building graphemes but we did not take into account the polyphonemy of graphemes, which is enormous in English. For example the letter occurs in 36 different graphemes but each of these graphemes can be used to represent different pho-
40
Fan Fengxiang and Gabriel Altmann
Table 7: Ranked weighted participation of letter in graphemes Letter
Rank x
fx
NPx (12)
e u o h a i t s r l c g y
1 2 3 4 5 6 7 8 9 10 11 12 13
169 110 100 86 85 67 48 46 44 36 35 33 23
169.89 112.03 90.44 77.36 67.95 60.56 54.45 49.21 44.60 40.48 36.75 33.32 30.15
Letter
Rank x
fx
NPx (12)
w p n d z m x b f k j q v
14 15 16 17 18 19 20 21 22 23 24 25 26
21 20 20 17 14 13 13 11 11 8 6 5 4
27.20 24.44 21.85 19.41 17.09 14.90 12.82 10.85 8.97 7.20 5.52 3.94 3.62
K = 2.8339, M = 0.6885, n = 26, DF = 22, X 2 = 14.64, P = 0.88
nemes, i.e. it can be polyphonemic. In this section we consider all occurrences of individual letters in graphemes, i.e. we compute the weighted participation of letters. The results are presented in Table 7. Again, the number of cases is too small (26) and, since almost all letters have different participation, no model could be set up. Hence we use, as above, the ranked weighted participation and as expected we obtain again the 1-displaced negative hypergeometric distribution. The result of fitting is graphically displayed in Figure 5. 200 190 180 170 160 150 140 130 120 110 100 90 80 70 60 50 40 30 20 10 0
f(x) NP(x)
2
4
6
8
10
12
14
16
18
20
22
24
Figure 5: Ranked weighted participation of letters in graphemes (NHG)
26
Graphemic representation of English phonemes
8
41
Letter utility
The last aspect we shall analyze here is the so-called letter utility. In the previous sections we considered the presence of a letter in different graphemes and its presence in all graphemes; here we consider its position in the grapheme. We take into account only different graphemes and ignore their polyphonemics. The position of a letter in a grapheme is a measure of its relevance. The earlier the letter appears in the cluster the more it contributes to its phonetic value. At least this can be assumed to hold in general (although it does not hold in each case). This is especially well expressed in French where the morphology represented by inflections is dying out and the first letter of a long grapheme is decisive for the phonetic form, e.g. <parlent> → /parl/. Let us illustrate the problem using the graphemes containing , namely < cq, cqu, q, qu, que > . Let nq be the set of graphemes containing and |n | the cardinal number of this set. Let wx be the weight of the letter <x> given by its position in the grapheme. We define first PP<x> =
∑
(13)
wx
x∈n<X>
as the sum of all weights (positions) of <x> in the graphemes of the set nx . For the letter we obtain from the above example nq = 5 and PP = 2 + 2 + 1 + 1 + 1 = 7. For comparative purposes we define PP<x> =
1
∑ |n<x> | x∈n <x>
wx .
(14)
In our example we obtain PP = 7/5 = 1.4. The results for all letters are given in Table 8, #G denoting the number of graphemes and MLU mean letter utility. Ordering the letters according to their utility in graphemes we obtain the order: . It would be possible to order the letters according to their absolute weight, too. The mean utility of all letters can be expressed as the ratio of the sum of the second column of Table 8 to the sum of the third column, i.e. 1231/687 =
42
Fan Fengxiang and Gabriel Altmann
Table 8: Letter utility in English graphemes Letter a b c d e f g h i j k l m
Weight
#G
MLU
Letter
59 17 40 19 201 17 44 100 57 5 10 87 21
38 13 27 15 97 13 27 45 36 4 8 37 14
1.5526 1.3077 1.4818 1.2667 2.0722 1.3077 1.6296 2.2222 1.5833 1.2500 1.2500 2.3514 1.5000
n o p q r s t u v w x y z
Weight
#G
MLU
31 81 26 7 103 75 69 91 4 24 10 19 14
21 54 19 5 45 39 41 47 3 14 5 11 9
1.4762 1.5000 1.3684 1.4000 2.2889 1.9231 1.6829 1.9362 1.3333 1.7143 2.0000 1.7272 1.5555
1.7918, or as the ratio of the sum of the second column to the number of letters, i.e. 1231/26 = 47.35. Comparing this last number with Italian where the ratio is 6.48, (cf. Bernhard & Altmann, this volume, pp. 13ff.), one sees that the diversification of graphemics in English is enormous.
9
Conclusions
English graphemics is a very complex matter. Some of the measures (indices) introduced here make it evident. They differ drastically from those in other languages. The loss of a unique phonetic value of a letter reduces letters to merely graphical signs obtaining a phonetic value only in a grapheme. The way to hieroglyphism is open. All indices introduced here can be analyzed further statistically. They have their sampling distributions, asymptotic tests can be set up, languages can be classified according to their graphemics, and there is a possibility to find interrelations among all these properties and also between graphemic and nongraphemic properties. The last aim is, of course, to find laws of graphemics and join them in a system of laws, i.e. in a theory. At present, any such enterprise would be premature.
Graphemic representation of English phonemes Table 9: English phoneme-grapheme correspondences Phoneme
Graphemes
Frequency
Examples
/AA/
<e> <ea> <eau> <eah>
1412 3 14 18 1 1 31 6 16 20 2 1 15 1 2427 7 2 4 8 4859 2 1 11 9 1 1 5028 1 7 20 36 10
(a)bo baz(aa)r y(ah) (al)mond arkans(as) baccar(at) astron(au)t (aw)ful s(er)geant wholeh(ea)rtedly bur(eau)cracy exh(au)stively (ho)nors l(i)ngerie abd(o)minal j(oh)n s(ol)der c(ou)gh ackn(ow)ledgement ab(a)ck, zigz(a)gging g(ah)n, p(ah) pl(ai)d beh(al)f, s(al)mon (au)nt, l(au)ghter y(eah) chop(i)n (a)bide, mad(a)m is(aa)c an(ae)sthesia, minuti(ae) abdall(ah), tor(ah) barg(ai)n, vill(ai)ns (au)gusta, nickl(au)s
/AE/
/AH/
(continued on next page)
43
44
Fan Fengxiang and Gabriel Altmann
Table 9 (continued from previous page) Phoneme
Graphemes
Frequency
Examples
<e> <ea> <eau> <ei> <eo> <eou> <er> ∈ ∈ ∈ ∈ ∈ ∈ ∈
3619 24 2 7 13 7 2 9 3 8 2379 185 36 1349 66 1 2680 1 5 2 1 17 317 1 2864 1 3 5 53 414 32 5 21 20 43 31
abandonm(e)nt, zab(e)l chang(ea)bl, veng(ea)nce bur(eau)crat, bur(eau)crats for(ei)gn, surf(ei)t bludg(eo)n, surg(eo)ns advantag(eou)s, right(eou)sness paraph(er)nalia, res(er)voir budd(ha), wind(ha)m ve(he)mence, ve(he)mently anni(hi)lation, pro(hi)bition abdom(i)nal, zoolog(i)st acac(ia), venet(ia)n anc(ie)nt, trans(ie)nt abduct(io)n, volit(io)n ambit(iou)s, vivac(iou)s belg(iu)m aband(o)n, zool(o)gy mendelss(oh)n conn(oi)sseur, tort(oi)se linc(ol)n, norf(ol)k m(on)sieur bl(oo)d, fl(oo)d adulter(ou)s, zeal(ou)s mccull(ough) abd(u)ction, y(u)m etiq(ue)tte br(uh)n, (uh) bisc(ui)t, circ(uit)s anal(y)ses, vin(y)l able, ambling babbled, bubbling subtler, subtly article buckled beadles, idling, haydn addle, huddling (continued on next page)
Graphemic representation of English phonemes Table 9 (continued from previous page) Phoneme
/AO/
Graphemes
Frequency
Examples
∈ ∈ ∈ ∈ ∈ <m> ∈ ∈ ∈ ∈ <s><m> ∈ <sc> ∈ <ss> ∈ <st> ∈ ∈ ∈ <m> ∈ ∈ <x> ∈ <ea> <eo>
7 13 24 12 10 9 42 18 146 3 1 25 19 44 5 251 2 24 344 1 24 1 257 1 132 3 1 5 7 2 2 1487 54 16 1 72
rifle baffle bedraggled ankle mcalister one ample apple activism muscle tussle apostle beetles battle logarithm, algorithm accum(u)lated axle dazzled (a)lbany, y(a)lta ut(ah) b(al)ked, w(al)kways extr(ao)rdinary appl(au)d, v(au)lts v(augha)n (aw), y(aw)ning (awe)some, dr(awe)rs s(ea)n g(eo)rgia, g(eo)rgetown ex(hau)st, inex(hau)stible ex(ho)rtations, ex(ho)rting dig(io)rgio, g(io)rgio abh(o)rrent, zl(o)tys ab(oa)rd, washb(oa)rd d(oo)r, outd(oo)rs f(ore)runner betanc(ou)rt, y(ou)r (continued on next page)
45
46
Fan Fengxiang and Gabriel Altmann
Table 9 (continued from previous page) Phoneme /AW/
/AY/
/EH/
Graphemes
Frequency
Examples
<ei> <ey> <eye> <e> <ea> <ee> <eh> <ei> <eo> <er> <ette>
1 6 32 3 486 3 219 2 1 21 6 2 80 5 10 2293 2 14 22 2 6 6 342 5 494 12 120 4 4955 257 1 2 4 5 2 1
unt(owa)rd l(ao), t(ao)ists aden(au)er, t(au)ssig (hou)r, (hou)rs ab(ou)nd, whereab(ou)ts b(ough), h(ough)s all(ow), y(ow) d(owe)r, h(owe) m(ae)stro al(ai), th(ai)land b(ay)ou, sant(ay)ana (aye), (aye)s alam(ei)n, z(ei)tler ch(ey)enne, m(ey)ers bug(eye)d, (eye)witness ab(i)des, wr(i)ting d(ia)mond, d(ia)monds l(ie), unt(ie) h(igh), th(igh) c(oy)ote, c(oy)otes beg(ui)led, disg(ui)se b(uy), sch(uy)ler acol(y)te, wr(y)ly b(ye), r(ye) actu(a)rial, y(a)rrow (ae)rial, kr(ae)mer ad(ai)r, volt(ai)re pr(ay)er, s(ay)s ab(e)d, z(e)st abr(ea)st, z(ea)lous k(ee)lson g(eh)rig, k(eh)le l(ei)sure, th(ei)rs j(eo)pardizing, l(eo)pards int(er)rogation, int(er)rogator pirou(ette) (continued on next page)
Graphemic representation of English phonemes Table 9 (continued from previous page) Phoneme
/ER/
/EY/
Graphemes
Frequency
Examples
<ey> <ar> <arr> <ear> <er> <ere> <err> <eur> | |