Showing posts with label genetics. Show all posts
Showing posts with label genetics. Show all posts

Tuesday, October 11, 2011

The Maasai of Kenya

The Maasai are an ethnic group of semi-nomadic people located in Kenya and northern Tanzania. They are among the best known of African ethnic groups, due to their distinctive customs and dress and residence near the many game parks of East Africa. I came across these four Maasai tribesmen in Nairobi National Park, Keyna.

Maasai DNA

Recent studies of Maasai DNA reveal a complex history of the Maasai people. In 2005, Elizabeth Wood, et al., sampled the Y-DNA of 26 Maasai tribesmen, and from this sample determined that 50% of the individuals had Y-DNA of Haplogroup E1b1b1 (M35), 27% had A3b2 (M13), 16% had E1b1a1 (M2), and 8% B2a (M150). An explanation of each, in the context of the history of the Maasai, follows.

Eastern Sudanic Ancestry

The largest Y-DNA found among the Maasai sample, E1b1b1, represents Eastern Sudanic ancestry (although it is also associated with certain other "Nilo-Saharan" populations). One would expect to see a high percentage of this haplogroup based on the fact that the Maasai speak a Nilotic (Eastern Sudanic) language. However, based on the fact that only 50% of their DNA is Eastern Sudanic, it would appear that the Maasai are remnants of an older culture whose language went extinct after prolonged contact and interbreeding with Nilotic peoples. Examining the other components of their Y-DNA is the key to understanding the origin of the Maasai.

Niger-Congo Ancestry

Haplogroup E1b1a1 represents a component of the Masaai ancestry common among most Sub-Saharan populations, indicating Niger-Congo ancestry. The most likely source of this haplogroup is the Bantu expansion, whereby Bantu-speaking peoples spread across Sub-Saharan Africa from East to West, generally. However, this haplogroup does not reveal the ancient origin of the Masaai people.

Pygmy and Khoisan Ancestry

It would appear that prior to the Bantu expansion and the southward migration of the Nilotic peoples, the Maasai were a combination of Pygmy and Khoisan people, assuming that haplogroup B is the African Pygmy modal haplogroup and A3b is the Khoisan or "Bushman" modal haplogroup. The deep components of Masaai DNA suggest an admixture between the most divergent branch of the Pygmies and the most divergent branch of the Khoisan. Most Pygmy populations exhibit a significant percent of the B2b clade, whereas the Maasai exhibit the B2a sister clade. Most Khoisan peoples exhibit a significant percent of the A3b1 clade, whereas the Maasai exhibit the A3b2 sister clade. It would appear that the distant ancestors of the early Maasai (excluding recent Supra-Saharan admixture) were a mixture of members of two ancient African hunter-gatherer cultures. The spread of Nilotic and Bantu culture wiped out the entire language families of the component populations, although the Maasai maintained much of their ancient nomadic culture.

Maasai DNA Reveals Origins of Sandawe People

Until recently, the Sandawe people, although living in Tanzania, were considered to be closely related to the Khoisan ethnicities of the Kalahari desert. Much of this presumption was based on the fact that like the Kalahari Bushmen, the Sandawe speak a click language. However, in 2007, Sarah Tishkoff, et al., conducted a study on the Y-DNA of 68 Sandawe people. The results were more complicated than even the Maasai. However, excluding the 56 individuals who tested for Supra-Saharan Y-DNA (Eastern Sudanic and Niger-Congo haplogroups described above), the remaining 12 individuals exhibited 72% haplogroup B2b (M112), 22% haplogroup A3b2 (M13) and 6% B2a (M150). Note the 22:6 ratio of A3b2 to B2a, comprared with the nearly identical 27:8 ratio among the Maasai. This suggests that the Sandawe were originally B2b pygmies, with subsequent admixtures from (not necessarily in this order) the Maasai, Bantus and Nilotics. Unlike the Maasai, however, the Sandawe retained their ancient click language. Since the Hadza people of Tanzania are the nearest B2b click-speaking tribe, one could presume that the language of the Sandawe is distantly related to the Hadza language, both languages perhaps descending from a common language spoken by the original Pygmy who underwent the B2b Y-DNA mutation. That would mean that the only extant "true" A3b1 Khoisan languages are those click-languages spoken in and around the Kalahari desert.

Implications Beyond Sub-Saharan Africa

According to the most recently accepted version of the mt-DNA phylogenetic tree, it is believed that the first split occurred separating the Khoisan L0 clade from the L1-6 superclade (representing the founding populations of the speakers of all non-Khoisan languages). In the non-Khoisan superclade (represented by haplogroup BT in the Y-DNA phylogenetic tree), the first node appears to separate the African Pygmies from the ancestor of the speakers of all non-Khoisan and non-pygmy languages. This split corresponds to the split between B and CT in the Y-DNA tree. It is interesting that both the haplogroup A Khoisan languages and the haplogroup B Pygmy languages are click languages, and that CT (which is downstream from BT) is the only clade originating in this time period not dominated by click languages. This is some evidence, although not conclusive, that there may have been at one time a Proto-World language that had clicks among its sound inventory, ancestral to all modern languages, including the CT languages (including Enlgish, for example). That is to say, perhaps around 75,000-100,000 years ago, all languages had clicks, and the Supra-Saharan CT branch lost its clicks. An analogy would be the English language having lost grammatical gender despite its Indo-European origin.

References

History of Click-Speaking Populations of Africa Inferred from mtDNA and Y Chromosome Genetic Variation. Tishkoff, Sarah A. et al 2007.

Contrasting patterns of Y chromosome and mtDNA variation in Africa: evidence for sex-biased demographic processes. Wood, Elizabeth T et al 2005.

Sunday, October 2, 2011

Basque Y-DNA

Basque Y-DNA

The above chart includes data from 162 male volunteers who submitted their Y-chromosomal DNA results to Family Tree DNA's Basque DNA project. Individuals who submitted their Y-DNA results claim to be of direct male Basque descent. Contributing volunteers included residents of Europe, Asia, and North and South America. Analysis of the data reveals that 71.6% of the participants in this study carry Y-DNA of haplogroup R1b1 and its subclades.

Basque People

The Basques as an ethnic group, primarily inhabit an area traditionally known as the Basque Country, a region that is located around the western end of the Pyrenees on the coast of the Bay of Biscay and straddles parts of north-central Spain and south-western France. Since the Basque language is unrelated to Indo-European, it is often thought that they represent the people or culture who occupied Europe before the spread of Indo-European languages there.

Y-DNA in the Field of Linguistics

Y-DNA haplogroup testing is a valuable tool in the study of historical linguistics. Y-DNA is carried from father to son, and mutates at a somewhat predictable rate. Haplogroups are clades of DNA types that share a distinct defining mutation or mutations. Each such mutation occurred in a single person at some point in the past. Since the person in which that mutation occurred necessarily spoke a language (at least in the time-frame of the past 50,000 years or so), and since a large percentage of children learn to speak the same native language as their father, one can use Y-DNA studies to track the historical evolution of languages, and uncover ancient relationships between living language families, to a surprising degree of accuracy.

When using Y-DNA as a tool to discover relationships between living language families, however, one must take into account the fact that there are several reasons why children may not learn to speak the native language of their fathers. The most obvious of such a situation is when the father either moves to a region that speaks a different language or has a child in a region where another language is dominant in addition to his native language (i.e., a more dominant local language is taught in schools, used in business, etc.), and rather than learning the native language of the father, children adopt the local language.

The goal in interpreting the data from this study is to determine which, if any, of the individuals whose Y-DNA first contained the defining mutations of the haplogroups, may have spoken a language ancestral to modern Basque, i.e. "Ancient Basque."

Haplogroup E Among Basque People

93.8% of those tested reported haplogroups of Eurasian origin, whereas 6.2% reported haplogroup E and its subclades. Haplogroup E is common among ethnic groups which originated along northern portions of the Nile River in Africa, including speakers of Nilo-Saharanm, Niger-Congo, Mande and certain Afro-Asiatic languages. The infusion of Y-DNA haplogroup E among the Basque population likely took place long after speakers of the ancestral Basque language arrived on the Iberian Peninsula, perhaps after the Afro-Asiatic speaking Moors invaded southern Europe. That is to say, the Basque language is unlikely to share a common origin with the Nilo-Saharan or Afro-Asiatic language families, although males of Northern African descent who migrated north to the Iberian Peninsula apparently interbred with women of the Basque population, perhaps influencing the local "Ancient Basque" language, but not replacing it.

Outliers in Haplogroups L, O and Q

Of the 162 individuals tested, there was one individual who carried haplogroup L, one who carried haplogroup O, and one who carried haplogroup Q. These haplogroups are of Eurasian origin, but are probably not associated with speakers of Ancient Basque. The individual who carried haplogroup Q resided in China and reported that his most distant known paternal ancestor resided in Mexico and had a Basque surname. It should be noted that Y-DNA haplogroup Q is the most common haplogroup among native (non-European) Mexicans, and not likely indicative of a Basque origin, despite the Basque surname. Haplogroups L and O generally correspond with South Asian and Asiatic languages, respectively. As the present-day Basque language shares little in common with members of these well-studied language families, the individuals likely represent a very small segment of the Basque population who descend from recent (less than 5000 years ago) immigrants to the Basque country from southern and eastern Asia.

Haplogroup R1a1 on the Iberian Peninsula

While haplogroup R1a1 appears in this sample at a percentage of 3.7%, that rate is similar to, if not less than, the occurrence of R1a1 in surrounding regions. R1a1 likely corresponds to DNA of native speakers of Indo-European languages, who settled the Iberian Peninsula and likely wiped out all recent branches of the Ancient Basque language with the exception of the languages of the Basque country. Haplogroup R1a1 is found in all locations where Indo-European languages are spoken, and the person in which its defining mutation occurred likely spoke a language ancestral to Indo-European, not to Basque.

Northwest Caucasian Haplogroup G

Y-DNA haplogroup G has not been definitively associated with any living language family, although it is common among speakers of Northwest Caucasian languages. The language of the progenitor of haplogroup G may only be manifested in the Northwest Caucasian substrate which differentiates the Northwest from the Northeast Caucasian languages. Since there are unlikely any modern surviving languages that descend directly from the language spoken by the progenitor of haplogroup G, it is difficult to rule out the haplogroup as corresponding to an Ancient Basque precursor. However, due to the overwhelming majority of haplogroup R1b1 (which shares the same lack of known modern surviving descendant languages) among the Basque population, it seems logical that many carriers of haplogroup G may have spoken a Vasconian (pre-Basque) language (perhaps since as long as 10,000 years ago), after native Vasconian speaking carriers of R1b1 dominated their native culture, perhaps shortly after the last ice age.

Cro-Magnon Haplogroup IJ

10.5% of the sample reported haplogroups of either I or J, both haplogroups that represent subsequent mutations from an earlier Cro-Magnon haplogroup IJ, which appears to have originated in the Caucasus. This is a large percentage of the sample that cannot be discounted. The progenitor of haplogroup J may have spoken a language ancestral to the Northeast Caucasian and Kartvelian language families. The progenitor of haplogroup I spoke a language belonging to an extinct family that may only have modern observable manifestation in the substrate of vocabulary found in the Germanic languages (approximately 1/3 of the lexicon) that is not traceable to Proto-Indo-European origin. While neither haplogroups I nor J can be definitively ruled out as corresponding to the Basque language, it appears that the Basque language does not share much in common with the Northeast Caucasian or Kartvelian languages, nor have I found any source that suggesting that its lexicon overlap the Proto-Germanic substrate.

Vasconian Haplogroup R1b1

Based on the data from this study in a vacuum, it seems very likely that if the progenitors of any of these haplogroups spoke a Vasconian language, it should be R1b1, as R1b1 accounts for 71.6% of the sample population. However, looking outside this study, R1b1 is equally common among most Indo-European speaking (Spanish and Portuguese, e.g.) populations of the Iberian peninsula, and almost as common in regions to the east where other Romance languages are spoken such as French and Italian. Perhaps remnants of the language spoken by the progenitor of R1b1 can be found by studying the differences between the Romance (Italic) brancih of the Indo-European languages from other Indo-European subfamilies. I suspect some of the differences may be accounted for by a Vasconian substrate that represents linguistic elements of other languages descended from Vasconian that may have been spoken by populations who assimilated into the western Indo-European culture, and adpoted Indo-European as their language. I believe this hypothesis is more sound than a IJ origin of the Basque language, based on geographic data on the present location of haplogroup R1b1 vs. haplogroups I and J. For example, haplogroup I is distributed widely in Scandinavia and in regions where Germanic languages are spoken. It seems likely that if there was a living descendant language of the language spoken by the progenitor of haplogroup I, it would have the highest likelihood of surviving in Germanic speaking Europe, not in the Pyrenees Mountains where R1b1 y-DNA is dominant among Italic speakers who (if they inherited their language from their ancestors rather than by assimilation) would be expected to have R1a1 Indo-European DNA.

Basque Language Isolate

One might ask, if the Basque language is associated with R1b1, and the Indo-European languages are associated with R1a1, (both clades of R1), why is the modern Basque language so different from all of its closest genetic relatives? However, consider, for example, the incredible difference between the English and Hindi languages (both Indo-European) which probably only diverged from their most common ancestor about 5,000 years ago. If not for available linguistic data from the numerous other languages in the Indo-European family, one might be highly skeptical about their common origin, especially due to the geographic distance where the two languages are spoken, and the differences in culture, appearance, religions, etc., between the populations by whom they are spoken. One must keep in mind that R1a and R1b diverged from their common y-DNA ancestor R1 approximately 18,000 years ago, and there are no intermediate languages on the R1b side that survived to modern times, with the possible exception of Basque. It should not be surprising that Basque seems completely foreign to the Indo-European languages, in this context.

Furthermore, even if the progenitor of R1a spoke an ancient Indo-European language and the progenitor of R1b spoke an ancient Vasconian language, that does not necessarily imply that the two ancient languages were closely related. The progenitor of R1 (ancestral to R1a and R1b) probably lived in Siberia some 25,000 to 30,000 years ago, approximately 10,000 years prior to the mutations that occurred that created the subclades of R1a and R1b. In those 10,000 years, the descendants of the progenitor of R1 may have come to speak many languages unrelated to the the native language of their ancestor by means of assimilation as they migrated across Asia. That is to say, while R1a and R1b carriers are undoubtedly genetically related, an inter-disciplinary approach including efforts by experts in the fields of anthropology, archaeology, linguistics, genetics, philology and other sciences, is required to prove or disprove any ancient relation between the Basque language and the Indo-European languages. Also, an examination of the Burushoski language spoken by (among others) descendants of R1's closest relative R2 may provide some evidence helpful to determining the origin of the Basque language. Unfortunately, Burushoski, spoken in portions of present-day Pakistan, is also a language isolate.

Related Reading

For what they were... we are: Linguistic musings: Basque and Proto-Indoeuropean

Diagram, research and analysis by Kevin Borland. Data provided by Family Tree DNA. Text of subsection "Basque People" derived from Wikipedia.