Sunday, November 29, 2015

Special Relativity and General Relativity Explained With A Ten Hundred Word Vocabulary

A quite lucid explanation of special relativity and general relativity with a limited vocabulary by Randall Munroe of xkcd fame has been reprinted at the New Yorker. A taste:
There once was a doctor with cool white hair. He was well known because he came up with some important ideas. He didn’t grow the cool hair until after he was done figuring that stuff out, but by the time everyone realized how good his ideas were, he had grown the hair, so that’s how everyone pictures him. He was so good at coming up with ideas that we use his name to mean “someone who’s good at thinking.”

Two of his biggest ideas were about how space and time work. This thing you’re reading right now explains those ideas using only the ten hundred words people use the most often. The doctor figured out the first idea while he was working in an office, and he figured out the second one ten years later, while he was working at a school. That second idea was a hundred years ago this year. (He also had a few other ideas that were just as important. People have spent a lot of time trying to figure out how he was so good at thinking.)

The first idea is called the special idea, because it covers only a few special parts of space and time. The other one—the big idea—covers all the stuff that is left out by the special idea. The big idea is a lot harder to understand than the special one. People who are good at numbers can use the special idea to answer questions pretty easily, but you have to know a lot about numbers to do anything with the big idea. To understand the big idea—the hard one—it helps to understand the special idea first.


Saturday, November 28, 2015

Ancient DNA from Neolithic Greece and Anatolia

Investigators have managed to get right to the source of the European Neolithic Revolution with ancient DNA from Greece and Anatolia, which is described in two new papers.

Like all other Early European Farmers, genetically, they are quite similar to modern day Sardinians and lack the European and Central Asian steppe contribution that is ubiquitous in most modern European populations.

The Y-DNA in the Neolithic Greeks and Anatolians was, unsurprisingly G2a2.  The mtDNA of two Mesolithic individuals was K1c and the mtDNA of two Neolithic individuals was K1a2.

Of course, their genes look quite unlike modern Greeks and Turks.

Monday, November 23, 2015

Ancient And Modern Ideograms

The Language Log blog highlights the research of Genevieve Von Petzinger into the use of symbols that may be ideograms as a form of written proto-language in Upper Paleolithic cave art as described in a linked TED talk that she has given.

In a somewhat related point, somewhere on my blogging "to do" list is the task of figuring out how many ideograms are in common used by readers of American English, ideally accompanied by examples of the use of ideograms to convey a sound associated with the word for which the ideogram stands (in the tradition of military use of words like Victor or Charlie to spell out words in situations where radio transmission quality is poor) in languages that are predominantly ideographic.

The point, of course, would be to illustrate that the line between phonetic writing systems and ideographic writing systems is one of degree, rather than being an all out either/or alternative.

Some of the common ideograms in American English include:

@ "at"
#  "pound
$ "dollar"
% "percent"
& "and"
~ "approximately"
> "greater than"
< "less than"
= "equals"
+ "plus"
- "minus"

There are many other ideograms familiar to readers of American English that can't be produced with a single keystroke.

There are thousands of Chinese ideograms in the proto-typical ideogram based writing system (Coptic hieroglyphics being another such language).  But, I'd guess that the number of ideograms widely understood by readers of American English.

Some fields, such as mathematics, make particularly heavy use of ideograms which are typically global in reach across the lines of the languages of the people who use them.

Saturday, November 21, 2015

Cosmology Based Limits On The Sum Of The Three Neutrino Masses Tighten

The Data

The sum of the three neutrino masses is less than or equal to 110 meV at a 95% confidence level, according to the latest effort to integrate multiple sources of astronomy data including the Planck satellite cosmic microwave radiation background data, the WiggleZ Dark Energy Survey, the Sloan Digital Sky Survey -Data Release 7 (SDSS-DR7) sample of Luminous Red Galaxies (LRG), and Baryon Acoustic Oscillation (BAO) data.

This is confirmed by an independent analysis of the Lyman-α power spectrum from BOSS cited in the paper, which limits the sum of the three neutrino masses to less than or equal to 120 meV at the 95% confidence level.

Background

In an inverted hierarchy of neutrino masses, the minimum sum of the three neutrino masses given current neutrino oscillation data is around 98 +/- 1 meV.

As previously noted at this blog, the state of the art measurements of the difference between the first and second neutrino mass eigenstate is roughly 8.66 +/- 0.12 meV, and the difference between the second and third neutrino mass eigenstate is roughly 49.5 +/- 0.5 meV, which implies that the sum of the three neutrino mass eigenstates cannot be less than about 65.34 meV with 95% confidence.

The hypothesis that there are more than three neutrinos that oscillate with each other has also been largely ruled out by experimental data.

Hypothesis Testing

Taken together, this data favors the normal neutrino mass hierarchy over the inverted neutrino mass hierarchy, even though the inverted neutrino mass hierarchy is not yet ruled out at a 95% confidence interval.

The body text of the paper does not expressly compare the relative likelihoods of the minimal mass normal hierarchy hypothesis to the minimal mass inverted hierarchy hypothesis, but, eyeballing the graphs with the pertinent data, it appears that the normal hierarchy is many times more likely than the inverted hierarchy.

Future Experiments

We need only about a 10% improvement in measurement precision to rule out the inverted neutrino mass hierarchy at the 95% confidence level.  We shouldn't be surprised that the neutrinos appear to have a normal mass hierarchy, as this is what we observe in the up-type quarks, the down-type quarks, and the charged leptons as well.

Knowing whether the neutrino mass eigenstates have a normal or inverted hierarchy also increases the certainty with which we can interpret neutrino oscillation data, for example, in an effort to determine the CP violating phase (if any) of neutrino oscillations.

The tight bound on the absolute neutrino masses discussed below also sets firm expectations regarding the expected amount of neutrinoless double beta decay in any particular model that has Majorana mass neutrinos.  Current experimental precision needs to improve by more than a factor of ten before it can meaningfully distinguish between Dirac and Majorana mass scenarios.

Bounds On Absolute Neutrino Masses Given What We Know

If the neutrinos do have a normal hierarchy, then the experimental bounds on the three neutrino mass eigenstates at the two sigma level based upon the latest data is:

v1: 0 meV to 12 meV              absolute precision +/- 6 meV
v2: 8.42 meV to 21.9 meV      absolute precision +/- 6.74 meV
v3: 56.92 meV to 72.4 meV    absolute precision +/- 7.74 meV

Any higher masses would violate the 110 meV upper bound on the sum of the three neutrino mass eigenstates.  So, while absolute neutrino mass has been called an "unsolved problem" we are tantalizingly close to determining it with a precision that is stunning in absolute terms and compares favorably with the precision with which we know the light quark masses in relative terms.

The bound between the minimum and maximum neutrino mass ranges in an inverted mass hierarchy is currently about 4.7 meV (i.e. +/- 2.35 meV for the lightest of the three).  If the neutrinos do indeed have an inverted mass hierarchy, the bounds upon the absolute masses are tight indeed.

An Unambitious Neutrino Mass Prediction

Given all of the data and the patterns that we see, I personally would be surprised to see anything other than a normal neutrino mass hierarchy with a lightest neutrino mass eigenstate of 2 eV or less, with a mass of 1 meV or less most favored. If that hypothesis is correct, then the neutrino masses would be:

v1: 0 meV to 2 meV          
v2: 8.42 meV to 11.9 meV    
v3: 56.92 meV to 62.4 meV

and the sum of the three neutrino masses would be not more than 76.3 meV (and not less than 65.34 meV).

Footnote On Inflation Constraints

In other news, the constraints on evidence of certain kinds of gravitational waves in the early universe which are predicted by inflation theories is also tightening considerably.  Generally speaking, this data tends to rule out many of the more elaborate inflation scenarios.

Hollywood Habits Reach Particle Physics

Usually, the abstract of a scientific journal article simply summarizes what the article says in a single paragraph (some longer, some shorter).  The main distinction among abstracts is between those that actually put their key conclusion in the lede, and those that bury the lede and tell you what the conclusion is about while forcing you to actually read the paper to get that result.

But, taking a cue from Hollywood, a researcher from the ALICE collaboration has gone one step further.  After the usual abstract telling you what the current paper says, she throws in one more sentence about coming attractions which are not actually included in the paper, stating: "Recent results obtained from these measurements will be presented and the measured cross sections will be compared to perturbative Quantum Chromodynamics calculations at next-to-leading order."

Now, admittedly, maybe she just means that they will present the results in this paper.  But, in context, it reads more to me like an announcement of an upcoming conference paper rather than a description of what is currently being presented.

Wednesday, November 18, 2015

New Denisovan DNA

Until now, we had one individual's Denisovan autosomal DNA and Denisovan mtDNA from two individuals.  Renanalysis of one of the teeth used the first time around and ancient DNA from a third morphologically similar tooth has given us new data.

We now have (partial) autosomal DNA from two more Denisovan individuals (the one who was the source of the mtDNA sample and a new individual) and a new set of Densiovan mtDNA (with a different haplogroup than the other two samples) from the same new individual.

This new DNA data confirm that all three teeth comes from individuals of the same Denisovan species of archaic hominin.

Denisovans appear to be a sister clade of archaic hominins to Neanderthals and all three samples come from a single cave in Siberia.  But, significant Denisovan admixture (in addition to the ordinary amount of Neanderthal mixture for East Eurasians) is present in modern humans with Australian Aboriginal ancestry or Papuan ancestry.  There may also be an independent source of Denisovan admixture in some Asian Negrito populations (e.g. in the Philippines, but not in the Andamanese) and possibly some very slight traces of Denisovan ancestry in modern humans in Southeast Asia and East Asia.  The introduction to the paper notes that:
In 2008, a finger phalanx from a child (Denisova 3) was found in Denisova Cave in the Altai Mountains in southern Siberia. The mitochondrial genome shared a common ancestor with presentday human and Neandertal mtDNAs about 1 million years ago, or about twice as long ago as the shared ancestor of present-day human and Neandertal mtDNAs. However, the nuclear genome revealed that this individual belonged to a sister group of Neandertals. This group was named Denisovans after the site where the bone was discovered. Analysis of the Denisovan genome showed that Denisovans have contributed on the order of 5% of the DNA to the genomes of present-day people in Oceania, and about 0.2% to the genomes of Native Americans and mainland Asians. 
In 2010, continued archaeological work in Denisova Cave resulted in the discovery of a toe phalanx (Denisova 5), identified on the basis of its genome sequence as Neandertal. The genome sequence allowed detailed analyses of the relationship of Denisovans and Neandertals to each other and to present-day humans. Although divergence times in terms of calendar years are unsure because of uncertainty about the human mutation rate, the bone showed that Denisovan and Neandertal populations split from each other on the order of four times further back in time than the deepest divergence among present-day human populations occurred; the ancestors of the two archaic groups split from the ancestors of present-day humans on the order of six times as long ago as present-day populations. In addition, a minimum of 0.5% of the genome of the Denisova 3 individual was derived from a Neandertal population more closely related to the Neandertal from Denisova Cave than to Neandertals from more western locations .
The abstract also notes that:
The mtDNA of Denisova 8 is more diverged and has accumulated fewer substitutions than the mtDNAs of the other two specimens, suggesting Denisovans were present in the region over an extended period. The nuclear DNA sequence diversity among the three Denisovans is comparable to that among six Neandertals, but lower than that among present-day humans. 
All of this is pretty much what we would expect from additional Denisovan DNA samples and none of them answer the big unsolved questions we have regarding the Denisovans, but it is still nice to have the additional data.

The one point I would add is that the more basal nature of the mtDNA from Denisovan 8 is used to argue that this tooth is much older and represents a prolonged occupations of the site.  This is not a necessary interpretation, or even, in my humble opinion, a likely one.

It is common for particular individuals in modern human populations living at the same time, to have both more basal and less basal mtDNA.  For example, in the same village in Nigeria, there might be one individual with mtDNA which most recently mutated 1,000 years ago, and another individual with mtDNA that most recently mutated 40,000 years ago.

This is an elementary inference from the apparent common mitochondrial origin of all hominins, and the fact that mutations happen with a low random frequency at each generation.  In any substantial sized population, the mtDNA sequence with the least recent mutation is likely to have last mutated many thousands of years earlier than the mtDNA sequence with the most recent mutation.

Given the archaeological context of the teeth, in similar layers of debris in a single cave, the likelihood that there was mtDNA diversity with both older and younger clades of mtDNA present seems more likely to me than a continuous occupation for thirty thousand or so years that managed to be deposited in such close proximity to each other.  One could estimate a predicted population size on this basis and compare it to the estimate using other methods.

The open access PNAS paper is here.  John Hawks has an analysis at his blog.

Tuesday, November 17, 2015

Connecting the Cultural Dots With Relics And Legends

Gruda Boljevića tumulus is one of the most important archaeological sites found recently in Europe. The reason why I believe that this tumulus is so important, is because it shows that the dolmen building, golden cross disc making culture which developed in Montenegro in the first half of the third millennium BC, has its direct cultural roots in Yamna culture of the Black Sea steppe. Why is this important?

I have already shown that the golden cross discs which appear in Ireland and Britain around 2500 BC have their predecessors in golden cross discs from Montenegro which were dated to 2700 BC (Mala Gruda) and some time between 3050 BC and 2700 BC (Gruda Boljevića). Considering that these golden cross discs first appear in Montenegro and then in Ireland and Britain and nowhere else in between suggests that this cultural trait could have been a result of a direct cultural transfer between Montenegro and Ireland and Britain. Irish archaeologists are reluctant to say whether this cultural influence was due to trade or missionary contacts, or whether it was a consequence of a migration of a group people into Ireland.

This is because Irish archaeologists don't read pseudo histories like the Irish annals. If they did they would have seen the old Irish annals tell us that right at the time when the metallurgy and the first golden cross discs appear in Ireland, a group of people, a tribe a clan lead by Partholón arrives in Ireland. Partholón and his people are credited with introducing cattle husbandry, plowing, cooking, dwellings, trade, and dividing the island in four and most importantly for this story, they are credited with bringing gold which before them was not used in Ireland. They bring the golden cross discs. But where did Partholón and his people come from? The Irish annals tell us that too. They tell us that Partholón arrived to Ireland from the Balkans via Iberia. The Lebor Gabála Érenn, an 11th-century Christian pseudo-history of Ireland, tells us more. It tells us that Partholón came to the Balkans from the Black Sea steppe, the land where at the beginning of the 3rd millennium BC we find Yamna culture...
From the Old European culture blog.

It has long been known that Iberians and remarkably similar genetically to the people of the British Isles, and that both populations are rich in Y-DNA R1b. It has also been recently discovered that the Yamna people had Y-DNA R1b, although more fine distinctions of sub-haplotypes of R1b muddy the connection between the Yamna people and Western Europeans.

It is much less well known, as observed in the comments to this post at the Eurogenes blog, the Croatians and the English, at least at a naive three ancestral component analysis level, seem to have very similar autosomal genetic makeups to each other.

Are the Irish legendary histories telling us the true tale of how the people bearing Y-DNA R1b came to arrive in Western Europe?

The Old European Culture blog, slowly and in individually intriguing and convincing installments is cumulatively making a convincing argument in that direction that can be corroborated with well dated archaeological relics. And, while this analysis doesn't name the Western European cultures involved, it does point to some very specific times and places to look for the culture that probably brought Y-DNA R1b to Europe along with a prominent role for cattle. This time and place turn out to be a pretty good fit to the Bell Beaker culture, while it is a rather poor fit to the megalithic culture that was already present when the Bell Beaker culture emerged.

The path suggested by Irish legendary history is notable, in part, because it is a match for one of several outstanding hypotheses for how Y-DNA R1b wound up in Western Europe, and also because legendary history from Ireland may be more reliable than in many other places because its position at an island isolated ocean frontier would have prevented it from being easily muddled with infusions of legendary histories from elsewhere.

This analysis also doesn't attach a definitively linguistic label to this population. Many scholars assume that the Yamna people of the steppe were Indo-Europeans, and there are some good reasons for that assumption. But, it is hardly definitive. While we have written Sumerian and Egyptian documents from this far back in history, we have no comparably old documents in any other languages.

The record is equally consistent with a hypothesis in which the Yamna people spoke a language that was in contact with Proto-Indo-European and borrowed words from it, but was itself non-Indo-European and related to one or more Caucasian languages. In this scenario, a non-Indo-European Yamna derived language arrives in Western Europe in the time frame, while Western Europe undergoes widespread language shift to Celtic or proto-Celtic languages much later, plus or minus a few centuries from Bronze Age collapse in most places. This scenario is attractive, because otherwise, the heavily Y-DNA R1b Basque people would have had to experience a language shift from an Indo-European language to Basque, which seems far less likely to be the case.

Thursday, November 12, 2015

Some Musings About Nomeclature, Mesons And Forces

Almost all of the ordinary matter in the universe (as opposed to "dark matter") is made up of protons and neutrons assembled together into atomic nuclei of atoms with one or more protons and sometimes with sometimes also with some neutrons each.

The atom is completed with one electron per proton in orbit around the nucleus (strictly speaking, "orbit" is a bit misleading as a classical approximation of the actual quantum physics behavior of electrons associated with an atomic nucleus, but it is good enough for the purposes of this post).  If the correspondence between protons and electrons in an "atom" is other than one to one, it is conventional to use the term "ion" rather than "atom" to describe it.

Protons and neutrons are composite particles made up of three quarks each, which are bound together by gluons which are emitted and absorbed by the quarks in the proton or neutron respectively (the generic name that encompasses both protons and neutrons is a "nucleon" often abbreviated "N").

It turns out that protons and neutrons are actually just the two most common examples of a larger class of composite particles made up of three quarks bound together by gluons which are call "baryons".

There is a still more general class of composite particles made up of quarks (not necessarily three) that are bound together by gluons which are called "hadrons" (there is also a hypothetical class of composite particles made up solely of gluons bound together in the absence of quarks often called "glueballs", but I'm not sure if they count as hadrons or not).

Hadrons made up of two quarks (or of blended combinations of two quark pairs) are called "mesons", a term that was coined in the 1930s when the need for a force carrying particle to bind protons in atomic nuclei was hypothesized many years before the first meson was actually observed.

Mesons made up of two different kinds of quarks are usually named based upon the heaviest quark in the mesons.  When that quark is a bottom quark (formerly also known as a beauty quark), the meson is, logically, known as a B meson.  But, most of the lighter mesons were discovered and named before the quark theory of the Standard Model was worked out and the quarks were assigned names.  So, their names are quite arbitrary.

If the heaviest quark in a meson is a charm quark it is usually known as a D meson (but prior to 1986, a meson with a charm quark and a strange quark was known as an F meson). It is also curious that the physics community managed to abolish the historical irregular name for the Ds meson, but not many other of the historical irregular names of other hadrons.

If the heaviest quark in a meson is a strange quark, it is usually known as a "kaon" abbreviated "K".

If we were renaming mesons today, knowing what we do about the Standard Model, they would probably have been called S mesons, C mesons, and B mesons, respectively. But, historical accident and the immense amount of education needed to do the physics that makes knowing these names relevant has allowed the irregular historical monikers to survive.

Different naming conventions apply to "quarkonia", in which a meson and antimeson of the same flavor are both present.

I'd welcome comments from anyone who can explain how it is that Kaons, D mesons, F mesons and any of the other irregular hardon names were assigned (baryon naming, for what it is worth, has fewer irregular hadron names, perhaps because more of them were discovered after the quark model was in place).

As in other areas of language, irregular names for hadrons seem to have persisted more strongly for the most common mesons, than for the rare ones.

We still have arbitrary meson names not precisely tied to quark content for scalar and axial vector mesons, whose quark content is not well understood.

There are several kinds of mesons that contain only up and down quarks (or are at least dominantly comprised of up and down quarks).  The lightest are the pi mesons, also known as pions.  Another kind of meson containing only (or at least dominantly) up and down quarks is the rho meson.

Pions and rhos bring us from the land of history and language in physics to the land of QCD (the physics of the strong force that binds quarks with gluons called "quantum chromodynamics").

It turns out that the force hypothesized back in the 1930s that binds protons and neutrons into atomic nuclei called the "nuclear strong force" is not a fundamental force of nature.  Instead, it is basically a second order effect of the fundamental strong force which is mediated by gluons to hold hadrons together "leaking" out of hadrons to bind them to other nearby hadrons.

The nuclear strong force is mediated not by gluons, which are fundamental in the Standard Model, but by mesons, which are composite particles, although, like gluons, mesons are bosons (a class of particles that has one kind of quantum mechanical behavior) rather than fermions (a class of particles including all baryons and quarks and leptons such as electrons that has another kind of quantum mechanical behavior).

What caught my eye as I was looking into the history of the odd nomenclature of mesons was that while the nuclear strong force is mediated primarily (virtual) pions, it is also mediated in part by other kinds of mesons, especially (virtual) rhos.

This made me think about different kinds of mediator particles for different forces. Electromagnetism is mediated by photons, which are mostly identical, but can differ from each other in frequency, polariziation or helicity.  The strong force is mediated by gluons, which are also mostly identical to each other, but can come in eight different combinations of color charges.

The weak force, in contrast, is mediated by both W bosons (which come in W+ and W- varieties that are antiparticles of each other), and Z bosons which differ in mass and charge from W bosons.  In this respect, the weak force is a bit like the nuclear strong force, which has more than one kind of mediating boson, although the non-fundamental and emergent nuclear strong force, in principle at least within the Standard Model, can have more kinds of potential mediator mesons than the weak force had mediator weak force bosons.

It is interesting to consider how the Standard Model might be subtly modified to reflect the existence of additional weak force bosons that appear much less frequently the W and Z bosons in much the way that rhos and other mesons mediate the nuclear strong force much less frequently than pions do.

There are hypothetical W' and Z' particles which have been searched for experimentally (and thus far, not discovered).  But, it isn't entirely clear that those models are sufficient to account for the kinds of properties we would predict if the collection of particles that mediate the weak force are analogous to the array of mesons that can mediate the nuclear strong force.

Also, the analogy of the nuclear strong force to the weak force suggests that the W and Z bosons, unlike photons and gluons, might be composite particles, rather than being fundamental. (Electroweak unification involves a concept of blending of more fundamental particles to create the W, the Z and the photon, as opposed to a true composite concept.)  This seems like an interesting line of conjecture to extend to see how far one could take it.

And, of course, we've omitted discussion of the Higgs boson, which shows every indication of being basically a part of the electroweak unification scheme that is not closely related to QCD at all, and gravity, which just plain doesn't play well with the Standard Model, but mostly manages to stay out of the way in circumstances where gravity and the Standard Model might clash with each other.

Tuesday, November 10, 2015

Earliest Evidence of Marijuana Is From Jomon Japan

The earliest evidence of marijuana use by humans (for its hemp fibers) comes from Jomon Japan.

[T]he earliest traces of cannabis in Japan are seeds and woven fibers discovered in the west of the country dating back to the Jomon Period (10,000 BC – 300 BC). Archaeologists suggest that cannabis fibers were used for clothes – as well as for bow strings and fishing lines. These plants were likely cannabis sativa – prized for its strong fibers – a thesis supported by a Japanese prehistoric cave painting which appears to show a tall spindly plant with cannabis’s tell-tale leaves.

Until World War II, it was an important commercial crop in Japan.

This changed after World War II, and Japan now has some of the most strict anti-cannabis laws in the world.

The Selfish Case For Sharing Data

Some scientists horde data, on the theory that this gives them an edge that other scientists in the community lack so that they can publish on something that no one else can publish upon.  Other scientists share data, on the theory that this allows more people to engage with their work and by doing so recognize it.

In academia, the first law of career advancement is publish or perish.  More publications are good. But, academia recognizes superstars based not only upon how many articles one publishes, but how many times those articles are cited by others.

Empirically, if you are a scientist who manages to publish at all (and many people in academia publish very little once they receive tenure), the best way to advance your citation count and thrust yourself into academic superstar status, is to share your data, according to a short preprint examining the question released Sunday, at least in astrophysics, but probably in a much broader array of other disciplines as well.

Overall, making your underlying data available will increase you citation count for a paper by 20% on average.

Monday, November 9, 2015

Waiting For The Next Big Paper

There have been a number of interesting papers on historical genetics and prehistory in the past couple of weeks that I haven't covered because I've been spending lots of days litigating cases in court and in arbitration forums for my clients.

But, a fair amount of interest in the field is devoted to the next big ancient DNA paper which a number of conference announcements discussed by Bell Beaker blogger suggest is coming sometime soon this autumn (in the Northern hemisphere) that will finally shed light on the genetics of the Western European Bell Beaker phenomena.

Of course, it could turn out that the blogging community had read too much into the tea leaves.  But, if such a paper is released, it could resolve a lot of the biggest remaining questions in historical genetics and prehistory.  Western European ancient DNA has been covered much less comprehensively than Central and Eastern Europe in recent autosomal ancient DNA studies, but the archaeology of Western Europe is pretty comprehensive and the technology can now do amazing things with lower quality remains, so there is good reason to be hopeful that the rumored new paper or two will be far more than hype.

Frustratingly, some of these conferences took place in October, but the content of the presentations has nonetheless apparently escaped the Internet's grasp.  But, this does argue against any major breakthrough papers.

This information, in turn, will also shed a lot of light on the question of when, how and from where the Indo-European languages and any predecessor languages arrived in Western Europe, although since pots are not people, new data can only strengthen or weaken arguments about historical linguistics, rather than resolve them definitively.

More On Math Genius Srinivasa Ramanujan

The collected works of the mystical Indian mathematical genius Srinivasa Ramanujan are still yielding valuable mathematical insights almost a century after his death.  The latest insight's involve a class of near counter-examples to Fermat's last theorem (which was subsequently proven to be true in my lifetime, four hundred years after it was proposed).

Monday, November 2, 2015

Razib's Riff On Eurasian Origins

Razib Khan has an extended meditation on the genetic origins of the Eurasians.

Some key points:

(1) Culturally driven punctuated change has been the norm for much of the history of humanity.

(2) One factor driving this population replacement trend has been the fragility of the economic subsistence basis of the conquered people.  Often they have been so fragile and with so little excess that exacting tribute from them and exploiting them as a ruling class has been impossible and instead they were replaced.

(3) The Y-DNA R1a and R1b branches have seen the explosion of the modern sublineages since ca. 3000 BCE, which correspond with the sweeping expansions of these patrilineages into Europe and South Asia, and to a less extent, into the Middle East.

(4) The Neolithic revolution and the subsequent explosion of steppe DNA beyond the steppe has resulted in the integration of populations that were more or less isolated from each other for tens of thousands years created a new homogenized blended population genetic community that had never existed until then.

Put more bluntly, European hunter-gatherers had an extremely different phenotype from the earliest farmers of Europe.  In other words, they were of very distinct races, at least as distinct as Chinese people are from French people today.

One major blending of these two populations gave rise to Early European Farmers to produce a mix much like modern day Sardinians.

Another major blending happened when population genetics from the steppe made its way to Europe sometime in the time frame of the late Copper Age to the early Iron Age to create more or less the modern European phenotype which didn't exist until then.  These steppe people were about half Early European farmer (itself a blend of hunter-gatherers and "basal European farmers", with additional contributions of Ancestral Northern European and Eastern Hunter-Gatherer).  Admixture percentages from substrate populations varied quite a bit in a systemic fashion with more admixture where the food production package brought be the conquering people was less well suited to local conditions.

(5) Y-DNA lineages have a tendency to extend beyond the tightly knit human communities that drive their initial expansions, and also show strong tendencies towards replacement in excess of that seen in autosomal and mitochondrial DNA.

(6) Selective fitness based selection has been ongoing into the modern era.

(7) Razib feels as it most of the puzzle pieces are in place already.  In contrast, I feel as if we have lots of the puzzle pieces together and a good understanding of some regions, but still do not understand the links between the genetic shifts and the historical events that caused them, nearly so well in Western Europe as we do in Eastern Europe and some other areas.

Friday, October 23, 2015

Deur et al make breakthrough link between non-perturbative and perturbative QCD

Alexandre Deur is a physicist at Jefferson Labs whom I have previous praised for his work on graviton self-interaction based alternative to dark matter analogous to the effects of gluon self-interaction in QCD (the part of the Standard Model of Particle Physics describing strong force interactions between quarks), which is his primary research area.

But, he's no slouch at his day job either, as a recent new paper that put him on the path to being a major innovator in the field rather than one cog of hundreds or thousands of physicists collaborating in conducting Big Science experiments (where he has most of his publications).

Deur and several colleagues posted a ground breaking preprint this week that links the fundamental scale constant of perturbative QCD (applied mostly to high energy collisions, a.ka. ultraviolet QCD, analytically (i.e. with equations)  to the fundamental scale constant of non-perturbative QCD, which is used mostly to study the properties of quarks confined in hadrons that are not interacting at such high energies a.k.a. infrared QCD.  This is a big deal because, in practice, ultraviolet QCD research is much more expensive than infrared QCD research.  This allows comparatively cheap low energy research (millions and tens of millions of dollar experiments in many cases for a related set of experiments and including lots of data that has already been collected) to boost expensive high energy research (costing in the billions and tens of billions of dollars for each sustain experimental program each of which is forging into unknown territory about which we have no prior data).

Both the strength of high energy strong force interactions, described by perturbative QCD, and the lion's share of the masses of hadrons comprised of lighter quarks (which in turn provides the lion's share of the mass of the ordinary matter (as opposed to dark matter and dark energy) in the universe, both ultimately flow from the strong force coupling constant and the strength of the strong force color charge of quarks (which is identical in magnitude for all quarks).

But, in practice, physicists use the physical constant "lamda_s" to do many of the strong force calculations in perturbative QCD, and use the physical constant "kappa", to do many of the strong force calculations involved in establishing hadron masses from first principles in non-perturbative QCD.  Both of these approaches are phenomenological approximations in their respective domains of applicability of the exact equations of QCD which are believed to be known, but are too mathematically intractable to calculate with directly.

This breakthrough means that experiments from perturbative QCD can now be used to provide the key QCD physical constants used to calculate hadron masses, while hadron mass measurement, in turn, can be used to determine the key QCD physical constant for making calculations in perturbative QCD.  Previously, the two constants had to be measured separately in practice, even though everyone knew that there must be some analytical relationship between them.

The results are largely consistent with experimental data from both regimes (although there is one calculation that has a two sigma tension between the theoretical prediction determined in this manner and the experimental data), and the uncertainties are predominantly due to issues involved in calculating a numerical approximation of equations with infinite numbers of terms.

The actual accuracy with which the physical constants involved are known, roughly 3%-6%, isn't terribly impressive.  But, in the long run, there is a clear path to reducing theoretical uncertainty by simply putting more resources from more powerful supercomputers on the problem so that the theoretical calculations can have less uncertainty by including far, far more terms in the calculation than anyone has been able to do to date with limited resources.  And, since the light hadron masses are known to far more accuracy than the extremely high energy measurements of perturbative QCD, this connection could ultimately use constants determined in the precisely measured low energy QCD regime to dramatically improve the accuracy of calculations in the high energy perturbative QCD regime that applies, for example, in the high energy particle accelerator collisions conducted at the Large Hadron Collider (LHC) by the ATLAS and CMS experiments.

This paper is also a critical intermediate step in linking both the perturbative QCD and nonperturbative QCD calculations done in the real world to the exact equations of QCD and the fundamental physical constant of the Standard Model which is the strong force coupling constant, from which  they are both in principle derived, thereby helping to make it possible to use first principles calculations using the actual exact equations of QCD of real world quantities.  The equations and constants of perturbative QCD and non-perturbative QCD are both informed by knowledge of what the exact equations of QCD look like but ultimately are phenomenological approximations of the exact equations, rather than being rigorously and exactly derived matheatically from the exact equations of QCD.

Popular accounts of QCD and the Standard Model often read as if this is a solved problem.  But while we believe that we know the exact equations of QCD, the quark masses and the strong force coupling constant with sufficient precision to make these calculations in principle, in fact, no one has yet managed to do it without major approximations, in practice.

The most precisely measured hadron masses (the proton, neutron and pion) are known to six significant digits, and even the least precisely determined ones (heavy hadrons with bottom quarks) are known to six significant digits.  But, the strong force coupling constant is known only to about 0.5% precision.  But, theoretically, it should be possible using only the light quark masses known to their current accuracy, the most precise several hadron masses, and the known exact equations of QCD, to calculate the strong force coupling constant to roughly 200 times as much accuracy as it is known today without conducting another experiment ever, if one has sufficient computational capacity.  This paper is a major intermediate step in that direction.

Moreover, one of the reasons for a significant amount of the uncertainties in the experimentally determined quark masses in the Standard Model is due to the uncertainties in the strong force coupling constant together with the accuracy lost in numerical approximations of the true equations of QCD.  So, improvement in measurement of the strong force coupling constant facilitated by this research has the potential to greatly improve the accuracy with which six other Standard Model fundamental constants are known using existing experimental data.  And, knowing both the strong force coupling constants and the quark masses with more precision, in turn, also makes it possible to greatly improved the statistical power of experiments done to determine the four CKM mixing matrix parameters.  This is because uncertainties regarding the Standard Model background predictions from QCD greatly reduce the statistical power of experiments measuring other Standard Model constants.

Finally, great precision in all of the physical constants going into QCD calculations which are used to determine Standard Model backgrounds in high energy particle accelerator experiments, in turn greatly improves the statistical power of experiments setting out to identify beyond the Standard Model physics.

For example, the primary decay path of the Standard Model Higgs boson is to quark-antiquark pairs of bottom quarks.  But, lots of other Standard Model processes also produce quark-antiquark pairs of bottom quarks.  The measurement of the Higgs boson signal in the bottom quark decay channel is determined by using perturbative QCD to make estimates of Standard Model bottom quark decay backgrounds from other processes, which have quite significant error bars of their own, and then to look at the total number of observed bottom quark decays observed to estimate the number of Higgs boson sourced bottom quark decays observed.  But, since quantum mechanics is stochastic, the number of Higgs boson bottom quark decays expected even with perfect backgrounds is a gaussian distribution around a most likely number of bottom quark decays for any given Higgs boson mass, the expected number of Higgs bosons produced is subject to further statistical variation, and the backgrounds with error bars (only some of which are irreducible statistical variation) that are large compared to the expected signal.  So, it is hard to see the Higgs boson in its main decay channel even when there are lots of Higgs boson bottom quark decays out there to be seen even at fairly low Tevatron energies.  But, if you can dramatically reduce the non-statistical errors in the Standard Model background prediction, it would be much easier to distinguish the signal of bottom quark decays from Higgs bosons from other Standard Model backgrounds, even with Tevatron data which is far inferior to the LHC in energy scale and total number of events observed.

Going forward, reducing error bar noise in Standard Model backgrounds in the current LHC experiments would significantly improve the ability of ATLAS and CMS to confirm that the Higgs boson seen at the LHC at 125 GeV or so has all of decays expected at the frequencies expected for a Standard Model Higgs boson of that mass, or in the alternative, to see statistically significant differences from the Standard Model Higgs boson expectation even if they are quite subtle differences.

Similarly, if new physics manifest at some characteristic energy scale "lamda_BSM", the energy scale at which the new physics can be detected experimentally could be reduced by an order of magnitude or two, if we were able to leverage our existing precision knowledge of light hadron masses into more precise values of the Standard Model fundamental constants of QCD using improved mathematical approximations of the exact equations of QCD which papers like this one are bringing closer to reality.

If BSM physics exist at some level greater than the electroweak scale of O(100 GeV) and the GUT scale of O(10^16 GeV), we might be able to find them using current experiments using current numerical QCD methods at the scale of O(1000-10,000 GeV) using the LHC with currently available technology and perturbative QCD calculation accuracies.  But, the kind of improvements that may be possible in QCD with much more accurately known QCD physical constants could stretch our experimental research to revealing or ruling out new physics up to scales of O(100,000-10,000,000 GeV) (i.e. 100-10,000 TeV).

There estimates may be a bit optimistic (because some Standard Model backgrounds have inherent statistical variation that is large relative to the expected signal even if the background is calculated perfectly, requiring experimenters to look at signals with little or no Standard Model background instead), but testing new physics up to the several hundred TeV scale with technology not involving any major technological breakthroughs not present at the LHC today is not unthinkable if we can make better progress on the math of QCD, which may be possible to achieve to a significant extent with nothing more than a big investment in spending on the supercomputers (without any advances in supercomputing technology itself from current levels) that are available to QCD physicists.

This paper is an important step in making these advances a function of our willingness to spend the money to allow our scientists to make absolutely inevitable and certain progress, as opposed to a gamble on whether no conceptual breakthroughs can be devised by physics geniuses, if those breakthroughs are even out there waiting to be discovered, which they might not be at some point.

Right now, most new physics scenarios have strong minimum energy scales, but their maximum energy scales are far in excess of the minimum energy scales at which they can be ruled out,  But, if the parameter space in which new physics can be sought expands enough, the entire parameter space of many BSM theories may be possible to confirm at particular values or rule out, if we can simply improve the statistical power of present day LHC technology experiments by using more precision knowledge of Standard Model fundamental constants to more precisely predict the expected Standard Model backgrounds.

For example, the non-detection of proton decay and neutrinoless double beta decay places an energy scale ceiling on many kinds of supersymmetry (SUSY) theories.  But, this ceiling is much higher than the minimum energy scales at which new physics from SUSY theories can be excluded using the LHC and other experimental data that is available.  Increased experimental power from a more precise knowledge of the fundamental constants of QCD, however, might make it possible to close that gap for many kinds of SUSY theories.  And since string theory almost universally assumes that its low energy approximation resembles fairly genetic versions of SUSY this could even make it possible to experimentally rule out immense swaths of the string theory landscape.

Lest I overhype too much, I do need to provide some perspective.  Physicists have known that what Deur and his colleagues did was possible in principle for half a century.  We knew already that this was a problem with a correct solution that was out there waiting to be found.  But, the fact that it took half a century to get from knowing that the answer to this intermediate result was out there, and actually discovering it, is also a testament to how non-trivial an effort this very lucid paper really is in fact, even if it seems deceptively simple.  The authors of this paper have not only reached an important intermediate result, but have also artfully make it look more much elementary and obvious than it actually was (much of the really hard stuff is hidden in results from QCD methods such as the light front method which are described only by bottom line result and citation in this paper).

Monday, October 19, 2015

Did Dogs Originate In Mongolia Or Tibet?

Dog genetics are most diverse in Mongolia and Tibet and show a roughly clinal trend towards less diversity with distance away from that area, suggesting that the domesticated dog may have originated there. But, there are conflicting indications from different kinds of data.
Dogs were the first domesticated species, originating at least 15,000 y ago from Eurasian gray wolves. Dogs today consist primarily of two specialized groups—a diverse set of nearly 400 pure breeds and a far more populous group of free-ranging animals adapted to a human commensal lifestyle (village dogs). Village dogs are more genetically diverse and geographically widespread than purebred dogs making them vital for unraveling dog population history. Using a semicustom 185,805-marker genotyping array, we conducted a large-scale survey of autosomal, mitochondrial, and Y chromosome diversity in 4,676 purebred dogs from 161 breeds and 549 village dogs from 38 countries. Geographic structure shows both isolation and gene flow have shaped genetic diversity in village dog populations. Some populations (notably those in the Neotropics and the South Pacific) are almost completely derived from European stock, whereas others are clearly admixed between indigenous and European dogs. Importantly, many populations—including those of Vietnam, India, and Egypt—show minimal evidence of European admixture. These populations exhibit a clear gradient of short-range linkage disequilibrium consistent with a Central Asian domestication origin.
L. Shannon et al., Genetic structure in village dogs reveals a Central Asian domestication origin, PNAS (Published online October 19, 2015).

Siberian Genetics

Siberia and Western Russia are home to over 40 culturally and linguistically diverse indigenous ethnic groups. Yet, genetic variation of peoples from this region is largely uncharacterized. We present whole-genome sequencing data from 28 individuals belonging to 14 distinct indigenous populations from that region. We combine these datasets with additional 32 modern-day and 15 ancient human genomes to build and compare autosomal, Y-DNA and mtDNA trees. Our results provide new links between modern and ancient inhabitants of Eurasia. Siberians share 38% of ancestry with descendants of the 45,000-year-old Ust-Ishim people, who were previously believed to have no modern-day descendants. Western Siberians trace 57% of their ancestry to the Ancient North Eurasians, represented by the 24,000-year-old Siberian Malta boy. In addition, Siberians admixtures are present in lineages represented by Eastern European hunter-gatherers from Samara, Karelia, Hungary and Sweden (from 8,000-6,600 years ago), as well as Yamnaya culture people (5,300-4,700 years ago) and modern-day northeastern Europeans. These results provide new evidence of ancient gene flow from Siberia into Europe.
Valouey et al., "Reconstructing Genetic History of Siberian and Northeastern European Populations" (2015).

Eurogenes has noted that while there is a close correspondence between Malta-like ancestry and Eastern European Hunter-Gatherer genetics, that there is not a good correspondence between this autosomal component and the Y-DNA N1c1 commonly found in Uralic and Scandinavian populations.  This maker may instead distinguish circumpolar from non-circumpolar populations.

Time depth is a tricky issue in Siberian populations.  There appear to have been multiple populations sweeps from East to West and then back from West to East again across the region, even in historical times, and this area was completely depopulated during the LGM with less clarity than in some other places about where the refugia from which it was repopulated were located.

There are Paleo-Siberian layers (pre-modern Siberian), modern indigenous Siberian layers (pre-Uralic), Uralic layers (ca. 35th to 25th centuries BCE), Proto-Indo-European layers (Tocharians ca. 20th century BCE to ca. 6th century CE), Turkic migrations (1st to 6th centuries CE), Islamic expansions (West to East starting in the 7th century CE), Mongolian migrations (East to West ca. 13th-14th centuries CE), and Russian migrations (West to East starting in the 17th century or so CE).  It is possible that instances of Y-DNA C in Europe in ancient DNA may represent traces of old East to West migration, but they are outliers and it is hard to say why they turn out. This list is illustrative only and surely contains mistakes and overlooks nuances.

Honestly, the extent to which Siberian ancestry is Paleolithic is remarkably high and probably reflects the fact that the region is largely ill suited to intense farming.  It isn't clear to what extent "Ust-Ishim ancestry" and "Malta ancestry" overlap from my initial glance at the paper.  But, it does appear that modern Siberia is more eastern influenced than western influenced.  It also isn't clear to which extent the chosen populations minimize some of the migrations known to have occurred historically.

Tuesday, October 13, 2015

Scores Of New Ancient European Genomes

The Mathieson, et al. (2015) paper on "Eight thousand years of natural selection" has few surprises when it comes to the title of the paper that weren't either widely known or strongly suspected before the study was published, but its barrage of new raw ancient DNA data, and ancient Y-DNA data in particular, is impressive.  The study has very little ancient DNA from the Atlantic region, however.

Saturday, October 10, 2015

Did The Yamnaya Die Or Run?

Razib Khan has tweeted the biggest story of the ASHG 2015 Conference (key points in bold):
@iosif_lazaridis revision of paper @mathiesoniain with more ancestry stuff on biorxiv soon
@iosif_lazaridis [Neolithic Anatolian] mtDNA look familiar to EEF. Y mostly G2a2. also J2 H and I at low frequency. C1 too
@iosif_lazaridis anatolian neolithic close to EEF on pca. but EEF shifted toward WHG #ASHG15
@iosif_lazaridis anatolian neolithic different from modern anatolian and se europe populations.
@iosif_lazaridis eurasian steppe, population transect done. 5,500 to 1,200 BC. author told me some R1a1a possibile stuff here yesterday
@iosif_lazaridis indo-european steppe = EHG + near eastern. new data eneolithic samara. 75% EHG ancestry. 25% "armenian" 5,200 to 4,000 BCE
@iosif_lazaridis poltavka people 3000 to 2200 BC basically like yamnaya. 50% EHG and 50% armenian-like. then srubnaya different.
@iosif_lazaridis srubnaya 2/3 yamnaya 1/3 middle neolithic european
@iosif_lazaridis yamnaya/poltavka went from R1b to R1a in the srubnaya period. z93 group found on bronze age steppe samara (s asian R1a)
@iosif_lazaridis there was back migration of EEF to the steppe after the initial yamnaya migration.
Via Eurogenes.

This is a huge new set of facts with potentially profound implications for how we understand European history.

In the span of a few centuries, starting around 2200 BC, a Y-DNA R1b dominated population, the Yamnaya and their descendants the Poltavka, were replaced in the southern part of the European steppe (or at least the men were) by a Y-DNA R1a dominated population with strongly overlapping autosomal genetic profiles, the Srubnaya.

One possibility is that the Yamnaya men were slaughtered by the Srubnaya men who may have assimilated some of the Yamnaya women, in a scenario mirroring that of the battles described in the Biblical Book of Numbers.

But, something remarkable happens in Western Europe right around the time that Y-DNA R1b men disappear from the southern part of the European steppe.  All of the sudden, Y-DNA R1b that was virtually absent from Western Europe rapidly becomes the predominant Y-DNA haplogroup of Western Europe and there are substantial shifts in the mtDNA mix of Western Europe.

The distribution of Y-DNA R1b sub-haplogroups in Europe and their phylogenetic relationship, suggests a route from East to West of Y-DNA R1b carriers from the European steppe to central France and from their spoke-like migrations in all directions.

The insight in today's talk provides a push factor - the Srubnaya (whom Davidski at Eurogenes describes as more militaristic and technologically advanced than the Yamnaya).  A collapse of Western European first wave Neolithic farming societies as a consequence of the 4.2 kiloyear event, meanwhile, may have left their societies in turmoil and collapse (including population collapse) leaving a political vacuum and slack in food production capacity once the event's harsh climate abated, into which the Yamnaya people, participating in a folk migration much like that of the "migration period" travels of Germanic tribes like the Goths, Visigoths, and Vandals in the Dark Ages, into Western Europe.

The Yamnaya were basically steppe pastoralists, which is to say, herders.  Faced with a potentially deadly military adversary, farmers stand their ground upon which they rely to survive, even if the consequences are dire.  But, herders in a culture that exults cattle and bulls rather than corn and wheat, don't have to suffer the consequences of standing and fighting against opponents who may be superior to them militarily, or just more determined.  They can run, as an entire community, taking the cattle and horses that provide the source of their wealth with them, at a lower cost that does not have to be paid in blood.

And, wouldn't it stand to reason that people in a cattle herding society would be more likely to have LP genes that allow them to drink cow's milk as adults which was gradually selected for over thousands of years, which they would bring with them in their genes to their new homeland, than a society of farmers would be to suddenly develop this gene through explosive and rapid natural selection?

And, given that these people could have been ancestral to Europe's Basque (who have high frequencies of Y-DNA R1b, has traditions that place an emphasis on cattle, who arrives in their current relict homeland in my view from France, have high levels of the LP genes, and speak a language that is distinct in being ergative, just like the language of the Georgians with whom the Yamnaya's non-Eastern Hunter-Gatherer autosomal genetic component shows strong affinity), it is highly plausible that their language (and hence Basque and other Vasconic languages) was an offshoot of the Kartvelian language family, possibly after creolization with Eastern Hunter-Gatherer languages that also contributed strongly to the Proto-Indo-European language, and with substrate influences from whatever first farmer Neolithic language was spoken in Western Europe before they arrived.

Razib recently hypothesized that Proto-Indo-European and Afro-Asiatic were hunter-gatherer substrate languages that were adopted by Early European farmers who admixed with them.

But, I don't think that scenario is plausible.  When hunter-gatherers and farmers collide, usually, the farmer's language prevails (see, e.g., Japan, where the rice farming Yayoi's language became the backbone of the Japanese language, but almost no words from the hunter-gather Jomon who spoke a language in the same language family as the Ainu made it into Japanese, despite the fact that something like 40% of the genetic ancestry of the Japanese is Jomon including a large share of the male Y-DNA), although sometimes there are some substrate influences that shape the superstrate language dialect that comes to be spoken in the blended community.

The alternative, which usually happens when the superstrate population is greatly outnumbered by a substrate population and the two populations have no linguistic common ground, is the development of a creole language, which Indo-European shows some elements of (or at a minimum simplification of the language driven by a large community of second language learners in the society), and which has been suggested on archaeological grounds by the mixed ethnicity communities that existed around the time and place that PIE came into being, long before ancient DNA tools were available.

The lexical similarities that Maju recently discovered between both Proto-Indo-European and Nilotic Nubian languages may very well be real, but he may have misapprehended the direction of the connection.  Perhaps, the lexical similarities in Nilotic Nubian may be the result of Neolithic migrants to Africa who arrive via the Sinai and the Nile bringing the words of their language, related to Kartvelian, with them, where the existing residents adopted them, rather than the other way around.

The Yamnaya folk migration hypothesis which I have just sketched out, which is strongly motivated by powerful ancient DNA evidence, has the potential to pull together myriad puzzle pieces of European prehistory in a single stroke.

It doesn't answer all of the questions.

What connection did the Bell Beaker culture have to the Yamnaya?

Is there any archaeological evidence to support this hypothesis?  And, if there was such evidence, what would we expect it to look like?  Is the lack of evidence of an apocalyptic war that destroyed the vast majority of Yamnaya men itself evidence favoring this hypothesis?  Before dismissing this conjecture for a lack of archaeological evidence, at a minimum, the archaeological evidence should be reviewed with fresh eyes informed by this hypothesis.

Did other members of a Yamnaya diaspora make their way to Western Anatolia (perhaps Troy I?) or Crete?

Finally, when did the migration(s) start?

Perhaps the Poltavka people are the Yamnaya who held out on the southern European steppe longer, but the transition from the Yamnaya culture to the Poltavka culture was a product of the disruption caused by the folk migration of the rest of the Yamnaya to Western Anatolia, Crete and Western Europe.  A 3000 BCE start date for these migrations is a better fit to the archaeological culture that could potentially reflect their arrival in their putative destinations.

This also illustrates the fact that dying and running are not necessarily mutually exclusive possibilities.  Perhaps some of the Yamnaya ran in various directions, giving rise to the various European cultures in which Y-DNA R1b is common, while other Yamnaya stood their ground, becoming the Poltavka, and ultimately died at the hands of Y-DNA R1a dominated peoples from the northern European steppe who slaughtered the Poltavka and took their land.

Fortunately, it is likely, given the stunning improvements that have been made in ancient DNA extraction, that these are questions to which the answer probably isn't, "we may never know." Instead, stay tuned.  More answers seem to lurk around every corner.

Thursday, October 8, 2015

A Brief History of Exponents

The Math With Bad Drawing blog has a nice little post explaining very lucidly who the notion of exponents of repeated multiplication was generalized in a way that is pretty much unique to allow for exponents that have values other than whole numbers.

This fact, typically first taught in middle school or high school algebra, has been well known for a long time. Euclid toyed with the idea a little.  Ancient Greek scientist Achimedes first generalized the concept and proved the law of exponents. A fairly efficient form of exponential notation was invented by Nicolas Chuquet in 1484. More than three hundred years ago René Descartes established the modern superscript notation for exponents in the late 1600s around the same time that Newton's law of gravity and motion were invented and around the same time that Newton and Leibniz invented calculus (the modern notation used in undergraduate calculus follows the practice of Leibniz and not Newton's much more awkward notation).

There has been one notable elaboration of a similar concept in mathematics since then, called the fractal dimension which was first formally defined using that name by the late Benoit Mandelbrot in 1967 and entered the upper level college mathematics curriculum in the late 1980s and early 1990s, around the time I was an undergraduate math major. This concept was also invented in Newton's day, but then consigned to the dustbin of history as a curiosity until the late 1800s when several mathematicians developed it some more, and then remained out of sight until Mandelbrot, more or less single handedly repopularized the concept in a way that actually stuck and found practical applications.

The fractal dimension generalizes the notion of a dimension in a manner similar to the way that the law of exponents generalizes the notion of repeated multiplication by relating change in detail to change in scale.  For example, the smaller the ruler you use to measure a shoreline, the longer the shore gets in ruler lengths, because the ragged pattern of a shoreline has a high fractal dimension, while a smooth shoreline would have a low fractal dimension and doesn't change in length at all based upon the length of the ruler used to measure it.

I probably wouldn't ordinarily have found any of the blog post on exponents notable at all. But, earlier just this week, I had been thinking about the precise issue of how the generalized notion of an exponent is so much more subtle than the naive repeated multiplication definition, in the context of thinking about Euler's formula and the Euler's number "e", which is equal to approximately 2.71828 and is a transcendental number that cannot be produced from the ratio of any two integers (something called a rational number). It felt remarkable to see in illustrated print found at random on the Internet, almost exactly the same line of thought.

I guess I still belong to the math tribe, even though I'm a lawyer now.

4500 Year Old Ethiopian Ancient DNA

UPDATE 3 (January 25, 2016): The portion of this paper pertaining to Eurasian admixture in people outside East Africa was due to an IT error and was retracted.  Some key inaccurate conclusions have been stricken below, but a careful reread is necessary to confirm the accuracy of all statements below.

UPDATE 2 (October 9, 2015): More figures at this tweet.

UPDATE: This is the first African autosomal ancient DNA sample that I am aware of, a remarkable technological feat, and it is paradigm shifting.  There is so much data in the whole genome of even a single individual, and the accumulated genomes of various modern and ancient populations is sufficiently significant already, that it is possible to make reliable and powerful inferences even with a sample size of just N=1 from a new population, ancient or modern, as this paper does.

ORIGINAL POST:

I had originally read this paper to imply that the 4500 year old Southwest Ethiopian male Mota in the sample was Eurasian admixed. Upon a more careful reading, it appears that this individual predates significant recent Eurasian admixture and can be used as a reference point to establish the levels of Eurasian back migration found in other African populations, since he has no measurable Eurasian admixture himself.
Characterizing genetic diversity in Africa is a crucial step for most analyses reconstructing the evolutionary history of anatomically modern humans. However, historic migrations from Eurasia into Africa have affected many contemporary populations, confounding inferences. Here, we present a 12.5x coverage ancient genome of an Ethiopian male (‘Mota’) who lived approximately 4,500 years ago. 
We use this genome to demonstrate that the Eurasian backflow into Africa came from a population closely related to Early Neolithic farmers, who had colonized Europe 4,000 years earlier. The extent of this backflow was much greater than previously reported, reaching all the way to Central, West and Southern Africa, affecting even populations such as Yoruba and Mbuti, previously thought to be relatively unadmixed, who harbor 6-7% Eurasian ancestry.
M. Gallego Llorente et al, "Ancient Ethiopian genome reveals extensive Eurasian admixture throughout the African continent" Science (October 8, 2015) DOI: 10.1126/science.aad2879

Hat tip to Dienekes.

Uniparental Haplogroups

Eurogenes notes that "this individual belongs to Y-haplogroup E1b1 and mtDNA haplogroup L3." These uniparental haplogroups, which are disclosed in the supplemental materials to the paper.

The mtDNA haplogroup is more specifically L3x2a which "is restricted to the Horn of Africa and the Nile Valley in modern Ethiopian samples, suggesting a degree of maternal continuity in Ethiopia over the past 4,500 years. . . . Mutation E-P2, present in Mota, represents the most widespread subclade of haplogroup E and has been found at high frequency in modern Ethiopians."

This individual also strengthens the case for an African origin of Y-DNA E, relative to a back migration hypothesis, because trace levels of Neanderthal ancestry found in other Africans can now be firmly attributed to recent Eurasian sources and are not present in this not really very old Y-DNA E African individual without Neolithic era Eurasian admixture.  This suggests that any back migration of Y-DNA E would have happened, if it did happen, prior to any Neanderthal admixture, which is present in all modern non-Africans.

The Source and Context of the Ancient DNA Sample
Mota Cave, situated 1,963 meters above sea level in the Gamo highlands of southwest Ethiopia, overlooks the Kulano River, a tributary of the Deme-Omo River. The cave was found in 2011 in collaboration with local Gamo elders and partially excavated in 2012. It measures 14 meters in width and 9 meters in depth and contains more than 60 centimeters of anthropogenic deposits and substantial rock fall. The cave’s deposits suggest at least seven different human occupations from the middle to late Holocene (c. 5295 BP to c. 300 BP), and contains the only middle Holocene burial known in southwest Ethiopia. This burial consists of a complete but fragmentary male adult skeleton dated via AMS radiocarbon to the fifth millennium BP (OxA-29631: 3997 ± 29 BP; 4524-4418 Cal BP).
This is part of an endorheic basin that flows into Lake Turkana in the Southwestern Ethiopia.

The context isn't well enough established to know if this individual was part of a hunter-gatherer, Neolithic, or metal age material culture, but there are hints that he might have been a hunter-gatherer because the only relics found with the body are "A geode and at least 27 obsidian, chert, and basalt flaked stone tools were found in the grave; such artefacts are characteristic of the Later Stone Age lithic tool assemblage present in much of the cave’s deposits.", and because of his genetic affinity to the Sandawe people, discussed below, who are a click speaking people who were a relict hunter-gatherer population of East Africa until about 150 years ago.

Genetic Affinities Of Linguistic Groups

This is right in the vicinity of the homeland of the Southern Omotic languages like Ari.  Cushitic languages are also spoken in the region, which has a high level of linguistic, religious and ethnic diversity in an area that is 90% rural.

The supplement also notes that "Principal component analysis shows that Ari and Sandawe are the closest contemporary populations to Mota.", and that Mota has no discernable Neanderthal component relative to modern African populations.
Mota was placed close to the Ethiopian samples, in between the clusters formed by the Ari and the Sandawe (but very close to an Ari individual that stands out from the rest of that group). The Ari can be split into two castes, Ari Cultivator and Ari Blacksmith, which share a common origin within the last 4,500 years. Since data on a larger number of SNPs are available for Ethiopian populations, we repeated the PCA using this higher quality dataset, which gave us 484,161 usable SNPs that could be called in Mota. Once again, Mota fell in between the Ari and the Sandawe cluster. . . . The Ari speak a language classified as Omotic, which is the most differentiated branch of the Afro-Asiatic languages. Gumuz, a population member of the Nilo-Saharan family (also an Afro-Asiatic language), also shows a high level of shared drift with Mota, but significantly less than the Ari. Sandawe, which are closer to Mota in the PCAs, do not show high shared drift with Mota in the f3, possibly because they are closer to the Khoisan populations than the other Eastern African populations.
Southwest Ethiopia is about as far from the place where any putative migration across the Gate of Tears would have taken place as one can be in Ethiopia and is in an area where Omotic languages are among the languages currently spoken.  The Sandawe people who also cluster with Mota currently live in central Tanzania but almost surely had a much larger geographic range in the past.  The highly tonal features of the Omotic languages may perhaps reflect a modified click language heritage.

The date is also a bit early for any Eurasian admixture to have an Ethio-Semitic source.  And, other recent studies have suggested that Eurasian admixture in non-Ethio-Semitic populations of Ethiopia (presumably arriving via the Blue Nile) took place at about the same time as Ethio-Semitic admixture.

Mota is not particularly close to Nilotic, Cushitic or Ethio-Semitic populations genetically.  Nor was Mota close to the Hadza people, another relict population of Paleo-Africans in East Africa.  Only the Sandawe and Omotic populations were reasonably close to Mota in the PCA analysis.  The study also tends to show that the Sandawe and Ari people of the Owo Valley cluster together rather closely genetically relative to other African populations, possibly shedding some light on the linguistic position of the Omotic language.  Both Cushitic and Ethio-Semitic populations deviate from the cluster that includes Mota in the same direction, while Nilotic and Hadza populations are essentially orthogonal to the Afro-Asiatic populations in the PCA.  The Omotic people, Sandawe people, and Mota are clustered together midway between the Afro-Asiatic, Nilotic and Hadza spokes at about 120 degree angles from each other.

Implications For Eurasian Ancestry In Other Africans
We used f4 ratio analysis to formally assess the extent of back-migration to Africa by West Eurasians . . . . Mota does not show any evidence of a West Eurasian component. . . . This contrasts in particular with the Ari, their closest contemporary relatives, which show large West Eurasian components (17.8%±1.0% and 14.9%±1.2% for Ari Cultivator and Ari Blacksmith, respectively). We confirmed that such a difference is not due to a comparison of a single individual to population estimates by recomputing the f4 ratio for each individual belonging to an Ethiopian population in our dataset. 
The absence of a West Eurasian component in Mota supports the dating of the backflow into Africa, which, at ~3.5kya, is younger than our ancient genome (dated to 4.5 kya). Given that Mota predates the backflow, it potentially provides a better unadmixed African reference than contemporary Yoruba. Thus, we recomputed the extent of the West Eurasian component in contemporary African populations using Mota . . . instead of Yoruba in our f4 ratio. By using this better reference, we estimated West Eurasian admixture to be significantly larger than previously estimated, with an additional 6-9% of the genome of contemporary African populations being of Eurasian origin. Importantly, this analysis shows that the West Eurasian component can be found also in West Africa, albeit at lower levels than in Eastern Africa. Importantly, a sizeable West Eurasian component is also found in the Yoruba and Mbuti, which are often used a representative of an unadmixed African population.
Ethiopians have more Eurasian admixture than other Africans, but essentially all modern Africans have significant levels of Eurasian admixture relative to Mota.

With respect to Neanderthal and Denisovan ancestry:
Given that Mota is our best example of an unadmixed African population, we used it as a reference to assess the affinity of a number of contemporary genomes with Neanderthals. We also investigated the effect of using Mota as a reference when estimating Denisovan introgression. We performed this analysis using the complete genomes (rather than a subset of SNPs as in earlier analyses), since a large number of SNPs is needed to obtain accurate estimates. . . . Both Yoruba and Mbuti were shown to have a small Neanderthal component, in line with their West Eurasian ancestry. As expected, estimates for French and Han were higher than for either of the two contemporary African genomes (from 0.21% in Mbuti to 2.96% in Han).
No evidence of any Denisovan ancestry was found in Mota or any of the other African samples tested with him as a reference for an unadmixed African genome.

The Nature of the Neolithic Eurasian Population Migrating To Africa

The fact that the inferred Eurasian component in other Africans determined with reference to Mota is similar to early Neolithic farmers in Europe is also notable.
Since we have in Mota an unadmixed African population, we can look for the origin of the West Eurasian backflow by modelling contemporary Ari as a mixture of Mota and possible source populations. 
We do this by using the admixture f3-statistics . . . from our global panel or a Eurasian ancient genome. For the latter, we used a representative of Mesolithic hunter-gatherers (Loschbour), and one of the Early Neolithic farmers (LBK, also known as Stuttgart); these two genomes were chosen for their high coverage, allowing us to use most of the SNPs available for contemporary populations and Mota.... 
LBK (an early Neolithic farmer) and Sardinians are the two most likely sources (showing the most negative admixture f3 values) for the Eurasian admixture in the Ari. A number of other analyses have shown Sardinians to be the closest contemporary population to early Neolithic farmers that came into Europe from the Near East, as contemporary populations from that region have been affected by large-scale populations movements in the last few millennia. Thus, the West Eurasian backflow originated from the direct descendants of the same early farmers who brought agriculture into Europe. Given that we have a putative source for the West Eurasian component, we can re-estimate its extent by using LBK as its source in our estimation of the f4 ratio . . . . without having to worry about West African ancestry in the source.

We next tested whether the West Eurasian component found in Yoruba, which had been previously suggested to be older than Mota [dated to 9.6k±1.8k yrs ago . . .], comes from the same source found for the Ari. We use the D statistics . . . . from our global panel or a Eurasian ancient genome. Sardinians and LBK were again found to be the most likely source of the West Eurasian component (giving the strongest positive values that indicate excess affinity between X and Yoruba compared to Mota). This result suggests that there was a single source for the West Eurasian component found throughout Africa.
The Copper and Bronze Age steppe and indigenous European hunter-gatherer population sourced admixtures that transformed the gene pool of early Neolithic Europe did not, by and large, extend to Africa.  But, given recent ancient DNA results from Neolithic Western Anatolia and the shifts that Near Eastern populations have seen genetically since then, I'm inclined to call a population that is similar to LBK and Sardinian individuals, Western Anatolian rather than Near Eastern.

This is consistent with previous studies showing that Eurasian admixture in East African and Khoisan people was more similar to Levantine people than to South Arabians and (except among highly West Eurasian Ethio-Semitic individuals with some South Arabian affinities) was of a uniform character throughout Africa similar in proportion and type to modern Omotic people. It would be interesting to see, however, if the Chadic people who live mostly in the Sahel between North Africa and Sub-Saharan Africa, have different autosomal Eurasian affinities to match their unique Y-DNA Eurasian affinities.

As long as the people with Early European Farmer type genetic began their migration that culminated in sub-Saharan Africa before the influx of Steppe-like people into Europe, this doesn't pose a paradox. As summarized in a blockbuster paper earlier this year:
By ~6,000-5,000 years ago, a resurgence of hunter-gatherer ancestry had occurred throughout much of Europe, but in Russia, the Yamnaya steppe herders of this time were descended not only from the preceding eastern European hunter-gatherers, but from a population of Near Eastern ancestry. Western and Eastern Europe came into contact ~4,500 years ago, as the Late Neolithic Corded Ware people from Germany traced ~3/4 of their ancestry to the Yamnaya, documenting a massive migration into the heartland of Europe from its eastern periphery. This steppe ancestry persisted in all sampled central Europeans until at least ~3,000 years ago, and is ubiquitous in present-day Europeans.
Allowing at least 500-1,500 years for a group of Early European Farmer-like people to migrate from Western Anatolia to Ethiopia before the resurgence of hunter-gatherer ancestry or the steppe ancestry had changed the European gene pool is not an unreasonable scenario. This trip involves a march of about 1200 miles more or less due South (although, obvious, the route would not be as the crow flies).

This is comparable to the time needed for Early European farmers to advance that far (i.e. to the Northern coast of Continental Europe and Southern Scandinavia) and with that much of a change in latitude in Europe during the first wave of the Neolithic revolution in Europe.

The time depth and distribution of Y-DNA T (which is present at relatively high levels on Omotic and Cushitic populations relative to Ethio-Semitic populations) suggests that this may have been an important Y-DNA haplogroup of the EEF-like Neolithic farmers whose autosomal DNA contributed to Africa's gene pool via the Levant, possibly with Y-DNA J mixed in (although the multiple possible historical events that could have spread Y-DNA J complicate the analysis).  But, Y-DNA T is too young to be a plausible candidate accompanying the spread of mtDNA M1 (as has been suggested by some) and U6 in their migrations back from Eurasia ca. 30,000 years ago, and is a poor fit to mtDNA clades that probably arrived in Africa via Iberia and then spread across North Africa to East Africa.

Y-DNA F* is basically absent from Africa, and Y-DNA I, while old enough, has a distribution that is to thin and patchy to be a very strong candidate for a companion to mtDNA M1 and U6.

Y-DNA J has about the right geographic spread in Africa to match mtDNA M1 and U6 as part of the same back migration, but it is hard to know how much of Y-DNA J is due to Semitic migration to Africa (Ethio-Semitic and Phoenician first, and then Arab later) in the last 4,000 years, how much is due to earlier Neolithic and Paleolithic migrations.  Another possibility is that a Y-DNA E population migrated to Iberia early in the Upper Paleolithic era (where it left genetic traces) and then back migrated to NW Africa ca. 30,000 years ago with mtDNA M1 and U6 women from Europe.

The Origins of African Herding And Farming

The apparent timing of the Eurasian Neolithic admixture in almost all modern Africans is also relevant to determining what role, if any, migrant people from food producing societies played in the conversion of wild African plants into domesticated crops that sustained early African-style farming.

Well dated plant remains can determine when domestication happened, and this can be compared to the apparent dates of admixture of people who resembled Early European Farmers genetically (probably about 1500 BCE - 2500 BCE with the Ethio-Semites arriving closer to 1500 BCE and the best guess for other farmers via the Nile closer to 2000 BCE).

Some of the data on the switch to food production is here. Cattle reached Egypt in the earliest part of the Neolithic revolution in Africa around 7000 BCE, the donkey was locally domesticated around 6000 BCE, and sheep and goats appear around 5000 BCE.

Fertile Crescent crops were mostly unsuited to sub-Saharan climates, so farming came much later to African than herding. Dillon (2007) argues that "The domestication of sorghum has its origins in Ethiopia and surrounding countries, commencing around 4000–3000 BC." Pearl millet cultivation became ca. 3200-2700 BCE in Africa and started in the West with a transfer to the East and to India by 1700 BCE. An early type of locally developed farming in Ethiopia was originally conducted primarily by Omotic people and was flourishing when the Ethio-Semites arrived, but was largely displaced and set aside when Ethio-Semites brought their more developed farming techniques to the Ethiopia ca. 1500 BCE.

The balance of the evidence, therefore, favors the development of African domesticated plants after Fertile Crescent pastoralist populations arrive in Africa, but before a major demic contribution of people with Early European Farmer genetics.

Of course, Mota, because he lived in such an isolated area, could have been one of the last unadmixed people of Africa when he died.  There is a decent chance that there would have been some Omotic farmers within a couple hundred miles or so of the place he died at the time.

Functional Traits

Functional traits discerned from Mota's genome include the following:
Skin colour could not be determined although Mota did not have common European variants associated with light skin colour (rs16891982 and rs1426654). Mota was determined to have had brown eyes (p-value = 0.997) and dark (p-value = 0.996), probably black (p-value = 0.843) hair. . . . Mota did not have any of the major alleles known to cause lactase persistence. . . . Mota . . . lived at high altitude and was . . . likely adapted to hypoxia.
Thus, Mota was probably black, brown eyed and dark haired, lacked lactase persistence associated with many herding and farming populations in Africa, and was genetically adapted to high altitudes.