Tuesday, July 28, 2026

The Linguistic Golden Age

The most notable findings of the study are that the number of languages in the world was already starting to decline five hundred years before the Columbian exchange and the Renaissance, and that small pre-agricultural populations placed an effective limit on the number of languages spoken in the pre-Holocene era. The peak number of languages was on the order of 10,000 to 35,000 according to the study under various model assumptions.

A linguistic “golden age” flourished between 1,000 to 3,000 years ago when tens of thousands of languages were spoken throughout the world, according to a new study coauthored by Yale linguist Claire Bowern that traces trajectories in global language diversity over the past 12,000 years.

The golden age was followed by a period of rapid decline in linguistic diversity that coincided with the rise of large states and multinational empires, such as the Roman Empire, the researchers found. The finding challenges a commonly held view that widespread language extinction began later, about 500 years ago, with the onset of European colonial expansion.

The languages of expanding states and empires were disproportionately likely to survive while those of absorbed, displaced, or declining populations disappeared, the researchers concluded. This means that the roughly 7,600 languages spoken or signed today represent a small and historically biased sample of the languages that once existed, which has important implications for the study of global patterns in language and culture, they said. . . .

The study, which was published July 23 in the journal Science, is the result of a long international collaboration between experts in language, demography, and evolutionary biology. . . .

Direct evidence of languages doesn’t exist before the appearance of writing about 6,000 years ago. So for the study, the researchers combined ethnographic information, estimates of prehistoric population size, and statistical and social-computational modelling to estimate linguistic diversity over the course of the Holocene — the current geological epoch that began about 12,000 years ago.

To gauge the likely distribution and sizes of populations before the advent of agriculture, the researchers used ethnographic data from 171 hunter-gatherer and fisher societies whose traditional subsistence and mobility had not been profoundly transformed by contact with food-producing populations. They combined these estimates with independent analyses suggesting that the global human population 12,000 years ago was between approximately 4.4 and 7 million. Assuming that the number of languages in the early Holocene corresponded with the number of distinct groups that the total human population could accommodate at the time, between 4,500 and 6,200 languages were spoken 12,000 years ago, the study’s models estimated.
The world immediately before agriculture was probably not exceptionally rich in languages,” said lead author Damián Blasi, an ICREA (Catalan Institution for Research and Advanced Studies) research professor based at the Center for Brain & Cognition at the Pompeu Fabra University in Barcelona, Spain. “There were most likely fewer languages than there are today. Linguistic diversity then grew alongside the human population for thousands of years.”

To capture fluctuations in linguistic diversity as populations expanded dramatically following the onset of agriculture and technological innovation and then contracted through wars, plagues, and other ecological factors, the researchers modelled thousands of possible trajectories connecting language estimates at the beginning of the Holocene with the present.

The models consistently produced patterns indicating that language diversity peaked between 1,000 and 3,000 years ago when upwards of tens of thousands of languages were being spoken worldwide, this period that the researchers call the world’s linguistic “golden age.”
That golden age ended with the rise of states and multinational empires as the dominant groups spread their languages, cultures, and pathogens to previously independent populations, the study shows. European colonialism intensified the steep decline in linguistic diversity, but did not cause it, according to the study.

The loss of so much linguistic diversity over the course of two millennia affected the development of today’s languages, the researchers said.

“The languages we see today are the survivors of a massive and highly selective historical bottleneck,” said study coauthor Russell Gray, director of the Department of Linguistic and Cultural Evolution at the Max Planck Institute for Evolutionary Anthropology in Leipzig, Germany. “A linguistic feature may be common not because it is inherently efficient or better for communication, but because it happened to be carried by populations that expanded. Extinction may have shaped linguistic diversity much more profoundly than we previously appreciated.”
"Study uncovers lost ‘golden age’ of languages: A new study coauthored by Yale linguist Claire Bowern suggests that tens of thousands of languages were spoken between 1,000 and 3,000 years ago." via a Yale University press release.

Conventional linguistic reconstruction methods struggle to reconstruction linguistic pre-history much before 4,000 to 8,000 years ago.

The 7,600 languages spoken today is rapidly falling. A very large share of all languages spoken today have a small number of speakers or are moribund and on the brink of extinction in the next couple of generations or sooner. 
The general consensus is that between 6,000 and 7,000 languages are currently spoken. Some linguists estimate that between 50% and 90% of them will be severely endangered or dead by the year 2100. The 20 most common languages, each with more than 50 million speakers, are spoken by 50% of the world's population, but most languages are spoken by fewer than 10,000 people. . . . More than 50% of the world's endangered languages are located in just eight countries: India, Brazil, Mexico, Australia, Indonesia, Nigeria, Papua New Guinea and Cameroon.

From Wikipedia. About 107 languages are spoken by 7 million people or more. A list of the number of languages per country can be found here. The median Native American language in the U.S. has 12 speakers. The median Australian Aboriginal language has 10 speakers. the median Brazilian indigenous language has 210 speakers. The median Papuan language has 1,315 speakers. The median Indonesian language has 3,500 speakers.

Reviving an extinct or endangered or purely liturgical language is not impossible but is very difficult and there have been only about twenty moderately successful attempts to do so. The most successful effort has been the Hebrew language.

The Model Dependence Of Cosmological Neutrino Mass Estimates

Even with greatly relaxed cosmology based bounds in neutrino mass that free it from the strong model dependence it has in the deeply flawed ΛCDM model, cosmology based limits on neutrino mass have a 90% confidence level upper bound that is about five times more strict than direct measurements of neutrino mass.

The relaxed less model dependent bound on neutrino mass still limits the lightest neutrino mass to about 60 meV in an inverted hierarchy scenario, and to about 73 meV in a normal neutrino mass hierarchy. In the restrictive and model dependent ΛCDM model estimation of the neutrino masses, there is a normal neutrino mass hierarchy and the lightest neutrino mass can't be much more than 2 meV which a best fit value that is much smaller than that.

Very low neutrino masses greatly limits the role that neutrinos can play to make up the gap between modified gravity theories that account for most, but not all, dark matter phenomena, such as MOND.
Neutrino oscillations establish that neutrinos are massive, providing the only laboratory detection of physics beyond the Standard Model. Direct kinematic experiments bound the electron-neutrino mass to m(νe) < 0.45 eV (KATRIN, 90% CL), implying ∑m(ν) ≲ 1.3 eV. 
Conversely, cosmology within ΛCDM is highly constraining: Planck CMB, CMB lensing, and DESI DR2 BAO yield ∑m(ν) < 0.056 eV (95% CL), in 2-3σ tension with the inverted-ordering floor (0.10 eV). However, this bound relies on ΛCDM, while data hint at an evolving dark energy. 

To determine the model dependence of cosmic neutrino mass bounds, we deconstruct each probe's sensitivity to late-time physics and pursue two robust routes to a ∑m(ν) bound: 
(i) The existing dark-energy-marginalized route, retaining all data and marginalizing over (w(0),w(a)), is shown to also be immune to flexible binned and cubic w(a) histories, yielding ∑m(ν) < 0.152 eV, sharpening to σ(∑m(ν)) ≈ 0.043 eV with Simons Observatory lensing and Spec-S5 BAO.
(ii) A new late-Universe-free route combines primary CMB, marginalizing over acoustic-peak smoothing via Alens, with the reconstructed lensing spectrum CκκL, removing late-time expansion dependence by construction. This yields ∑m(ν) < 0.41 eV today, tightening to 0.31 eV (Simons Observatory) and 0.28 eV (cosmic-variance limit) across all tested dark-energy models. These relaxed bounds trade statistical power for model independence. Interestingly, they land in the sensitivity range targeted by next-generation laboratory experiments like Project 8 (m(νe) ∼ 0.1 eV), motivating vital synergies between future cosmological and terrestrial neutrino measurements.
Frank J. Qu, et al., "Measuring Cosmic Neutrino Masses Independently of Dark Energy" arXiv:2607.24742 (July 27, 2026).

Monday, July 27, 2026

The Pre-Greek Substrate

A nice Facebook reel discusses the pre-Greek substrate of the Indo-European Greek language for which we have about 1000 Greek words that don't appear to have Indo-European origins as evidence.


The geographic extent of the pre-Greek substrate

It notes the geographic extent of the substrate (which is similar to that of modern Greece plus the west coast of Turkey), and the subject matter of the substrate words including animals, plants, toponyms, terrain types, large sea related words, other natural phenomena, proper names, and particularly notably, most of the gods in the Greek pantheon. It also notes distinctively non-Indo-European phonemes found in many pre-Greek substrate words like "nth" and "ss" which were suffixes in the pre-Greek substrate.

My intuition is the the pre-Greek substrate had a larger than average impact on Greek compared to other Indo-European languages, but probably less than the Indo-Anatolian languages. The fact that many god names were borrowed into Greek also suggests that this may be a general feature also shared by Indo-Aryan (i.e. Sanskrit derived) and Indo-Iranian language, and perhaps by Germanic languages as well.

Some substrate source words:



Davidski At Eurogenes On The Indo-European Homeland

Davidski at Eurogenese has provided some thoughts on the Indo-European homeland in a teaser post with an updated large data set of ancient and modern DNA from Europe. He states:

(Interactive version here)

- Yamnaya is almost certainly derived from Serednii Stih (aka Sredny Stog). I've been talking about this for years and you can probably pick it up without too much trouble in the PCA above. Finally, even Iosif Lazaridis, David Reich, Nick Patterson and friends are now on board with this idea as per their latest paper on the topic (see Lazaridis et al. 2025).

- However, I'm not convinced by the Lazaridis et al. hypothesis that the Caucasus–lower Volga (CLV) cline is intimately linked to the Indo-Anatolian and proto-Anatolian expansions. That's because the CLV cline is an artifact of isolation-by-distance working for thousands of years across a vast and highly diverse landscape, and it includes a wide range of populations that usually have nothing to do with each other.

- The question of who actually spoke Indo-Anatolian first will, in all likelihood, never be solved to everyone's satisfaction. But the number of true candidate groups is actually quite small, and I reckon they're almost certainly all highlighted in my PCA above. Wink, wink, nudge, nudge.

My longstanding hypothesis is that Indo-Europeans established themselves in Anatolia close in time to the first historical attestation of their presence in the 18th century BCE, and that the Indo-Anatolian languages are more divergent than other Indo-European languages due to a stronger substrate effect from the Hattic language and related pre-Indo-European languages from Anatolia, because the substrate culture that they conquered was a vital and functioning Copper Age culture rather than the nearly collapsed Neolithic cultures that other Indo-Europeans encountered as a substrate. I believe that conventional wisdom in the field of historical linguistics attributes too much of the divergence between the Anatolian languages and other Indo-European language to the time depth from their common ancestor dialect.

New Combined Standard Model Constant Measurements

A new study tries to extract several Standard Model Constant measurements from the same data set and obtains results generally consistent with prior efforts to measure the same constants, but with greater uncertainty than the state of the art measurements of these quantities.

X17 Hypothesis Less Constrained Following Reanalysis of Data

The conclusion of a new reanalysis of X17 relevant experimental data loosens the constraints greatly. But, I remain deeply skeptical of its existence.

Also, why use units of 10^-3 GeV instead of MeV in your key chart?

[T]he revised exclusion limits no longer extend to the X17 mass of 16.88 MeV, leaving a sizeable region of parameter space for the vector-boson interpretation of the anomaly, bounded by the NA64 and Orsay constraints. Interestingly, the remaining allowed interval, 6.5 × 10^−5 ≲ εe ≲ 1.1 × 10^−4, is compatible with the recent preliminary measurement of the X17 lifetime. We further observe that, for these coupling values, the X17 decay length for a 100 GeV/c e− beam-dump experiment is of the order of few meters, resulting to an invisible signature for a missing-energy setup such as NA64-e [51]. The X17 parameters space not any longer constrained by E141 could thus be explored in the near future by NA64 through the invisible-mode dataset accumulated so far by the experiment.

Since its first observation by the ATOMKI experiment in 2018, the ∗Be anomaly has attracted considerable interest within the dark sector community as it may indicate the existence of a new fundamental particle with a mass of about 16.9 MeV, the X17. However, the minimal model describing X17 as a new vector boson is severely constrained by null results from legacy beam-dump experiments. Among these, the E141 experiment at SLAC places stringent limits on the X17 coupling to electrons εe in the 5.1 × 10^−5 ≲ εe ≲ 1.7 × 10^−4 region. 
This excludes the possibility of a long-lived X17, potentially in contrast with the preliminary estimate of the particle lifetime recently reported by the ATOMKI collaboration. The E141 limits commonly adopted in the literature rely on reinterpretations of the original analysis under solid but simplifying assumptions. While these studies provide a reliable estimate of the experiment's reach, they rely on approximations that were well justified when the boson mass was largely unconstrained. With the X17 mass now confined to a narrow region by recent experimental observations, a more refined treatment of the E141 sensitivity becomes necessary. 
In this work, we revisit the E141 exclusion limits in the X17 scenario by performing a dedicated reanalysis that incorporates a more accurate treatment of the experimental setup and signal prediction. We quantify the impact of these refinements on the excluded parameter space and discuss their implications for the compatibility between the E141 constraints and the vector boson interpretation of the ATOMKI anomalies.
A. Celentano, A. Marini, L. Marsicano, "Updated E141 constraints on a long-lived X17 vector boson" arXiv:2607.22102 (July 24, 2026).

Another recent X17 paper highlighted by @neo is as follows:

The X17 particle has been proposed to explain the invariant mass anomalies observed in electron-positron pairs during nuclear transitions at the Atomki experiment. Motivated by recent observations of 8B solar neutrinos induced coherent elastic neutrino-nucleus scattering (CEνNS), we present the first comprehensive analysis of the hypothetical boson using data from multi-ton dark matter direct detection facilities. 
We consider the new particle as a light Z′ mediator arising from a spontaneously broken U(1)′ symmetry, featuring both vector and axial-vector couplings to leptons. By evaluating the latest datasets from XENONnT, PandaX-4T, and LUX-ZEPLIN, we derive stringent limits on the effective vector coupling utilizing marginalization procedures. Our global analysis provides competitive constraints that meaningfully narrow the allowed parameter space of the model, while exhibiting a clear sensitivity to the tau-flavor coupling.
M. F. Mustamin, M. Demirci, M. Deniz, "Signatures of X17 through Coherent Elastic Solar Neutrino-Nucleus Scattering in Direct Detection Searches"arXiv:2607.12691 (July 14, 2026).

The conclusion of this article states:
Motivated by the first observation of coherent elastic neutrino-nucleus scattering induced by 8B solar neutrinos, we have derived robust constraints on the effective couplings of a light Z′ gauge boson, interpreted here as the X17 particle. Initially conjectured to explain the invariant mass anomalies observed by the Atomki experiment, this hypothetical mediator can be handled within a U(1)′ gauge symmetry framework. Crucially, we have explicitly accounted for the flavor-dependent nature of the couplings, which naturally arise from solar neutrino flavor transitions during their propagation to Earth. 
Our analysis utilized the latest datasets from direct detection searches at XENONnT, PandaX-4T, and LUX-ZEPLIN experiments. While these multi-ton direct detection facilities were fundamentally designed to search for weakly interacting massive dark matter particles, their recent milestone in detecting solar neutrino-induced nuclear recoils offers a novel and powerful avenue for probing BSM physics. We systematically evaluated the flavor-dependent effective couplings—Cνe eff, Cνµ eff, and Cντ eff—that quantify the X17 interactions with the xenon target. By rigorously marginalizing over the relevant effective neutrino couplings, we have derived comprehensive 1 dof and 2 dof limits of the effective vector couplings using each dataset released by the experiments. Our global analysis reveals that PandaX-4T and LUXZEPLIN provide relatively similar behavior, while the combined datasets effectively break individual experimental degeneracies. Most importantly, we have found that the SM hypothesis remains completely protected and consistent within the 90% CL allowed regions across all flavor combinations. Overall, our derived bounds are highly competitive with the allowed parameter spaces previously mapped by IceCube and COHERENT+reactor studies. We also highlight the sensitivity of these detectors to the tau-flavor coupling, a direct consequence of the oscillated solar neutrino flux.
A non-lepton universal coupling seems particularly dubious.