Rome and the Northern Catch Up Reanalysed

In July 2026 I published Rome Was Ahead Then the Barbarians Caught Up. https://substack.davidepiffer.com/p/rome-was-ahead-then-the-barbarians The article compared genetic scores for educational attainment in ancient north-central Italy and northern Europe. It found an early Italian lead followed by convergence in the Imperial and post-imperial centuries.

The result was provocative, but the northern sample had an obvious weakness. It contained 427 people, nearly four in five from Denmark, and none from Germany. A claim about northern Europe rested overwhelmingly on Scandinavia.

This follow-up puts the claim to a harder test. The source cohort now contains 1,916 eligible ancient people, including 1,556 from Denmark, Norway, Sweden, Germany, Belgium and the Netherlands. After the primary imputation-quality screen, the analysis retains 1,454 people: 1,231 from continental northern Europe and 223 from north-central Italy.

The main conclusion survives, although in a more qualified form. The larger sample still supports an Italian lead in the Iron Age, and the narrower Republican Rome-region comparison points in the same direction. Northern scores rise relative to Italy across the full span. The early contrasts are stable when I change the quality filter, but the size and statistical strength of the long-run catch-up are more sensitive.


A much larger northern sample

The previous article’s 427 northern Europeans included 337 people from present-day Denmark, 65 from Norway and 17 from Sweden, with only eight from Lithuania and Finland. The new continental northern group is larger and geographically broader. Germany, Belgium and the Netherlands now appear alongside Scandinavia, while Britain remains reserved for a separate analysis.

The combined source cohort covers 5000 BCE to 950 CE. It includes 1,556 eligible continental northerners and 360 eligible Italians north of 41.5 degrees latitude. Sardinia and explicitly Longobard-labelled Italian samples are excluded. These are geographical categories, not ethnic labels. Continental north does not mean that every person was Germanic, and north-central Italy does not mean that every Italian was Roman.

The narrower Republican comparison therefore restricts the Italian side to the Rome region, corresponding to present-day Lazio. Even there, most retained burials are associated with Etruscan contexts rather than residents of the city of Rome. The point is to test whether the broader Italian result survives a more historically focused geographical restriction, not to treat the regional sample as a direct sample of Rome’s urban population.

The expanded dataset draws partly on the 2026 collection published by Akbari and colleagues and partly on additional ancient genomes in my existing collection. Duplicate representations were consolidated, and dates were checked against the source studies. The primary analysis retains 1,454 people after the imputation-quality screen explained below. Figure 1 shows their distribution.

Figure 1 Geographical distribution of the primary sample


How the reanalysis works

The outcome is an EA4 polygenic score, meaning a genetic score for educational attainment based on the Okbay and colleagues 2022 genome-wide association study. It records the weighted frequency of modern education-associated alleles across the available ancient genotypes.

Ancient genomes are incomplete, so many genotypes must be imputed. Instead of converting each imputed genotype into a single hard call, I calculate the score from its posterior dosage: the expected number of copies of the scored allele given the sequencing data, nearby variants and the reference panel. A genotype assigned with modest confidence therefore contributes differently from one assigned with near certainty.

Posterior dosages preserve uncertainty, but they do not make poor imputation harmless. I therefore apply a sample-level imputation-quality score, abbreviated IQS only from this point onward. In this analysis it measures the average posterior confidence of heterozygous calls across the 3,876 variants used for EA4. The primary sample requires an EA4-panel IQS above 0.9. This is a score-panel version of the quality measure used by Akbari and colleagues, not an exact reproduction of their genome-wide statistic. A separate sensitivity analysis instead retains people with documented genome coverage above 0.1 times.

The ancestry adjustment has also changed. The earlier approach used principal components, or PCs, which reduce genome-wide variation to a few major axes. Here I replace that PC adjustment with a genetic relationship matrix, or GRM, built from 1,989 neutral markers. The GRM models pairwise genetic similarity across the whole sample rather than asking a small number of PCs to summarize it.

I report two kinds of result. The primary site-clustered model adjusts for date and source but does not remove genome-wide ancestry differences, because ancestry turnover may be part of the historical change being studied. The GRM model asks the narrower question of what remains after conditioning on genetic similarity. It is an ancestry-conditional sensitivity, not the uniquely correct answer.

This distinction matters because a GRM can overcontrol for ancestry in a historical comparison. If migration or population turnover changed the ancestry composition of Italy or northern Europe, and that ancestry carried different EA4 scores, the GRM may absorb part of the historical change itself. A smaller GRM estimate is therefore not automatically a less biased estimate of the total population change. It is an estimate conditional on genome-wide genetic similarity. The opposite limitation applies to the site-clustered model: it preserves ancestry-associated change but cannot distinguish migration from selection, residual stratification or other causes. I show both because they answer different questions.


What the larger sample looks like

Figure 2 plots the individual posterior-dosage scores. The curves summarize the raw observations and show why a single early-versus-late contrast is inadequate. The samples are uneven through time, and the two regions do not follow parallel trajectories.

Figure 2 Individual EA4 scores through time


The Iron Age Italian lead

The common comparison windows are 5000–4000 BCE, 4000–1000 BCE, 1000–28 BCE, 27 BCE–475 CE and 476–950 CE. These are shared bins for comparison, not a claim that every region passed through the same archaeological transition on the same date.

Table 1 gives the continental-north-minus-Italy contrast in each period and the narrower Republican comparison with the Rome region. Negative estimates mean higher Italian scores. In the 1000–28 BCE window, the primary estimate is minus 0.368 standard deviations, with a 95 percent interval from minus 0.629 to minus 0.108 and p=.0063. The corresponding q value is .0127 within the primary model family.

The GRM estimate is minus 0.281, with an interval from minus 0.569 to +0.006 and p=.0549. Both estimates favour Italy, but the ancestry-conditional interval just includes zero. The result is therefore clearest as a total observed difference in this pooled sample and weaker after conditioning on genome-wide genetic similarity.

Table 1 Continental north minus Italy and Rome region comparisons

Figure 3 presents adjusted means from the primary site-clustered models. Both groups are evaluated at the same mean date and source mixture inside each period. The connecting lines are visual guides, not a separate temporal model.

Figure 3. Adjusted EA4 means across the five periods


The Republican Rome region comparison

The narrower Republican check covers 509–28 BCE and restricts the Italian side to the Rome region. After the IQS filter, it compares 27 people from this region with 33 northerners. The primary north-minus-Rome-region estimate is minus 0.507 standard deviations, with a 95 percent interval from minus 0.769 to minus 0.245 and p=.00043.

The GRM estimate is minus 0.488, with an interval from minus 0.896 to minus 0.080 and p=.020. The Italian-side lead therefore survives the ancestry-conditional model in this narrower comparison. The Rome-region analysis overlaps the broader Iron Age sample and is not an independent replication.

The headline contrast is only the beginning. The full analysis asks whether the apparent catch-up reflects northern rise, Italian decline, or both—and whether the result survives stricter genetic and data-quality checks. Upgrade to continue reading.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论