From Hunter-Gatherers to Romans: Testing Northern and Central Italian DNA

Breaking a modern genome into ancient ancestries is easy. Knowing when to trust the result is much harder.

Online ancestry reports love a dramatic cast list: Viking, Celt, Roman, Western Hunter-Gatherer, Anatolian Farmer. The percentages look wonderfully precise. The trouble is that an analyst can keep swapping ancient populations until a seductive combination appears. If the failed attempts vanish and only the prettiest bar survives, the result tells us as much about the search as it does about the genome.

PifferPilfer is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

I chose modern Italians as my test case, but the same procedure can be used for any population. I begin with sources that make historical sense, test deep, prehistoric and historical ancestry separately, and add components one at a time. When a new source fails to improve the simpler model, I stop.

I wanted to ask those questions in order. First: which deep ancestry components werThanks so much for supporting my work!e actually needed? Second: which Neolithic, Bell Beaker, Corded Ware or Yamnaya-era populations could model the prehistoric layer? Only then did I ask whether a northern historical source improved an Iron Age Italian baseline. The source set changed at each step, so a Roman population never shared an ancestry bar with WHG or Anatolian farmers.

Figure 1. Digging through the layers of ancestry

qpAdm, in plain English: a statistical test that asks whether a target genome can be represented by one or more ancient populations chosen in advance. It compares patterns of shared genetic variation against a separate reference panel and reports both model fit and source weights.A model with p < .05 is rejected. A model above that threshold is merely not rejected: it is not automatically true, unique or historically best. Its percentages are weights assigned to the chosen genetic sources.

Layer 1: Deep ancestry

I began with the deepest ancestry layer: Western and Eastern hunter-gatherers, Anatolian Neolithic farmers, Iranian Neolithic ancestry and a western Yamnaya-related proxy. These are broad genetic building blocks, not historical peoples, and no Etruscan, Roman, Celtic or Germanic source appears in this part of the analysis.

Figures 2 and 3 compare like with like: five northern individuals—one each from Ferrara, Cremona, Domodossola, Taio and Treviso—and five Central Italian individuals from Bolsena, Ascoli Piceno, Montepulciano, Perugia and Chiusi. The regional pools appear only in the historical section.

All ten individuals retained a two- or three-source deep model. Anatolian Neolithic and western Yamnaya-related ancestry appeared throughout. Iranian Neolithic ancestry was also retained for the Emilia-Romagna and Veneto representatives and all five Central Italian individuals; it was not needed for the representatives from Lombardy, Piedmont or Trentino–Alto Adige. WHG and EHG were available to the search but did not earn a separate place in the retained models.

Figure 2. Deep ancestry models for northern and Central Italy


Layer 2: Prehistoric ancestry

I then moved one rung closer in time and changed the source set completely. The northern targets were tested with Neolithic Italian, Bell Beaker, Corded Ware and regional Yamnaya sources. For the Central Italian targets I declared a separate, period-matched menu: Central Italian Neolithic and Bronze Age groups from Lazio and Sicily, Adriatic-Balkan groups from Cetina and Mokrin, and Aegean Late Bronze Age groups from Aegina, Koukounaries and Chania. Historical populations were excluded.

This layer was much less forgiving. Two northern individuals retained a passing, feasible model: the Lombardy representative was estimated as 53% Central Italian Neolithic plus 47% Corded Ware [model p = .309], while the Trentino–Alto Adige representative was estimated as 32% Central Italian Neolithic plus 68% Bell Beaker [model p = .775]. With the separate Central-Mediterranean menu, Bolsena also passed: 44.4% Mokrin Early Bronze Age plus 55.6% Aegina Late Bronze Age [model p = .0714]. The other seven individuals had no passing, feasible prehistoric model.

Figure 3. Prehistoric ancestry models for northern and Central Italy


Layer 3: Historical ancestry

Only at the final layer did I introduce Etruscans and Roman-Republic Italians, then ask whether Celtic- and Germanic-period sources improved those Iron Age Italian baselines. The figures that follow answer that narrower historical question. Deep and prehistoric sources do not appear in these bars.

Historical models: the northern regional pools

I first pooled modern AADR individuals by region: Emilia-Romagna (n = 7), Lombardy (n = 3), Piedmont (n = 4), Trentino–Alto Adige (n = 2) and Veneto (n = 3). Pooling gives a regional signal, although the last two pools are small and should not be mistaken for population surveys.

Exactly what each pool contains: Emilia-Romagna: 7 individuals; Lombardy: 3; Piedmont: 4; Trentino–Alto Adige: 2; Veneto: 3—19 people in total. Every individual row later in the article represents one person (n = 1), not another pool.

Figure 4. Historical additions retained for five northern Italian regional pools

Emilia-Romagna, Lombardy and Piedmont retained no northern addition in any of the four historical families. That result does not deny Celtic or Germanic ancestry in those regions. It says that, with these sources and this statistical panel, none of the additions improved an already viable Iron Age Italian baseline strongly enough to clear the stopping rule.

Trentino–Alto Adige was different. One model assigned 72.9% to the Roman-Republic baseline and 27.1% to a Hungary La Tène proxy [model p = .658; Holm p = .0227]. A separate, later comparison assigned 81.5% to late Lazio Etruscans and 18.5% to an early Roman-period group from Alken Enge in Denmark [model p = .158; Holm p = .0306]. These are alternative temporal models, not components to be stacked together.

Veneto also retained a Roman-period northern proxy: 75.0% late Lazio Etruscan and 25.0% Alken Enge [model p = .0515; Holm p = .000125]. Its early-medieval family produced 61.4% Roman-Republic and 38.6% Saxon [model p = .780; Holm p = .0227], but that bar deserves an asterisk because the early-medieval calibration behaved abnormally.

What happens when every pool member is tested?

The only fair way to interpret a pooled result is to test everyone inside it. I therefore repeated all four historical families for each of the 19 pool members—912 additional qpAdm models under the same SNP panel, source rotations and corrected stopping rule. No person was represented by a regional stand-in.

The negative pools were internally consistent. None of the seven Emilia-Romagna individuals and none of the three Lombardy individuals retained a northern addition. In Piedmont, three of four individuals—including Cuneo—retained nothing, but the Domodossola individual retained 24.8% Alken-Enge-related ancestry [model p = .480; Holm p = .0157]. A 42.0% Saxon-like early-medieval model also passed narrowly [model p = .0530; Holm p = .0434], but belongs to the technically unreliable early-medieval family.

Trentino now makes sense. The Roman-period pool estimated 18.5% Alken-Enge-related ancestry. Nogarè (Nei27) retained 23.9 ± 8.1% [model p = .0804; Holm p = .0422], whereas Taio (Nei18) estimated 14.6 ± 8.1% but did not improve the Italian baseline after correction [Holm p = .683]. The pooled Roman-period result is therefore consistent with the stronger signal in Nogarè.

The Trentino La Tène result is a different case. Nogarè estimated 29.1 ± 11.1% and Taio 25.3 ± 13.0%, but neither cleared Holm correction alone [Holm p = .0796 and .148]. The two-person pool estimated a similar 27.1 ± 8.9%; its smaller uncertainty allowed the addition to pass [Holm p = .0227]. Pooling did not invent a different coefficient—it made a shared but noisy signal precise enough to retain.

Veneto supplied the cleanest replication. All three members retained a Roman-period northern proxy: Bassano del Grappa 24.6%, Treviso 30.5% and Sappada 27.5%. Their pooled estimate was 25.0%. The pool’s 38.6% Saxon-like early-medieval result was not retained by any member separately and remains provisional because of the early-medieval calibration problem.

Figure 5. Northern Italian regional pools and all 19 people inside them

Historical models: Central Italy

The Central Italian comparison used five modern individuals from Bolsena, Ascoli Piceno, Montepulciano, Perugia and Chiusi. Four retained no northern historical addition at all. Their results were compatible with the sampled Iron Age Italian baselines without requiring Hallstatt, La Tène, Roman-period northern or early-medieval sources.

Bolsena was the sole exception, and only barely in ancestry terms: 98.9% late Lazio Etruscan plus 1.1% Hungary Longobard-like [model p = .335; Holm p = 1.44 × 10⁻⁹]. Alternative early-medieval models put the northern coefficient between 0.6% and 0.9%. A statistically sharp distinction can therefore correspond to a biologically tiny coefficient.

I would not build a migration story around that one percent. The synthetic early-medieval tests printed zero jackknife standard errors in every replicate, an obvious warning that this family was overconfident. The honest summary is that the five Central Italian genomes were almost entirely compatible with Iron Age Italian baselines, with one tiny and technically fragile exception.

Who were the Etruscans? They lived in city-states centred on Etruria—roughly modern Tuscany, northern Lazio and western Umbria—from about the eighth to third centuries BCE. They spoke a non-Indo-European language, strongly influenced early Rome and were gradually incorporated into the Roman Republic.


What the comparison actually shows

The cleanest contrast is not ‘Roman Central Italians versus Germanic Northern Italians.’ It is uniformity versus heterogeneity. The Central Italian individuals were consistently explained by Iron Age Italian baselines. Across the 19 northerners, robust retained additions were concentrated in Domodossola, Nogarè and all three Veneto localities, while fourteen individuals retained none.

Figure 6. Historical tests for five Central Italian individuals

*Early-medieval estimates are shown for completeness but treated cautiously because the synthetic calibration returned zero printed standard errors.

The retained northern weights were substantial where they appeared: about 18–30% in the most reliable La Tène and Roman-period results. Closely related ancient sources can sometimes substitute for one another, so each source label is best read together with its period and comparison family.

The sensitivity experiment was reassuring for the Roman-period family: simulated 5–30% northern mixtures were usually detected after correction, and no false northern component was retained at 0%. Hallstatt and La Tène models were less stable because some simulated mixtures failed the overall model test. The early-medieval family remains provisional because of the zero-standard-error problem.


Why the stopping rule matters

If I had simply displayed every passing two-way model, nearly every target could have acquired a colourful Celtic, Danish, Saxon or Longobard slice. A high p-value would then look like discovery. The nested test asks the more useful question: did that slice improve the simpler model enough to justify the extra story?

That rule removed most additions. It also preserved Roman-period northern signals in Domodossola, Nogarè and all three Veneto individuals, plus the corresponding Trentino and Veneto pools. A method that sometimes says yes and often says no is more informative than one designed to produce an ancestry cocktail for everyone.

So how Roman is Italy today? The cautious answer is that these modern genomes remain strongly compatible with sampled Iron Age Italian ancestry, especially in Central Italy. Northern Italy shows additional northern historical affinity in some places and people, but not as a universal layer.


Who is hiding in your DNA?

An Etruscan? A steppe herder? A suspiciously persistent Longobard?

Most ancestry websites produce a colourful cocktail of Romans, Vikings and hunter-gatherers, then quietly neglect to mention how many alternative recipes they tried first. I take the opposite approach: I test your genome against real ancient populations, discard models that fail, and stop adding ancestors when the evidence stops improving.

To celebrate this project, I am offering five new yearly subscribers a complimentary Personal Ancient-Ancestry Modelling Report.

Your report will explore three distinct chapters of your genetic history: your deepest hunter-gatherer and early-farmer foundations; your Neolithic, Bronze Age and steppe-related ancestry; and your affinities to Iron Age, Roman and early-medieval populations.

You will receive the models that worked, the ones that failed, and a plain-English explanation of the results. The same stopping rule used in this article will be applied before the report is written.

Become a yearly subscriber and apply: email pifferdavide@gmail.com with the subject ‘Ancient DNA report’ and the address used for your subscription. I will reply with private upload instructions. Please do not attach your genome to the email.

Raw DNA files will be submitted privately, used only to prepare the report and deleted afterward. The analysis concerns population history—not health, paternity or medical risk.

Subscribe to keep digging through the layers of human ancestry, and become a yearly subscriber for a chance to receive your own qpAdm report.


Methods note

To limit model fishing, I kept the 173,185-SNP AADR v66.HO-aligned panel, right-population set, source definitions and 200 bootstrap replicates fixed within each declared model family. Deep, prehistoric and historical layers remained separate, and source groups generally required at least four unrelated, quality-controlled individuals. Figures 2 and 3 use ten single-person targets: five northern and five Central Italian individuals. Each faced 31 deep combinations. The northern prehistoric menu contained 1,023 combinations per person; the revised Central-Mediterranean menu contained 47 predeclared combinations per person, or 235 models in total. Its groups contained four Central Italian Neolithic individuals, four Lazio Bronze Age, eight Sicily Early Bronze Age, eight Cetina Middle Bronze Age, twelve Mokrin Early Bronze Age, four Koukounaries Late Bronze Age, four Aegina Late Bronze Age and seventeen Chania Late Bronze Age individuals. The historical section then switches explicitly to five northern regional pools, each screened across 98 historical combinations, before testing all 19 constituent individuals separately. That individual extension was historical only: 912 models and 1,444 nested comparisons.

Expanded historical models were retained only when the overall fit was p ≥ .05, every coefficient lay between zero and one, the sources were distinguishable, and the improvement over the matching one-source Italian model survived both Benjamini–Hochberg and Holm correction. In Figures 4 and 6, a full red bar shows the retained one-source Iron Age Italian baseline; a blue segment appears only when a northern addition clears the full rule. Figure 5 uses a dash for a northern addition that was not retained.


References

Antonio, M.L., Gao, Z., Moots, H.M. et al. (2019). Ancient Rome: A genetic crossroads of Europe and the Mediterranean. Science, 366(6466), 708–714. https://doi.org/10.1126/science.aay6826

Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045. https://doi.org/10.1093/genetics/iyaa045

Mallick, S., Micco, A., Mah, M., Ringbauer, H., Lazaridis, I., Olalde, I., Patterson, N. & Reich, D. (2024). The Allen Ancient DNA Resource (AADR) a curated compendium of ancient human genomes. Scientific Data, 11, 182. https://doi.org/10.1038/s41597-024-03031-7

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论