Am I 100% Republican Roman? It depends.
After testing ancient ancestry models of Italian genomes in my September 3 article, I wanted to put my own DNA through the same exercise. I had my 23andMe data, ancient genomes to compare it with, and a fairly clear idea of what I expected to find: Iron Age Italian ancestry, with some Celtic or Germanic input.
The interesting part, for me, was how much of that northern contribution would show up. Would it be a small addition to an otherwise local Italian profile, or would it account for a substantial part of the result? And could the comparison tell me anything about whether that affinity was closer to the Celtic or Germanic side of my expectations? I wanted to see how far the data would take me.
Then I got a result I had not anticipated. My genome fitted an ancient Roman/Italic source on its own, without needing an added northern source to pass the test. Expressed as an ancestry bar, that is a rather arresting 100% Roman. Hence the question in the title.
I had gone looking for a mixture, so this was an unexpected place to start. Where was the Celtic or Germanic contribution I had expected? Would it emerge in other comparisons, or would those also place me close to the ancient Italians? Before deciding what to make of my result, I wanted to know how often the same thing happened to other northern Italians. A result shared by many of them would be a different story from one peculiar to me.
So this post follows my genome through those comparisons, with other northern Italians alongside it. I start with where I sit among living populations, then turn to the ancient samples and the ancestry models. The question that interests me is how much of the history I had in mind can actually be recovered from my DNA, and what happens to that surprising Roman result when I look more closely.
How I made the comparison
The baseline Roman/Italic source contains four people from Castel di Decima, Palestrina, Ardea and Boville Ernica. They represent Roman, Latin and Volscian or Hernician contexts, spanning roughly 900–200 BCE. Some lived before the Republic, so I use the broader name Roman/Italic. Etruscans are tested separately.
For this article’s PCA, all reference genotypes come from the Allen Ancient DNA Resource (AADR), which includes modern as well as ancient people. I fitted the axes to a selected subset of 693 modern references in 90 groups, using 61,521 filtered autosomal SNPs, then projected my genome and the 19 study individuals onto those fixed axes. None of us helps define the coordinate system. [2]
In the first qpAdm comparison, I tested the four-person Roman/Italic pool and a separate ten-person Etruscan sample from Tarquinia, both alone and with each of nine alternative sources. Etruscans provide a neighbouring central Italian comparison. Repeating the twenty-model comparison with the six-person Roman/Italic pool gives 40 models per person and 800 across all 20 genomes. Each mixture contains one Italian source and one additional source; I do not fit Gaulish and Germanic contributions simultaneously. All memberships stay fixed between targets.
I use p ≥ 0.05 as the model-fit threshold. For an added source, I also require valid proportions, distinguishable source populations and an improvement over the corresponding Italian-only model after multiple-testing correction. The appendix gives the sample lists and settings.
A second comparison adds Imperial Italian communities and eastern Mediterranean sources. Here I keep the twelve comparison populations fixed for every model, rather than rotating unused candidate sources into that set. This gives the four-person, six-person and Imperial source models a common basis for comparison. Figures 1–2 use the first design; Figures 3–4 use this fixed-panel test.
For model selection, I start with each Italian source alone and test the permitted additions one at a time. I retain an addition only when the resulting model passes (p ≥ 0.05), its proportions are valid, its sources are distinguishable, and the nested improvement test has Holm-corrected p < 0.05. I stop a branch when no tested addition qualifies; if its model fails, that branch supplies no adequate model. When several alternatives qualify, I report them separately. A larger overall model p-value is not enough to retain a source. Figure 3 applies this rule; Figure 4 separately shows the highest-p-value candidates. [1]
Who the comparison populations were
The Gaulish reference consists of 40 people from Bucy-le-Long in northern France, a La Tène community of Iron Age Gaul. The Gauls were Celtic peoples; this group tests the Celtic side of my expectation using an actual ancient community, rather than modern French people. It represents northern Gaul, not every Celtic population. [9, 10]
Alken Enge (14), Asnæs (10), Lille Vadsby (9) and Simonsborg (19) are Roman-period groups from Denmark. I use them as Scandinavian proxies for Germanic-related ancestry, keeping the sites separate to see whether the result depends on one community. Wielbark (17), from two sites in northern Poland, adds a Baltic comparison. Its archaeological culture is commonly associated with the Goths, although that does not establish the ethnic identity of every person buried there. [10, 12]
The Longobards, or Lombards, were a Germanic people whose migration from Pannonia into Italy in 568 CE makes them especially relevant here. The Hungarian references come from their associated cemetery at Szólád. I retain its main northern/central-European group (13) and two additional AADR groups (7 and 4) separately. “Core”, “1” and “2” identify those genetic groupings, not different tribes. The community contained people of differing ancestry, so a percentage assigned to a Szólád group is not automatically a percentage of northern ancestry. [11]
Every source is a group of people, not a single ancient genome. Together, these nine alternatives contain 133 distinct individuals. They were chosen for their historical relevance, archaeological context and available qualifying genomes, not for the highest fit to my DNA.
My genome is closest to northern Italian samples, particularly Bergamo and Torino.
Migration brings ancestry into a population; later gene flow and drift change how that population relates to its neighbours. PCA and FST summarise the resulting differences. They do not separate Celtic, Roman and Germanic contributions. To ask about those historical sources, I need an explicit model.
What the single Roman bar means
With one source, qpAdm fixes its coefficient at 100%. The test is whether that source can account for my genetic rFigure 1. Every model of my genomeelationships with the comparison populations. My four-person Roman/Italic model passes at p = 0.2257. The Etruscan-only model fails at p = 0.0080. [1]
Adding the main Szólád Longobard group to the four-person Roman/Italic baseline gives a Longobard-source estimate of 2.4% ± 13.4 percentage points and p = 0.1463. The Gaulish alternative produces a negative coefficient, −5.7%, so it is not a valid ancestry mixture. Figure 1 shows every model, including the rejected and invalid fits.
Figure 1. Every model of my genome
Including the Roman outliers
Figure 1 compares the full model sets. The six-person pool adds two individuals whom the metadata identify as eastern-Mediterranean-shifted outliers, one from Palestrina and one from Ardea. They belong to Roman/Latin archaeological contexts but differ genetically from the four-person baseline. Including them tests a broader version of the ancient reference, while keeping Etruscans outside it.
With this pool, my one-source p-value is 0.1122, which still passes. Adding the Gaulish group gives 18.2% ± 13.6 percentage points and p = 0.1293. The main Szólád group gives 14.8% ± 10.4 percentage points and p = 0.1081. Both mixtures pass, but neither added source survives the corrected improvement test. The one-source model remains sufficient under this comparison.
Including eastern-Mediterranean-shifted individuals raises the estimated contribution from several northern alternatives. A northern source can partly compensate for that change in the Italian reference. Yet the uncertainty is large enough that the added contribution is not required. The difference between a passing mixture and evidence for an additional source matters here: I get the former, but not the latter.
The two additional Szólád groups give negative personal coefficients and, in several fits, extremely large errors. Those unstable estimates are not usable ancestry percentages. Keeping the groups separate makes this visible instead of hiding the cemetery’s diversity inside one Longobard average.
The Etruscan-only model fails in both analyses, as do all eighteen Etruscan-plus-additional-source fits for my genome. This result concerns the particular Tarquinia sample and comparison populations used here. Ancient-DNA studies also find substantial genetic similarity between Etruscans and neighbouring central Italians, so the archaeological labels should not be read as sharply separated genetic categories. [3, 4]
Below, I compare my result with 19 other northern Italians and test what happens when Imperial-era genomes replace the Iron Age references. Upgrade to read the full results, including the Celtic and Longobard-associated models and how the choice of ancient samples changes the answer.