Education and intelligence polygenic scores, thirteen years later
Thirteen years ago, we finally had three genome-wide-significant genetic hits for education. Three! I used their frequencies across 14 populations from the 1000 Genomes project in my 2013 paper to investigate whether education-associated alleles followed a pattern across populations that tracked IQ. [1, 2]
Since then, the GWAS samples have grown enormously. EA3 arrived in 2018, followed by EA4 in 2022. And now, in 2026, we finally have a much larger study of fluid intelligence. [3, 4, 5] I wanted to see what happened when those original three SNPs met their successors.
There is also a new Norwegian education GWAS in the mix. Kvalvik and colleagues published it in September 2026, using the Norwegian Mother, Father and Child Cohort Study, usually shortened to MoBa. Here, “Norway” means their years-of-education score, based on 120,527 genotyped adults. The paper also studied educational milestones and twins; the score used here comes from the adult MoBa GWAS. [6]
For a comparison with the older intelligence results, I also include the Pan-UK Biobank fluid-intelligence score. Its weights come from the European-ancestry analysis of the earlier UKBB test data, covering 135,088 participants. The Pan-UKBB resource paper appeared in 2025. [7]
They agree remarkably well. Across 46 population samples, the old score correlates 0.939 with EA3, 0.779 with EA4 and 0.874 with the new fluid-intelligence score. It also correlates 0.810 with the population IQ estimates, across 39 samples. After thirteen years, those three variants are still keeping up.
From three SNPs to thousands: subscribe for more on the genetics of education and intelligence.
Meet the GWAS
Table 1A puts the education studies on a timeline, and Table 1B does the same for intelligence. The growth is quite something: from three education hits in 2013 to thousands. The sample sizes belong to the discovery studies, not to the population samples in our scatterplots.
Table 1A. The education GWAS behind the scores
Table 1B. The intelligence GWAS behind the scores
The scores used in this post retain EA1’s 3 SNPs, 2,589 from EA3 MTAG, 2,933 from EA4, 42 from MoBa, 69 from Pan-UKBB, 454 from new FI and 502 from FI plus COGENT.
Back when three SNPs were a big deal
Rietveld and colleagues studied 126,559 people to find those three hits. [2] A genome-wide association study, or GWAS, searches for genetic variants associated with a trait. The three SNPs passed the significance threshold for years of education or college completion. In 2013, that was enough to get started.
A polygenic score adds up the associated alleles, giving each one a weight based on its estimated effect. With population allele frequencies, we can calculate the average score for each sample. For this comparison, I took the original three variants and weighted them using Rietveld’s published effects on years of education. I call that score EA1.
My original paper examined the three significant SNPs as well as a broader set of ten variants. The main 1000 Genomes analysis covered 14 populations, with another analysis of 11 HapMap populations. ALFRED extended the education analysis to 50 populations using seven markers, including proxies for all three lead SNPs. [1]
Eight of those ALFRED population groups are represented in today’s full 51-sample roster: Yoruba, Italians/Tuscans, Dai, Han Chinese, Japanese, Koreans, Mongolians and Papuans.
Three SNPs meet a few thousand
Start with Figure 1. Each dot is a population sample, with EA1 on the horizontal axis and the newer education score on the vertical axis. The EA3 plot is striking: r = 0.939. A score built from three variants follows one containing 2,589 variants remarkably closely.
EA4 agrees too, at r = 0.779, although the points are more spread out. A larger GWAS does not automatically give the highest correlation in every comparison. There is also a difference in what these scores measure: the EA3 panel combines education with cognitive and mathematical traits, whereas EA4 comes from the later education analysis.
Figure 1. Three original SNPs versus EA3 and EA4
None of the three original SNP identifiers appears in either of these newer score files. That rules out the simplest explanation, that we are just adding the same three SNPs back into a bigger sum. Nearby variants can still be inherited together, so some of the newer markers may track the same genetic regions.
Now intelligence joins in
The new paper by van den Berg, Malawsky, Huang and colleagues gets much more out of UK Biobank by combining measured fluid-intelligence tests with estimates for missing scores. Fluid intelligence means reasoning and solving unfamiliar problems. Their combined analysis reaches 455,402 people; adding the Cognitive Genomics Consortium, COGENT, brings it to 490,700. [5]
Figure 2 is the comparison I was most interested in. EA1 correlates 0.874 with the new UKBB intelligence score and 0.846 with FI plus COGENT. The first three education hits still track a much larger intelligence GWAS thirteen years later.
Figure 2. Three original SNPs versus the new intelligence scores
The FI scores use reconstructed panels of significant variants: 454 SNPs for FI and 502 for FI plus COGENT.
UK Biobank was not in the first study
This is where the new FI result becomes especially interesting. UK Biobank was absent from Rietveld’s 2013 study. I checked the original cohort list, and the 2016 education paper explicitly used UKBB as an independent replication sample. [8] Another UKBB study that year also replicated all three original education hits, including two for verbal–numerical reasoning. [9]
So the UKBB-only FI result is an independent replication with respect to the GWAS discovery sample. The genetic associations were estimated in people from a cohort that did not contribute to EA1. Yet the population scores correlate 0.874.
Adding COGENT changes that answer. Its GWAS includes Helsinki, Lothian 1936 and the Minnesota twin and family cohort, which also contributed to EA1. [2, 10] An earlier COGENT paper actually removed Helsinki and Lothian to avoid overlap in its independent comparison. [11] FI plus COGENT therefore has some cohort overlap with the first study. The published lists do not tell us exactly how many individuals were shared.
EA3 and EA4 also reuse earlier education cohorts, and UKBB contributes to both. Their agreement is useful evidence, with that shared history taken into account. UKBB-only FI gives us the clearer comparison with an independent discovery cohort.
And EA2? I started with the EA3 and EA4 population scores from the previous FI post, so I have not calculated an EA2 score here. The 2016 study was an important step between EA1 and EA3, expanding the discovery sample to 293,723 people and finding 74 significant loci. Its absence from these plots reflects the scores available for this comparison. [8]
The old hits still fit the newer scores
Table 2 puts all seven scores and population IQ together. I also included the Norwegian MoBa education score [6] and the older Pan-UKBB FI score. [7] The Norwegian score tracks EA1 very closely, at r = 0.895. Its correlation with EA3 is even higher, 0.948.
Table 2. The population correlation matrix
Figure 3 shows EA1 against the Norwegian MoBa education score and the older Pan-UKBB FI score. Pan-UKBB has the weakest correlation with EA1 among these six comparisons, at 0.390. The new FI result is much closer to the pattern in the original education hits.
Figure 3. Three original SNPs versus older FI and Norwegian education
For IQ, we need to compare the scores on the same populations. Table 2 uses every available pair, so the number of samples varies. Table 3 fixes that by using the same 39 samples for all seven scores.
Table 3. IQ correlations on the same 39 populations
New FI leads at 0.870, followed by EA3 at 0.855 and FI plus COGENT at 0.849. Our three-SNP EA1 score gives 0.810, and Norway 0.792. EA4 is lower, at 0.745, while the older Pan-UKBB FI score gives 0.313.
The newer FI scores fit these IQ estimates much better than the older one. EA1 does surprisingly well for such a tiny score. The ordering is descriptive: I have not tested whether each gap in Table 3 is statistically significant, and these correlations concern population averages rather than prediction in individuals.
The IQ association is still there
Figure 4 returns to the question behind my 2013 paper. The correlation between EA1 and population IQ is 0.810 across 39 samples. Giving all three increasing alleles equal weight produces almost the same result, 0.816. The pattern hardly changes when we remove the effect-size weights.
Figure 4. Three original SNPs versus population IQ
These IQ estimates were assembled before this study, mainly from Lynn and Vanhanen, David Becker’s National IQ dataset version 1.3.5 and other publications. They combine national and subgroup estimates collected in different ways. That is worth bearing in mind when reading the plots.
I also checked the other IQ measures used in the previous post. Table 4 shows that the Parra and Kirkegaard compilation gives 0.822, or 0.803 when the Ashkenazi estimate is added. [12] The online BRGHT and IIQTEST measures give lower correlations, 0.698 and 0.688. They cover different populations and rely on self-selected test takers, so we should expect some variation.
Table 4. EA1 against the five IQ measures
For the Parra and Kirkegaard comparison, the U.S. estimate goes with a demographic composite of the African-American, Mexican-American and Utah White scores. Mexico uses Mexico_overall. I retain 110 for the Ashkenazi estimate, based on the historical data discussed in Dunkel and colleagues. [13]
The Dutch sample is an unusual point in several plots. Removing it raises the EA1–IQ correlation from 0.810 to 0.833, so it is certainly not responsible for making that association look strong.
Enjoy seeing the new GWAS put to the test? Subscribe for more comparisons, plots and discussion.
What ancestry changes
There is a substantial ancestry pattern in these plots. To see how much of the agreement remains within the broad groups, I subtracted each group’s average from both measures and correlated what was left. Table 5 gives the results.
Table 5. How much agreement remains within broad ancestry groups
EA1 still correlates 0.665 with EA3 and 0.465 with new FI after that calculation. Its correlation with IQ falls much more sharply, from 0.810 to 0.205. Much of the raw IQ association therefore lies between the broad ancestry groups. Those groups are coarse and some are small, but the drop is substantial and deserves attention.
In 2013, I proposed using population allele-frequency patterns to investigate recent polygenic selection. [1] The newer scores strengthen the evidence that the pattern recurs across GWAS. Establishing its cause takes more than these correlations: shared ancestry and the way genetic markers track one another can also produce agreement, while environmental differences affect the IQ estimates.
What happens when we account for LD
There is a catch with GWAS hits: the SNP we score may simply be travelling with the variant that affects the trait. We call it a tag SNP. Its estimated effect then partly reflects how well it tracks that causal variant in the discovery sample. This tendency for alleles to be inherited together more often than chance predicts is linkage disequilibrium, or LD.
Imagine a tag allele that usually accompanies an intelligence-increasing causal allele in Britain. Elsewhere, the same tag might accompany that allele much less often, or even accompany the opposite allele. Applying the British weight unchanged can then overstate its contribution or get its direction wrong, even if the causal effect itself is identical. A tag that works well in Britain can be a poor guide elsewhere. [14]
For population averages, there is another wrinkle: a difference in tag-allele frequency need not mirror a difference in causal-allele frequency. Summing tags with British weights can therefore distort comparisons between populations. Different LD patterns can also reduce individual prediction accuracy. These are related concerns, and an adjustment that changes population-score correlations still needs testing in individuals.
For each SNP, I compared its local LD profile in Britain with its profile in each of the other 25 populations in 1000 Genomes. The correlation between those profiles became a population-specific multiplier on the GWAS effect. This is an experimental weighting rule, rather than an established correction that makes a score “LD-free”.
One detail matters for reading Figures 5 and 6: these weighted scores are deviations from the British reference, GBR. For each population, I subtract a matched British score calculated with exactly the same retained SNPs and the same LD weights.
A positive value means a higher weighted sum than the matched British reference; a negative value means a lower one, and zero means no net weighted difference. These are score units, not IQ points. Because the weights and retained SNPs can differ between populations, each deviation uses its own scoring rule. Subtracting the matched reference makes the comparison explicit; it does not turn these into one identical score definition for everyone.
Our weighting asks how similar each SNP’s surrounding LD pattern is to Britain’s, using nearby variants as a practical proxy for the unknown causal links. It does not identify the causal variant or directly measure how well the score SNP tags it. That is why I treat this as an exploratory sensitivity analysis.
The three old hits are still very much in the game. Figure 5 shows EA1 correlating 0.874 with LD-weighted FI and 0.906 with LD-weighted FI plus COGENT, across 25 population samples. Table 6A adds the other scores: the weighted FI plus COGENT score correlates 0.943 with EA3 and 0.909 with EA4. These are strong agreements, although they use a smaller population set than the earlier figures.
Figure 5. The first three education hits versus LD-weighted intelligence
Table 6A. LD-weighted intelligence scores meet the older scores
Does the weighting improve the fit to IQ? Mostly, it lowers the correlations; FI plus COGENT against BRGHT is the small exception. Figure 6 shows correlations of 0.740 for FI and 0.824 for FI plus COGENT with the original population IQ compilation. Table 6B puts the weighted and original versions side by side, using exactly the same samples for each comparison.
Figure 6. LD-weighted intelligence scores versus population IQ
Table 6B. Does LD weighting improve the IQ correlations
The corrected IQ_P&K comparison is less flattering to the weighting: FI falls from 0.815 to 0.771, and FI plus COGENT from 0.814 to 0.804. That comparison keeps the U.S. demographic composite and counts it once. Mexico_overall has no matching phased reference in this subset, so Mexico contributes no national point here. MXL remains a component of the U.S. composite.
So LD weighting leaves the strong agreement with EA1 intact, especially for FI plus COGENT. It does not give us a general improvement across the IQ measures. The original and weighted versions belong side by side; the weighted score still needs validation as a method of improving prediction.
What stands out to me is how well those first three hits have held up. EA3, EA4 and the Norway educational attainment GWAS agree with them. The new UKBB intelligence GWAS gives us a strong result from an independent discovery cohort. Thirteen years and vastly larger samples later, the original population pattern is still there. That is a satisfying result to come back to.
This article is free to read. If you’d like to support more analyses like this, consider becoming a paid subscriber.
Methods appendix
The collection contains 51 named population samples. EA1 is calculated as twice the sum of the increasing-allele frequencies multiplied by the published combined-sample EduYears effects: A at rs9320913, +0.101; G at rs11584700, +0.095 after reversing the reported allele; and T at rs4851266, +0.082. These use the three-decimal precision printed in Rietveld’s supplement. The three SNPs reached genome-wide significance across the EduYears and College analyses, rather than all three for EduYears alone. [2]
In the 2013 paper, Table 4 examined the three significant SNPs. The broader ten-variant analysis produced a first rotated principal component correlated 0.90 with IQ across 14 populations. The present weighted three-SNP score and its 39-population IQ comparison use a different calculation and larger population set. [1]
EA1 is complete for 46 samples. CHINA HAN, HONG KONG, KOREAN, MONGOL and TURKISH lack a recoverable frequency for rs9320913 after searches of the available original releases and coordinate checks. They remain in the data with a missing score. Partial two-SNP scores and imputed frequency values were not substituted.
The frequency refresh replaces the whole Japanese population block with ToMMo 61KJPN v20250616 (61,424 people, GRCh38, autosomes 1–22). For Taiwan, all requested frequencies are replaced by the Taiwan View Illumina WGS release (1,492 people, GRCh38): 7,155 exact marker allele pairs were recovered across the project panels; 835 remain unavailable from the public pages. Missing current-release frequencies are not filled with older Taiwan values. Other population frequency data, the IQ compilations and the GWAS effect estimates are retained. Marker intersections are recalculated separately for each score, so expanding or changing a panel can change scores even in populations whose allele frequencies were unchanged.
The EA3 MTAG panel combines education, cognitive performance and mathematical phenotypes. The refreshed 51-population frequency intersection retains 2,589 of its 3,269 candidate variants and 2,933 of 3,925 EA4 candidates. Norway uses 42 MoBa EduYears lead variants, accession GCST90832184; Pan-UKBB FI uses 69 SNPs. These are fixed restricted panels, rather than full genome-wide prediction scores.
The authors reported 550 FI and 620 FI plus COGENT lead SNPs using FUMA. The resources checked here—the supplementary workbook, public summary-statistics deposit and code repository—did not contain separate FUMA lead lists. I reconstructed panels from the summary statistics using P < 5 × 10⁻⁸, PLINK clumping with a European reference at r² < 0.1 within 1,000 kb, and the complete 51-population frequency intersection. This retained 454 FI and 502 FI plus COGENT SNPs. None of the exact three EA1 identifiers appears in EA3, EA4 or either new FI panel. Nearby correlated variants may still tag the same regions.
UKBB was absent from the EA1 cohort roster. COGENT adds 35,298 participants and has at least three shared named cohorts with EA1: Helsinki, Lothian 1936 and Minnesota. The cohort counts do not establish the exact number of shared individuals. EA3 reused 59 cohort result files from the preceding education GWAS and added 12 new or updated files; EA4 reused EA3 statistics for 69 cohorts, alongside larger 23andMe data and a revised UKBB analysis. [2, 3, 4, 10, 11]
The new FI study uses education among the measurements informing estimated intelligence scores. Its discovery cohort is independent of EA1, while its phenotype incorporates some education-related information. The population comparisons also reuse the existing reference samples and IQ compilations. [5]
The IQ_P&K analysis pairs U.S. national IQ with the previously specified demographic composite of ASW, MXL and CEU scores, rather than assigning that national estimate separately to those samples. Mexico_overall represents Mexico; MXL is a Mexican-American sample. The Ashkenazi value of 110 is an estimate based on historical Wisconsin data, rather than a representative modern survey of all Ashkenazi Jews. [12, 13]
The scatterplots and matrices use Pearson correlations of population means. Plotted scores are standardized within each comparison; fitted lines summarize the association. Eighteen boxed labels are used consistently in Figures 1–4, selected to cover the seven population groups and the spread of standardized score profiles. Larger standalone plots label all 46 samples in each PGS comparison and all 39 in the IQ plot. The label selection leaves the plotted points and correlations unchanged. Figures 5 and 6 label all 25 comparison populations using the standard 1000 Genomes codes.
Table 5 subtracts the mean of each measure within the existing groups: African, European, East Asian, South Asian, admixed American, Middle Eastern and North African, and Other. It retains all paired rows, including small groups. This removes group means but leaves finer ancestry, linkage disequilibrium and environmental differences unresolved.
Conventional confidence intervals use Fisher’s transformation, and p-values assume independent population rows. Shared ancestry and repeated country IQ estimates can violate that assumption. The accompanying CSV files retain full precision, pairwise sample counts, rank correlations, panel information and missing-data reasons. The larger supplementary matrix contains the 34 named panels from the previous subset analyses plus EA1 and IQ, for 36 measures; the 20,000 random resamples are not separate named scores.
The LD sensitivity uses the original 430- and 476-SNP FI panels. Query SNPs require MAF ≥ 0.01 in GBR and the comparator. For each query, tags within ±500 kb with phased r² > 0.01 in both populations are matched by variant identity, and the query itself is excluded. Pearson r across the two tag-r² profiles is calculated when at least three common tags are available. LD uses the authoritative Phase 3 set of 2,504 people; frequencies use the 3,202-person release. Each score is 2Σβᵢrₚᵢ(fₚᵢ − fGBR,ᵢ), a contrast with Britain using the identical retained markers and multipliers. The matched GBR baseline is 2ΣβᵢrₚᵢfGBR,ᵢ, using the comparator population’s rₚᵢ rather than GBR self-comparison weights of one. The reported value is a weighted deviation from GBR, with no division by SNP count, sum of weights or standard deviation before plotting. Undefined profile correlations are excluded. Signed negative profile correlations are retained and reverse a marker’s multiplier under this experimental rule. This is an anticorrelation between tag-r² profiles, not evidence for a reversal of signed tag–causal LD or a causal effect. Profile similarity is also insensitive to a uniform rescaling of all r² values, so a high profile correlation does not establish equally strong tagging. The primary weighted sets contain 390–430 FI markers and 413–476 FI plus COGENT markers per population. Britain supplies the zero reference and is excluded from correlations. The original LD marker panels are retained after the frequency refresh. Table 6A matches population labels to the refreshed scores, including subgroup proxies, while Table 6B recalculates the original comparison from the identical frequency source. Alternative minimum-tag thresholds, full-precision correlations, conventional confidence intervals and dependent-correlation tests are saved in the analysis files. LD weighting does not establish causal effects or individual predictive validity.
The eight ALFRED matches are between population groups; shared individuals have not been verified. The present EA1–IQ comparison includes 39 population samples.
For MoBa, the authors’ 54 lead SNPs become 42 because 12 lack a usable allele-matched frequency in at least one of the 51 population samples. Every population is scored using the same retained markers within each panel.
The frequency refresh also shows how sensitive a small panel can be to the markers available in every population. On the same 39 IQ samples, Pan-UKBB’s correlation falls from 0.586 with the earlier 63-SNP set to 0.313 with the refreshed 69-SNP set. If I hold the 61 SNPs shared by both versions fixed, it barely moves: 0.537 with the old frequencies and 0.536 with the new ones. The change comes mainly from which markers enter the score. That is another reason to keep track of the exact panel, rather than treating its name as a guarantee of stability.
The Dutch EA1 allele frequency was checked against the original source. Removing the Dutch sample raises the EA1–IQ correlation from 0.810 to 0.833.
The LD profiles were calculated locally from phased genomes with PLINK. [15]
Figures 5 and 6 standardize the GBR deviations across the displayed populations. Zero on those plotted axes marks the comparison-set average, rather than GBR. In the Yoruba comparison, both Yoruba and British allele frequencies use the Yoruba–GBR multipliers; GBR self-comparison correlations of one are not used for this matched reference.
References
[1] Piffer, D. (2013). Factor Analysis of Population Allele Frequencies as a Simple, Novel Method of Detecting Signals of Recent Polygenic Selection: The Example of Educational Attainment and IQ. Mankind Quarterly, 54(2), 168–200. Study link
[2] Rietveld, C. A., Medland, S. E., Derringer, J., et al. (2013). GWAS of 126,559 individuals identifies genetic variants associated with educational attainment. Science, 340(6139), 1467–1471. Study link
[3] Lee, J. J., Wedow, R., Okbay, A., et al. (2018). Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Nature Genetics, 50, 1112–1121. Study link
[4] Okbay, A., Wu, Y., Wang, N., et al. (2022). Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nature Genetics, 54, 437–449. Study link
[5] van den Berg, D. M., Malawsky, D. S., Huang, W., et al. (2026). Imputation of fluid intelligence scores reduces ascertainment bias and increases power for analyses of common and rare variants. Nature Genetics. Study link
[6] Kvalvik, E. H., Wang, Y., Walhovd, K. B., Lyngstad, T. H., & Røgeberg, O. (2026). Beyond years of schooling: Genetic associations across educational milestones in two Norwegian cohorts. PLOS Genetics, 22(9), e1012310. Study link
[7] Karczewski, K. J., Gupta, R., Kanai, M., et al. (2025). Pan-UK Biobank genome-wide association analyses enhance discovery and resolution of ancestry-enriched effects. Nature Genetics, 57, 2408–2417. Study link
[8] Okbay, A., Beauchamp, J. P., Fontana, M. A., et al. (2016). Genome-wide association study identifies 74 loci associated with educational attainment. Nature, 533, 539–542. Study link
[9] Davies, G., Marioni, R. E., Liewald, D. C., et al. (2016). Genome-wide association study of cognitive functions and educational attainment in UK Biobank (N = 112,151). Molecular Psychiatry, 21, 758–767. Study link
[10] Trampush, J. W., Yang, M. L. Z., Yu, J., et al. (2017). GWAS meta-analysis reveals novel loci and genetic correlates for general cognitive function: a report from the COGENT consortium. Molecular Psychiatry, 22, 336–345. Study link
[11] Trampush, J. W., Lencz, T., Knowles, E., et al. (2015). Independent evidence for an association between general cognitive ability and a genetic locus for educational attainment. American Journal of Medical Genetics Part B: Neuropsychiatric Genetics, 168B(5), 363–373. Study link
[12] Parra, L., & Kirkegaard, E. O. W. (2025). National IQs: measurement and defense. OpenPsych. Table A1 supplies the national estimates used here. Study link
[13] Dunkel, C. S., Woodley of Menie, M. A., Pallesen, J., & Kirkegaard, E. O. W. (2019). Polygenic Scores Mediate the Jewish Phenotypic Advantage in Educational Attainment and Cognitive Ability Compared With Catholics and Lutherans. Evolutionary Behavioral Sciences, 13(4), 366–375. Study link
[14] Wang, Y., Guo, J., Ni, G., Yang, J., Visscher, P. M., & Yengo, L. (2020). Theoretical and empirical quantification of the accuracy of polygenic scores in ancestry divergent populations. Nature Communications, 11, 3865. Study link
[15] Chang, C.C., Chow, C.C., Tellier, L.C.A.M., Vattikuti, S., Purcell, S.M. & Lee, J.J. (2015). Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience, 4, 7. Phased LD calculations follow the PLINK 2 LD documentation. Study link · PLINK LD documentation