Cognitive Sanctuaries Or: Manifesto of the Department Chair

[This is a guest post by Jess Werk. This blog post was initially written in a different file format and converted using AI. — T.]
Hello, I am Jess Werk, professor and chair of Astronomy at the University of Washington. Astronomy may interest mathematicians right now because our field has already been reshaped by supercomputers and survived to tell the tale. Our magneto-hydrodynamic simulations, built on the Navier-Stokes equations (plus lots of other “subgrid” physics), are run for over a hundred million core-hours, and produce emergent astrophysics that takes us years to understand and verify. They do not replace analytic theory; they complement it. The Rubin Observatory’s nightly stream of sky survey data, petabytes over the life of the project, would be unusable without advancements in cloud storage and AI algorithms. Astronomers were early adopters of machine-learning techniques because we have a whole Universe’s worth of beautiful data, and our space telescopes are built at the edge of what is technically feasible (e.g. the James Webb Space Telescope, and the Habitable Worlds Observatory now being designed). We are generally a technology-forward field, but a growing number of us see AI-enabled workflows as a risk to the profession. I will tell you what I think that risk is, and then put on my department chair hat and tell you what I am doing about it. My optimism is intact because it has to be.
As scientists, we design our experiments around an unattainable ideal: an objective and mindless observer from nowhere, studying a reality that exists separately from our subjective understanding of it. As creatures who study the cosmos from a relatively tiny rock, 93 million miles from the nearest star, embedded within the roiling gases of the Galactic interstellar medium produced by hundreds of billions of gasping stars, some fraction of which violently explode, we astronomers appreciate the physical impossibility of the observer from nowhere (and yet we endeavor to achieve it!). Thomas Nagel writes about the paradox of The View from Nowhere and argues later, in Mind and Cosmos, that an intelligible natural order which produces minds capable of understanding it is itself something that requires explanation. Intelligibility is the biological mind’s responsibility; every experiment we design must be understood by someone because that is what gives science its purpose. Generative AI attempts to embody the mindless observer from nowhere. Its achievements in both mathematics and science reveal the incompleteness of this long-held ideal and underscore the importance of human scientists and mathematicians, whose minds make results and proofs meaningful.
The marketing for and media coverage of generative AI invite a belief that the speed and volume of its output render slow human understanding meaningless. Thankfully, mathematicians have already begun to reject this idea. In astronomy, the result plays the role of the theorem: a paper is judged on what it found, how significant it was, and whether it found it first, far more than on what the finding means. The result serves as a proxy for understanding, and that proxy is now a problem (Kra, B., 2026). Across academia, incentive structures have long rewarded productivity over understanding, and generative AI now makes productivity easy to manufacture. Together, the two effects undermine the perceived value of doctoral-level study. Rather than swiftly implementing field-wide changes to our systems, I argue that protecting the Ph.D. requires simple, department-level fortifications that we can all make now. Together, these structural fortifications make up a model I have started calling the cognitive sanctuary.
Under the current productivity-weighted model, the technical debt that Henry Cohn describes in his blog post is borne by the most vulnerable members of our community. Ph.D. students compete for postdocs on the number of their publications; postdocs compete for faculty jobs on their h-indices; early-career faculty are judged on the dollar value of their grants (funders, in turn, count our “research products” which are ingested into a database) and the quantity of their scholarly output (including the number of Ph.D. students trained). These metrics do not measure or reward scientific merit. Citations accrue to papers that are already well-cited (e.g. Merton 1968, The Matthew Effect in Science), so the h-index inherits and magnifies biases present in the field (e.g. Kelly and Jennions 2006, Caplar, Tacchella and Birrer 2017). In the meantime, the literature keeps growing: astronomy arXiv submissions rose 14% from 2024 to 2025 (Lewis, Shah and Alfred 2026) and more than half of the papers posted in 2025 are written with language-model assistance (Saad and Ting 2026). Only about one paper in 66 that uses these AI tools declares it, a gap Saad and Ting attribute to failing disclosure norms and I would also attribute to perceived stigma. The people with the strongest incentive to use these tools and hide that they have done so are the same ones under the most pressure to publish: our students and early-career scientists. Unfortunately, they pay twice. Those who are still building new skills are the ones most likely to suffer “cognitive debt” from an overreliance on an LLM (e.g. Kosmyna et al. 2025; Bastani et al. 2025) and they will also face the steepest professional consequences for failing to declare its use (e.g., forced retractions and publication bans).
Scientific and mathematical thought relies on our collective judgment being tested, then re-tested and held up continually against alternative explanations. Our workflows, increasingly built on technology, have been streamlined to enable first discoveries and to decorate them with our own names. The slower, iterative processes of refining theories and interpreting data are relegated to methods sections that a few people might read, if they appear at all in the publication. A paper is rarely dedicated to showing how a particular idea turned out to be wrong after a careful analysis. “Nobody has time for that,” we think. The struggle of doing science is the struggle of learning, and it is also where the joy of discovery lives. A cognitive sanctuary, therefore, must reward the effort and celebrate a process that includes failure. It is a space that encourages unhurried thinking with plenty of room for rabbit holes and the occasional mad hatter.
Unlike a mad hatter, agentic AI creates a chain of probabilistic decisions and steers users away from the improbable. Some of my colleagues have envisioned a future in which student learning centers on individualized, guided conversations with an “Agentic Professor” (Cornillon and Prochaska 2026). In this future, the ability to derive and manipulate equations becomes secondary, and it is assumed that the capacity to interrogate a hypothesis, a necessary Ph.D.-level skill, can be built without the practice of trying and failing and trying and failing and trying and eventually succeeding. Sam Altman envisions a personal AI team for everyone and a virtual tutor for every child (Altman, S 2024). Taken to its end, the vision leaves no Ph.D. advisor to serve as a witness to the student’s judgment, and no one to vouch that the student can be trusted to produce and judge knowledge. I reject this future. The Ph.D. student cannot supervise mathematics without being able to produce it themselves through “active shaping of experience performed in the pursuit of knowledge” (Polanyi, M. 1966). Practice builds a tacit element of understanding that my field calls physical intuition. Nothing yet shows that physical intuition can be developed through closed-loop agentic AI conversations, and recent evidence points in the opposite direction (Bastani et al. 2025). Interrogation is best practiced among other scientists who can offer surprising alternatives and explanations (sometimes incorrect!), while an AI agent’s alternatives are drawn from the distribution of what has already been written. Our system of knowledge depends upon the Ph.D. and the Ph.D. depends upon iterative judgment of people who already have it.
And the people who have that judgment already work down the hall from you. Academic departments bring together kindred spirits and support many of the structures a cognitive sanctuary needs: we host seminars, discussions, hack-a-thons, and community events designed to honor thinking minds. What I see as a department chair is that the senior faculty with the most recognition are often the ones who use these structures least. Nobody wins an award for being a department chair, as I sadly have discovered, and the work of showing up to preprint discussions carries no service credit. When you are leading national committees and international collaborations, it is easy to justify skipping colloquium or a lunch talk for the hour gained in research productivity. Faculty are expected to do too much, and being overachievers, we do even more. Held against all that external and often unpaid labor, taking time to honor a half-formed thought process looks like an unaffordable luxury. Meanwhile, Ph.D. students are craving opportunities to interact with the senior faculty who are so often absent from department life. A department that works as a cognitive sanctuary gives senior faculty the cover to decline some external work, and it counts showing up for junior colleagues as critical service.
At our Astronomy faculty retreat on September 17, we discussed a scenario modeled on what is happening in mathematics. Briefly, OpenAI posts a preprint, press release, and public decision log of 40,000 steps reporting a five-sigma detection of an evolving dark energy equation of state, inconsistent with a cosmological constant, with error bars a factor of 2.5 tighter than anything our community could achieve from the same public data, drawn from a future survey in which our department has already invested heavily. None of the faculty were especially fazed and none thought the scenario to be implausible. Although they are all using AI in their research to different degrees, the faculty broadly agreed on four ideas:
- Generative AI is changing who has access to discovery, and early-career scientists face the steepest barriers.
- The interpretation of a discovery is more valuable than the discovery itself, and how we figure something out matters more than what we figure out.
- Scientific writing is best when it is slow and iterative because the writing refines the science (Gopen and Swan 1990).
- Verifying someone else’s work is less rewarding than doing your own, so a field that is about to depend on verification more than ever must reward it deliberately.
A cognitive sanctuary protects the thought process, for everyone in the department, but above all, for Ph.D. students. The Harvard Summit on PhD Math Education in the Age of AI met on the same day as our astronomy faculty retreat, and reached many of the same conclusions about assessment and AI use; it does not address faculty incentives, which is where much of the department’s leverage lies. Below are several practical suggestions for how to build a cognitive sanctuary in your own department. They fall into three groups: Ph.D. processes, faculty reward systems, and department community.
Ph.D. Processes
- Develop detailed rubrics for both the dissertation and the defense that distinguish a pass from a conditional pass from a failure. One dimension, for example, is whether the student can state the competing interpretations of the result and say what measurement would distinguish them. Require committee members to score the rubric independently before any closed-door discussion. Following a successful defense, require the committee to sign a short report stating what the student demonstrated. The report can go into recommendation letters, and it gives hiring committees something other than a publication list to read. Post the rubric and the evaluation process on the department website.
- Design any evaluation checkpoints (e.g. qualifying exams, required committee meetings) to be conducted live, preferably in person, and unassisted.
- Remove any requirement that a student publish a paper to advance to candidacy. Do not push students to submit publications early, before interpretations have been discussed and worked out. Encourage faculty and Ph.D. students to submit short write-ups of ideas that did not pan out.
- Ask Ph.D. candidates to explain their work at the board, in group meetings and once a year to the whole department. Explaining an idea to a room of scientists helps refine it, and explaining how an idea turned out to be wrong is part of learning. Ask faculty to provide feedback on the substance of the explanation, not the delivery.
- Set clear boundaries on AI use in Ph.D. research to honor the learning process. These boundaries apply to advisors and external collaborators as well as to students. Agree that advisors will not send students feedback that is written with generative AI. Agree that students will not send advisors text or analysis generated by AI.
- Teach students to interrogate AI model outputs on well-understood problems and build these skills in coursework, and group and one-on-one meetings.
Faculty Reward Systems
- Define unit-level promotion and tenure criteria that consider Ph.D. student mentoring to be at least as important as a prestigious award or large grant.
- Publicly post faculty department service assignments, and weight Ph.D. student advising and committee work more heavily than external service (including university-level service). Provide service credits for faculty who commit to regularly show up to key department events (e.g. chalkboard talks, colloquium).
- Develop a teaching policy that includes credit for Ph.D. student advising and committee work. For example, a faculty member who sits on three or more Ph.D. committees in each of two consecutive years is offered a quarter of teaching relief once every 2 years.
- Set an expected external service load for faculty that they may cite when they decline invitations. Encourage faculty to report declined external service alongside accepted service in their annual merit reports.
Department Community
- Create opportunities for meaningful connection among senior faculty and students, e.g. department coffee hours, potlucks, celebrations of student milestones.
- Thoughtfully design colloquium and seminars with time for discussion, including dedicated time for Ph.D. students to meet with speakers. Commit to showing up yourself, even when speakers are discussing a topic outside of your area of expertise.
- Set a department-wide standard that no one uses AI-generated text in research publications, except for disclosed, light use for editing, grammar, or translation.
- Develop department-specific policies on generative AI use and have regular discussions with faculty on whether these policies are achieving their stated purpose.
None of these suggested structural fortifications presents a case against generative AI, a remarkable tool that can improve the practice of science and mathematics and reduce the burden of some administrative tasks. Cognitive sanctuaries encourage its use in the open, especially where nothing the student is supposed to be learning is at stake. The suggestions above do not address structural problems in our fields that will be amplified by AI. They will not themselves generate much-needed funding for basic science and mathematics at the national level. They cannot solve emerging issues in grade-school education that increasingly relies on AI.
Departments, and the Ph.D. students they educate, sit where all the above challenges meet. They absorb the funding cuts and receive the students that grade-school education produces. As we create microcosms of joy and learning in our departments, I hope that we can gradually steer universities toward rewarding process. Cognitive sanctuaries build the minds that make discoveries mean something, the “deployable intellectual reserve” that Amit Sahai envisions. In an optimistic future where AI-driven discoveries and innovations must be understood and tested, academia is intentionally restructured around process, collaborations that demonstrate understanding win the highest awards, and Ph.D. students find joy as their engaged human professors challenge them over and over again until what they know expands, and they can vouch for all of it.
Acknowledgements
Thank you to Tatiana Toro for the introduction to Terry, and to Terry for hosting this piece. Thank you also to Terry for pointing me to the report of the Harvard Summit, which I read only after finishing a full draft of this piece. I was genuinely heartened to find how much our two fields had converged on their own. Thank you to the faculty in the UW Department of Astronomy who inspire me with their thoughtfulness, brilliance and musical talent. Thank you to Professor Xavier Prochaska, my postdoc mentor who frequently had me go to his chalkboard with my ideas. He came to me over a year ago with a claim that, under his guidance, Claude could write a Ph.D. thesis in a month that would be better than that of an average Ph.D. student in Astronomy. I did not believe him at the time. I do now, but I have decided that the thesis is not the point of the Ph.D. The student is.
I found a strange comfort in reading philosophy this summer, particularly The View from Nowhere and Mind and Cosmos, both by Thomas Nagel, after my guitar teacher was killed in an accident on June 1st. He taught me how repetition builds skill, always turned on the metronome, and taught me intricate finger patterns. One day I would be all tangled up in a finger pattern, and the next it would fall into place. Our lessons were a microcosm of joy, a cognitive sanctuary, while I was suffering from the burnout of my first years as department chair. Our time together meant more to me than I can say, and what I learned from him will stay with me forever.
AI disclosure: I drafted every paragraph of this piece by hand, then typed and edited it. I used Claude (Fable 5.1) to check my sentences against Gopen and Swan 1990, to find and verify citations, and to argue with. A handful of sentences began as its suggested wording and were rewritten by me; the arguments are mine, though some were sharpened in the arguing. No text in this piece was generated and pasted.
References
Altman, S. 2024. “The Intelligence Age.” Blog post, September 23, 2024. https://ia.samaltman.com/
American Astronomical Society. 2026. “Author Guidelines for Use of AI and LLMs in Manuscript Preparation.” AAS Journals, posted September 9, 2026. https://journals.aas.org/author-llm-guidelines
Avila, A., et al. 2026. “A Severe Misalignment of AI in Mathematics.” Zenodo, September 11, 2026. https://doi.org/10.5281/zenodo.22737751
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., and Mariman, R. 2025. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences 122 (26): e2422633122. https://doi.org/10.1073/pnas.2422633122
Caplar, N., Tacchella, S., and Birrer, S. 2017. “Quantitative evaluation of gender bias in astronomical publications from citation counts.” Nature Astronomy 1: 0141. https://doi.org/10.1038/s41550-017-0141
Cohn, H. 2026. “The technical debt of AI-generated mathematics.” Guest post, What’s new (Terence Tao’s blog), September 15, 2026. https://terrytao.wordpress.com/2026/09/15/the-technical-debt-of-ai-generated-mathematics/
Cornillon, P., and Prochaska, J. X. 2026. “The Agentic Professor: Exploring GenAI-Supported Futures in Higher Education.” EDUCAUSE Review, September 14, 2026. https://er.educause.edu/articles/2026/9/the-agentic-professor-exploring-genai-supported-futures-in-higher-education
Gopen, G. D., and Swan, J. A. 1990. “The Science of Scientific Writing.” American Scientist 78 (6): 550–558.
Kelly, C. D., and Jennions, M. D. 2006. “The h index and career assessment by numbers.” Trends in Ecology & Evolution 21 (4): 167–170. https://doi.org/10.1016/j.tree.2006.01.005
Kosmyna, N., et al. 2025. “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task.” arXiv:2506.08872. https://arxiv.org/abs/2506.08872
Kra, B. 2026. “Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought. AI has broken this system.” Guest post, What’s new (Terence Tao’s blog), September 13, 2026. https://terrytao.wordpress.com/2026/09/13/deep-theorems-were-scarce-and-difficult-and-so-became-an-effective-mechanism-to-identify-deep-thought-ai-has-broken-this-system/
Lewis, R., Shah, H., and Alfred, A. 2026. “Astrophysics Wrapped 2025.” arXiv:2602.12303. https://arxiv.org/abs/2602.12303
Merton, R. K. 1968. “The Matthew Effect in Science.” Science 159 (3810): 56–63. https://doi.org/10.1126/science.159.3810.56
Nagel, T. 1986. The View from Nowhere. New York: Oxford University Press.
Nagel, T. 2012. Mind and Cosmos: Why the Materialist Neo-Darwinian Conception of Nature Is Almost Certainly False. New York: Oxford University Press.
Polanyi, M. 1966. The Tacit Dimension. Garden City, NY: Doubleday. Reissued 2009, Chicago: University of Chicago Press.
Saad, S. M., and Ting, Y.-S. 2026. “More than half of recent astronomy papers are written with language-model assistance.” arXiv:2609.10664. https://arxiv.org/abs/2609.10664
Sahai, A. 2026. “We’re gonna need a lot more mathematicians.” Guest post, What’s new (Terence Tao’s blog), September 24, 2026. https://terrytao.wordpress.com/2026/09/24/were-gonna-need-a-lot-more-mathematicians/
Sanderson, G. 2026. “If math is more than proof, we need to better celebrate the rest of it.” Guest post, What’s new (Terence Tao’s blog), September 18, 2026. https://terrytao.wordpress.com/2026/09/18/if-math-is-more-than-proof-we-need-to-better-celebrate-the-rest-of-it/
Summit on PhD Math Education in the Age of AI. 2026. Report of Summit on PhD Math Education in the Age of AI, September 17–18, 2026. Harvard Center of Mathematical Sciences and Applications. https://cmsa.fas.harvard.edu/media/2026/09/Summit-on-PhD-Math-Education-in-the-Age-of-AI.pdf