Becoming a Benchmark

Becoming a Benchmark 图片 1

[This is a guest post by Talia Ringer. This blog post was initially written in a different file format and converted using AI. — T.]

When I was wrapping up graduate school in computer science in Spring 2021, I was given access to a curious little programming model on OpenAI’s “playground.” The model, called Codex, took in natural language text and generated programs from that text. The task of automatically generating programs given an expression of programmer intent had been one of many major questions in my field of study—programming languages—for decades. This was the first time I had seen a neural model, with little to no input from anyone in my field, actually succeed at that task to any degree. This was more than a year before ChatGPT was released, and yet I knew that everything was about to change.

On one hand, I was glad that programming was becoming more accessible. My mom came to visit shortly after that, and I pulled out my laptop and told her she could now program. I had her make a little video game. It was not perfect, sure, but my mom was programming, and that was wild. I had always wanted programming to become accessible to everyone.

On the other hand, I had so many worries. At a technical level, would the software of the future be full of AI-introduced bugs? Would people run a bunch of unit tests on AI-produced programs and think that they are OK, but miss out on edge cases? Would AI tools overfit to the tests they are given access to? Would people sometimes feel more productive using these tools, but actually get less done? (Yes, yes, yes, and yes.)

Those technical worries were overshadowed by a larger existential dread. My specific focus within programming languages research had been on using classic programming languages techniques to make it easier to write formal, machine-checkable proofs using proof assistants like Rocq and Lean. Was my entire field of research about to die? To be swallowed by AI? I was not worried about my field actually being fully “solved,” but I was very worried about the prospect of AI researchers claiming to fully solve my research area, and of the general public actually believing them. I had seen this play out before in linguistics and natural language processing.

Worse, if AI swallowed my field, would AI’s culture leak into my field’s culture? I had known AI’s culture to involve all sorts of things I find distasteful and immoral in research, like “scooping,” competition, and secrecy. My field, by contrast, was (and thankfully still is) a lot more like mathematics in this regard. We value communication, collaboration, and openness. I was scared that AI would rot this culture from without.

I have come to understand what I went through as the AI grief cycle, the one that starts when one’s life’s work becomes a benchmark for AI companies. And as I worked through this cycle of grief, I realized that I had to communicate both to AI companies and to the general public that my work will not be replaced by these tools, but that it will change. I had to understand that myself, come to terms with it, and move with it in my own work. And above all, I had to make sure incentives stayed aligned with that reality.

This started with communicating what I actually do, both to the AI community and to those outside of it who set incentives. This was the most exhausting part of it, since I was coming from a small research community, while the AI community is large and powerful. So that means I really had to immerse myself and learn to speak their language. In lieu of that, it would have been too difficult to honestly and reputably assess the limitations and impacts of what AI tools do. (In doing so, I accidentally nerd-sniped myself into actually doing AI work as part of my broader research portfolio, but it is probably possible and fine to do this in a way that still keeps one’s research far away from AI.)

On an individual level, funders and the AI community alike have since come to respect me. And also, my field of programming languages has come out pretty OK so far. I will never know how much of an impact I have had on that. But I have found that the whole community has reacted in ways that are pretty aligned with what I have done. We have engaged honestly with the changes, and we really have assessed both the capabilities and limitations of the tools that had infringed on our field so suddenly. When relevant, now, we use neural techniques to improve our own work. But we have also found ways that our techniques are strictly complementary to those techniques, and we have figured out how to communicate that. Our culture has come out intact, too.

But I think one thing we have going for ourselves in programming languages is that basically nobody has ever heard of our field, and most people do not care about what we do. People do care a lot about math. We don’t even get the “oh, I hate math” response that mathematicians get; hatred means that people care.

Culturally, at least in the US, math is simultaneously revered and hated. Peak “intelligence” is often culturally associated with mathematical ability—people even bring up Terry Tao as an example of this. Poor performance at math, or anxiety around such poor performance, is often coupled with statements about not being “smart enough” to do math. Mathematicians, then, come to represent the intellectual elite, with all of the scapegoating such a label carries. And leaders of AI companies come to believe that if they can “solve math,” they can “solve everything,” whatever that means. (Sorry, Jesse; we are still friends.)

Because of this dual reverence and hatred, people are paying way more attention to AI infringing on math than they did to AI infringing on my field. This is good and bad. It does mean that mathematicians’ statements are getting lots of coverage; they have a real chance to communicate to the entire world what it is they actually do. But it also means that their legitimate sour gripes are being misread as sour grapes. There is a perception that they are gatekeeping. Just like, in 2021, if people had paid this much attention to my field, they might have wrongly concluded that I was just being bitter because I did not want my mom to be able to program. (I did!)

This makes mathematicians’ jobs harder than mine was. Their audience can at times be actively adversarial. Just learning the language of AI folks will not be enough.

Still, I think mathematicians are on the right track. The work many are already doing of addressing the public is even more important than addressing the AI industry. Yes, the AI industry’s goals are misaligned with those of the math community. But it is actually OK if that remains true, so long as the rest of society recognizes that misalignment and continues to value and incentivize the kind of work that mathematicians value. AI companies will always want to use whatever field is hot at the time to show that their models are the “best.” There are ways to help them better understand what “best” should mean, but it’s also pretty OK if they never do understand that, as long as the general public does. So math—the process, not the benchmark—will need a PR campaign that lasts for as long as math the benchmark is relevant to AI companies. That means deliberate and consistent engagement with news outlets, social media, education systems, policymakers, and funders.

This work of engagement is exhausting and unending, but it does get easier as society and AI companies alike move on. Remember when art and writing were the main targets of these companies? Two things seem to have happened: First, society seems to have collectively developed a distaste for AI art and writing, and even for human art and writing that vaguely resembles AI art and writing. Second, AI companies seem to have reached the point at which showing off their tools’ art and writing results no longer proves them to have the “best models,” so they have largely moved on to other fields.

Have art and writing been negatively impacted? Absolutely. And the fight is still ongoing. But at least the fight is no longer all-consuming. (Collective action like unionizing, or like the Hollywood writer’s strike, might also have to do with that, though. I do think mathematicians should consider unionizing.)

And what if mathematicians actually do want AI tools that help with math the process, and not just math the benchmark? (I do. I’m super excited about what math could be like in such a world.) In that case, I think it is probably best to look for collaborations with smaller companies and academics, especially those that have a participatory model where mathematicians get to co-design the tools to reflect their own values and use-cases. (I have one such collaboration with Emily Riehl already, and I am hungry for more!) It also helps to be upfront about expectations around credit, especially where the culture might clash.

In any case, a friend who is a professor of linguistics told me in 2021 that, in ten years’ time, my research would be different in ways I could not possibly predict ahead of time. And that still, it would be OK. I took great comfort in that. It really will be OK.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论