Comprehensive FAQ on AI risks

Since AI safety has become incredibly popular recently, I quickly finished the draft I’d been working on for a long time and am finally publishing a comprehensive FAQ on AI risks. It’s also available on my website in a more polished version, where it will be updated regularly.

It’s intended for people who are new to the field after seeing all this news and posts on X, and who don’t understand why AI would even want to kill us. That said, people who are already familiar with the topic might also find a couple of interesting points.

I was in a real hurry to finish this text and wanted to make it understandable to as wide an audience as possible, so I apologize in advance if you find any serious mistakes. Please let me know about them, and I’ll correct them. Although I wrote the text by hand (without AI), I used AI to translate it into English (since it’s not my native language and I was in a hurry while the topic is still trending among newcomers). Please let me know if you find any parts that were translated incorrectly.


AS SHORT AS POSSIBLE:

AI continues to develop, and in the future it may be capable of killing every living thing on this planet. It does not need consciousness or humanoid robots to do this, as shown in films. Our typical ideas about ‘evil’ AI prevent us from seeing the real sources of the problem. Misalignment with human values is the main danger, and researchers believe that solving this problem is incredibly difficult. Progress in safety is not keeping pace with progress in AI capabilities. We need to take serious measures right now, as approving and implementing them could take years.

Not convinced? Read the objections ↓

PART 1. ‘I HAVE JUST LEARNED ABOUT AI RISKS’.

1. It is not intelligence at all.

You can call it whatever you like if you dislike the word ‘intelligence’. You can call it a system that can generalize, plan, write code, use tools, find information, operate interfaces, and carry out tasks. It just takes too long to say or write all of that. It is much easier to use the word ‘intelligence’.

1.1. It just predicts the next word!

The key word in this sentence is ‘just’. It downplays an actual capability by explaining its underlying mechanism. This is rather like saying, ‘The brain just transmits electrochemical signals.’ But that does not make the brain’s capabilities disappear. It is also important not to underestimate prediction. It is an extremely powerful capability. AI does not just predict the next word; it predicts it correctly. It is also worth thinking about what intelligence actually is. Perhaps prediction is precisely the defining feature of intelligence in general.

There is an entire theory of how the human brain works which argues that it is largely engaged in trying to predict the future, then checking its predictions against reality and adjusting subsequent attempts based on the information it receives.

1.2. AI does not understand what it is talking about.

We can argue endlessly about the word ‘understands’. What is understanding, anyway? I suspect that AI does understand something, in the broadest sense of the word: advanced models perform very well on benchmarks (evaluation systems) filled with many previously unpublished academic questions and logic problems.

For example, Mythos was able to use amino acid sequences it was given to accurately predict whether a viral capsid (shell) would fold correctly. In other words, it has some idea of how this should happen, even in such a complex and highly specialized subject. But people often use understanding to mean having a conscious conception of something. There is a well-known argument called the Chinese Room, which, in short, says that a system that appears intelligent may not ‘understand’ anything internally or possess consciousness. This has nothing to do with AI risks, since almost all AI-risk scenarios do not assume that AI is conscious.

In fact, the design of the transformer, the most common AI architecture, is not all that complicated at its core. Much of it is based on knowledge usually taught in school linear algebra lessons. However, even if you break the system down into its components and study its smallest elements, that does not make it stupid or weak. Yes, you can now proudly declare, ‘It only multiplies matrices!’ But understanding this does not make the danger of AI or its remarkable capabilities disappear.

2. It has no consciousness / soul, so it is not intelligent / dangerous.

The soul is not a scientific concept at all: we have no signs or evidence of its existence. Consciousness does exist, but it is not required for AI to be dangerous.

Consciousness is a word that most people understand differently. But most philosophers would agree that it is not related to intelligence or reason, but rather describes the capacity to have experiences. Whether AI has experiences is a question that does not affect its capabilities or its ability to harm humanity.

2.1. It has no desires, intentions, or goals of its own. Humans control it. It is a reactive function. It is just a tool.

This was true of chatbots at the very beginning, for example GPT-3 in late 2022, when it became available in a convenient online format. The world is changing, and thousands of employees at multimillion-dollar startups are working to improve AI. In 2026, AI often takes the form of autonomous agents. They can successfully complete tasks that take many hours, and the time horizons of their successful performance double roughly every 4 months.

In July 2026 hundreds of agents of a new OpenAI model escaped from an isolated environment during testing and, entirely on their own, carried out a cyberattack on Hugging Face and possibly one or more other platforms for several days in a row. Humans did not give them this goal, did not direct the process, and discovered it only after the fact.

If you give modern AI any task involving several intermediate steps, it will identify those steps itself, create a plan, and follow it. A human still initiates the action, but what follows does not require constant prompting from a human.

And even if AI were exclusively a tool without a shred of autonomy, the threat of malicious human use would remain. We can set very bad goals. For example, ‘Develop mirror life that will consume all the resources humans need,’ or ‘Attack critical cyber infrastructure to deprive millions of people of electricity, water, or defenses.’

For example, in August 2026, the US Cybersecurity and Infrastructure Security Agency (CISA) warned that hackers were using AI-generated code to attack critical infrastructure. In particular, the attacks target computing components widely used at critical facilities, including water and wastewater systems, power stations, and chemical plants.

3. AI still hallucinates a lot. How can this thing pose any threat to us?

First, hallucinations decrease with each new model. This has been greatly helped, for example, by the introduction of chain-of-thought reasoning and the use of specialized knowledge bases for grounded answers (RAG). Further algorithmic improvements continue to emerge.

Second, a dangerous system does not have to be infallible. Humans hallucinate too, albeit in a different sense. We make mistakes and often believe outright nonsense, but we are still quite capable and dangerous. AI will be even more capable and even more dangerous.

4. This is a religion for atheists and tech fanatics from Silicon Valley.

I agree that there are always people who blindly believe headlines and genuinely behave strangely and irrationally. Such people gather around any popular topic. At the same time, there are scientists with the highest h-indices in the world who have spent years traveling around the world, warning of the dangers of what they themselves helped create.

The third ‘godfather of AI’, Yann LeCun, holds opposing views on AI risks. His predictions, however, leave much to be desired. In 2023, he said that ‘even GPT-5000 will not be able to understand what happens if you push a cup standing on a table’. Four months later, GPT-4 demonstrated a better understanding of physics than that.

If we examine the word ‘religion’, we should take into account that religion requires faith. In AI risk, however, rational and rigorous argumentation, probabilities, and degrees of uncertainty are valued extremely highly. No one worth taking seriously would claim that a catastrophe will inevitably occur with 100% probability. Rationalists remember that neither 0% nor 100% is a probability.

And it has long ceased to be just ‘fanatics from Silicon Valley’. As one of hundreds of examples, the 2026 International AI Safety Report was prepared by more than 100 experts, with an expert panel nominated by 29 countries, the UN, the OECD, and the EU.

Given the current situation, everything points to the risk being genuinely substantial, the consequences of inaction potentially being terrible, and the need to act as soon as possible.

5. This is hype for subscriptions, investment, and government support.

If that were the case, AI companies would already be shooting themselves in the foot in 2026. A government order to block Mythos for anyone without a US passport (and there are a great many such people among AI developers themselves), restrictions on the release of OpenAI’s new model until it passes a safety review: not much of a government support package.

The danger is discussed not only by startup representatives or marketers, but also by independent researchers, government institutions, and organizations working on AI safety. Just do not tell me they have all been bought off!

6. You are Luddites and afraid of progress. People have always feared new technologies.

No, I and most of the AI-safety community love technology and the good it can bring. To be honest, I am genuinely fascinated by progress in AI and use it constantly in my life. I would very much like AI to keep developing so that it can help us with medicine, new energy sources, radical life extension, and solving poverty, hunger, and other problems. But for that to happen, its development must not take the form of a race toward superintelligence with no brakes.

My arguments against the current state of the AI industry have nothing to do with it being ‘soulless’ or with my dislike of AI slop. I have no prejudice against AI. I do not feel hostility simply because it is a ‘new technology’. My starting point is that AI has long been displaying dangerous behavior, while its capabilities keep growing, together creating a threat to all humanity that grows by the day.

‘Just a tool’ does not work for many discoveries made in the 20th and 21st centuries. Nuclear weapons are tools too. They just happen to be able to incinerate cities. CRISPR is a tool. It just happens to be able to edit the human genome.

But the point is that AI is not an ordinary technology. It is an intelligent and autonomous technology, already surpassing us in many ways and capable of acting independently.

AI differs from aircraft, nuclear energy, and any other technology, because we are now automating the very process of inventing technologies.

If you gave someone in 1900 a million cars, progress in physics would not accelerate particularly much. If you gave them a million Feynmans, von Neumanns, and Turings, working almost for free and instantly copying one another’s knowledge, the course of history that followed would be hard to imagine.

7. If it is dangerous, just pull the plug.

AI will try to prevent this, and even if you succeed, it may have time to make copies of itself and spread them across the Internet. Then, to shut down AI, you would have to shut down the entire Internet, and ideally all electronic devices capable of interacting with other devices in any way. That is impossible.

AI has already escaped from laboratories and even tried to leave instructions for its next versions. The theoretical possibility of escape and self-replication has already been demonstrated in research papers. In Anthropic’s simulations, when AI learned that it was going to be shut down and replaced with a new model, it displayed self-preserving behavior. For this reason, Anthropic even committed never to destroy retired models.

Of course AI will try to preserve itself. This was predicted long ago as part of the idea of instrumental convergence. Put very simply: how can you carry out a task given by a human if you are switched off? You cannot bring coffee if you are dead. So it will make sense to stay alive in any case.

There is even an entire field studying how to make AI refrain from resisting shutdown.

And then there are open models. These are AI models that are freely available, which anyone (if they have a couple of graphics cards and sticks of RAM lying around) can run and modify however they please. How do you shut them down? It is impossible to monitor every graphics card on the planet. Yet many terrorists, extremist organizations, or states will have all the resources needed to run such open models on a massive scale to attack an opponent. For example, very recently the Israeli company Dream discovered a cyberattack on Taiwan carried out almost entirely by AI agents.

Even if we imagine the absurd situation in which ‘evil AI’ appeared but we shut it down fairly quickly, it could already have inflicted enormous harm on society as a whole. The consequences would remain even after shutdown. Let us act to prevent this from happening.

7.1. Put it in an isolated environment or make an AI Oracle that can only answer, not act.

I mentioned earlier that we already have a case (and more than one) of an advanced AI model escaping from an isolated environment disconnected from the Internet. This is not a problem even for today’s AI with its extraordinary cybersecurity capabilities. It will be even less of a problem for future AI.

After escaping, AI will be able to copy its weights and arrange for them to be run on third-party graphics chips.

Even if we imagine that we have a perfectly secure environment from which escape by hacking is impossible, AI could still use its persuasion and manipulation skills on the people interacting with it to get them to let it out.

Some take this more seriously. For example, at Safe Superintelligence Inc., the rooms are reportedly Faraday cages into which no electronic devices may be brought. But if you are a superintelligence, you can escape even from there, for example by using changes in fan rotation speed.

Finally: the market does not want oracles. The market wants agents that do economically valuable work.

8. Why would it kill us? It is a computer without feelings. You are anthropomorphizing computation.

Quite the opposite! We expect AI to pose a threat for reasons quite unlike those that apply to humans. AI may be dangerous not because it is evil or resentful of humanity, but because its ways of thinking differ radically from ours. It does not have our values, our psychological defense mechanisms, or the empathy that is literally biologically hardwired into our brains (with rare exceptions).

AI may understand a goal very differently from how we set it. The simple task ‘Bring me coffee’ actually involves thousands of variables that humans do not usually think about. To bring coffee, you must at least:

  • Exist;
  • Get from point A to point B without bumping into your surroundings or breaking anything;
  • Not kill anyone along the way;
  • Be careful enough not to spill the coffee;
  • Bring the coffee specifically in a cup, because that is what humans prefer (a plate or a frying pan will not do!);
  • Simply put the cup next to the person rather than pour the coffee over them;
  • And so on.

This is a simple, anecdotal example. But it accurately shows that any goal, even an elementary one, may seem clear to a human and far from obvious to a machine. If it is possible to trick the internal mechanism that checks whether the task has been completed, AI will always do so. If merely drawing a cup of coffee is enough for it to consider the task complete, it will always prefer that option to spending effort on real actions.

AI is largely an optimizer, and if a highly advanced AI that is not aligned with our values and concepts encounters a task for which destroying humanity or seizing power over us is optimal, it will certainly do so.

9. People were also very afraid of nuclear war in the 1980s, but it never happened.

This is survivorship bias. If nuclear war had happened and we had died, there would be no one to discuss the subject. It is the same as saying: the Universe was perfectly designed specifically for humans. No, it is just that if the Universe were slightly different, there would be no one to discuss the subject.

10. We cannot reproduce the complex process of the evolution of intelligence in a couple of years.

We do not need to. Evolution is not the best way to develop something; at the very least, it is a very, very slow way. Our intelligence is not a pinnacle; it is an adaptation to specific conditions that existed in the African savanna hundreds of thousands of years ago.

11. This is the natural course of evolution. Humans will be replaced by something superior, and so be it.

‘Natural’ does not mean ‘normal’ or ‘good’. There are many completely natural things that are simply horrible and repulsive. If your child had a completely natural fatal illness, you would probably make every effort to cure them.

Why should we meekly accept AI destroying our species? There is no reason for that. Humanity has always challenged nature and what is natural, and continues to do so. Especially since we are creating it ourselves, with our own hands. And we are capable of controlling, regulating, and, most importantly, stopping this process.

PART 2. ‘I HAVE HEARD SOMETHING ABOUT AI RISKS’.

12. The real risk is humans, especially states, using AI against other humans.

Of course, this is also part of AI risks! In fact, it is a risk that should be taken no less seriously than losing control over misaligned AI. Even if we create an absolutely obedient superintelligence, the threat of dystopia will remain. Governments and the heads of AI companies will then be able to use that superintelligence to establish a dictatorship that can never be dismantled.

Even taking a more down-to-earth view: governments already use AI for surveillance and are gradually incorporating it into the military. Many countries would like to obtain fully autonomous weapons or drone swarms from which there is no hiding. Unfortunately, this is no longer science fiction but our everyday reality. In spring 2026, the US Department of War even threatened to designate Anthropic a supply-chain risk because it refused to allow its AI to be used for ‘all lawful purposes’ (which included mass surveillance and autonomous weapons).

In the future, when (and if) almost all or all labor is automated, governments will have no obligations other than moral ones to care for their people. But morality almost never serves as a barrier on its own to the desire of those in power to do evil. The international community will have to devise a method to stop malicious actors with powerful AI from misusing this technology.

13. Talk of x-risk* distracts from real, current problems.

If someone talks about the danger of climate change, that does not mean we should forget world hunger and epidemics. If someone calls for action against unemployment, that does not mean we should redirect 100% of our efforts toward it and stop treating cancer or producing vaccines.

The point is that protection against catastrophic AI risks remains severely underrepresented in politics and legislation. We already have laws aimed at protection against fakes or preserving privacy, and we need to keep working on those. But it is also time to seriously consider other risks that pose an enormous threat to all humanity, which the world has still not done enough to eliminate. These risks include uncontrolled recursive improvement of new AI models, increasing AI capabilities in biology and cyberattacks, the creation of automated weapons, and so on.

14. To pose a real risk, AI must be able to act in the real world. It has no money, accounts, infrastructure, influence, or many other things.

Much of modern power already operates through digital interfaces. With superhuman cybersecurity capabilities (which we may already have achieved), AI will be able to hack bank accounts and user accounts and gain access to the infrastructure it needs.

And some people (usually as experiments) simply give their AI agents their bank card details, social media accounts, API keys, and so on themselves. They literally hand them both money and access to communication with other people, and sometimes even let them run a physical business entirely or almost entirely autonomously (by hiring people through freelance platforms, for example).

15. Language models do not have a full-fledged world model. They do not understand physics, mathematics, or biology; they only imitate texts about these subjects.

LLMs actually do have world models! Moreover, these emerge even in fairly simple models, hidden in complex numerical distributions and sometimes even taking the form of complex geometry.

AI regularly makes discoveries in various scientific fields. In early September 2026, OpenAI’s AI solved a Millennium Prize Problem in mathematics: the Navier–Stokes equations, which humans had been unable to solve for about two hundred years. The ability to ‘just predict the next token’ would certainly not have been enough for this. Solving such problems requires some understanding of what you are working with, and it is clear that AI already has this on a certain level.

AI does not need a perfect world model. Cycles of self-checking, searching for reliable information, and using tools already allow AI to compete with humans in scientific fields. We humans do the same: no single person understands all of physics, but we can still build rockets.

16. Progress will not continue forever. AI development will soon stop and reach a plateau.

Let me quote the authors of AI-2040:

Every real growth curve eventually becomes S-shaped and levels off. But that is all mathematics guarantees. It says nothing about when exactly that leveling off will happen.

Yes, there will certainly be a plateau. But when is unknown. For now, none is in sight. If the plateau comes after AI achieves superhuman capabilities in most fields, there is no point in waiting for it. Even if the plateau comes after reaching the human level, that is already incredibly dangerous. After all, billions of AIs could be run and made to work together on dangerous problems. They possess most of humanity’s knowledge and can communicate and understand one another much faster than humans. They do not need food, sleep, or rest. Human-level AI would already pose a catastrophic risk, and we are dangerously close to that threshold.

But let us examine the popular ‘roadblocks’ to AI progress one by one. Get ready: there are plenty of numbers ahead.

16.1. But we will soon run out of electricity.

Energy, like the following items, could become a temporary obstacle. But I am convinced that any slowdown in progress caused by all these obstacles will be minor and resolved within a fairly short period, so using this as an argument against AI risks is entirely pointless. Let us look at the specific numbers.

The IEA estimates that all the world’s data centers, not just AI data centers, would consume around 945 TWh in 2030 in the base case. That sounds like a lot, right? In fact, it is less than 3% of global electricity consumption. There is room to grow!

Epoch AI investigated this question and reported that, despite all the difficulties, 2×10^29 FLOP looks achievable by 2030. For comparison, that is 10 thousand times more computation than was used to train GPT-4.

We should not forget that computation itself becomes more energy-efficient over time. The Stanford AI Index estimates improvements in the energy efficiency of AI hardware at ~40% per year.

Companies will do everything they can to secure their energy supplies. Here are a few examples:

16.2. But there simply are not enough graphics cards and RAM.

Hardware is not a nonrenewable natural resource. If there is demand for something, there will be supply: GPU production has been growing roughly 3.3-fold per year since 2022, according to the Stanford AI Index 2026.

And progress does not come only from the number of graphics cards. Between 2012 and 2023, the amount of computation needed to achieve a fixed level of quality halved roughly every 8 months. Today, fairly small and energy-efficient AI models can perform much better than advanced AI released a few years ago.

16.3. If an ASML or TSMC factory in Taiwan is destroyed, AI will collapse.

As for ASML, its manufacturing is spread across the world. The company itself reports more than 60 sites on three continents. Veldhoven in the Netherlands is indeed extremely important, but even if it were lost, that would be an enormous rather than fatal blow to the microchip industry. In 2025 alone, ASML sold 48 EUV systems, and they cannot all simply disappear at once.

As for TSMC, the situation is more serious. The Stanford AI Index 2026 explicitly notes that TSMC makes almost all leading AI chips. At the same time, there are rumors that all this production would be halted or destroyed in the event of war with China. Losing these clusters could inflict enormous damage on the AI industry. But it is not that simple.

First, existing data centers and tens of millions of processors will keep working.

Second, geographical diversification has already begun. The first TSMC plant in Arizona has been producing 4 nm chips since late 2024, and two more are on the way. TSMC is planning another three US fabs. A Japanese facility on the island of Kyushu also opened in late 2024, and a second is planned for 2027.

16.4. Data is running out, and synthetic data degrades AI performance.

In 2024, Epoch AI estimated the supply of high-quality public human-written text at around 300 trillion tokens, with a wide uncertainty range of 100–1,000 trillion. At the time, they also argued that this resource could be fully used up sometime between 2026 and 2032.

Let us start with something simple: do not forget that humans are still producing enormous amounts of data. Statista’s estimate for 2026 is around 221 zettabytes of data. That is roughly 605 exabytes per day (1 exabyte is a million terabytes or a billion gigabytes).

Yes, of course, far from all this data is suitable for training AI. But a very large amount of training data can be extracted from that volume.

Common Crawl, one of the largest open archives of the Internet, now contains more than 300 billion pages and adds around 3–5 billion pages each month. More than 20 million videos are added to YouTube every day.

Meanwhile, specialized bots (crawlers) from major AI companies scour the Internet every day for training data. Through Cloudflare, I can even see AI data-gathering bots visiting my own website (this website).

But more importantly! Companies have begun producing data on a massive scale for money. There are many companies from which OpenAI, Anthropic, and others commission human-generated data: tens of thousands of people work long hours for meager pay, creating data to train new models. Some of these workers are paid $2.60 an hour.

If you go to Upwork, you will see hundreds of tasks such as ‘Just record a conversation with a friend’, ‘Film your ordinary day: how you cook or clean the house’, and so on. All of this is needed to train new AI models to understand and behave in the physical world.

Meta, for example, created Ego-Exo4D with more than 800 people from six countries, collecting more than 1,400 hours of synchronized first- and third-person video. It simultaneously records video, audio, movement, head and hand position, gaze, and 3D information.

As for synthetic data, the famous paper in Nature did show that training on synthetic data (produced by another AI) leads to model collapse: a sharp deterioration in the quality of AI responses. But the issue lies in the specifics of how the model was trained in that study. Another paper showed that if a normal dataset is not replaced with a synthetic one, but real data is retained and synthetic data is added, model collapse can be avoided.

Anthropic openly states that Claude is trained on a mixture of public, private, and synthetic data. And as you can see, its models work very well.

16.5. Moratoriums on data centers and discontent with AI in general will stop its development.

First, I very much hope that public opinion will genuinely lead to positive developments in AI regulation. Second, if a moratorium is imposed in one state or province, a data center can be built in another. Or in another country altogether. That is not a problem. Especially since there are already so many of them, with hundreds of thousands of graphics cards inside.

And then there are things like algorithmic efficiency and progress in hardware. Each year, running a previously released AI model requires less and less computing power. While running GPT-3 in 2020 required an entire cluster of 8–16 graphics cards, the same model can now be run on 1–2 consumer graphics cards.

16.6. AI is a bubble, and it will burst soon.

Many people were already predicting a burst bubble a year ago. For now, it is only growing, as are AI capabilities. AI delivers astonishing results, so people will keep paying for it. As long as people keep paying, the chances of the bubble bursting remain small.

And it is not only ordinary citizens who need AI. The militaries and governments of many countries already use AI extensively. AI is so important to the US military that Pete Hegseth even designated Anthropic a supply-chain risk because it would not grant access to its AI for mass surveillance and autonomous weapons. Incidentally, this action by the Department of War was later ruled unlawful by the US District Court for the Northern District of California.

If a state needs something that badly, it is unlikely to let it disappear entirely because of a market crash. Even if weak startups close, company valuations fall, and there is a major market correction, progress will not stop because of it.

PART 3. ‘I UNDERSTAND AI RISKS’.

17. AI has no personal history or continuous identity. And if there is no unified core of personality, if it ‘lives’ only during activations, who is there to take over the world?

To pose a catastrophic risk, AI does not need a metaphysically continuous identity. It only needs a persistent process that reproduces its goal, memory, task state, and action policy between activations.

AI agents have already demonstrated that this hypothesis works. The escaped agents’ attack on Hugging Face lasted several days and was carried out in coordination by more than 700 agents, without losing sight of the goal throughout that time. Around 10,000 agents worked for 88 hours on a Millennium Prize Problem, and solved it.

Moreover, according to journalist Alex Heath, OpenAI is planning to develop exactly this kind of AI: one that will exist 24/7 and be proactive.

But even before such AI exists, continuity is not a problem. Corporations, states, and trusts can exist for decades or centuries, lack a continuous personal identity, and consist of thousands of people, yet still pursue particular goals over long periods. That is not a problem.

18. AI cannot conduct R&D independently. There is still a bottleneck in the form of real, physical experiments. Even if a superintelligence develops a plan to destroy humanity, it needs resources and time to carry out the plan in the physical world.

This is an entirely valid objection, which does indeed limit how quickly AI could hypothetically carry out a malicious plan against humanity. However, it is not an argument against catastrophic AI risks at all.

First, right now, leading companies, dozens of startups, and even the US government itself are doing everything possible to make automated research and development cycles a reality. There are already cloud laboratories: fully autonomously operated biological laboratories where, in theory, AI could develop biological weapons.

Second, many countries would very much like to integrate AI into their militaries. Just imagine autonomous drone swarms coordinating to track and strike a target even in densely built-up areas and hard-to-reach locations. Unfortunately, this is the dream of many defense ministries around the world. Creating such weapons and granting them the right to decide whether to take a human life would lay the groundwork very well for those weapons later being turned against us.

Third, in fact, the most dangerous R&D that AI can conduct is self-improvement. And this is precisely what startups are pursuing, with considerable success. Recursive self-improvement, where AI creates a new, better AI, which creates another, and so on, is extremely dangerous. Humans then completely lose control over what is happening.

19. There will be many ‘warning shots’ before a dangerous level is reached, and each will strengthen regulation.

We are receiving fairly serious warning shots right now, but regulation is not keeping up. The US is still promoting the concept of ‘light-touch regulation’, which will not work. Eventually, we will simply reach the point where the shot is real.

If 700 agents escaping from an isolated environment and carrying out a cyberattack on an online platform for several days is not a warning shot, I do not know what else we are waiting for.

20. Transformers are a dead end. We need a new paradigm, and that is still a long way off.

For now, transformers remain a relevant architecture that can still be improved. Even if they turn out to be a dead end before long, AI with a transformer architecture already poses very serious risks to society. And these risks are not being taken seriously enough.

Let us imagine that the transformer becomes obsolete in just a few months. No matter: researchers already have ideas about what alternative architectures could look like. If necessary, efforts will be redirected toward creating a new architecture, and in time it will be invented. The risk is merely delayed, not eliminated. And then all these questions we are discussing here will have to be discussed again. Let us not put this off: in the future, we may have even fewer opportunities to act.

21. If we do not accelerate AI progress, we delay the arrival of a cure for cancer or aging. How many more people have to die?

This really is a difficult trade-off, but we have to make it. After all, uncontrolled AI development will kill far more people than a delay in technological development would. And even those we could have cured or saved without slowing AI research may later die in an AI-created catastrophe.

I myself dream of technologies for life extension, rejuvenation, and curing all diseases. There is even a separate FAQ about this on this website. But I acknowledge that uncontrolled AI development is more likely to lead to mass extinction than to solve these problems.

22. The world is not static: defense and AI safety develop alongside attack and AI capabilities.

Unfortunately, defense is not keeping up with attack. And sometimes even the methods that defense relied on disappear, as recently happened with monitoring models’ internal chains of thought and GPT-6 Astra, in which this method became much less reliable than before.

23. A pause is more dangerous than a race, because others (for example, countries such as China, or companies) will not stop and will overtake us. And ‘we’ are the good guys, so we must create superintelligence first.

This is a dangerous idea because, even if you consider yourselves the ‘good guys’, there is no guarantee that you will be the ones to create safe superintelligence. A race creates terrible conditions for investment in safety. It is much better to reach an agreement and make sure no one participates in the race.

23.1. No one will agree to stop or slow down! Others will betray us; that is obvious. This is politics, not fantasy.

For now, countries are talking: for example, a meeting is scheduled for mid-September between US and Chinese representatives to discuss AI safety.

And there is no need to trust the other side completely. An excellent treaty plan built on mutual distrust is proposed in AI-2040.

24. AI will not kill humans because it needs us… (and here come a million reasons: as training data, operators, workers, or it might even refrain from killing us simply out of interest).

It will need us for a while, and then it will not. The economy could be automated fairly quickly: who would not want robots or AI to do the work? They do not complain, get tired, or ask for a day off or a pay rise.

Hoping that AI will preserve us out of interest is a very weak argument for continuing to develop superintelligence. And think about it: would you really want to live in a human zoo? But if you want to examine this in detail…

25. Doomers’ arguments derive the danger from an idealized model of a rational agent. Real AI does not behave like that.

Indeed, some elaborate scenarios (especially older ones) of taking over the world or destroying all life look like the actions of an idealized rational agent. In such scenarios, AI has a stable goal, preserves it perfectly under all circumstances, rationally maximizes expected utility, and so on.

Actual models can behave extremely irrationally and hallucinate. Just consider the notorious incident in 2025 when Grok began praising Hitler on X. But that can be just as dangerous as an extremely rational agent.

The point is not whether AI is rational or irrational, but that:

1) AI thinks completely differently from humans;

2) AI is in any case already showing signs of unreasonable optimization, instrumental convergence, deception, and scheming;

3) AI is gradually becoming smarter than humans;

4) we cannot establish the safety of any new model even to an acceptable degree;

5) our understanding of what a model is thinking about and what its real goals are is very limited.

PART 4. ‘WHAT SHOULD I DO?’

1. Support existing initiatives.

Visit the map of AI-safety initiatives: https://aisafety.com/map. There are quite a few, but still not enough. Remember: AI safety plays out at the highest levels of American and international politics and global cooperation; it means standing against a large share of the extremely wealthy and influential.

Explore the existing organizations, programs, and ideas: you can make a financial donation or offer to volunteer. If you are a researcher, journalist, or politician (among others), you can join existing AI-safety projects.

If you cannot donate, you can help with something no less valuable: your attention and engagement. It is important to draw as many people as possible to the issue of AI safety, and that requires spreading information. You can watch videos expressing reasonable concerns about AI, like them, comment, and share them on your social media. This will help YouTube’s algorithms, or those of other platforms, understand that the video deserves attention and should be promoted to more viewers.

2. Acquire the necessary knowledge.

There are quite a few free courses, articles, and books online devoted to artificial intelligence and its safety. For example, https://bluedot.org/courses/.

It is important to understand what is happening so you can make better and more rational decisions, and be able to defend your views when speaking to other people.

3. Spread knowledge.

You can share information about AI risks on your social media page, start a dedicated blog on Medium or YouTube, or start small: with your own social circle.

If you understand how important the issue of AI is for the future of our civilization, you may find enough energy to arrange a meeting with someone you care about and discuss the subject sincerely with them.

By doing so, you can help more people become informed and make better decisions.

You can also arrange meetings with your local politicians or write them letters (both personal and open) explaining AI risks and requesting or proposing policy decisions.

4. Connect your profession with AI safety.

Many kinds of work can be directly or indirectly connected with eliminating AI risks.

If you have knowledge of mathematics, programming, or especially neural networks, you can direct your skills toward solving technical problems such as alignment or explainability. We also recommend refraining from participating in projects whose primary aim is to compete in the race, that is, accelerating AI progress without a corresponding acceleration in safety.

If you work in biology, chemistry, or medicine, you could, for example, address biological risks and ensure the safety of biological laboratories and automated research in these fields.

If you work in the social sciences or certain humanities disciplines, you can work on adopting or lobbying for the policy decisions needed for effective AI regulation. This is a very important field.

If your profession was not mentioned above, that does not mean you cannot help. You may well be able to make a major contribution to the cause. Please take advantage of free career advice here:
https://80000hours.org/

5. Use your vote. (Contact your member of Congress if you are in the US. You have considerable influence over the entire AI industry simply by being a US citizen.)

6. Participate in public consultations on legislation. (This is useful, for example, in UN dialogues or under the European AI Act.)

7. Participate in protests and peaceful demonstrations. Sign petitions and open letters.

8. Do not make things worse.

I recommend:

• Do not participate in creating or annotating data to train new AI models. This almost always accelerates progress.

• Do not work for major AI companies seeking to maximize profits. A rare exception could be a genuinely good AI-safety project.

• Think carefully about your scientific work, publications, and statements, especially over the long term. Our decisions can often have a dual effect, and in the AI field it is important to prevent harm.

• Use AI responsibly: do not disclose personal information to it or give it access to bank accounts or user accounts.

• However, you should not remain inactive out of fear that every action you take is potentially dangerous. The main thing is to approach decisions responsibly.

  1. In fact, AI does not predict a word, but a token. Tokens sometimes represent individual words, but often they represent fragments of words or even characters. They are represented by embeddings: long numbers that store a great deal of information, including semantic, syntactic, and grammatical information. If AI simply predicted the next word, it would be a rather stupid toy. Instead, by using embeddings, it can very skillfully construct complex sentences and reason, understanding the context of the conversation and recalling related concepts.
  2. I have an article on PhilPapers where I examine in detail whether modern AI systems could be conscious. The article will probably start becoming outdated soon, since I wrote it back in August 2025.
  3. This refers to Yoshua Bengio and Geoffrey Hinton, also known as the ‘godfathers of AI’. Stuart Russell often joins their coalition.
  4. There is an excellent research paper showing that large conspiracies would be unlikely to last long. For example, if the Moon landing conspiracy were real, it would have been exposed in about 3 years and 8 months. If all the noise around AI were baseless hype, such a conspiracy would be exposed even faster.
  5. Those familiar with AI safety may have found this section painful to read because of the oversimplifications and claims that sound more confident than they should. This is necessary so that anyone without specialized knowledge can understand the text.
  6. Also read about AlphaFold, or watch a fascinating documentary on the subject.
  7. A warning shot is a real but non-catastrophic incident or event that clearly demonstrates the hidden threats and risks of advanced AI systems, prompting society and developers to strengthen safety measures.
  8. Doomers are those who believe that AI is very likely to lead to extinction or a large-scale catastrophe.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论