What Really Is the Best AI Tool in 2026?

My Medium friends can read this over on Medium.

"Which AI is the best?" Is there really an answer for that question?

ChatGPT started it all. Claude went nuts in coding. Gemini looked great for a while.

So who wins? 2026?

I have five on the list, plus five more that are somewhat unique, and I can't name a winner right away. Not for coding, not for writing, not for research. But I can rank and share what my opinion is.

The models got really close. The products drifted apart somewhat. That's basically the story here, and it's why this piece is long.

So, let's look at all of them.

Model or product

Two questions hide inside "which AI is best."

The model decides how well it reasons, how good the code is, how much context it holds, how fast it answers, how well it… does things you want.

The product on the other hand decides how (well) you work with it every day. Search, files, images, your mail and your docs, agents, what it can touch on your machine, and who gets to see your data.

A great model in a weak product loses to a decent model that sits right next to your work. That’s a crucial piece of the puzzle.

The deeper a tool sits in your system, the more it can do for you. That's what pushed me from the browser tab to the terminal.

Keep both questions in mind. The five below win different ones.

1. Coding

I start with that because coding is my most-used AI use case.

And yes, Claude Code is still my default. At home and at work. Has been for a while. It runs in the terminal, holds context across a big, messy project, and makes fewer dumb mistakes than anything else I use. At least, that’s how it used to be. But over the last few weeks, especially with the release of Opus 5, I feel like Claude is making more and more mistakes. Not a good look.

It's also the most expensive of the bunch. A big catch, and I'll get to the numbers.

Codex from OpenAI gets close. Very close. Fast, accurate, (much) cheaper. I mainly use it to review what Claude Code did. And it’s dang good.

Codex and the Mac app have gotten so good, I am really tempted to switch to OpenAI entirely (for a while) and see how that goes.

Gemini's coding tool is Antigravity, a VS Code fork with the Google suite next to your code. Version one crashed on me every twenty minutes. Version two works (better). I went through all three of those recently.

It’s still lagging behind Claude and Codex. Mistakes, crashes, weird logic, strange behaviors. For coding, Gemini isn’t great.

Then there’s China.

Kimi K3 is relatively new. Moonshot AI shipped it on July 16, the open weights followed on July 27. 2.8 trillion parameters, the biggest open-weight model so far I’d say. Moonshot's own report puts it behind Claude Fable 5 and GPT-5.6 Sol overall, but level with Fable 5 on agentic coding tests like SWE Marathon, and ahead of GPT-5.6 Sol there. For an open model, that's remarkable. Then it burns about twice the tokens getting there. More on that in the money section.

DeepSeek is a pretty cheap engine. Not a coding product, a model you plug into your own tools. Solid code for the price, and the price is the story: V4 Pro costs $0.66 per million input tokens off-peak. Fable 5 costs $10. I mean that’s quite the difference.

So in short: For serious code, Claude Code. Still. But basically tied with Codex models. If money matters more, definitely Codex, or even DeepSeek behind a tool you already have. Kimi K3 if you want open weights and a frontier-level coder in one.

2. Research

Few people talk about this, at least in my space. But research is such an important and somewhat underrated topic in AI, because it’s the backbone of good, correct answers.

ChatGPT has search and Deep Research built in. Ask it to compare five things and hand you a table with sources, and it does. Pretty solid. Not perfect, but great. Getting better with each release.

Gemini has Google Search underneath, which sounds like an unfair advantage and sort of is. Plus Gemini Notebook, the tool formerly known as NotebookLM. Upload your sources, ask questions, get answers grounded in those sources only. Since the rename in July, every notebook also gets a small cloud computer that writes and runs code against your files. In practice, though, I don’t see much difference between Gemini research and ChatGPT research. For now.

Claude searches too. But its strength is reading what you hand it. Long PDFs, a folder of notes, a spec, a codebase. For many questions I don't open Google anymore. I ask Claude. The answers are good. But it's not the best at research by any means. And again, Claude keeps making more and more mistakes with the latest models.

Kimi and DeepSeek depend on the setup. Kimi's app has agent modes that search and browse. DeepSeek is mostly a model behind an API. You bring your own search.

For anything that gets published: check the sources anyway. A convincing answer and a correct answer are two different things. Still.

3. Images and video

Media is getting more and more important, because the outputs get better and better each week.

ChatGPT generates images natively, and they're good. Video was Sora. Was. OpenAI shut the Sora app down on April 26, and the API goes on September 24. So, for now, ChatGPT is an image tool, not a video tool. That changes quickly. Images are great. They used to suck. Now they’re way up there.

Gemini has the biggest media package of them all. Image generation, and since I/O in May, Gemini Omni Flash makes and edits video right in the Gemini app, for AI Plus, Pro, and Ultra subscribers. You describe the change, it edits the clip. Short clips, for now. For image and video, Gemini is still hard to beat.

Claude makes no images and no video. None. It reads them fine. For a writer that's no problem. For someone producing social visuals all day, it's a gap. And I don’t understand why they don’t already do it. At least, Claude has great integration with media AI tools like Higgsfield. MCP, quick, well-integrated. Works like a charm. But that’s another hefty price tag on top of the already pricey Claude subscription.

Kimi K3 takes images and video as input; generating them goes through plugins in the Kimi app. DeepSeek is a text model. Bring your own image model.

4. Writing

I use AI to write a lot of drafts, ideas, plans, and such. And I use it to research, organize, and clean up. And writing stays one of the biggest limitations of AI models. They all don’t sound great. Without good (and long) instructions, you can’t produce anything remotely creative.

But that can change. It will. And with the right prompts and skills, AI gets the job done much better. For that:

Claude is good for long, structured text. Blog posts, briefings, concepts, editing. With great prompting and well organized skills for tone, voice, structure, it can produce fine drafts. Nothing to publish immediately, but closer than most.

ChatGPT for volume. Twenty hook variants, social posts, repurposing, templates. Many small pieces, high frequency. None sound great. But it’s quick, relatively cheap, and easy to automate.

Geminiis for when the piece starts with research that already lives in your Drive, Gmail, Docs, or wherever. Combined with Google’s searches or Gemini Notebook content, it’s a powerhouse. But that’s quite the setup and highly depends on your Google ecosystem integration.

Kimi K3 for the odd case where you don't want a chat answer but a finished file. It produces editable .docx, .xlsx, .pptx and .pdf right in the app. Claude can do that through Skills. Kimi does it out of the box.

DeepSeek for bulk. First drafts, summaries, classification, a thousand product descriptions. A human or a better model does the last pass.

And it doesn't matter who wrote it, if it's a good read. I said that before.

5. Agents

Here the products drift apart the most. Intentionally.

ChatGPT has Codex, an agent mode, Voice, GPTs, Actions, and the biggest consumer ecosystem of the five. It is, in my opinion, the best all-rounder of the bunch.

Claude has Claude Code, Cowork for the non-coders, and Skills that load themselves when a task fits. Less plugin store, more reliable execution. It’s great for a coding and orga workflow setup that doesn’t require any media gen capabilities.

Gemini's agent is Spark (if you get access to it), a background agent that keeps working with your laptop closed, deep in Gmail, Calendar, Drive, Docs, Sheets. Since the end of July it's on the $20 Pro plan in most of the world. Not the EU, not the UK, not Switzerland. Not for me, then.

Kimi has an Agent mode and Agent Cluster. Cluster coordinates up to 300 sub-agents in parallel, for big searches, batch processing, very long documents. For a batch job, interesting.

DeepSeek is the building block. Model via API, tools via MCP or function calling, the rest you build. And maintain.

6. Unique features

Feature lists overlap a lot by now. Search, files, code, voice, everybody has it. The few things that only one of them has matter more to most people.

Claude: Claude Design. Since April, as a research preview, included in Pro and up. A prompt, a document, or a whole codebase in; slides, one-pagers, prototypes, and landing pages out, in your own design system, exportable to Canva (MCP), PDF, PPTX, or plain HTML. Anthropic calls it a Labs product. Better than what ChatGPT or Gemini have for that kind of work. It’s pretty great. For agency work, social media, and creative workflows, Claude Design is a huge plus. What hinders it a lot, is again the lack of image gen and video gen capabilities of Claude.

ChatGPT: Sites. New this summer, in public beta for Plus, Pro, and workspaces. Describe a website, a web app, or a game in chat, ChatGPT builds it, hosts it, and hands you a URL. Hosting, access controls, storage, database, all included. No deployment, no hosting bill. That’s a great tool set for any web designer & developer or anyone who needs a quick web page. Hard to beat.

Gemini: Gemini Notebook and the Workspace depth. Nobody else sits inside Gmail, Docs, Sheets, Drive, and Search the way Gemini does, and Notebook is the one research tool that only answers from what you gave it. If your life is in Google, this is hard to beat.

Kimi K3: open weights at frontier level, plus files. You can download the model. Editable Office files fall out of the chat. And Agent Cluster.

DeepSeek: the price. And the off-peak discount, half price outside weekday peak hours, all weekend. As far as I know, nobody else prices by the clock.

7. Context and money

Context windows are boring now. Kimi K3 and DeepSeek V4: 1 million tokens. Gemini: a million or more, per Google's docs. Claude and ChatGPT: depends on the model and the plan.

Subscriptions start at $20. ChatGPT Plus, Claude Pro, Google AI Pro, all $20 a month. Claude Pro includes Claude Code, Cowork, and Design. ChatGPT Pro is $200.

The APIs show the gaps. Per million tokens, input and output:

  • Claude Fable 5: $10 / $50
  • Kimi K3: $3 / $15
  • DeepSeek V4 Pro: $0.66 / $1.98 off-peak, double at peak
  • DeepSeek V4 Flash: $0.22 / $0.66 off-peak

Price per token isn't what you pay, though. Price per finished task is. In Artificial Analysis' test suite, K3 generated about twice the tokens the field needed for the same tasks. A third of the price times double the tokens is not a third anymore. I ran that math when K3 launched.

DeepSeek's clock trick: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else, including the whole weekend, is half price. Run the batch job on Sunday.

8. Downsides

Cost first. Claude is the best of the five for hard work and the most expensive, and heavy Claude Code use adds up fast. I know.

Then the data question. Moonshot and DeepSeek are Chinese companies. Open weights mean you can host the model yourself or through a European provider, and your prompts never leave. That's the good part. Using their apps means your prompts go to their servers. Where those are, and who reads along, check before you paste a client contract in. Kimi's weights also come under their own "Kimi K3 License", not plain MIT or Apache. Read it before you build a business on it.

Then the EU thing. Gemini Spark: not available here. The most interesting product features ship US-first and reach us months later. If at all. That's regulation and lawyers, not the model.

Claude: no images, no video, and Claude Design is a research preview. Previews change or disappear. Sora just showed how fast.

Gemini: 3.5 Pro was promised for June and still isn't out. And which Gemini you get depends on the plan and the app. Confusing.

ChatGPT: it does everything, for everyone, and OpenAI is the company behind it. You have to want to buy from them. Many do. I'm not sure.

Kimi K3: slow output, 62 tokens per second when it launched, below the median of its price tier, and the token appetite. DeepSeek: no product to speak of. You're the product team.

The "others"

Five more AI tools that deserve a paragraph here:

Grok. xAI shipped Grok 4.6 on August 12, and it lives inside X. A plus for some, a minus for others. Grok 4.7 is supposed to follow within weeks. Musk said so, so… we'll see. I haven’t tried Grok much, so I can’t give an honest verdict.

GLM. z.ai released GLM-5.3 on August 14 and calls it the top open-weight coding model. The full weights were promised two weeks later and hadn't landed when I wrote this; the smaller GLM-5.3-Flash did, on August 26, under MIT. On Artificial Analysis' open-weight list it sits third right now, behind Kimi K3 and Alibaba's Qwen3.8. Three Chinese models at the top of the open list. All three.

Mistral. The European one. Mistral Medium 3.5, 128 billion parameters, 256k context, open weights, a chat product called Le Chat and a coding product called Vibe. Smaller than the others. But the data stays in the EU, and for some of us that's the whole point. Another story, though.

Perplexity. No frontier model of its own. Pro lets you pick between GPT, Claude, and Gemini underneath, and Perplexity sells the product on top: search with sources, a browser called Comet that's free since late 2025, and since February an agent called Perplexity Computer that runs a whole set of models for one task. Everything I said about product over model applies here. It's all product. And Perplexity is a good one.

Cursor. The other way around. An IDE with the frontier models inside, and its own: Composer 2.5, released in May, trained on top of Kimi K2.5 with Cursor's own reinforcement learning. Artificial Analysis puts it third on their Coding Agent Index, behind Claude Opus 4.7 and GPT-5.5, at ten to sixty times lower cost per task. Only inside Cursor, though. No API. Pro is $20 a month, like everyone else.

The Bottom Line

Ten tools, no winner. I don't think that changes anytime soon. It’s too close.

What I do: Claude Code for the work that matters. Codex as the reviewer. A browser tab with ChatGPT or Claude when I'm away from the Mac. Gemini when I want the Google world. The Chinese models I read about and run the numbers on. Nothing switched. Yet.

Two paid tools, maybe three. Not one. The hard, expensive stuff goes to the best model, the bulk to the cheap one.

How many do you pay for?

If you're not already, follow me on Medium.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论