I wanna begin with conclusion first, because that's the only honest way to start:
1. This is a snapshot, not a verdict. Things move fast — new models drop almost every week, and keeping up feels less like learning and more like chasing a train that keeps accelerating. So everything here reflects where things stand as of mid-2026. Check the date before you trust the take.
2. The names change, the leaders don't. Tools launch every single day, but most of them are wrappers — a pretty frontend with a backend that just pipes your data off to someone else's model. Strip that away and you're left with the players who actually build the engines: OpenAI (ChatGPT), Anthropic (Claude), and Google (Gemini). So, lets go straight to the source. Skip the middleman.
3. Nobody wins every category. There is no model that tops every column. Each one trades strengths for weaknesses. So please — stop announcing that one model is "the best." That sentence is almost always wrong.
And here's the whole point: the question isn't which AI is best. It's which AI is best for you.
A quick word on the others people always bring up — Perplexity, Grok, DeepSeek. DeepSeek is banned in the US, so that settles that. Perplexity… I'm not sold on where they're headed. And Grok — look, it's backed by the same orbit as SpaceX, and I'd never bet against that crowd long-term. But right now? It's not there yet. The moment it is, that's its own video.
Sooo, lets begin…
One thing to get straight before we compare anything: ChatGPT, Gemini, and Claude are interfaces — the app you open, the chat box you type into. Behind each one sits a rotating lineup of models doing the actual work, and that lineup changes constantly.
Each also comes in tiers, from a few dollars a month all the way up to $200+. Comparing a $200 plan against a free one tells you nothing. So to keep this fair, I'm comparing the tier most people actually pay for: the ~$20/month plan, which lands in roughly the same place across all three.

Source: https://chatgpt.com/pricing/

Source: https://claude.com/pricing
Here's how we'll judge them — ten rounds:
Video generation (spoiler: 🥇 Gemini )
Music generation (spoiler: 🥇 Gemini )
Image generation (spoiler: 🥇 ChatGPT )
Chat/Writing (spoiler: 🥇 Claude )
Code generation (spoiler: 🥇 Claude )
Agentic AI (spoiler: 🥇 OpenAI (Codex) and Anthropic (ClaudeCowork)
Third-party tools (spoiler: 🥇 Claude and ChatGPT )
Other features (spoiler: 🥇 Gemini )
Benchmarks — (spoiler: 🥇 it depends )
1. Video generation.
I start with video generation, since this is the most obvious category…
Claude: no video model. Nothing to see here.
OpenAI: gone, and I'm a little sad about it.
Sora was genuinely great — Sora 2 could pull off stuff that used to be impossible, like a gymnast's routine or a triple axel with a cat clinging on.. or doing such things - my example - https://www.instagram.com/p/DRAR2HNjjmq/?
But it's over. The Sora app shut down on April 26, 2026, and the API follows on September 24, 2026. What killed it wasn't quality — it was money. Video eats far more compute than text or images, and OpenAI couldn't keep the costs in check; at its peak it cost an estimated $1 million a day to run while users drifted away. A shame. I liked that one. OpenAI + 4
Gemini: the last one standing — and it's strong. This is now Google's game by default. Their video model just got a new name: Gemini Omni, which replaced Veo in the Gemini app. It animates photos or builds video from basically any input, and you shape it step by step just by talking to it — the editing-by-conversation thing is the part that feels genuinely new. Under the hood it's the Veo line, which is the model people kept calling the most realistic, with native audio — dialogue, lip-sync, and sound effects generated in the same pass instead of bolted on after. Gemini + 2
So for video, the choice basically makes itself:
1st place🥇: Gemini (with its model Omni)
2nd place 🥈: —
3rd place 🥉: —
2. Music generation
Music generation — easy call, and Gemini's the only one that even shows up.
Neither Claude nor ChatGPT generates music. Gemini's the only one of the three that does, and the model is called Lyria 3. Describe a genre, mood, or upload a photo, and it builds a track — with lyrics, vocals, and instrumentals, plus auto-generated cover art. Google
And here's the part that matters if you're not just messing around: on the $20 plan, this is actually usable for real projects. The Plus and Pro tiers unlock Lyria 3 Pro, which generates full songs up to three minutes long — not just a 30-second loop, but a real structure with intro, verse, chorus, bridge, and outro that you can actually direct. Touted use cases are exactly the creator stuff — custom tracks for vlogs, podcasts, and tutorial videos. 9to5Google9to5Google
So for music, the answer's simple: it's Gemini or it's a separate tool entirely.
1st place🥇: Gemini (with its model Omni)
2nd place 🥈: —
3rd place 🥉: —
3. Image generation
With image generation that’s also pretty simple to compare…
Lets begin with Claude. Claude can't generate images. Period. That's not a knock — it's just not what Claude is built for. So if image generation is your thing, look elsewhere; this isn't Claude's lane. What Claude can do is the reverse: feed it a screenshot or a photo and it reads what's inside — text, charts, UI, handwriting. That's OCR (Optical Character Recognition — tech that turns text inside an image into text the machine can actually read). Great for pulling info out of an image. Useless for making one. Don't confuse the two.
Gemini. Not long ago Google's Nano Banana model broke the internet — for a stretch it was the one everyone was posting. But the landscape moved.
ChatGPT is, by a wide margin, the best right now. In April 2026 OpenAI dropped ChatGPT Images 2.0 (the model's called gpt-image-2), and it's not a close race. The thing that used to instantly out an AI image — garbled text, menus full of "enchuita" and "burrto" — it now nails. Near-perfect text rendering, 4K output, and you can edit an image across multiple turns while it keeps everything else consistent. For most real-world use, it's the one to beat. OpenAI + 2
Now the part that actually matters day-to-day: the watermark.
⚠️ Heads up — this is where the common wisdom is now wrong. It used to be true that ChatGPT images were clean and Gemini stamped a watermark. The visible part still holds: Gemini brands every downloaded image with a sparkle logo in the corner, and there's no setting to turn it off — your only outs are the API, an editor, or a removal tool. ChatGPT doesn't stamp a visible logo. AI Toolbox
But here's what changed: as of 2026, OpenAI now applies Google's invisible SynthID watermark — plus C2PA content credentials — to every image generated through ChatGPT and its API. So both now carry an invisible "made by AI" fingerprint baked into the pixels. And that one's the stubborn kind: SynthID is embedded during pixel generation and can't be removed without wrecking the image — it survives crops, compression, and editing. GPT Watermarker + 2
So if your beef is the visible sparkle, that's still Gemini's problem, and yeah — you can wipe the visible Gemini logo with a free tool in seconds. (I'm lazy; I resent every minute spent on watermark cleanup. But it's a minute, not an hour.) Just know the invisible mark is now on both, and that one isn't coming off either platform.
So my pick for Image generation:
1st place🥇: ChatGPT (image-2)
2nd place 🥈: Gemini (nano-banana-2)
3rd place 🥉: —
4. Chat, writing & text generation
…And here's where it gets genuinely hard.
I'll be straight with you: this is the toughest category to call. All three are good enough at writing. And writing — especially copywriting — is subjective in a way image quality just isn't. Tone of voice, the canvas/editing experience, citations, how often it hallucinates, how it holds an angle — different people weight all of these differently, so there's no clean scoreboard here.
But I'll give you my honest take, and it comes from working alongside actual copywriters — people with 10+ years in, the kind who feel every word, every shift in mood, every angle. Across the board, they say the same thing: Claude is the best writer, and it's not close.
That's not a benchmark. It's the verdict of people whose whole job is sensing the difference between writing that's technically fine and writing that actually lands. Take it for what it's worth — but when the pros consistently point the same direction, I pay attention.
So instead of pretending I can crown a winner objectively, I'll do something more useful: a simple test, run across all three, so you can see each tool's personality with your own eyes and decide which one fits your ear.
Create an email where I communicate like an alien making first contact with humanity.
One Prompt, Three AIs: ChatGPT vs Gemini vs Claude
And in my opinion, the quality of this reply is dominated by Claude. There are also two features I want to underline here:

Claude providedme with three takes on first contact, depending on the vibe I want:

Claude shows to me ‘Send email’ in one click… as you remember,I am lazy, so I love such 1-click integrations
ChatGPT has now deprecated its editing feature — Canvas, the side panel you used to collaborate on a draft. Which surprises me, because it was a genuinely great feature. As of the GPT-5.5 update, OpenAI replaced it with inline "writing blocks" right in the chat — not the same thing. The good news: if you're a paid user, Canvas still works in the older models like GPT-5.4, at least until those get sunset. (Worth noting — Canvas was OpenAI's answer to Claude's Artifacts in the first place, and Claude kept theirs. So if a dedicated editing canvas matters to you, that's now a point in Claude's column.)

You can edit document inside and ask questions
And Gemini? Gemini has it all. Its Canvas combines the strengths of both — Claude's side-panel editing surface and the kind of in-place document workspace ChatGPT just walked away from. Select Canvas, type your prompt, and you get a document or coding workspace alongside the chat. The highlight-and-refine part is the standout — select any chunk of text and tell it to shorten, rephrase, or rewrite without leaving the document, plus sliders to dial tone from casual to formal and length from short to long. And here's the kicker: Canvas is free for everyone — no $20 tier required to touch it.

Bottom line for this toughest category to call:
1st place🥇: Claude Opus 4.8
2nd place 🥈: Gemini 3.5 and ChatGPT 5.5
3rd place 🥉: —
5. Code generation
Here's why this category matters more than most people think: code is just automation of any digital process. And if you're not a developer, that arguably makes it more important, not less — because AI can now fill a gap you didn't even know you had. The thing you assumed required hiring someone? You might be able to just... describe it and get it.
For judging this one, lean on the benchmarks — Arena's coding board plus a few others — because coding is one of the rare creative-AI categories where the differences are actually measurable. And the data here is unusually clear: Claude dominates. On Arena's coding leaderboard, the top cluster is almost entirely Claude Opus variants — four of the top five. The coding lead has been consistent across multiple Claude releases, and the gap to GPT in Code Arena has run as wide as ~90 Elo points — which, in benchmark terms, is not a rounding error. Gemini sits in the contender pack, and ChatGPT trails on the pure-coding boards. Arena Leaderboard
But — and this is the nuance that actually matters — "coding" isn't one thing. It splits into very different jobs:
Building a website or web app — front-end, visual, "make me a landing page." This is its own arena (Arena even breaks out a WebDev board), and it's where Claude's lead is most visible.
Automating a process — scripts, data pipelines, connecting tools together. This is probably what you'd actually use it for, and it's a different skill than making things look pretty.
Agentic coding — where the AI doesn't just write the code but runs it, tests it, fixes its own errors, and iterates. (That's its own category — we'll get there.)
So when someone asks "which AI codes best," the honest answer is: best at what kind of coding? On the raw benchmarks across the board, though, the trophy's Claude's right now.
Bottom line for this category:
1st place🥇: Claude Opus 4.8
2nd place 🥈: Gemini 3.5 Flash (WebDev) and ChatGPT 5.5 (SWE-Bench)
3rd place 🥉: Gemini 3.5 Flash (SWE-Bench) and ChatGPT 5.5 (WebDev)
6. Agentic AI
First, what "agentic" actually means, because the word gets thrown around loosely. A normal chatbot answers. You ask, it replies, done. An agent acts. You hand it a goal — "go research these ten companies and build me a spreadsheet" — and it breaks that into steps, uses tools, runs code, checks its own work, fixes its mistakes, and comes back when the whole thing is done. The shift is from "tell me how" to "go do it." That's the difference that matters. We have separate amazing article about that, so be sure to checkit out >here.
Why care, especially if you're not technical? Because this is where AI stops being a smarter Google and starts being something closer to a junior employee you can delegate to. That's the leap.
Anthropic (Claude): this is their home turf. Claude's agentic products are Claude Code (the developer one — it lives in your terminal and can build, run, and debug across a whole codebase) and Claude Cowork, the version aimed at non-developers — give it a multi-step knowledge-work task and it goes off and does it. Anthropic has leaned hardest into agents of the three, and it shows: the same coding strength from the last section is what makes an agent reliable, because an agent that writes broken code just fails faster.
OpenAI (ChatGPT): strong, and arguably more polished for everyday users. Their offering is ChatGPT agent — it can browse the web, fill forms, click through sites, and complete tasks on your behalf inside a controlled environment. It's well-integrated and approachable, which for a non-technical person is a real advantage. Both Claude Cowork and ChatGPT agent are the products Google was openly chasing when it built its own. OpenAI Help Center
Google (Gemini): the agent exists — but not for you on $20. Google launched Gemini Spark at I/O 2026, a 24/7 agent that runs multi-step tasks in the background, and reviewers called it "shockingly good." Sounds great — except here's the catch for this comparison: it's rolling out only to Google AI Ultra subscribers, 18+, in the US. Ultra is $100/month — five times our price point. So on the $20 plan we're actually comparing, Spark isn't on the table. Google's deepest agent advantage (it already owns your Gmail, Calendar, Drive, and Chrome) is real, but it's locked behind the premium tier today. Wikipedia + 3
Bottom line for agents on the $20 tier: it's a two-horse race — Claude and ChatGPT.
1st place🥇: Claude Code (if you are Dev) or Codex (if you are non-Dev)
2nd place 🥈: Claude CoWork (if you are non-Dev)
3rd place 🥉: —
But — and I want to be fair to Google here — this is a matter of time. Spark already exists and reviewers say it's strong; it's just gated behind the premium tier right now. My bet is that within a month or two, we'll see Spark land at the $20 level and go head-to-head with the likes of Claude Cowork and ChatGPT's agent. Google moves fast, and they're clearly all-in on the agentic push. So treat this round as a snapshot, not a verdict — the gap here is about access and pricing, not capability.
7. Third-party tools (MCP)
Third-party tools & ecosystem — and here's where the gap is widest.
Quick context on why this matters: an AI that can only talk to you is useful, but an AI that can reach into your Gmail, your calendar, your Notion, your CRM — and actually do things there — is on a different level. The plumbing that makes this work is called MCP (Model Context Protocol — basically a universal adapter that lets an AI plug into outside apps, the same way USB-C connects to anything).
Claude is built around this. Claude and Claude Cowork connect to a huge and growing list — Gmail, Calendar, Notion, Zapier, and on and on — through MCP connectors, and there's a public directory where you just browse and connect. Anthropic having a public connector directory is specifically what competitors are measured against. This is arguably Claude's strongest practical advantage for everyday workflows. Manus
ChatGPT is right there too. Through its Apps and connectors system, ChatGPT plugs into a similarly broad set of third-party services. So on integrations, Claude and ChatGPT are genuinely competitive — both treat "connect to your other tools" as a core feature.
Gemini is the odd one out — and this is the important nuance. It's not that Gemini can't connect to outside tools — it's where and how. On the consumer $20 app, you mostly get Google's own world: Gmail, Docs, Sheets, Drive, Calendar, Maps, YouTube. The consumer Gemini app does not support custom MCP connections for personal accounts — it's a polished assistant, but it's not built for third-party developer integrations. Real third-party MCP support exists, but it lives in the developer and enterprise tooling — Antigravity, the CLI, Gemini Enterprise — not the consumer app.
But, I hope, I am sure and I want to be fair to Google here — this is a matter of time and Google will laucnh it very very soon.
So the honest summary: on the $20 consumer plan, Claude and ChatGPT let you wire your AI into the tools you already use; Gemini mostly keeps you inside Google's ecosystem.
1st place🥇: Claude
2nd place 🥈: ChatGPT
3rd place 🥉: —
8. Other features
Other perks — and this is where Google just runs away with it. This isn't close. The $20 plan for ChatGPT and Claude buys you... the AI. That's it. Good AI — but that's the whole package. Google's $20 plan (Google AI Pro) is a bundle, and once you add up what's inside, the value gap is almost unfair.
Here's what rides along with that one subscription:
Google Flow — their AI creative studio for generating cinematic video and images. You get 1,000 monthly Flow credits to make clips and scenes.
NotebookLM (premium tier) — the research tool that turns your own documents into summaries, study guides, and those AI podcast-style audio overviews. The paid access here alone is a real tool, not a toy.
5 TB of cloud storage across Drive, Gmail, and Photos — the Pro plan bundles 5 TB, which on its own is worth roughly half the subscription price.
YouTube Premium (Lite) — ad-free YouTube bundled in (varies by country).
Google Home Premium — the Advanced home/security plan, a ~$20/mo value, included free — what used to be Nest Aware.
Gemini built into Gmail, Docs, Sheets, and Slides — the AI baked right into the apps you already work in.
So on raw value-for-money, Gemini wins this round and it isn't a debate. The catch — and you already know it from the last two rounds — is that all this value only pays off if you actually live in Google's world.
So, 1st place🥇: Gemini.
10. Benchmarks — where they land on LLM Arena
Main stop: the benchmarks.
If you ever need quickly check what models are good at different areas (text, code, image, video, etc) - the one worth knowing is Arena (you'll still hear people call it LMArena or Chatbot Arena — it rebranded to "Arena" in January 2026). It's the closest thing we have to a neutral scoreboard, and the way it works is what makes it trustworthy.
It's blind testing. You type a prompt. Two anonymous models answer side by side. You don't know which is which. You pick the better one. After millions of these comparisons, you get a leaderboard built entirely on what real people actually preferred.
Why does "blind" matter so much? Because the moment you can see a model's name, you lean toward the brand you already trust. Stripping the labels off kills that bias. No marketing, no logos — just which answer was better. And it's run out of UC Berkeley, not by any of the model makers themselves, which keeps it honest.
But — and this matters — know the limits. Arena tells you which model people preferred on average, on their prompts. That's not the same as which model is best for your prompts. It narrows the field fast, but it won't tell you which one works best for what you do — only your own testing can. It also leans toward answers that feel good in a quick side-by-side, which isn't always the same as the answer that's most correct or most useful for real work. Treat it as a starting map, not the final word.

Antropic and OpenAI dominate this chart, Anthropic has Claude Cowork, OpenAI has Codex. Agentic benchmark shows how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability.

View overall rankings across various AI models in text-to-text tasks across math, coding, creative writing, and other open-ended domains.

View overall rankings across LLMs with integrated web search.

View overall rankings across AI models on front-end web development tasks, including agentic coding workflows that require multi-step reasoning and tool use.

View overall rankings across text to image AI models.

View overall rankings across text to video AI models.
So — which one?
Remember the rule we started with: this is a snapshot of mid-2026, not a verdict. Things move fast, and even the benchmarks only tell you which model people preferred on average, on their prompts — which isn't the same as which one's best for yours.
But you didn't watch this far for a shrug, so here are my actual picks:
If you just want to play with AI for fun — image, video, music, asking it stuff — go Gemini or ChatGPT. Gemini especially, if you want the most toys for your $20 (remember that perks round).
If you want to automate your work — connect your tools, hand off real tasks — I'd point you to ChatGPT (its agent, and Codex for code-flavored automation) or Claude (Claude Cowork for non-developer workflows). Both let your AI actually do things, not just describe them.
If you're a developer — it's Claude Code, full stop. The coding benchmarks back it up, and it's where Anthropic is strongest.
And that's the whole point of this video. The honest answer was never going to be a single name on a trophy. The question isn't which AI is best. It's which AI is best for you. Pick the one that fits the work you actually do — and don't be afraid to switch when the next snapshot rolls around.


