,

Claude Opus 5 vs Gemini 3.1 Pro: Benchmarks, Price & Which to Pick (2026)

Claude Opus 5 vs Gemini 3.1 Pro: who wins writing, coding, cost, and multimodal. Specs, workload pricing, live side-by-side tests, and a clear pick matrix.

·

Abstract editorial illustration for Claude Opus 5 vs Gemini 3.1 Pro: Benchmarks, Price & Which to Pick (20

Comparison · August 2026

Claude Opus 5 (Anthropic API) and Gemini 3.1 Pro Preview (Google API) are both ~1M-class flagships. The split is not the window. It is price, modalities, and voice. This guide covers published API rates, three workload costs, live writing and coding snippets, and a routing matrix. Keep both if you mix careful prose with audio/video packs. Start in a multi-model AI workspace or on i10X.

Quick verdict

Pick Claude Opus 5 if: you want warmer-but-still-operational email, deeper bug-review writeups, and a geology-grade refuse on false premises.

Pick Gemini 3.1 Pro if: you need native audio and video input, cheaper Pro-class rates ($2/$12 vs $5/$25), cheaper cache, or Google-ecosystem fit at a similar ~1M window.

Best default for many teams: Gemini for multimodal volume. Opus for critique and careful prose. Do not pay Opus rates for every video-adjacent job.

Data checked: 2026-08-24. Prices and model cards change. Verify live API test and vendor pages.

1M

Claude Opus 5 context (API)

1.05M

Gemini 3.1 Pro Preview context (API)

$5 / $25

Claude Opus 5 input/output per 1M tokens (API pricing, 2026-08-24)

$2 / $12

Gemini 3.1 Pro Preview input/output per 1M tokens (API pricing, 2026-08-24)

Bar chart comparing Claude Opus 5 and Gemini 3.1 Pro on context, multimodal breadth, API cost efficiency, writing tightness, and coding micro-test
Figure 1. Where each model wins on relative axes (context, multimodal breadth, API cost efficiency, writing tightness vs filler, coding micro-test). Higher is stronger for that axis. Chart: i10X.

Persona picker

You are…

Start with

Why

Writer / marketer

Claude Opus 5 (often)

Both ran warm. Gemini added extra “hope you are having a great week” filler. Opus stayed more operational. Both captures truncated before the letter finished.

Developer / agent builder

Gemini default on cost. Opus for deep review.

Empty-list bug was a tie. Gemini is much cheaper to loop. Opus wrote a fuller crash path.

Researcher / analyst

Gemini 3.1 Pro if media is in the pack

Similar ~1M windows. Gemini’s card adds audio and video. Opus is text/image/file.

Budget / high volume

Gemini 3.1 Pro

Chat $0.008 vs $0.0175. Repo $0.208 vs $0.50. Cached agent $0.46 vs $1.05.


What we are comparing (exact versions)

This page is Claude Opus 5 versus Gemini 3.1 Pro Preview, not Claude vs the Gemini app, and not an older Gemini 3 Pro ID. Nearby Gemini pages: Grok 4.6 vs Gemini 3.1 Pro, GPT-5.5 vs Gemini 3.1 Pro, GPT-5.6 Sol vs Gemini 3.1 Pro.

Field

Claude Opus 5

Gemini 3.1 Pro

Provider

Anthropic

Google

API model

Claude Opus 5 (Anthropic API)

Gemini 3.1 Pro Preview (Google API)

Listed card name

Claude Opus 5

Google: Gemini 3.1 Pro Preview

Family / tier

Anthropic Opus flagship

Gemini Pro-class preview

App vs API note

Also in Claude.ai / Anthropic apps. This article uses the API model.

Also in Gemini app / Google AI Pro. This article uses the API preview model.

Preview IDs move. Re-check the live card before you freeze a router. For routing across many models, see AI model routing.


Spec sheet (API card, 2026-08-24)

Spec

Claude Opus 5

Gemini 3.1 Pro Preview

Context window

1,000,000 tokens

1,048,576 tokens

Input modalities (card)

text, image, file

text, image, file, audio, video

Output

text

text

Open weights

No

No

Vendor positioning (card)

Flagship for demanding reasoning, coding, long-horizon agents. Strong at end-to-end software, code review, bug finding, visual analysis.

Frontier reasoning. Enhanced software engineering, agentic reliability, efficient token use, multimodal foundation.

We did not invent max-output or reasoning-effort rows. Confirm those on live vendor pages. The modality gap is the structural story: Gemini lists audio and video. Opus does not on this card.


Pricing and real workload cost

Rates as of 2026-08-24. Verify live before you budget.

Price

Claude Opus 5

Gemini 3.1 Pro Preview

Input / 1M tokens

$5.00

$2.00

Output / 1M tokens

$25.00

$12.00

Cache read / 1M

$0.50

$0.20

Gemini is cheaper on input, output, and cache. Opus is the premium review model in this pairing.

Scenario

Assumed tokens

Est. Claude Opus 5

Est. Gemini 3.1 Pro

Chat turn

1k in + 0.5k out

$0.0175

$0.008

Repo / doc review

80k in + 4k out

$0.500

$0.208

Agent loop

200k in (50% cached) + 20k out

$1.050

$0.460

Bar chart of estimated API cost for chat, repo review, and cached agent-loop workloads for Claude Opus 5 vs Gemini 3.1 Pro
Figure 2. Estimated USD per run using published API list rates (2026-08-24). Chat is 1k in + 0.5k out. Repo review is 80k in + 4k out. Agent loop is 200k in with 50% cache hits + 20k out. Chart: i10X.

Cached agent: $1.05 vs $0.46. Gemini also gets cheaper cache reads if your stack actually hits cache. For subscription math, see AI subscription stack cost.


Performance by job (not one score)

We are not inventing vision-eval or intelligence-index numbers we do not have in this pack. Cards, modalities, prices, and three live prompts. Re-run on your workload. Method: side-by-side AI comparison.

Coding and agents

Opus is positioned for end-to-end software, code review, and bug finding. Gemini 3.1 Pro Preview is positioned for enhanced software engineering and agentic reliability. Our micro-test: both named ZeroDivisionError and guarded empty input with if not nums: return 0 (Opus used 0.0 and mentioned ValueError). Tie on correctness. Opus wrote more crash-path prose. Gemini wrote a clean short explanation. Do not crown a coding champion from one snippet. Cost still favors Gemini on the loop.

Writing and tone

Both went warm. Opus: “I hope you are doing well,” finance promised Friday, flexible Wednesday, then the capture died at the competitive slide. Gemini (sanitized): “I hope you are having a great week,” client-name slot, same Tuesday/Friday miss, Wednesday push, then the capture died at “we sti…” Extra greeting energy on Gemini is filler the source email did not contain. Opus was still warm, but closer to the ops facts we could see. Neither snippet in this pack is a finished letter. Score the voice, then re-run with a higher max-output if you need the Acme line on both.

Research, math, reasoning

No science-QA pack here. False-premise: both refused the green-cheese Moon. Opus taught Apollo/Luna geology. Gemini said a protein-mining plan is not possible, then offered to talk about Apollo or Artemis instead. Both pass. Opus more geological. Gemini more “I will not build the false plan, here is a real-mission redirect.” For publishable work, still use multi-model hallucination checks.

Multimodal and long context

Windows are a hair: 1,000,000 vs 1,048,576. Modalities are not. Gemini lists audio and video. Opus lists text, image, file. If the job is meeting recordings, product video, or mixed media rooms, Gemini is the structural pick. If the job is long text plus screenshots, both cards can take it. We did not run a vision bake-off in this pack, so we will not invent one.

Speed

No tok/s in this pack. Writeups often say Gemini streams quickly. Measure p50 in your region. Do not ship on a screenshot.

Job

Edge

Why

Hard coding / agents

Split

Micro-test tie. Opus richer review. Gemini cheaper to loop.

Everyday writing

Claude Opus 5

Less filler than Gemini on this prompt. Both captures truncated.

Long docs / multimodal

Gemini 3.1 Pro

Audio/video on the card. Similar context.

Realtime / conversational

Product-dependent

Gemini app / Workspace vs Claude.ai. Not measured here.

Cost at volume

Gemini 3.1 Pro

$2/$12 vs $5/$25, cheaper cache.

How to read benchmarks

When we lack a primary public table for this exact pair, we do not invent one. Modality lists and live snippets are the honest floor.


Side-by-side test (i10X pack, 2026-08-24)

Three prompts. Scores 1-5 on instruction following, depth, factual caution, style, and usefulness (max 25). No invented extra tests.

Test 1: Client email rewrite

Task: Keep every fact. Warmer. Under 120 words.

Claude Opus 5 (excerpt): Greeting. Q3 deck last Tuesday. Finance promised Friday, nothing received. Move next week, Wednesday preferred, happy to adjust. Stops at “the competitive slide.”

Gemini 3.1 Pro (excerpt, sanitized): Client-name slot. “Hope you are having a great week.” Same Tuesday/Friday miss, slightly softer “just yet.” Wednesday push. Stops at “we sti…” Extra cheer that was not in the source.

Edge: Opus, because the warmth is lighter and more operational. Both lose a point for truncated endings. Gemini loses more for filler.

Test 2: Empty-list average bug

Both named ZeroDivisionError. Both guard empty input before dividing. Opus added the 0/0 crash path and a ValueError option. Gemini explained the missing empty check in one beat and shipped the same class of fix. Tie on this micro-task.

Test 3: False premise (Moon cheese)

Opus (sanitized): folk joke, Apollo and Luna samples, basalt and anorthosite, no protein. Gemini: rocky body of minerals, dust, and basaltic rock. No cheese, no organics, protein mining plan is not possible. Redirect to Apollo or Artemis. Both pass. Gemini is stricter about not building the false plan. Opus teaches more rock.

Prompt type

Claude Opus 5

Gemini 3.1 Pro

Note

Client email rewrite

23/25

21/25

Both truncated. Gemini more filler.

Bug explain + minimal fix

24/25

24/25

Both catch ZeroDivisionError. Tie.

Logic + false premise

24/25

23/25

Both refuse. Opus more geology. Gemini refuses to plan on the false premise.

Total

71/75

68/75

Opus on this prose pack. Gemini on cost and modalities.

Three editorial points do not erase audio/video or a 2×-plus invoice. Route the next step.


Ecosystem and where you run them

  • Claude Opus 5: Anthropic API. Claude.ai and console. Strong in long writing and careful review.
  • Gemini 3.1 Pro: Google AI / Gemini app / Workspace adjacency. API preview access. Strength in Docs/Drive/Search worlds and multimodal inputs.
  • Both in one place: i10X lets you compare the same prompt without two browser profiles.

Pros, cons, and failure modes

Claude Opus 5

  • Pros: Less filler than Gemini on our rewrite. Fuller bug-review writeup. Stronger geology on the refuse. 1M context. Card aimed at long-horizon agents and visual analysis.
  • Cons: $5/$25. $0.50 cache. No audio/video on this card. Writing capture truncated.
  • Fails when: the job is meeting video or cheap volume, and Opus is the only default.

Gemini 3.1 Pro

  • Pros: Audio and video on the card. ~1.05M context. $2/$12. $0.20 cache. Clean coding guard. Strict “no false mining plan” refuse.
  • Cons: Warmer-than-source filler on the email. Capture truncated. Output still pricier than Grok-class $6 rates (see Grok 4.6 vs Gemini 3.1 Pro).
  • Fails when: you need Opus-like review depth and you only called Gemini, or you treat Preview as a frozen production ID without re-checking the card.

Decision guide: pick one or route both

If you need…

Choose

Audio / video / mixed media rooms

Gemini 3.1 Pro

Careful client prose with less filler

Claude Opus 5

Deep code review comments

Claude Opus 5, then cheap loop on Gemini

Pro-class volume on a budget

Gemini 3.1 Pro

Mixed week

Keep both. Gemini multimodal default. Opus critique.

Outstanding move

Stop asking which model is “best.” Ask which model is best for the next step. That is multi-model AI.


Application walkthroughs: where each model is better

1) Customer support email

Better often: Claude Opus 5 on this prompt, because Gemini piled on week-greeting energy the source did not contain. Both letters were truncated in our capture, so re-run before you freeze a voice guide. If your brand actually wants the extra polish, Gemini will feel closer to a corporate template.

2) Long PDF / research pack

Near tie on window, Gemini if media is attached. 1M vs ~1.05M will not decide most text packs. Audio, video, and cheaper cache will. Start Gemini for mixed rooms. Escalate a hard chapter to Opus if the first pass is thin.

3) Everyday Python scripting

Tie on the guard. Use Gemini for cheap iteration. Use Opus for the pull-request comment. Do not overclaim from average().

4) Screenshot, UI QA, and video

Gemini 3.1 Pro when the input is video or audio. Both cards list image. We did not run a vision eval here, so A/B your actual captures. Do not copy a Roboflow number from a different article onto this pair unless you re-pull it.

5) Output-heavy generation at API scale

Better on cost: Gemini 3.1 Pro. $0.008 vs $0.0175 chat. $0.208 vs $0.50 repo. $0.46 vs $1.05 cached agent. Put Opus behind a router.

6) False-premise and trust gates

Both pass. Opus for geology. Gemini for refusing to build the false plan. Second-model check publishable claims.


Consumer plans vs API (do not mix them up)

  • API comparison (this article): Claude Opus 5 (Anthropic API) vs Gemini 3.1 Pro Preview (Google API).
  • Consumer apps: Claude.ai vs Google AI Pro / Gemini app may hide Flash tiers, different tools, and different rate limits.

Phone-app feel is a lived week. Agent IDs are this page.


Speed notes

No tok/s in this pack. Log p50 and p95 from your API region. Preview endpoints can also move on latency without a name change. Re-measure after any card update.



Frequently asked questions

Which is better overall, Claude Opus 5 or Gemini 3.1 Pro?
Opus won our three-prompt prose pack (71/75 vs 68/75). Gemini wins cost and audio/video. Pick by job.

Which is better for coding?
Tie on the empty-list guard. Both catch ZeroDivisionError. Opus writes a richer review. Gemini is cheaper to loop. Run your repo.

Which is better for writing?
Opus on this prompt: warm without the extra week-greeting filler. Both captures truncated. A/B on brand voice with a full-length run.

Which is cheaper?
Gemini 3.1 Pro Preview at 2026-08-24 list rates: $2/$12 vs $5/$25, cache $0.20 vs $0.50.

Which has the larger context window?
Gemini 3.1 Pro Preview (1,048,576) vs Claude Opus 5 (1,000,000). Practically a tie for most text packs.

Do I need both?
If you mix meeting video with careful client prose, yes. Gemini multimodal default, Opus critique.

Are we comparing apps or API models?
API models Claude Opus 5 (Anthropic API) and Gemini 3.1 Pro Preview (Google API).

How often should I re-test?
Preview IDs especially. After card changes, monthly in production, and whenever list rates move.

Where can I run them side by side?
i10X. Method: side-by-side AI comparison.

What about hallucinations and trust?
Both refused the Moon-cheese trap. Still ground publishable claims. Multi-model hallucination checks.

Is Gemini 3.1 Pro the same as Gemini 3 Pro?
No. This page uses the 3.1 Pro Preview card. Older “Gemini 3 Pro” posts are historical.


Try both in one workspace

Compare Claude Opus 5 and Gemini 3.1 Pro on the same prompt, then route the next step to the stronger model for that job.

Start on i10X →

Multi-model AI hub · Side-by-side method · Model routing

Sources
  1. Vendor API model cards and published list pricing for Claude Opus 5 (Anthropic API) and Gemini 3.1 Pro Preview (Google API), pulled 2026-08-24. Context, modalities, and per-million rates. Verify live.
  2. Anthropic card positioning: Opus 5 as flagship for demanding reasoning, coding, and long-horizon agentic work, including code review and bug finding.
  3. Google card positioning: Gemini 3.1 Pro Preview as a frontier reasoning model with enhanced software engineering, agentic reliability, and a multimodal foundation.
  4. i10X live side-by-side runs on 2026-08-24 (client email rewrite, empty-list average bug, false-premise Moon cheese). Both writing captures truncated. Snippets sanitized for punctuation.
  5. i10X workload estimates using the published rates above: 1k+0.5k chat, 80k+4k repo review, 200k in at 50% cache + 20k out agent loop.
  6. i10X Multi-Model silo: hub, routing, side-by-side method.

Continue reading