HEAD TO HEAD • COMPARATIVE EVALUATION

gpt-oss-120b vs Gemini 3 Pro: What You Can Actually Compare in 2026

Gemini 3 Pro is shut down. We compare gpt-oss-120b against Gemini 3.1 Pro on price, benchmarks and 22 hosting providers — with the numbers and sources.

gpt-oss-120b vs Gemini 3 Pro: What You Can Actually Compare in 2026


[Verified 22 September 2026]
[Verified 22 September 2026 Comparison Dataset]

Finding 01 • Pricing Arithmetic

The ~70× Price Chasm
66.7× in / 70.6× out
Cheapest gpt-oss-120b input is $0.030/1M vs Gemini 3.1 Pro Standard at $2.00/1M ($2.00 / $0.030 = 66.7×). Cheapest output is $0.17/1M vs Gemini at $12.00/1M ($12.00 / $0.17 = 70.6×). On a standard linear axis, gpt-oss-120b is an invisible sliver, making a logarithmic scale mandatory across all visual pricing comparisons.

Finding 02 • Test Comparability

Reasoning Effort Trumps Model Choice
4.4× Swing on Coding
Reasoning effort changes gpt-oss-120b more than switching models on complex tasks. On Terminal-Bench Hard, high reasoning resolves 23.5% vs 5.3% on low — a 4.4× delta. On Humanity’s Last Exam (HLE), high effort yields 19.6% vs 5.9% (3.3×). GDPval-AA yields 4.8% vs 0.0%. A benchmark table that doesn’t state reasoning effort is meaningless.

Finding 03 • Host Discrepancy

Same Weights, 22 Different Outcomes
5.6 Pt Spread • 40× TPS Delta
GPQA Diamond by provider: AkashML 76.6%, Bedrock EU 76.1%, Parasail 75.9%, Bedrock 74.4%, Phala 71.5%, DigitalOcean 71.0% — same weights, 5.6 points of spread. Throughput spreads 40× from SiliconFlow (17 tps) to Cerebras (688 tps). Uptime spreads from 64.54% to 100%. “Which model” is the wrong question for an open-weight model — “which model on which host” is the real one.

Finding 04 • Asymmetric Cliffs

Opposite Directional Cost Cliffs
200k Token vs 12× Routing Cliff
Gemini 3.1 Pro doubles input price ($2.00 to $4.00) and adds 50% to output ($12.00 to $18.00) above 200,000 tokens. gpt-oss-120b has no context-length cliff, but its cheapest host ($0.030) is 11.7× (~12×) cheaper than its dearest ($0.350). Your Gemini bill depends on prompt length; your gpt-oss bill depends entirely on routing.

What It Costs
Sources: OpenRouter • Google API Pricing

The economic difference between open-weight inference and closed proprietary API hosting cannot be evaluated with single average price tags. Gemini 3.1 Pro enforces strict tiered pricing with a sharp penalty cliff at 200,000 tokens, whereas gpt-oss-120b offers 20+ independent hosting providers with rate variance exceeding 1,100%. Use the live interactive calculator below to evaluate your exact production workload.






↑ 200,000-Token Pricing Cliff Triggered: Your input exceeds 200,000 tokens. Gemini 3.1 Pro input rate doubled to $4.00/1M and output rate increased by 50% to $18.00/1M.

gpt-oss-120b (Open Weights)
Provider: AkashML (Fixed Rate)
$5.80
Input: $2.40 (80M tokens) • Output: $3.40 (20M tokens)
Unit Cost: $0.00058 per request

Gemini 3.1 Pro (Preview)
Google Cloud Vertex / AI Studio
$400.00
Input: $160.00 ($2.00/1M) • Output: $240.00 ($12.00/1M)
Unit Cost: $0.04000 per request

Gemini 3.1 Pro costs 69.0× more
gpt-oss-120b saves $394.20 per month (98.6% cost reduction)

Monthly Volume: 100M Total Tokens

Frontier Price Disparity per 1M Tokens (Input vs Output)

Chart Metric: USD per 1M Tokens (Logarithmic Axis)

• Why a logarithmic scale is mandatory: On a linear vertical axis, gpt-oss-120b’s $0.030 input cost is 1.5% of Gemini’s $2.00 standard rate, rendering open-weight pricing invisible. Gemini output at $12.00 and cliff output at $18.00 dwarf all local and decentralized hosting options.

Gemini 3 Pro is Gone — What Replaced It
Source: Google AI Model Archive

Developers evaluating historical benchmarks often encounter citations for gemini-3-pro-preview. As of 22 September 2026, that checkpoint is officially archived and decommissioned. Requests dispatched to the endpoint return deprecation errors.

Google’s designated successor is gemini-3.1-pro-preview. However, engineering teams must note its operational constraints:

  • Preview Status: Gemini 3.1 Pro remains in Preview without commercial SLA uptime guarantees. Google explicitly reserves the right to modify backend checkpoint weights without version incrementation.
  • No Free Tier: Unlike Gemini 1.5/2.0 Flash models, Gemini 3.1 Pro has no free request allowance in Google AI Studio. Every call is metered.
  • Flash Tier Monopoly on Stability: As of late September 2026, the only production-stable Gemini models offered by Google Cloud are Flash-tier models (Gemini 2.5 Flash / Gemini 3 Flash). The frontier Pro tier remains experimental.

The Benchmarks, with Reasoning Effort Shown
Source: Artificial Analysis (Third-Party Telemetry)

In accordance with our standing benchmark methodology, AI coding models cannot be compared as single-number scalar entities. gpt-oss-120b features native dynamic reasoning effort. Varying reasoning tokens fundamentally alters model capability, expanding coding resolution rates by more than 400%.

gpt-oss-120b Benchmark Scores by Reasoning Effort


Third-party measurements independently recorded by Artificial Analysis. Default view enforces simultaneous display of High and Low reasoning effort to eliminate selective reporting bias.

Evaluation Harness High Effort Low Effort Reasoning Delta Harness Focus & Methodology
Terminal-Bench Hard 23.5% 5.3% +343% (4.4×) Sandboxed Linux CLI, git conflicts, and multi-file bash execution.
Humanity’s Last Exam (HLE) 19.6% 5.9% +232% (3.3×) Frontier multimodal academic reasoning across multidisciplinary domains.
GDPval-AA 4.8% 0.0% +4.8 pt Econometric analysis and professional domain logic synthesis.
Coding Index 30.4 21.2 +43.4% Composite programming metric across code editing, syntax, and unit tests.
GPQA Diamond 78.2% 67.2% +11.0 pt Expert PhD-level biology, physics, and chemistry multiple-choice QA.
τ²-Bench Telecom 65.8% 45.0% +46.2% Telecommunications protocol compliance and RFC structural parsing.
IFBench (Instruction Following) 69.0% 58.3% +18.4% Verifiable constraint following (word counts, JSON formats, negative constraints).
AA-LCR (Long Context Reasoning) 52.0% 46.0% +13.0% Needle-in-a-haystack retrieval and cross-document reasoning at 100k+ tokens.

Single-Configuration Benchmarks
Verified Third-Party Metrics (Artificial Analysis)
• Intelligence Index: 11.6
• Agentic Index: 3.7
• SciCode: 34.0%
• CritPt (Critical Points): 1.1% (High) vs 0.0% (Low)
• AA-Omniscience Accuracy: 21.8% (High) vs 19.8% (Low)
• Non-Hallucination Rate: 9.2% (High) vs 8.6% (Low)

Design Arena Elo Ratings
Frontier WebDev & Visual Code Generation
• Game Development: 1,016 Elo
• Data Visualization: 1,002 Elo
• Websites: 979 Elo
• Code Categories (Overall): 977 Elo
• UI Components: 943 Elo
• 3D Generation: 930 Elo

The Same Model, 22 Different Hosts
Source: OpenRouter Telemetry

Unlike proprietary closed models hosted solely within a single vendor’s cloud boundary, gpt-oss-120b can be hosted by any datacenter with NVIDIA H100 hardware. The OpenRouter benchmark telemetry reveals extreme divergence in latency, throughput, reliability, and algorithmic accuracy across 22 verified hosts.

Cost vs Speed Frontier: 22 Providers Benchmarked

Log X (Blended $/1M) vs Linear Y (TPS)

Pareto efficiency points highlighted in emerald. Blended pricing calculated using standardized 3:1 input-to-output token weighting (0.75 × in + 0.25 × out). Also review our free OpenRouter endpoints guide.

Availability: 99.43% with dynamic routing • 88.43% single-host

Provider ▴▾ Input / 1M ▴▾ Output / 1M ▴▾ Prompt Cache ▴▾ TTFT Latency ▴▾ Throughput (TPS) ▴▾ Uptime ▴▾ GPQA Diamond ▴▾
AkashML Cheapest $0.030 $0.17 $0.030 0.79s 30 99.96% 76.6%
CoreWeave Cheapest $0.030 $0.17 $0.030 0.41s 38 99.95% —
DekaLLM Cheapest $0.030 $0.18 — 0.73s 29 99.86% —
DeepInfra (bf16) $0.037 $0.17 — 0.56s 32 99.65% —
Crusoe Low Latency $0.050 $0.25 $0.050 0.24s 183 99.56% —
Mancer $0.050 $0.30 — 0.82s 40 98.65% —
NovitaAI $0.050 $0.25 — 1.62s 30 95.63% —
DigitalOcean $0.060 $0.42 $0.012 0.63s 37 99.98% 71.0%
Google Vertex $0.090 $0.36 — 2.87s 38 78.81% —
Baseten Low Latency $0.100 $0.50 $0.100 0.26s 144 99.99% —
Parasail $0.100 $0.75 $0.055 0.59s 109 99.99% 75.9%
SambaNova $0.140 $0.95 — 0.69s 227 97.66% —
Groq $0.150 $0.60 $0.075 0.36s 292 99.98% —
Nebius Token Factory $0.150 $0.60 — 0.29s 246 99.4% —
Amazon Bedrock $0.150 $0.60 — 0.53s 96 99.68% 74.4%
DeepInfra (Turbo) $0.150 $0.60 — 0.46s 141 99.97% —
Together Low Latency $0.150 $0.60 — 0.26s 74 97.76% —
Phala $0.150 $0.60 — 1.53s 75 95.4% 71.5%
SiliconFlow $0.150 $0.60 $0.075 1.37s 17 93.39% —
MARA $0.150 $0.75 — 1.14s 190 82.26% —
DeepInfra (fp8) $0.200 $0.95 — 2.48s 90 64.54% —
Cerebras Fastest (688 tps) Low Latency $0.350 $0.75 $0.350 0.24s 688 99.97% —

The 200,000-Token Cliff
Source: Google Cloud Pricing Schedule

Large codebase evaluations reveal opposite pricing dynamics between Google Cloud and open-weights marketplaces.

Google enforces a severe 200,000-token pricing cliff. For small tasks (prompts ≤ 200k tokens), Gemini 3.1 Pro Standard bills at $2.00 input and $12.00 output per million. However, if your context window includes large repository dependencies or comprehensive documentation exceeding 200,000 tokens:

Gemini 3.1 Pro Service Tier Prompts ≤ 200k Tokens Prompts > 200k Tokens (Cliff) Price Escalation Rate
Standard Input $2.00 / 1M $4.00 / 1M +100% (Input Price Doubles)
Standard Output $12.00 / 1M $18.00 / 1M +50% Escalation
Batch / Flex Input $1.00 / 1M $2.00 / 1M +100% Escalation
Batch / Flex Output $6.00 / 1M $9.00 / 1M +50% Escalation
Context Cache Write $0.20 / 1M $0.40 / 1M +100% + $4.50/1M/hour storage fee

Conversely, gpt-oss-120b has no prompt-length cliff. A 120,000-token prompt on Crusoe or AkashML is billed at the exact same $0.030 or $0.050 rate as a 50-token query. Instead, gpt-oss-120b presents a provider cliff: dispatching requests to Cerebras costs $0.350 input (11.7× more than AkashML), while NovitaAI latency (1.62s) is 6.7× slower than Crusoe (0.24s). Consequently, developers using open weights must implement dynamic proxy routing rather than worrying about token counts.

Specs Side by Side
Source: Hugging Face Model Card • Google Dev

Engineering specifications contrasted directly. Parameters and architecture for Gemini 3.1 Pro are omitted by Google and marked explicitly as not disclosed.

Architecture & System Specification gpt-oss-120b (OpenAI) Gemini 3.1 Pro Preview (Google)
Vendor / Creator OpenAI Google
Release Status Production Open Weights (Released 2025-08-05) Preview Checkpoint (Unstable / Experimental)
License & Weights Open Weights (Self-hostable, weights downloadable) Proprietary Closed Source (API-only access)
Total Parameters 117 Billion not disclosed
Active Parameters (per token) 5.1 Billion (Sparse MoE) not disclosed
Underlying Architecture Mixture-of-Experts (MoE) not disclosed
Native Quantization MXFP4 (Microscaling FP4) not disclosed
Hardware Deployment Requirement Single NVIDIA H100 (80GB VRAM) Google Cloud TPU v5e/v5p Infrastructure
Context Window 131,072 Tokens (128k) 2,000,000 Tokens (2M)
Maximum Output Tokens 131,072 Tokens 65,536 Tokens
Knowledge Cutoff Date June 2024 (Over 2 years stale) March 2026
Pricing Structure $0.03 – $0.35 in • $0.17 – $0.95 out (per 1M) $2.00 / $12.00 (≤200k) • $4.00 / $18.00 (>200k)
Grounding & Live Web Access None native (Must be implemented via client tool) 5,000 free Google Search calls/mo, then $14/1k
Training Data Usage Policy Zero external data retention on self-hosted instances Not used to improve products on paid API tiers

Which Should You Use?

There is no universal victor between a 117B open-weight model and a proprietary trillion-token frontier cloud service. The optimal deployment hinges on context length, cost elasticity, and knowledge staleness.

Scenario 01 • Economic Scale

CI/CD Pipelines, Batch Synthesis & Testing
Recommended: gpt-oss-120b

When generating 100M+ tokens monthly for unit test creation, automated PR linting, and synthetic dataset bootstrapping, Gemini’s $2.00/$12.00 pricing creates severe financial drag ($1,400/mo). Routed gpt-oss-120b delivers identical reasoning for under $25/mo. Stale 2024 knowledge is irrelevant when all necessary context is supplied directly via git diffs.

Scenario 02 • Large Context & Recency

Full-Repository Audits & Modern Frameworks
Recommended: Gemini 3.1 Pro (Preview)

Caveat: Gemini 3.1 Pro is still in Preview. However, if your codebase requires loading 500,000+ tokens of multi-repo dependencies into active attention, gpt-oss-120b’s 131k context window cannot execute the task. Furthermore, coding against libraries released after June 2024 requires Gemini’s newer weights or Google Search grounding.

Scenario 03 • Interactive Latency

Real-Time IDE Autocomplete & Terminal Agents
Recommended: gpt-oss-120b on Cerebras/Groq

For interactive developer tools where waiting 2 seconds destroys concentration, Cerebras streams gpt-oss-120b at 688 tokens per second with 0.24s Time-to-First-Token. Groq delivers 292 tokens per second. Proprietary cloud APIs like Google Vertex (2.87s TTFT, 38 tps) cannot compete with wafer-scale inference hardware.

What We Could Not Compare

Rigorous engineering requires clearly defining the boundaries of available empirical data:

  • Parameter Efficiency: Because Google does not disclose parameter counts, active mixture counts, or quantization mechanisms for Gemini 3.1 Pro, FLOPs-per-token efficiency cannot be compared.
  • Post-Cutoff Coding Correctness: gpt-oss-120b has a hard knowledge cutoff of June 2024. Benchmarking both models on 2025/2026 programming frameworks (e.g., Next.js 16 or PyTorch 2.6) measures training recency rather than fundamental reasoning capacity.
  • Self-Hosted Hardware Performance: Because Gemini 3.1 Pro cannot be self-hosted, local GPU memory footprints, NVLink bandwidth scaling, and cold-start weights loading could only be measured for gpt-oss-120b.
  • Preview Checkpoint Drift: Gemini 3.1 Pro Preview is actively updated by Google without fixed weight hashes, whereas gpt-oss-120b weights are mathematically frozen and verifiable via SHA-256 signatures on Hugging Face.

Frequently Asked Questions

Is Gemini 3 Pro still available to use?
No. Gemini 3 Pro (gemini-3-pro-preview) has been shut down and listed under “Previous models” on ai.google.dev as of 22 September 2026. The live comparison model is Gemini 3.1 Pro Preview.
How does gpt-oss-120b run on a single 80GB H100?
gpt-oss-120b uses native MXFP4 quantization with a sparse Mixture-of-Experts (MoE) architecture. Out of 117 billion total parameters, only 5.1 billion are active per token, enabling it to fit comfortably within the 80GB VRAM footprint of a single NVIDIA H100 GPU without multi-node tensor parallelism.
Why do different cloud providers score differently on GPQA Diamond with the same weights?
Different providers deploy different inference engines (vLLM, TensorRT-LLM, SGLang, or custom ASICs), quantization kernels (BF16 vs FP8 vs INT4 KV-cache), system prompt defaults, and temperature presets. This produces a 5.6 percentage point spread on GPQA Diamond (76.6% on AkashML vs 71.0% on DigitalOcean) on the identical weights.
What is the 200,000-token pricing cliff on Gemini 3.1 Pro?
Google charges $2.00 per 1M input tokens and $12.00 per 1M output tokens for requests up to 200,000 tokens. Once a prompt exceeds 200,000 tokens, the input price doubles to $4.00 per 1M and the output price increases by 50% to $18.00 per 1M.
Does gpt-oss-120b require high reasoning effort for programming tasks?
Yes. On Terminal-Bench Hard, gpt-oss-120b resolves 23.5% with high reasoning effort compared to only 5.3% with low reasoning effort — a 4.4× performance difference. Benchmark figures must always state reasoning effort to be comparable.
Can I run gpt-oss-120b offline without internet access?
Yes. gpt-oss-120b is released under open weights and can be downloaded from Hugging Face for air-gapped on-premise execution with zero data transmission to external APIs.

Abdullah Zulfiqar
Abdullah Zulfiqar Founder & Technical Editor

Abdullah Zulfiqar is the founder and editor of Vibe Coder Journal, an independent publication that benchmarks AI coding tools. He verifies every figure against primary sources — official documentation, real release files and live leaderboards — rather than repeating secondary reporting. His work has corrected widely-circulated errors in Terminal-Bench scores, Ollama's official uninstall instructions and Anthropic's documented install commands. Vibe Coder Journal accepts no sponsorships or affiliate commissions.

Related Head-to-Head Comparisons & Alternatives

comparison 17 min read

Qwen 2.5 vs 3.5: Benchmarks, VRAM & Real Tests (2026)

Qwen 3.5 beats Qwen 2.5 on every benchmark that matters — a 4B 3.5 model outscores the 72B 2.5 on MMLU-Pro. Compare VRAM needs, context, thinking mode and Ollama sizes.