What It Costs
Sources: OpenRouter • Google API Pricing
The economic difference between open-weight inference and closed proprietary API hosting cannot be evaluated with single average price tags. Gemini 3.1 Pro enforces strict tiered pricing with a sharp penalty cliff at 200,000 tokens, whereas gpt-oss-120b offers 20+ independent hosting providers with rate variance exceeding 1,100%. Use the live interactive calculator below to evaluate your exact production workload.
Unit Cost: $0.00058 per request
Unit Cost: $0.04000 per request
Chart Metric: USD per 1M Tokens (Logarithmic Axis)
Gemini 3 Pro is Gone — What Replaced It
Source: Google AI Model Archive
Developers evaluating historical benchmarks often encounter citations for gemini-3-pro-preview. As of 22 September 2026, that checkpoint is officially archived and decommissioned. Requests dispatched to the endpoint return deprecation errors.
Google’s designated successor is gemini-3.1-pro-preview. However, engineering teams must note its operational constraints:
- Preview Status: Gemini 3.1 Pro remains in Preview without commercial SLA uptime guarantees. Google explicitly reserves the right to modify backend checkpoint weights without version incrementation.
- No Free Tier: Unlike Gemini 1.5/2.0 Flash models, Gemini 3.1 Pro has no free request allowance in Google AI Studio. Every call is metered.
- Flash Tier Monopoly on Stability: As of late September 2026, the only production-stable Gemini models offered by Google Cloud are Flash-tier models (Gemini 2.5 Flash / Gemini 3 Flash). The frontier Pro tier remains experimental.
The Benchmarks, with Reasoning Effort Shown
Source: Artificial Analysis (Third-Party Telemetry)
In accordance with our standing benchmark methodology, AI coding models cannot be compared as single-number scalar entities. gpt-oss-120b features native dynamic reasoning effort. Varying reasoning tokens fundamentally alters model capability, expanding coding resolution rates by more than 400%.
| Evaluation Harness | High Effort | Low Effort | Reasoning Delta | Harness Focus & Methodology |
|---|---|---|---|---|
| Terminal-Bench Hard | 23.5% | 5.3% | +343% (4.4×) | Sandboxed Linux CLI, git conflicts, and multi-file bash execution. |
| Humanity’s Last Exam (HLE) | 19.6% | 5.9% | +232% (3.3×) | Frontier multimodal academic reasoning across multidisciplinary domains. |
| GDPval-AA | 4.8% | 0.0% | +4.8 pt | Econometric analysis and professional domain logic synthesis. |
| Coding Index | 30.4 | 21.2 | +43.4% | Composite programming metric across code editing, syntax, and unit tests. |
| GPQA Diamond | 78.2% | 67.2% | +11.0 pt | Expert PhD-level biology, physics, and chemistry multiple-choice QA. |
| τ²-Bench Telecom | 65.8% | 45.0% | +46.2% | Telecommunications protocol compliance and RFC structural parsing. |
| IFBench (Instruction Following) | 69.0% | 58.3% | +18.4% | Verifiable constraint following (word counts, JSON formats, negative constraints). |
| AA-LCR (Long Context Reasoning) | 52.0% | 46.0% | +13.0% | Needle-in-a-haystack retrieval and cross-document reasoning at 100k+ tokens. |
• Agentic Index: 3.7
• SciCode: 34.0%
• CritPt (Critical Points): 1.1% (High) vs 0.0% (Low)
• AA-Omniscience Accuracy: 21.8% (High) vs 19.8% (Low)
• Non-Hallucination Rate: 9.2% (High) vs 8.6% (Low)
• Data Visualization: 1,002 Elo
• Websites: 979 Elo
• Code Categories (Overall): 977 Elo
• UI Components: 943 Elo
• 3D Generation: 930 Elo
The Same Model, 22 Different Hosts
Source: OpenRouter Telemetry
Unlike proprietary closed models hosted solely within a single vendor’s cloud boundary, gpt-oss-120b can be hosted by any datacenter with NVIDIA H100 hardware. The OpenRouter benchmark telemetry reveals extreme divergence in latency, throughput, reliability, and algorithmic accuracy across 22 verified hosts.
Log X (Blended $/1M) vs Linear Y (TPS)
| Provider ▴▾ | Input / 1M ▴▾ | Output / 1M ▴▾ | Prompt Cache ▴▾ | TTFT Latency ▴▾ | Throughput (TPS) ▴▾ | Uptime ▴▾ | GPQA Diamond ▴▾ |
|---|---|---|---|---|---|---|---|
| AkashML Cheapest | $0.030 | $0.17 | $0.030 | 0.79s | 30 | 99.96% | 76.6% |
| CoreWeave Cheapest | $0.030 | $0.17 | $0.030 | 0.41s | 38 | 99.95% | — |
| DekaLLM Cheapest | $0.030 | $0.18 | — | 0.73s | 29 | 99.86% | — |
| DeepInfra (bf16) | $0.037 | $0.17 | — | 0.56s | 32 | 99.65% | — |
| Crusoe Low Latency | $0.050 | $0.25 | $0.050 | 0.24s | 183 | 99.56% | — |
| Mancer | $0.050 | $0.30 | — | 0.82s | 40 | 98.65% | — |
| NovitaAI | $0.050 | $0.25 | — | 1.62s | 30 | 95.63% | — |
| DigitalOcean | $0.060 | $0.42 | $0.012 | 0.63s | 37 | 99.98% | 71.0% |
| Google Vertex | $0.090 | $0.36 | — | 2.87s | 38 | 78.81% | — |
| Baseten Low Latency | $0.100 | $0.50 | $0.100 | 0.26s | 144 | 99.99% | — |
| Parasail | $0.100 | $0.75 | $0.055 | 0.59s | 109 | 99.99% | 75.9% |
| SambaNova | $0.140 | $0.95 | — | 0.69s | 227 | 97.66% | — |
| Groq | $0.150 | $0.60 | $0.075 | 0.36s | 292 | 99.98% | — |
| Nebius Token Factory | $0.150 | $0.60 | — | 0.29s | 246 | 99.4% | — |
| Amazon Bedrock | $0.150 | $0.60 | — | 0.53s | 96 | 99.68% | 74.4% |
| DeepInfra (Turbo) | $0.150 | $0.60 | — | 0.46s | 141 | 99.97% | — |
| Together Low Latency | $0.150 | $0.60 | — | 0.26s | 74 | 97.76% | — |
| Phala | $0.150 | $0.60 | — | 1.53s | 75 | 95.4% | 71.5% |
| SiliconFlow | $0.150 | $0.60 | $0.075 | 1.37s | 17 | 93.39% | — |
| MARA | $0.150 | $0.75 | — | 1.14s | 190 | 82.26% | — |
| DeepInfra (fp8) | $0.200 | $0.95 | — | 2.48s | 90 | 64.54% | — |
| Cerebras Fastest (688 tps) Low Latency | $0.350 | $0.75 | $0.350 | 0.24s | 688 | 99.97% | — |
The 200,000-Token Cliff
Source: Google Cloud Pricing Schedule
Large codebase evaluations reveal opposite pricing dynamics between Google Cloud and open-weights marketplaces.
Google enforces a severe 200,000-token pricing cliff. For small tasks (prompts ≤ 200k tokens), Gemini 3.1 Pro Standard bills at $2.00 input and $12.00 output per million. However, if your context window includes large repository dependencies or comprehensive documentation exceeding 200,000 tokens:
| Gemini 3.1 Pro Service Tier | Prompts ≤ 200k Tokens | Prompts > 200k Tokens (Cliff) | Price Escalation Rate |
|---|---|---|---|
| Standard Input | $2.00 / 1M | $4.00 / 1M | +100% (Input Price Doubles) |
| Standard Output | $12.00 / 1M | $18.00 / 1M | +50% Escalation |
| Batch / Flex Input | $1.00 / 1M | $2.00 / 1M | +100% Escalation |
| Batch / Flex Output | $6.00 / 1M | $9.00 / 1M | +50% Escalation |
| Context Cache Write | $0.20 / 1M | $0.40 / 1M | +100% + $4.50/1M/hour storage fee |
Conversely, gpt-oss-120b has no prompt-length cliff. A 120,000-token prompt on Crusoe or AkashML is billed at the exact same $0.030 or $0.050 rate as a 50-token query. Instead, gpt-oss-120b presents a provider cliff: dispatching requests to Cerebras costs $0.350 input (11.7× more than AkashML), while NovitaAI latency (1.62s) is 6.7× slower than Crusoe (0.24s). Consequently, developers using open weights must implement dynamic proxy routing rather than worrying about token counts.
Specs Side by Side
Source: Hugging Face Model Card • Google Dev
Engineering specifications contrasted directly. Parameters and architecture for Gemini 3.1 Pro are omitted by Google and marked explicitly as not disclosed.
| Architecture & System Specification | gpt-oss-120b (OpenAI) | Gemini 3.1 Pro Preview (Google) |
|---|---|---|
| Vendor / Creator | OpenAI | |
| Release Status | Production Open Weights (Released 2025-08-05) | Preview Checkpoint (Unstable / Experimental) |
| License & Weights | Open Weights (Self-hostable, weights downloadable) | Proprietary Closed Source (API-only access) |
| Total Parameters | 117 Billion | not disclosed |
| Active Parameters (per token) | 5.1 Billion (Sparse MoE) | not disclosed |
| Underlying Architecture | Mixture-of-Experts (MoE) | not disclosed |
| Native Quantization | MXFP4 (Microscaling FP4) | not disclosed |
| Hardware Deployment Requirement | Single NVIDIA H100 (80GB VRAM) | Google Cloud TPU v5e/v5p Infrastructure |
| Context Window | 131,072 Tokens (128k) | 2,000,000 Tokens (2M) |
| Maximum Output Tokens | 131,072 Tokens | 65,536 Tokens |
| Knowledge Cutoff Date | June 2024 (Over 2 years stale) | March 2026 |
| Pricing Structure | $0.03 – $0.35 in • $0.17 – $0.95 out (per 1M) | $2.00 / $12.00 (≤200k) • $4.00 / $18.00 (>200k) |
| Grounding & Live Web Access | None native (Must be implemented via client tool) | 5,000 free Google Search calls/mo, then $14/1k |
| Training Data Usage Policy | Zero external data retention on self-hosted instances | Not used to improve products on paid API tiers |
Which Should You Use?
There is no universal victor between a 117B open-weight model and a proprietary trillion-token frontier cloud service. The optimal deployment hinges on context length, cost elasticity, and knowledge staleness.
When generating 100M+ tokens monthly for unit test creation, automated PR linting, and synthetic dataset bootstrapping, Gemini’s $2.00/$12.00 pricing creates severe financial drag ($1,400/mo). Routed gpt-oss-120b delivers identical reasoning for under $25/mo. Stale 2024 knowledge is irrelevant when all necessary context is supplied directly via git diffs.
Caveat: Gemini 3.1 Pro is still in Preview. However, if your codebase requires loading 500,000+ tokens of multi-repo dependencies into active attention, gpt-oss-120b’s 131k context window cannot execute the task. Furthermore, coding against libraries released after June 2024 requires Gemini’s newer weights or Google Search grounding.
For interactive developer tools where waiting 2 seconds destroys concentration, Cerebras streams gpt-oss-120b at 688 tokens per second with 0.24s Time-to-First-Token. Groq delivers 292 tokens per second. Proprietary cloud APIs like Google Vertex (2.87s TTFT, 38 tps) cannot compete with wafer-scale inference hardware.
What We Could Not Compare
Rigorous engineering requires clearly defining the boundaries of available empirical data:
- Parameter Efficiency: Because Google does not disclose parameter counts, active mixture counts, or quantization mechanisms for Gemini 3.1 Pro, FLOPs-per-token efficiency cannot be compared.
- Post-Cutoff Coding Correctness: gpt-oss-120b has a hard knowledge cutoff of June 2024. Benchmarking both models on 2025/2026 programming frameworks (e.g., Next.js 16 or PyTorch 2.6) measures training recency rather than fundamental reasoning capacity.
- Self-Hosted Hardware Performance: Because Gemini 3.1 Pro cannot be self-hosted, local GPU memory footprints, NVLink bandwidth scaling, and cold-start weights loading could only be measured for gpt-oss-120b.
- Preview Checkpoint Drift: Gemini 3.1 Pro Preview is actively updated by Google without fixed weight hashes, whereas gpt-oss-120b weights are mathematically frozen and verifiable via SHA-256 signatures on Hugging Face.