EMPIRICAL BENCHMARK • INDEPENDENT EVALUATION

Best Free OpenRouter AI Models for Programming

Empirical benchmark of all 22 active free models on OpenRouter (:free) across N=100 availability telemetry probes and 25 deterministic coding challenges.

INDEPENDENT BENCHMARKS FOR AI CODING TOOLS • LAST TESTED: 2026-09-19 • NEXT RUN: 2026-09-26

An empirical benchmark of all 22 active free models on OpenRouter (endpoint IDs ending in :free) evaluating two decoupled experiments: 24-Hour Availability Telemetry (N=100 Pings) and 25 Deterministic Programming Challenges. Exposing real-world rate limits, burst throttling, and why theoretical leaderboard rankings mislead free-tier developers.

Date Last Tested
2026-09-19
Next: 2026-09-26 (Weekly)
Models Evaluated
22 Models
100% of Active Free Pool
Coding Suite
25 Tasks
Deterministic Unit Tests
Availability Probe
N=100 Pings
24h Diurnal Sampling
Key Metric
Availability %
429 & 503 Drops Tracked

Availability Rate vs. Coding Pass Rate (Quadrant Analysis)

X-Axis: Tasks Passed (/25) • Y-Axis: Availability Rate (% of N=100 Probes). Top-right quadrant contains production-viable sweet spot models.

Sweet Spot (≥80% Avail, ≥18 Pass)
Throttled Geniuses
Fast Utility
Unreliable

★ THE USABLE SWEET SPOT (HIGH AVAILABILITY & HIGH ACCURACY) ⚡ THROTTLED GENIUSES (HIGH ACCURACY, FREQUENT 429/503 DROPS) FAST UTILITY (BOILERPLATE / LIGHT TASKS) UNRELIABLE / REFUSAL TIER 0% 25% 50% 75% 100% 0/25 5/25 10/25 15/25 20/25 25/25 Coding Tasks Passed (out of 25 Deterministic Challenges) → Availability Rate (% of N=100 Probes across 24h Window) → DeepSeek V4 Flash 0731 (deepseek/deepseek-v4-flash-0731:free) Probe Availability: 84% (N=100) Tasks Passed: 23/25 Median Latency: 1.84s Burst Limit: 11 reqs DeepSeek V4 Flash Qwen 3.8 27B (qwen/qwen3.8-27b:free) Probe Availability: 92% (N=100) Tasks Passed: 22/25 Median Latency: 1.42s Burst Limit: 14 reqs Qwen 3.8 27B Cohere North Mini Code (cohere/north-mini-code:free) Probe Availability: 96% (N=100) Tasks Passed: 21/25 Median Latency: 0.88s Burst Limit: 18 reqs Cohere North Code Google Gemma 4 31B (google/gemma-4-31b-it:free) Probe Availability: 88% (N=100) Tasks Passed: 21/25 Median Latency: 1.65s Burst Limit: 12 reqs Gemma 4 31B Google Gemma 4 26B A4B MoE (google/gemma-4-26b-a4b-it:free) Probe Availability: 92% (N=100) Tasks Passed: 20/25 Median Latency: 1.15s Burst Limit: 15 reqs Poolside Laguna S 2.1 (poolside/laguna-s-2.1:free) Probe Availability: 84% (N=100) Tasks Passed: 20/25 Median Latency: 1.72s Burst Limit: 10 reqs Laguna S 2.1 NVIDIA Nemotron 3.5 Lightning (nvidia/nemotron-3.5-lightning:free) Probe Availability: 88% (N=100) Tasks Passed: 18/25 Median Latency: 1.05s Burst Limit: 13 reqs Nex AGI Nex-N2.5-Pro (nex-agi/nex-n2.5-pro:free) Probe Availability: 84% (N=100) Tasks Passed: 18/25 Median Latency: 1.58s Burst Limit: 11 reqs NVIDIA Nemotron 3 Nano Omni Reasoning (nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free) Probe Availability: 80% (N=100) Tasks Passed: 20/25 Median Latency: 2.1s Burst Limit: 9 reqs Poolside Laguna XS 2.1 (poolside/laguna-xs-2.1:free) Probe Availability: 92% (N=100) Tasks Passed: 17/25 Median Latency: 0.94s Burst Limit: 16 reqs Thinking Machines Inkling (thinkingmachines/inkling:free) Probe Availability: 76% (N=100) Tasks Passed: 17/25 Median Latency: 2.3s Burst Limit: 8 reqs Z.ai GLM 5.2 (z-ai/glm-5.2:free) Probe Availability: 80% (N=100) Tasks Passed: 16/25 Median Latency: 1.75s Burst Limit: 10 reqs Nex AGI Nex-N2.5-Mini (nex-agi/nex-n2.5-mini:free) Probe Availability: 92% (N=100) Tasks Passed: 15/25 Median Latency: 0.98s Burst Limit: 16 reqs Thinking Machines Inkling Small (thinkingmachines/inkling-small:free) Probe Availability: 84% (N=100) Tasks Passed: 14/25 Median Latency: 1.25s Burst Limit: 12 reqs inclusionAI Ling 3.0 Flash VL (inclusionai/ling-3.0-flash-vl:free) Probe Availability: 80% (N=100) Tasks Passed: 13/25 Median Latency: 1.95s Burst Limit: 9 reqs Dots Studio Dots3-Note Preview (dots-studio/dots-3-note-preview:free) Probe Availability: 72% (N=100) Tasks Passed: 12/25 Median Latency: 2.45s Burst Limit: 8 reqs LiquidAI LFM2.5-2.6B (liquid/lfm-2.5-2.6b:free) Probe Availability: 100% (N=100) Tasks Passed: 11/25 Median Latency: 0.62s Burst Limit: 20 reqs LiquidAI (100% Avail, 0.6s) inclusionAI Ling 3.0 Flash Fin (inclusionai/ling-3.0-flash-fin:free) Probe Availability: 76% (N=100) Tasks Passed: 11/25 Median Latency: 2.1s Burst Limit: 8 reqs inclusionAI Ling 3.0 Flash Sante (inclusionai/ling-3.0-flash-sante:free) Probe Availability: 72% (N=100) Tasks Passed: 9/25 Median Latency: 2.2s Burst Limit: 7 reqs NVIDIA Nemotron 3 Super (nvidia/nemotron-3-super-120b-a12b:free) Probe Availability: 68% (N=100) Tasks Passed: 22/25 Median Latency: 2.85s Burst Limit: 7 reqs Nemotron Super (68%) NVIDIA Nemotron 3 Ultra (nvidia/nemotron-3-ultra-550b-a55b:free) Probe Availability: 48% (N=100) Tasks Passed: 24/25 Median Latency: 4.6s Burst Limit: 4 reqs Nemotron Ultra (48% Avail) NVIDIA Nemotron 3.5 Content Safety (nvidia/nemotron-3.5-content-safety:free) Probe Availability: 60% (N=100) Tasks Passed: 2/25 Median Latency: 0.9s Burst Limit: 6 reqs



Model ID ↕ Context Window ↕ Availability (N=100 Pings) ↕ Median Latency ↕ Tasks Passed (/25) ↕ Burst Limit (Reqs) ↕ Date Tested ↕
DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731:free 1,024k 84% (84/100) 1.84s 23/25 (92.0%) 11 reqs 2026-09-19
Qwen 3.8 27Bqwen/qwen3.8-27b:free 256k 92% (92/100) 1.42s 22/25 (88.0%) 14 reqs 2026-09-19
Cohere North Mini Codecohere/north-mini-code:free 256k 96% (96/100) 0.88s 21/25 (84.0%) 18 reqs 2026-09-19
Google Gemma 4 31Bgoogle/gemma-4-31b-it:free 256k 88% (88/100) 1.65s 21/25 (84.0%) 12 reqs 2026-09-19
Google Gemma 4 26B A4B MoEgoogle/gemma-4-26b-a4b-it:free 256k 92% (92/100) 1.15s 20/25 (80.0%) 15 reqs 2026-09-19
Poolside Laguna S 2.1poolside/laguna-s-2.1:free 256k 84% (84/100) 1.72s 20/25 (80.0%) 10 reqs 2026-09-19
NVIDIA Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning:free 1,000k 88% (88/100) 1.05s 18/25 (72.0%) 13 reqs 2026-09-19
Nex AGI Nex-N2.5-Pronex-agi/nex-n2.5-pro:free 256k 84% (84/100) 1.58s 18/25 (72.0%) 11 reqs 2026-09-19
NVIDIA Nemotron 3 Nano Omni Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 256k 80% (80/100) 2.1s 20/25 (80.0%) 9 reqs 2026-09-19
Poolside Laguna XS 2.1poolside/laguna-xs-2.1:free 256k 92% (92/100) 0.94s 17/25 (68.0%) 16 reqs 2026-09-19
Thinking Machines Inklingthinkingmachines/inkling:free 1,048k 76% (76/100) 2.3s 17/25 (68.0%) 8 reqs 2026-09-19
Z.ai GLM 5.2z-ai/glm-5.2:free 32k 80% (80/100) 1.75s 16/25 (64.0%) 10 reqs 2026-09-19
Nex AGI Nex-N2.5-Mininex-agi/nex-n2.5-mini:free 256k 92% (92/100) 0.98s 15/25 (60.0%) 16 reqs 2026-09-19
Thinking Machines Inkling Smallthinkingmachines/inkling-small:free 1,048k 84% (84/100) 1.25s 14/25 (56.0%) 12 reqs 2026-09-19
inclusionAI Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl:free 256k 80% (80/100) 1.95s 13/25 (52.0%) 9 reqs 2026-09-19
Dots Studio Dots3-Note Previewdots-studio/dots-3-note-preview:free 512k 72% (72/100) 2.45s 12/25 (48.0%) 8 reqs 2026-09-19
LiquidAI LFM2.5-2.6Bliquid/lfm-2.5-2.6b:free 64k 100% (100/100) 0.62s 11/25 (44.0%) 20 reqs 2026-09-19
inclusionAI Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin:free 256k 76% (76/100) 2.1s 11/25 (44.0%) 8 reqs 2026-09-19
inclusionAI Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-sante:free 256k 72% (72/100) 2.2s 9/25 (36.0%) 7 reqs 2026-09-19
NVIDIA Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b:free 256k 68% (68/100) 2.85s 22/25 (88.0%) 7 reqs 2026-09-19
NVIDIA Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b:free 1,000k 48% (48/100) 4.6s 24/25 (96.0%) 4 reqs 2026-09-19
NVIDIA Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety:free 128k 60% (60/100) 0.9s 2/25 (8.0%) 6 reqs 2026-09-19

The Availability Trap: Why Most AI Model Rankings Are Useless for Free Tiers

Almost every AI benchmark published online assumes a sterile, enterprise-funded environment: a tier-4 API account with generous rate limits, dedicated routing pools, and high-availability SLAs. In that hypothetical world, leaderboard scores translate directly into developer performance.

However, when building software with local AI code editors like Cursor, Continue.dev, or Cline, developers looking to avoid monthly API bills frequently route requests through OpenRouter free models (endpoint IDs ending in :free). This is where traditional benchmark rankings fall apart.

The Key Finding: On OpenRouter’s free tier, Availability Rate is the single most important metric for developer productivity. A model that achieves 96% coding accuracy on paper but fails on 52% of real-world connection attempts yields an effective productivity score of near zero. You cannot write software when your editor constantly freezes with network timeouts.

Two Decoupled Experiments, One Unified Benchmark

A critical flaw in naive benchmark reporting is conflating network availability with code reasoning IQ. If an endpoint drops a connection due to upstream server overload, scoring that drop as an algorithmic failure would distort the model’s actual coding aptitude. Conversely, testing code quality only when an endpoint happens to be alive conceals whether a developer can reliably use it during a work session.

To eliminate this arithmetic paradox, our methodology conducts two decoupled experiments:

  1. Experiment 1: 24-Hour Availability & Latency Probe (N=100 Pings):
    We sample each endpoint with 100 standardized, lightweight requests distributed across a 24-hour window (hourly batches) to capture diurnal global traffic peaks. This measures true Availability Rate (percentage of calls returning HTTP 200 with valid completions vs. HTTP 429 rate limits, HTTP 503 capacity errors, or gateway timeouts), median latency, and burst limits.
  2. Experiment 2: Deterministic Code Synthesis Suite (25 Tasks):
    Independently, models are evaluated against our 25 unit-tested coding challenges. During this evaluation, network-level transport drops are retried so that every model attempts all 25 challenges under temperature 0.0 and our verbatim zero-shot system prompt.

This decoupling explains why NVIDIA Nemotron 3 Ultra passed 24 out of 25 programming tasks when responsive, yet exhibited an alarming 48% availability rate (48 out of 100 probes succeeded) across the 24-hour observation period. It is an intellectual giant trapped behind an exhausted infrastructure pipeline.

Does OpenRouter Have Rate Limits? Understanding OpenRouter Rate Limits & Free Quotas

A primary query among developers setting up open-source coding assistants is: does OpenRouter have rate limits? The answer is an unambiguous yes. Understanding OpenRouter rate limits is essential for developers configuring AI tools like Cursor, Continue, and Cline. OpenRouter enforces strict, multi-tiered limits on free models to preserve compute infrastructure and prevent bot abuse.

Official OpenRouter Policy Citation:
“Free model endpoints (:free) on OpenRouter are subject to a standard rate limit of 20 requests per minute (RPM) and 200 requests per day (RPD), alongside upstream provider concurrency caps.”
— Source: OpenRouter Documentation: Limits (Retrieved September 19, 2026).

1. The 20 Requests Per Minute (RPM) Ceiling

On all endpoints ending in :free, OpenRouter enforces a baseline velocity cap of 20 requests per minute per user account. When you use an agentic IDE tool (such as Cline running autonomous file modifications or Cursor multi-file indexing), the editor can easily fire 15 to 30 requests in under 45 seconds. The moment your client crosses 20 RPM, OpenRouter returns an immediate HTTP 429 Too Many Requests with an active backoff header.

2. The 200 Requests Per Day (RPD) Free Model Limit

Beyond the per-minute limit, the OpenRouter free model limit caps free tier usage at 200 completions per rolling 24-hour day across all free endpoints combined. Once you hit 200 requests within a 24-hour window, all free model calls fail automatically until the rolling quota resets.

3. Burst Throttling and Concurrency Ceilings

Even if your average request rate remains well below 20 RPM, rapid consecutive calls (sub-300ms intervals) trigger burst throttling on many free endpoints. In our burst testing, smaller models like cohere/north-mini-code:free sustained 18 consecutive burst requests before throttling, whereas large-parameter models like nvidia/nemotron-3-ultra-550b-a55b:free choked after just 4 calls.

Why Free Models Fail on OpenRouter: The Three Real Failure Modes

When an API call to a free model fails inside your IDE, it almost always stems from one of three distinct failure architectures:

  1. Upstream Provider Capacity Exhaustion (HTTP 503): OpenRouter aggregates compute from third-party hosting partners (such as DeepInfra, Fireworks, Together AI, Lepton, and Cloudflare). During peak global hours, these providers prioritize paid enterprise traffic and de-prioritize free queues. Your request never reaches the neural network; it is rejected at the provider gateway with an instant HTTP 503 Service Unavailable.
  2. Account-Level OpenRouter Free Limits (HTTP 429): When you cross the 20 RPM or 200 RPD boundary, OpenRouter’s edge gateway intercepts and blocks the call prior to routing.
  3. Cloudflare Edge Challenges & Socket Hang-Ups: Shared free routing tunnels occasionally trigger Cloudflare bot mitigation challenges, dropping the connection before receiving completion tokens.

Best Free Model on OpenRouter for Programming: In-Depth Verdicts

Based on our empirical two-stage evaluation suite, here are the verified results for the top free models on OpenRouter:

1. Cohere North Mini Code (cohere/north-mini-code:free) — The #1 Overall Free Model

Cohere North Mini Code emerged as the single most dependable free model for software development. Delivering a 96% availability rate (96/100 probes) and a lightning-fast 0.88-second median latency, it passed 21 out of 25 programming tasks (84% accuracy). It adhered strictly to the zero-shot system prompt without conversational commentary, and sustained 18 consecutive burst requests before rate limits. It is our top recommendation for autocompletion, refactoring, and code explanation.

2. Qwen 3.8 27B (qwen/qwen3.8-27b:free) — The High-Accuracy Powerhouse

With 22/25 tasks passed (88%) and a 92% availability rate (92/100) at 1.42s latency, Qwen 3.8 27B provides the best balance of deep algorithmic reasoning and high availability. It succeeded on difficult challenges including LRU Cache eviction and Rotated Binary Search, making it an ideal primary model for complex coding sessions.

3. DeepSeek V4 Flash 0731 (deepseek/deepseek-v4-flash-0731:free) — The Deep Reasoning Titan

Solving 23/25 challenges (92%), DeepSeek V4 Flash posted the highest coding score in the usable tier. With a 1-million token context window, it easily handled multi-step algorithm synthesis and SQL window queries. However, high developer demand pushes its probe availability down to 84% (84/100), meaning roughly 1 in 6 requests may fail during peak US working hours.

4. Google Gemma 4 31B & 26B MoE (google/gemma-4-31b-it:free)

Google’s 2026 open models performed exceptionally well on polyglot scripting tasks (Python, JavaScript polyfills, and SQL). The 26B A4B MoE variant stood out with 92% availability and 1.15s latency, making it a reliable secondary fallback.

The “Throttled Genius” Warning: Beware of nvidia/nemotron-3-ultra-550b-a55b:free. While it scored an incredible 24/25 on coding tests, its 48% availability rate (52 of 100 probes failed) and 4.6s latency mean more than half your requests will timeout or fail with HTTP 503. It should never be used as a primary model in an interactive IDE.

Practical Developer Guide: Setting Up Resilient Fallback Chains

Because free endpoints can experience sudden capacity dips, relying on a single free model will eventually freeze your workflow. The professional solution is a cascading fallback chain. When your primary model returns HTTP 429 or 503, your client automatically redirects the request to your backup.

Here is an optimal fallback configuration for Continue.dev (add to ~/.continue/config.json):

{
  "models": [
    {
      "title": "Free Primary (Cohere North Mini Code)",
      "provider": "openrouter",
      "model": "cohere/north-mini-code:free",
      "apiKey": "YOUR_OPENROUTER_API_KEY",
      "contextLength": 256000
    },
    {
      "title": "Fallback 1 (Qwen 3.8 27B)",
      "provider": "openrouter",
      "model": "qwen/qwen3.8-27b:free",
      "apiKey": "YOUR_OPENROUTER_API_KEY",
      "contextLength": 262144
    },
    {
      "title": "Fallback 2 (Google Gemma 4 MoE)",
      "provider": "openrouter",
      "model": "google/gemma-4-26b-a4b-it:free",
      "apiKey": "YOUR_OPENROUTER_API_KEY",
      "contextLength": 262144
    }
  ],
  "tabAutocompleteModel": {
    "title": "Inline Autocomplete",
    "provider": "openrouter",
    "model": "cohere/north-mini-code:free",
    "apiKey": "YOUR_OPENROUTER_API_KEY"
  }
}

Evaluation Methodology & Published System Prompt

Our benchmark evaluates deterministic execution rather than subjective aesthetic preference:

  • Active Free Pool Enumeration: We queried the official OpenRouter Models API (GET https://openrouter.ai/api/v1/models) on September 19, 2026 at 14:00 UTC, filtered on model IDs ending in :free, and identified exactly 22 active endpoints across 9 providers. 100% of these active endpoints were evaluated.
  • 25 Deterministic Challenges: Algorithms, data structures (LRU Cache, Trie), regex parsing, log analysis, token buckets, and polyglot tasks (Python, JS, SQL).
  • Zero-Shot Single Attempt (No Retries): Each coding challenge receives 1 prompt execution. No cherry-picking, no best-of-N retries. If an endpoint outputs syntactically invalid code or fails test assertions, it receives 0.
  • Fixed Temperature: Locked at 0.0 to eliminate stochastic variability.
  • Verbatim System Prompt:
    You are an expert programming assistant. Respond ONLY with valid, clean code matching the requested function signature. Do NOT include markdown code blocks, conversational greetings, explanations, or commentary.

Limitations of This Benchmark

  • Temporal Snapshot: Free tier availability is dynamic. Measurements represent our 24-hour observation cycle on September 19, 2026. Upstream providers frequently adjust free tier compute allocations based on global demand.
  • Sample Size: 25 coding challenges provide strong signal on syntax, algorithmic reasoning, and basic system utilities, but do not capture full multi-repository software architecture.
  • Geographic Routing: Requests originated from US-East gateway clusters. Latency profiles may vary in other regions.

Download Benchmark Data & Public Repository

In accordance with our open science standards, the complete benchmark harness, 25 tasks, and raw CSV results are published with stable URLs and on GitHub:


↓ Download Raw Results (CSV)


GitHub Repository (vcj-openrouter-free-bench) ↗


Frequently Asked Questions (FAQ)

What are the best free OpenRouter AI models for programming?
The best overall free model is Cohere North Mini Code for speed and 96% availability (96/100 probes), followed closely by Qwen 3.8 27B for complex algorithmic tasks (88% pass rate, 92% availability). For 1M token context reasoning, DeepSeek V4 Flash leads in accuracy (92% pass rate), though with slightly higher throttling during peak hours (84% availability).
Does OpenRouter have rate limits on free models?
Yes. OpenRouter enforces a rate limit of 20 requests per minute (RPM) and a hard cap of 200 requests per rolling 24-hour day across all free models (openrouter.ai/docs/limits). Rapid burst requests can also trigger temporary 429 throttling after 8 to 15 calls.
Why do free models on OpenRouter return HTTP 503 or 429 errors?
HTTP 429 indicates that your account has exceeded the 20 RPM or 200 daily request limit. HTTP 503 indicates upstream capacity exhaustion: the underlying cloud provider hosting the open weights has prioritized paying customer traffic and dropped the free queue.
Can I use OpenRouter free models inside Cursor, Continue, or Cline?
Yes! All three tools support custom OpenAI-compatible endpoints with OpenRouter base URLs. To avoid being blocked by sudden rate limits, configure a cascading fallback list with at least two or three free models in sequence.

For head-to-head metrics on open weights against closed commercial models, see our gpt-oss-120b vs Gemini 3 Pro comparison.

Abdullah Zulfiqar
Abdullah Zulfiqar Founder & Technical Editor

Abdullah Zulfiqar is the founder and editor of Vibe Coder Journal, an independent publication that benchmarks AI coding tools. He verifies every figure against primary sources — official documentation, real release files and live leaderboards — rather than repeating secondary reporting. His work has corrected widely-circulated errors in Terminal-Bench scores, Ollama's official uninstall instructions and Anthropic's documented install commands. Vibe Coder Journal accepts no sponsorships or affiliate commissions.

Related AI Coding Benchmarks