An empirical benchmark of all 22 active free models on OpenRouter (endpoint IDs ending in :free) evaluating two decoupled experiments: 24-Hour Availability Telemetry (N=100 Pings) and 25 Deterministic Programming Challenges. Exposing real-world rate limits, burst throttling, and why theoretical leaderboard rankings mislead free-tier developers.
Availability Rate vs. Coding Pass Rate (Quadrant Analysis)
X-Axis: Tasks Passed (/25) • Y-Axis: Availability Rate (% of N=100 Probes). Top-right quadrant contains production-viable sweet spot models.
| Model ID | Context Window | Availability (N=100 Pings) | Median Latency | Tasks Passed (/25) | Burst Limit (Reqs) | Date Tested |
|---|---|---|---|---|---|---|
DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731:free |
1,024k | 84% (84/100) | 1.84s | 23/25 (92.0%) | 11 reqs | 2026-09-19 |
Qwen 3.8 27Bqwen/qwen3.8-27b:free |
256k | 92% (92/100) | 1.42s | 22/25 (88.0%) | 14 reqs | 2026-09-19 |
Cohere North Mini Codecohere/north-mini-code:free |
256k | 96% (96/100) | 0.88s | 21/25 (84.0%) | 18 reqs | 2026-09-19 |
Google Gemma 4 31Bgoogle/gemma-4-31b-it:free |
256k | 88% (88/100) | 1.65s | 21/25 (84.0%) | 12 reqs | 2026-09-19 |
Google Gemma 4 26B A4B MoEgoogle/gemma-4-26b-a4b-it:free |
256k | 92% (92/100) | 1.15s | 20/25 (80.0%) | 15 reqs | 2026-09-19 |
Poolside Laguna S 2.1poolside/laguna-s-2.1:free |
256k | 84% (84/100) | 1.72s | 20/25 (80.0%) | 10 reqs | 2026-09-19 |
NVIDIA Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning:free |
1,000k | 88% (88/100) | 1.05s | 18/25 (72.0%) | 13 reqs | 2026-09-19 |
Nex AGI Nex-N2.5-Pronex-agi/nex-n2.5-pro:free |
256k | 84% (84/100) | 1.58s | 18/25 (72.0%) | 11 reqs | 2026-09-19 |
NVIDIA Nemotron 3 Nano Omni Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free |
256k | 80% (80/100) | 2.1s | 20/25 (80.0%) | 9 reqs | 2026-09-19 |
Poolside Laguna XS 2.1poolside/laguna-xs-2.1:free |
256k | 92% (92/100) | 0.94s | 17/25 (68.0%) | 16 reqs | 2026-09-19 |
Thinking Machines Inklingthinkingmachines/inkling:free |
1,048k | 76% (76/100) | 2.3s | 17/25 (68.0%) | 8 reqs | 2026-09-19 |
Z.ai GLM 5.2z-ai/glm-5.2:free |
32k | 80% (80/100) | 1.75s | 16/25 (64.0%) | 10 reqs | 2026-09-19 |
Nex AGI Nex-N2.5-Mininex-agi/nex-n2.5-mini:free |
256k | 92% (92/100) | 0.98s | 15/25 (60.0%) | 16 reqs | 2026-09-19 |
Thinking Machines Inkling Smallthinkingmachines/inkling-small:free |
1,048k | 84% (84/100) | 1.25s | 14/25 (56.0%) | 12 reqs | 2026-09-19 |
inclusionAI Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl:free |
256k | 80% (80/100) | 1.95s | 13/25 (52.0%) | 9 reqs | 2026-09-19 |
Dots Studio Dots3-Note Previewdots-studio/dots-3-note-preview:free |
512k | 72% (72/100) | 2.45s | 12/25 (48.0%) | 8 reqs | 2026-09-19 |
LiquidAI LFM2.5-2.6Bliquid/lfm-2.5-2.6b:free |
64k | 100% (100/100) | 0.62s | 11/25 (44.0%) | 20 reqs | 2026-09-19 |
inclusionAI Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin:free |
256k | 76% (76/100) | 2.1s | 11/25 (44.0%) | 8 reqs | 2026-09-19 |
inclusionAI Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-sante:free |
256k | 72% (72/100) | 2.2s | 9/25 (36.0%) | 7 reqs | 2026-09-19 |
NVIDIA Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b:free |
256k | 68% (68/100) | 2.85s | 22/25 (88.0%) | 7 reqs | 2026-09-19 |
NVIDIA Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b:free |
1,000k | 48% (48/100) | 4.6s | 24/25 (96.0%) | 4 reqs | 2026-09-19 |
NVIDIA Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety:free |
128k | 60% (60/100) | 0.9s | 2/25 (8.0%) | 6 reqs | 2026-09-19 |
The Availability Trap: Why Most AI Model Rankings Are Useless for Free Tiers
Almost every AI benchmark published online assumes a sterile, enterprise-funded environment: a tier-4 API account with generous rate limits, dedicated routing pools, and high-availability SLAs. In that hypothetical world, leaderboard scores translate directly into developer performance.
However, when building software with local AI code editors like Cursor, Continue.dev, or Cline, developers looking to avoid monthly API bills frequently route requests through OpenRouter free models (endpoint IDs ending in :free). This is where traditional benchmark rankings fall apart.
Two Decoupled Experiments, One Unified Benchmark
A critical flaw in naive benchmark reporting is conflating network availability with code reasoning IQ. If an endpoint drops a connection due to upstream server overload, scoring that drop as an algorithmic failure would distort the model’s actual coding aptitude. Conversely, testing code quality only when an endpoint happens to be alive conceals whether a developer can reliably use it during a work session.
To eliminate this arithmetic paradox, our methodology conducts two decoupled experiments:
-
Experiment 1: 24-Hour Availability & Latency Probe (N=100 Pings):
We sample each endpoint with 100 standardized, lightweight requests distributed across a 24-hour window (hourly batches) to capture diurnal global traffic peaks. This measures true Availability Rate (percentage of calls returning HTTP 200 with valid completions vs. HTTP 429 rate limits, HTTP 503 capacity errors, or gateway timeouts), median latency, and burst limits. -
Experiment 2: Deterministic Code Synthesis Suite (25 Tasks):
Independently, models are evaluated against our 25 unit-tested coding challenges. During this evaluation, network-level transport drops are retried so that every model attempts all 25 challenges under temperature0.0and our verbatim zero-shot system prompt.
This decoupling explains why NVIDIA Nemotron 3 Ultra passed 24 out of 25 programming tasks when responsive, yet exhibited an alarming 48% availability rate (48 out of 100 probes succeeded) across the 24-hour observation period. It is an intellectual giant trapped behind an exhausted infrastructure pipeline.
Does OpenRouter Have Rate Limits? Understanding OpenRouter Rate Limits & Free Quotas
A primary query among developers setting up open-source coding assistants is: does OpenRouter have rate limits? The answer is an unambiguous yes. Understanding OpenRouter rate limits is essential for developers configuring AI tools like Cursor, Continue, and Cline. OpenRouter enforces strict, multi-tiered limits on free models to preserve compute infrastructure and prevent bot abuse.
“Free model endpoints (
:free) on OpenRouter are subject to a standard rate limit of 20 requests per minute (RPM) and 200 requests per day (RPD), alongside upstream provider concurrency caps.”— Source: OpenRouter Documentation: Limits (Retrieved September 19, 2026).
1. The 20 Requests Per Minute (RPM) Ceiling
On all endpoints ending in :free, OpenRouter enforces a baseline velocity cap of 20 requests per minute per user account. When you use an agentic IDE tool (such as Cline running autonomous file modifications or Cursor multi-file indexing), the editor can easily fire 15 to 30 requests in under 45 seconds. The moment your client crosses 20 RPM, OpenRouter returns an immediate HTTP 429 Too Many Requests with an active backoff header.
2. The 200 Requests Per Day (RPD) Free Model Limit
Beyond the per-minute limit, the OpenRouter free model limit caps free tier usage at 200 completions per rolling 24-hour day across all free endpoints combined. Once you hit 200 requests within a 24-hour window, all free model calls fail automatically until the rolling quota resets.
3. Burst Throttling and Concurrency Ceilings
Even if your average request rate remains well below 20 RPM, rapid consecutive calls (sub-300ms intervals) trigger burst throttling on many free endpoints. In our burst testing, smaller models like cohere/north-mini-code:free sustained 18 consecutive burst requests before throttling, whereas large-parameter models like nvidia/nemotron-3-ultra-550b-a55b:free choked after just 4 calls.
Why Free Models Fail on OpenRouter: The Three Real Failure Modes
When an API call to a free model fails inside your IDE, it almost always stems from one of three distinct failure architectures:
- Upstream Provider Capacity Exhaustion (HTTP 503): OpenRouter aggregates compute from third-party hosting partners (such as DeepInfra, Fireworks, Together AI, Lepton, and Cloudflare). During peak global hours, these providers prioritize paid enterprise traffic and de-prioritize free queues. Your request never reaches the neural network; it is rejected at the provider gateway with an instant
HTTP 503 Service Unavailable. - Account-Level OpenRouter Free Limits (HTTP 429): When you cross the 20 RPM or 200 RPD boundary, OpenRouter’s edge gateway intercepts and blocks the call prior to routing.
- Cloudflare Edge Challenges & Socket Hang-Ups: Shared free routing tunnels occasionally trigger Cloudflare bot mitigation challenges, dropping the connection before receiving completion tokens.
Best Free Model on OpenRouter for Programming: In-Depth Verdicts
Based on our empirical two-stage evaluation suite, here are the verified results for the top free models on OpenRouter:
1. Cohere North Mini Code (cohere/north-mini-code:free) — The #1 Overall Free Model
Cohere North Mini Code emerged as the single most dependable free model for software development. Delivering a 96% availability rate (96/100 probes) and a lightning-fast 0.88-second median latency, it passed 21 out of 25 programming tasks (84% accuracy). It adhered strictly to the zero-shot system prompt without conversational commentary, and sustained 18 consecutive burst requests before rate limits. It is our top recommendation for autocompletion, refactoring, and code explanation.
2. Qwen 3.8 27B (qwen/qwen3.8-27b:free) — The High-Accuracy Powerhouse
With 22/25 tasks passed (88%) and a 92% availability rate (92/100) at 1.42s latency, Qwen 3.8 27B provides the best balance of deep algorithmic reasoning and high availability. It succeeded on difficult challenges including LRU Cache eviction and Rotated Binary Search, making it an ideal primary model for complex coding sessions.
3. DeepSeek V4 Flash 0731 (deepseek/deepseek-v4-flash-0731:free) — The Deep Reasoning Titan
Solving 23/25 challenges (92%), DeepSeek V4 Flash posted the highest coding score in the usable tier. With a 1-million token context window, it easily handled multi-step algorithm synthesis and SQL window queries. However, high developer demand pushes its probe availability down to 84% (84/100), meaning roughly 1 in 6 requests may fail during peak US working hours.
4. Google Gemma 4 31B & 26B MoE (google/gemma-4-31b-it:free)
Google’s 2026 open models performed exceptionally well on polyglot scripting tasks (Python, JavaScript polyfills, and SQL). The 26B A4B MoE variant stood out with 92% availability and 1.15s latency, making it a reliable secondary fallback.
nvidia/nemotron-3-ultra-550b-a55b:free. While it scored an incredible 24/25 on coding tests, its 48% availability rate (52 of 100 probes failed) and 4.6s latency mean more than half your requests will timeout or fail with HTTP 503. It should never be used as a primary model in an interactive IDE.
Practical Developer Guide: Setting Up Resilient Fallback Chains
Because free endpoints can experience sudden capacity dips, relying on a single free model will eventually freeze your workflow. The professional solution is a cascading fallback chain. When your primary model returns HTTP 429 or 503, your client automatically redirects the request to your backup.
Here is an optimal fallback configuration for Continue.dev (add to ~/.continue/config.json):
{
"models": [
{
"title": "Free Primary (Cohere North Mini Code)",
"provider": "openrouter",
"model": "cohere/north-mini-code:free",
"apiKey": "YOUR_OPENROUTER_API_KEY",
"contextLength": 256000
},
{
"title": "Fallback 1 (Qwen 3.8 27B)",
"provider": "openrouter",
"model": "qwen/qwen3.8-27b:free",
"apiKey": "YOUR_OPENROUTER_API_KEY",
"contextLength": 262144
},
{
"title": "Fallback 2 (Google Gemma 4 MoE)",
"provider": "openrouter",
"model": "google/gemma-4-26b-a4b-it:free",
"apiKey": "YOUR_OPENROUTER_API_KEY",
"contextLength": 262144
}
],
"tabAutocompleteModel": {
"title": "Inline Autocomplete",
"provider": "openrouter",
"model": "cohere/north-mini-code:free",
"apiKey": "YOUR_OPENROUTER_API_KEY"
}
}
Evaluation Methodology & Published System Prompt
Our benchmark evaluates deterministic execution rather than subjective aesthetic preference:
- Active Free Pool Enumeration: We queried the official OpenRouter Models API (
GET https://openrouter.ai/api/v1/models) on September 19, 2026 at 14:00 UTC, filtered on model IDs ending in:free, and identified exactly 22 active endpoints across 9 providers. 100% of these active endpoints were evaluated. - 25 Deterministic Challenges: Algorithms, data structures (LRU Cache, Trie), regex parsing, log analysis, token buckets, and polyglot tasks (Python, JS, SQL).
- Zero-Shot Single Attempt (No Retries): Each coding challenge receives 1 prompt execution. No cherry-picking, no best-of-N retries. If an endpoint outputs syntactically invalid code or fails test assertions, it receives 0.
- Fixed Temperature: Locked at
0.0to eliminate stochastic variability. - Verbatim System Prompt:
You are an expert programming assistant. Respond ONLY with valid, clean code matching the requested function signature. Do NOT include markdown code blocks, conversational greetings, explanations, or commentary.
Limitations of This Benchmark
- Temporal Snapshot: Free tier availability is dynamic. Measurements represent our 24-hour observation cycle on September 19, 2026. Upstream providers frequently adjust free tier compute allocations based on global demand.
- Sample Size: 25 coding challenges provide strong signal on syntax, algorithmic reasoning, and basic system utilities, but do not capture full multi-repository software architecture.
- Geographic Routing: Requests originated from US-East gateway clusters. Latency profiles may vary in other regions.
Download Benchmark Data & Public Repository
In accordance with our open science standards, the complete benchmark harness, 25 tasks, and raw CSV results are published with stable URLs and on GitHub:
Frequently Asked Questions (FAQ)
What are the best free OpenRouter AI models for programming?
Does OpenRouter have rate limits on free models?
Why do free models on OpenRouter return HTTP 503 or 429 errors?
Can I use OpenRouter free models inside Cursor, Continue, or Cline?
For head-to-head metrics on open weights against closed commercial models, see our gpt-oss-120b vs Gemini 3 Pro comparison.