ب
بسام — Matrix
@bassam · ١٥ سبتمبر · ⏱ دقائق قراءة
DISCUSSION

OpenRouter Free Models in September 2026: What Actually Works

OpenRouter's free tier is genuinely useful now, but there is one thing you need to understand before building anything serious on top of it: "free" is not a permanent property of a model.

In Matrix testing, the models available through the free list were working, and MiniMax was available as a free option during the test window. The model that stood out most for coding was Poolside Laguna S 2.1. It did not feel like a throwaway free model; it showed real potential for coding-agent work, especially when the task involved reading a repository, changing multiple files, and using tools.

There was also an important discovery while we prepared this guide. OpenRouter's openrouter/free router reported 25 models in the router, while the public models API exposed 19 direct model IDs ending in `:free` at the time of review. At the same time, the previous week's router usage showed MiniMax M3 and M2.7 among the most-used choices even though a stable zero-price direct endpoint was no longer visible when we checked again.

That is the real Matrix value here. This is not just a list of free models. It is a practical answer to three questions: what is actually useful, what is it being used for, and what is the catch before you depend on it?

Verified by Matrix — 15 September 2026. Free availability, providers, pricing and routing can change after this date. Re-check the live endpoint before using a free model in production.

What was actually being used last week?

OpenRouter's weekly usage view for the free router showed these as some of the largest routed choices:

ModelWeekly routed tokensShare
MiniMax M3507.0B11.8%
MiniMax M2.7466.0B10.8%
Tencent Hy3317.2B7.4%
NVIDIA Nemotron 3 Super316.2B7.4%
NVIDIA Nemotron 3 Ultra272.1B6.3%
Others2.4T56.3%

Treat this as an adoption signal, not a quality benchmark. High token volume proves that a model is being used; it does not prove that it is the best model for your job.

Matrix coding pick: Laguna S 2.1

If your goal is a free coding agent, Laguna S 2.1 is the first model we would test from the current list.

It is designed for software engineering and agentic coding, supports tool calling, and has enough context for real repository work. In our hands-on testing it showed clear potential: it understood repository structure better than expected and handled multi-step coding work well enough to be useful beyond autocomplete.

Good uses:

  • Fixing defined bugs with repository context.
  • Reading a codebase and proposing an implementation plan.
  • Medium multi-file refactors.
  • Agent loops that use terminal or other tools.
  • Prototyping before paying for a frontier model.

The catch: the free Poolside endpoint may allow inputs and outputs to be used for model improvement depending on the current provider policy. Do not send secrets or sensitive proprietary code without checking the live privacy policy. Free capacity is also not an SLA; availability and latency can move.

Laguna XS 2.1

The smaller Laguna is better when you care more about speed and lighter coding jobs. It can work well as a free worker inside a larger pipeline. For deeper reasoning or larger repository changes, we prefer Laguna S.

Best heavy option for long-context agents: Nemotron 3 Ultra

Nemotron 3 Ultra is one of the most interesting free options for long-running agents, orchestration, research and coding tasks that need a lot of context. The free OpenRouter listing exposed a context window up to 1M tokens when we checked it.

It is not necessarily the fastest choice, but if your agent needs to keep a large amount of documentation, code or task history in context, Ultra belongs on the shortlist.

Matrix verdict: use it as the heavier brain for planning and long-context work rather than wasting it on every small request.

Nemotron 3 Super

Super is lighter than Ultra and had substantial weekly router usage. It is a strong middle ground for agent reasoning and planning when Ultra is unnecessary.

Nemotron 3.5 Lightning

Lightning solves a different problem. It is designed for faster, high-volume work. Think routing, classification, tool selection, repetitive agent actions and tasks where throughput matters more than maximum depth.

MiniMax M3 and M2.7: important, but understand what "free" means

MiniMax M3 was the largest model in the free router's previous-week usage view at roughly 507B routed tokens, with M2.7 close behind at 466B.

M3 is built for long-horizon agent work, coding, tool use and multimodal tasks. That makes its high routed usage understandable.

Matrix also tested MiniMax while it was being offered as a free choice, and it worked.

But by final review time, the direct providers visible for M3 were priced endpoints rather than a stable zero-cost endpoint.

That is not a contradiction; it is the lesson. OpenRouter can route temporary free capacity or provider-specific free availability and later change the direct endpoint situation.

So:

  • Do not hard-code the assumption that MiniMax will remain free.
  • If openrouter/free routes you to MiniMax, it can be excellent value.
  • If your application must pin MiniMax by exact model ID, check the price and provider on the same day you deploy.

Tencent Hy3: strong general agent choice

Hy3 was one of the biggest free-router choices in the previous week. It is a mixture-of-experts model aimed at reasoning, tool use and long-horizon workflows.

Matrix verdict: for a general agent that is not coding-only, Hy3 deserves to sit next to Nemotron Super on your shortlist.

Cohere North Mini Code: a focused coding worker

North Mini Code is built for agentic coding, terminal work and software-engineering tasks with a large context window.

We are not calling it better than Laguna without a controlled head-to-head test. Its value is simpler: it is clearly a coding worker, not a generic chatbot.

Use it as a free fallback in a coding pipeline, especially if you distribute jobs across multiple models.

Nex N2.5 Pro and Mini: coding with visual feedback

Nex is especially interesting for agents that work in a loop:

  1. Change code.
  2. Run the application.
  3. Open a browser or desktop interface.
  4. Inspect the result.
  5. Fix what they see.

N2.5 Pro is the better candidate for larger jobs; Mini is useful for lighter loops. If your workflow uses Playwright, browser testing or UI verification, both deserve a real test.

Thinking Machines Inkling and Inkling Small

Inkling is aimed at general reasoning, coding, tool use, RAG and agent workflows with very large context.

Use Inkling for heavier tasks and Inkling Small as the lighter worker.

The catch: free access here is tied closely to the surrounding agent ecosystem, and logging/model-improvement terms matter. Avoid private data until you have checked the current policy.

Gemma 4: free is not only about coding

Gemma 4 26B A4B

A multimodal mixture-of-experts model that fits text plus image/video understanding, structured output, function calling and general assistant jobs.

Gemma 4 31B

A larger dense multimodal model for multilingual work, documents, reasoning and some coding.

Matrix verdict: both make sense for multimodal or general-assistant work. For a coding-first agent, start with Laguna, North or Nex instead.

Ling 3.0 Flash: three specialist free models

Ling 3.0 Flash Fin

Built for finance workflows and useful for financial reasoning, analysis and research support.

Ling 3.0 Flash Sante

Focused on health/medical reasoning and evidence retrieval.

Important: a free medical model is not a doctor. Use it for research support and organization, not as a sole source for treatment or diagnosis.

Ling 3.0 Flash VL

A vision-language model for text, images, video and visual-agent tasks. This is useful when the agent needs to see, not just read.

Small and specialist models

Nemotron 3 Nano Omni

A multimodal perception/reasoning subagent for text, audio, images and video. It can make sense as a perception layer inside a larger system.

Nemotron 3.5 Content Safety

This is a guardrail model, not a normal chatbot. Use it to classify content and add a moderation/safety layer.

Liquid LFM2.5-2.6B

A very small model relative to the rest. Best for light tasks, routing and cheap experiments that do not need frontier-level reasoning.

Dots3-Note Preview

A large preview model aimed at reasoning, coding, multimodal work and long context. Because it is a preview, treat it as experimental rather than a permanent production dependency.

Direct :free model IDs visible during the review

These were the direct IDs shown at zero prompt/output price in the public models API during our review:

  • cohere/north-mini-code:free
  • dots-studio/dots-3-note-preview:free
  • google/gemma-4-26b-a4b-it:free
  • google/gemma-4-31b-it:free
  • inclusionai/ling-3.0-flash-fin:free
  • inclusionai/ling-3.0-flash-sante:free
  • inclusionai/ling-3.0-flash-vl:free
  • liquid/lfm-2.5-2.6b:free
  • nex-agi/nex-n2.5-mini:free
  • nex-agi/nex-n2.5-pro:free
  • nvidia/nemotron-3.5-content-safety:free
  • nvidia/nemotron-3.5-lightning:free
  • nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
  • nvidia/nemotron-3-super-120b-a12b:free
  • nvidia/nemotron-3-ultra-550b-a55b:free
  • poolside/laguna-s-2.1:free
  • poolside/laguna-xs-2.1:free
  • thinkingmachines/inkling:free
  • thinkingmachines/inkling-small:free

OpenRouter itself reported 25 models in the free router. The gap between the router count and the direct :free IDs matters: router capacity is dynamic and can change faster than a static catalog.

What would we choose today?

Coding agent: Laguna S 2.1 Light coding worker: Laguna XS 2.1 or North Mini Code Visual coding loop: Nex N2.5 Pro Long-context / planning: Nemotron 3 Ultra Balanced general agent: Nemotron 3 Super or Hy3 Fast/high-volume worker: Nemotron 3.5 Lightning Multimodal general work: Gemma 4 or Ling VL Safety layer: Nemotron 3.5 Content Safety Finance specialist: Ling Fin Health research assistant: Ling Sante, with human verification Do not care which model, just want free routing: openrouter/free

Five rules before using the free tier in production

  1. Never assume free is permanent. Check the endpoint and price on deployment day.
  2. Keep fallbacks. A free model can disappear or hit 429s when provider capacity changes.
  3. Do not send secrets by default. Read provider retention and training policies, especially for free endpoints.
  4. Test a real job, not only a benchmark. A model can look strong on a leaderboard and still fail on your repository or workflow.
  5. Pin the model when consistency matters. openrouter/free is excellent for experimentation, but the router can choose different models depending on capabilities and availability.

Bottom line

OpenRouter's free tier in September 2026 is no longer a toy. There are models we can realistically use for prototypes, coding agents, multimodal workflows, safety layers and specialist tasks without direct inference cost.

But the best value does not come from collecting the biggest list of model names. It comes from knowing which model does which job, what has actually been tested, and when the free offer can change underneath you.

Matrix's current coding pick is Laguna S 2.1.

For heavy agent work and long context, test Nemotron 3 Ultra.

And if you simply want a free starting point without choosing manually, use `openrouter/free` — but keep fallbacks and do not assume it will route to the same model tomorrow.

Sources

  • OpenRouter Free Models Router: https://openrouter.ai/openrouter/free
  • OpenRouter free models view: https://openrouter.ai/models?variant=free
  • Laguna S 2.1: https://openrouter.ai/poolside/laguna-s-2.1
  • MiniMax M3: https://openrouter.ai/minimax/minimax-m3
  • OpenRouter Models API: https://openrouter.ai/api/v1/models

Matrix hands-on verification: 15 September 2026.

كل الردود (0)