OpenRouter logo

OpenRouter

Unified API for accessing 500+ LLM models from 60+ providers

OverviewArticles

From the OpenRouter blog

Published by OpenRouter, not by us. Every card opens the original post.

Confidence Thresholds for Model Escalation Routing

Confidence-based escalation keeps most requests on a cheap model and sends only the ones it is unsure about to a stronger one. This guide covers forcing a numeric confidence field with structured outputs, setting the threshold from error rates on your own traffic, tuning it…

openrouter.ai

Cost vs. Quality Tradeoff Framework for Agent Models

The cheapest model that clears your quality bar is usually not the model at the top of a leaderboard. This framework sets the bar for one task, measures cost per quality point across a cheap, a mid-tier, and a frontier model on your own examples, and picks the cheapest one that…

openrouter.ai

How to Gate Pull Requests on LLM Evals in CI

A one-line prompt change can ship an agent that tells customers the wrong refund window, and nothing in a normal CI pipeline checks what the model says. This guide builds a fixed eval set for a support agent, a script that exits non-zero below a measured threshold, and a GitHub…

openrouter.ai

AI Agent Regression Testing After a Prompt or Model Change

An agent's behavior can change when you edit a prompt, swap a model, change a tool schema, or change what retrieval returns. This guide covers the locked case set, the per-case behavioral contract, and how to run the same suite against two concrete model slugs through…

openrouter.ai

Building a Golden Eval Dataset from Production Traffic

A golden eval dataset is a curated set of production inputs with reviewed expected outputs, versioned in Git and run before every deploy. This guide covers the five steps to build one from live traffic and how to run the same set against many candidate models through one API.

openrouter.ai

How to Test Tool-Calling Accuracy in AI Agents

An agent can call the wrong tool, or call the right tool with the wrong arguments. This guide covers three ways to test each failure mode, a Python harness that grades both, and how to run the same test cases against several tool-capable models through OpenRouter.

openrouter.ai

Image-to-Video AI Models Compared: Cost, Resolution, and Control

If you already have the image a video should start from, the model choice comes down to what has to happen after that frame. This post compares the Veo 3.1, Seedance, Kling, and Grok Imagine Video lines on duration, resolution, first-frame and last-frame control, generated…

openrouter.ai

Best Embedding Models in 2026

An embedding model decides what your retrieval system can find. We shortlisted the embedding models in our catalog for English RAG, multilingual retrieval, code search, text-and-image retrieval, and low-cost indexing, sent live requests to each one, and recorded their prices,…

openrouter.ai

How to Use Jev: Moderation with the Jev API in TypeScript

A method for putting Jev to work on a new problem, worked end to end on marketplace listing moderation: what stays in code, what Jev sees, how to phrase the questions, and how to turn its probabilities into publish, hold, or reject.

openrouter.ai

Batch API: half-price inference by bundling requests

Send a whole workload in one POST, collect the results within 24 hours, and typically pay half the per-token price. Across 230k+ batches that completed over our two week beta period, the median finished in 7 minutes.

openrouter.ai

Is Jev as Accurate as Frontier Models at Classification?

Claude Opus 5 leads OpenRouter's classification task ranking by spend. We sent the same 3,080 Banking77 utterances to it and to Jev 1.13 through the Decisions API. Opus scored 84.4% to Jev's 81.0%, and Jev answered in 175 ms at $0.11 per thousand requests against 2.3 seconds…

openrouter.ai

What Is Nemotron 3.5 Lightning

Nemotron 3.5 Lightning is NVIDIA's open-weight 30B mixture-of-experts model with about 3B active parameters per token, built for the high-volume execution calls in an agent run. This post covers what the architecture means, how the model compares with Nemotron 3 Ultra, what…

openrouter.ai