
Sail Research
The most efficient inference for long-horizon agents
About Sail Research
Sail Research is built on the observation that agent workloads have a completely different cost curve from chat, and almost nobody prices for it. A chat request needs an answer now. An agent grinding through a codebase overnight does not, and that slack is worth money. Sail turns it into a first class control: you declare a completion window (asap, balanced or flex) describing how much latency you can tolerate, and pay materially less for the patience, with published savings ranging from 5 to 80 percent depending on the model and window.
The second half of the product is the environment. Long horizon agents need somewhere to live that outlasts a request, so Sail provides persistent VM sandboxes called Sailboxes where compute keeps running indefinitely rather than being torn down between calls. The serving stack covers open weight models including GLM-5.2, DeepSeek V4 Flash, Kimi-K2.6, Qwen3.6, gpt-oss-120b, Gemma 4 and Nemotron 3, with published per million token pricing for every window. Endpoints are drop in compatible with both OpenAI and Anthropic clients, so adopting it is a base URL change rather than a rewrite. LoRA fine tuning and reinforcement learning rollouts run on the same platform, keeping the whole long horizon loop in one place.
Point your existing OpenAI or Anthropic client at the Sail endpoint, since the Responses, Chat Completions and Messages APIs are all supported. Choose a completion window per request. Use asap when a human is waiting, balanced for background work with a soft deadline, and flex for anything that can finish whenever, which is where the deepest discounts sit; some models are offered on flex only, at a fraction of standard rates. Pick a model from the open weight catalogue, with input, cached input and output priced separately per million tokens for each window, so the cost of a workload is calculable in advance rather than discovered on the invoice. Prompt caching is implicit and based on prefix matching, so repeated agent scaffolding is charged at the cached rate without you managing anything. For agents that need to persist, launch a Sailbox, a full VM that keeps running between calls. Fine tune with LoRA or run reinforcement learning rollouts on the same stack. Billing is pay as you go with monthly free credits, and no sales conversation is required to start.
- •Completion Windows - asap, balanced and flex tiers that trade latency for cost, with published savings of 5 to 80 percent
- •Drop-in Compatible Endpoints - OpenAI and Anthropic compatible Responses, Chat Completions and Messages APIs, so switching is a base URL change
- •Sailboxes - Persistent VM sandboxes that let long horizon agents keep compute running indefinitely
- •Open Weight Catalogue - GLM-5.2, DeepSeek V4 Flash, Kimi-K2.6, Qwen3.6, gpt-oss-120b, Gemma 4 and Nemotron 3
- •Implicit Prompt Caching - Prefix matching handled automatically and billed at a separate cached input rate
- •LoRA Fine-tuning - Serve your own fine tunes on the same inference stack
- •Reinforcement Learning Rollouts - Train agents on the platform rather than exporting the loop elsewhere
- •Batch and Background Requests - Built for large workloads rather than interactive turns
- •Self-serve Pay As You Go - Monthly free credits and no sales call required to scale
Teams running agents that work for hours or days rather than answering in seconds, where inference is the dominant line item and latency is negotiable. Overnight batch analysis, automated code review across a large repository, and long running research agents are the shapes that benefit most, because flex pricing turns tolerance for delay into a direct discount. The OpenAI and Anthropic compatible endpoints make it cheap to trial, since you can point an existing codebase at it and measure. Two caveats deserve weight. The catalogue is open weight only, so any team committed to frontier Claude or GPT models cannot use it at all. And every efficiency claim here is a self reported vendor benchmark with no independent verification, in a capital intensive market where price is the product and margins compress continuously.
Pricing
Usage based, no flat monthly plan
Usage-based pricing per million tokens with no seat or subscription fee, and $5 of free credits monthly - rates vary by model and completion window, for example GLM-5.2 runs $0.80 input and $3.00 output per million tokens on asap, dropping to $0.40 and $1.80 on flex - cached input is billed separately - enterprise volume pricing is custom
From the vendor pricing page, 2026-09-19














