Sail Research logo

Sail Research

The most efficient inference for long-horizon agents

Overview

About Sail Research

Sail Research is built on the observation that agent workloads have a completely different cost curve from chat, and almost nobody prices for it. A chat request needs an answer now. An agent grinding through a codebase overnight does not, and that slack is worth money. Sail turns it into a first class control: you declare a completion window (asap, balanced or flex) describing how much latency you can tolerate, and pay materially less for the patience, with published savings ranging from 5 to 80 percent depending on the model and window.

The second half of the product is the environment. Long horizon agents need somewhere to live that outlasts a request, so Sail provides persistent VM sandboxes called Sailboxes where compute keeps running indefinitely rather than being torn down between calls. The serving stack covers open weight models including GLM-5.2, DeepSeek V4 Flash, Kimi-K2.6, Qwen3.6, gpt-oss-120b, Gemma 4 and Nemotron 3, with published per million token pricing for every window. Endpoints are drop in compatible with both OpenAI and Anthropic clients, so adopting it is a base URL change rather than a rewrite. LoRA fine tuning and reinforcement learning rollouts run on the same platform, keeping the whole long horizon loop in one place.

Pricing

Usage based, no flat monthly plan

No free tier

Usage-based pricing per million tokens with no seat or subscription fee, and $5 of free credits monthly - rates vary by model and completion window, for example GLM-5.2 runs $0.80 input and $3.00 output per million tokens on asap, dropping to $0.40 and $1.80 on flex - cached input is billed separately - enterprise volume pricing is custom

From the vendor pricing page, 2026-09-19

Details

Data from: Website•Updated: Aug 21, 2026
autonomous-agentagent-frameworkinfrastructureinferencellm-api