Prime Agent
A self-improving RLM harness for long-running coding and research work
About Prime Agent
Prime Agent is an open-source coding and research agent from Prime Intellect, built for work that runs longer than a chat window. Its design rests on two ideas that set it apart from the usual agent loop. The Recursive Language Model treats context as variables - a prompt is a value you can hold and pass around - and treats tools and subagents as function calls inside a persistent Python REPL. The Continual Harness stores supplemental prompts, memories, skill descriptions and reusable subagent specifications as durable state the agent can refine over time, local to the session by default.
The practical consequence is that everything is programmatic. The REPL is the model's built-in tool, so file operations, shell commands, tool use, subagent spawning and context management all happen as code rather than as a fixed menu of tool calls. rlm(...) spawns real child agents for parallel or background work and hands their results back as values. /refine reviews the current trajectory and applies small, evidence-backed updates to the harness state, never touching the immutable base system prompt, with recorded snapshots so a bad refinement can be rolled back. One warning belongs at the top because the project puts it there itself: Prime Agent executes model-generated Python and project commands with your user permissions, and its worker and kernel processes improve lifecycle isolation and recovery but are explicitly not a security sandbox. Run it against a disposable clone, a clean worktree or a restricted environment, not a checkout you cannot restore. There is a paper behind the design (arXiv 2608.23552), and the agent and TUI are built on top of pi by earendil-works, which the project credits.
Install it on macOS or Linux with a single script that downloads a versioned release, verifies its SHA-256 checksum and prepares the Python runtime. Change into the project you want worked on and run prime-agent. On first launch, /login picks a subscription or API-key provider - the agent brings no inference of its own.
From there it behaves less like a chat and more like a process you manage. Sessions are daemon-backed, so they keep running when the terminal disconnects and can be reattached later with prime-agent attach or browsed with prime-agent agents. /goal keeps an objective and its progress alive across turns until you clear it. /heartbeat and prime-agent schedule re-enter a session periodically or at a set time. /autonomous continues within configured turn, token and time budgets and can run quality gates you define, and the documentation is careful to note that a passed gate checks only what that gate verifies and that hitting a limit does not imply the task succeeded.
- •Programmatic Everything - A persistent Python REPL is the primary model tool, so file edits, shell commands, tool use and context management are written as code rather than selected from a fixed tool list
- •Built-In Subagents -
rlm(...)spawns real child agents for parallel or background work and returns their results programmatically, and running agents can message and steer one another directly - •A Harness That Improves Itself -
/refinepersists focused, reviewable lessons as supplemental prompts, memories, skill descriptions or subagent specs, with refinement history and snapshot rollback, and without rewriting the base prompt - •Executable Skills - Skills are importable Python packages, and a built-in skill creator turns a recurring workflow into a project or personal skill
- •Daemon-Backed Continuity - Active sessions, REPL state, schedules and subagents survive terminal disconnects and can be reattached
- •Long-Horizon Controls - Automatic compaction, persistent goals, heartbeats, schedules and a bounded autonomous mode with turn, token and time budgets plus user-defined quality gates
- •Headless Modes - JSON mode and RPC mode for automation and integration into other systems
Researchers running long evaluations, and engineers with tasks measured in hours rather than prompts - migrations, large refactors, work that needs to survive a closed laptop. If you already think in Python, the RLM model is the appeal: composing subagents and manipulating context as values is far more expressive than a fixed tool schema, and skills as importable packages means your workflows are real code you can test.
It is the wrong tool for several people. If you want autocomplete or a quick in-editor edit, this is much heavier than you need. If you cannot run it against a disposable checkout or inside a container, the project's own warning about the absence of a sandbox should stop you. And if you are choosing your first terminal coding agent, something more conventional will be a gentler start - the value here is concentrated in the long-running and self-refining behaviour, which only pays off on work that actually runs long.
Pricing
No pricing page found
- Free (MIT, bring your own provider)$0/mo
Checked 2026-09-30














