
Dograh
The open source Vapi alternative for voice agents
About Dograh
Dograh is a voice agent platform built around a single constraint: your call data never leaves infrastructure you control. Voice agents handle phone conversations, which in most industries means the recordings contain the most regulated material a company holds, and the dominant hosted platforms require sending all of it to a third party. Dograh is the self-hosted answer, published under BSD 2-Clause and deployable on-premises or inside your own VPC.
The architecture is deliberately modular rather than opinionated. Speech to text, the language model, text to speech and telephony are each swappable, and you can run a speech-to-speech pipeline instead if latency matters more than control over each stage. Open models like Whisper and Kokoro can run entirely on your own hardware, so a deployment can be built with no external API calls at all. A hybrid voice mode mixes pre-recorded human clips with generated speech in the same voice, which is the practical trick for making the fixed parts of a call sound human while keeping the dynamic parts generated. Support runs to more than 70 languages, and agents can be built through Model Context Protocol from Claude Code, Cursor or another IDE agent rather than only through a web builder.
Decide first where it runs, because that is the decision the product is organised around. Self-hosting is free forever under BSD 2-Clause, a managed cloud is available to try without setting up infrastructure, and private VPC deployments are arranged directly with the team. Then assemble the stack: pick a speech-to-text engine, a language model, a text-to-speech voice and a telephony provider, or collapse those into a single speech-to-speech pipeline for the lowest latency. Build the conversation flow in the visual builder, or drive it over Model Context Protocol from a coding agent if you would rather define agents in your editor. Where a call has fixed segments, such as a disclosure or a greeting, record them with a human voice and let the same voice generate the rest, so the seam is not audible. Deploy, and the audio and transcripts stay inside your boundary.
- •BSD 2-Clause and Self-hosted - Free forever on your own infrastructure, with the full source published
- •Swappable Stack - Choose your own speech-to-text, language model, text-to-speech and telephony, or run speech-to-speech
- •Data Sovereignty - Deploy on-premises or in your VPC so call audio never crosses your boundary
- •Self-hosted Models - Run Whisper, Kokoro and other open models locally rather than calling an API
- •Hybrid Voice - Mix pre-recorded human clips with generated speech in the same voice
- •Real-time Audio - Speech-to-speech pipeline built for ultra-low latency
- •Build via MCP - Define agents from Claude Code, Cursor or another IDE agent rather than only a web UI
- •70+ Languages - Multi-language support across the pipeline
- •Managed Cloud Option - Self-serve credits from $5 if you would rather not host it yourself
Teams building phone agents in regulated industries, where sending call recordings to a hosted vendor is either a compliance problem or an outright blocker. Healthcare, finance and legal are the obvious cases, but so is any company whose legal team has opinions about where customer audio lives. It suits engineers who want to choose their own models rather than accept a platform default, and who are comfortable operating infrastructure. The MCP integration makes it unusually approachable for developers already working inside a coding agent. If you want a voice agent running this afternoon with no infrastructure work at all, a fully hosted competitor will get you there faster, and that is a legitimate trade.
Pricing
Usage based, no flat monthly plan
- Self-serve pay-as-you-goNot listed
- Committed VolumeContact sales
- EnterpriseContact sales
- Open Source$0/mo
From the vendor pricing page, 2026-09-22
















