Speech-to-Text.co
Convert speech to text: upload or record audio
About Speech-to-Text.co
Speech-to-Text.co is a browser-based transcription workspace that converts recorded speech into an editable transcript. You upload an audio or video file, or record a voice note with the microphone, and the file is sent securely to a remote transcription service that returns text you can correct, copy and export. It addresses a common problem with spoken material: interviews, lectures, meetings, podcasts, voicemails and WhatsApp voice notes are easy to record but hard to search, quote or skim. Nothing needs to be installed, and recordings of up to 5 minutes can be tried without an account.
The tool uses a file-based workflow rather than live dictation, so the microphone records first and transcription starts only after you submit the file. It accepts MP3, WAV, M4A, FLAC, OGG, AIFF, WMA, OPUS, AAC, MP4, MOV, AVI, MKV and WebM, can detect the spoken language. Once a transcript exists, three AI tools work on it: an assistant that answers questions grounded in the text, a summary tool, and translation into 13 target languages. The interface itself is localized in 13 languages. Pricing is freemium: a free account includes 15 minutes of transcription and 10 AI actions per day with TXT and JSON exports, while Pro costs $9/month or $84/year (about $7/month) and removes both daily caps and adds SRT, VTT, PDF and DOCX exports. A Teams plan is currently waitlist only.
- •Record or upload - Record a voice note in the workspace or choose a supported audio or video file from your phone, computer, messaging app or camera.
- •Submit for processing - Choose Transcribe to send the complete file to the remote transcription service and follow its progress.
- •Review and edit - The returned transcript opens in an editable workspace, where you correct uncertain names, numbers and specialist terms against the audio.
- •Work with the text - Ask the assistant questions, generate a summary or translate the finished transcript into another language.
- •Export - Copy the result or export it as TXT or JSON, or as SRT, VTT, PDF or DOCX on Pro.
- •Upload or record - Accepts common audio and video containers, including phone formats such as M4A, OGG and OPUS, plus a browser voice recorder for new notes.
- •Editable transcript - Results open in a workspace where text can be corrected and formatted before it is copied or exported.
- •Transcript assistant - An AI chat panel answers questions about the transcript, such as which decisions were made, which dates were mentioned or what follow-up tasks came up.
- •Summaries and translation - Condense a transcript into its main ideas, or translate it into English, German, Spanish, French, Italian, Portuguese, Russian, Chinese, Arabic, Japanese, Polish, Dutch or Vietnamese.
- •Subtitle and document exports - TXT and JSON on every account; Pro adds SRT and VTT for timed subtitles plus PDF and DOCX documents.
Speech-to-Text.co suits people who need a text version of a finished recording rather than live captions. Podcasters and interviewers can use it to find quotations, outline episodes and draft show notes. Students and researchers can turn lectures and interviews into searchable text and summaries. Teams can transcribe recorded meetings to locate decisions, owners and follow-up tasks. Video creators who need captions can export timed SRT or VTT files on Pro. Anyone who receives long voice messages or voicemails can convert them to text to read quietly and search later. Because files are processed by a remote service rather than in the browser, users handling confidential or regulated audio should check the published privacy terms and their own policies before uploading.













