Voz a texto
Transcribe voz a texto online y gratis: en vivo desde tu micrófono, o sube un archivo de audio o video para una transcripción real sin conexión.
Your microphone audio isn't uploaded to Docomint. It's handled by your browser's own speech-recognition engine, which — depending on your browser — may send it to that vendor's service to transcribe.
0 words • 0 characters
What is a Voz a texto?
Two ways to turn speech into text, both entirely in your browser. Record mode is voice typing: your browser's built-in speech recognition transcribes your microphone live, word by word, as you talk — useful for quick dictation instead of hunting for keys. Upload Audio/Video mode is a real audio-to-text transcriber: it runs an offline speech-recognition model (Whisper, via WebAssembly) on an audio or video file after a one-time model download, entirely on your device — no file is ever sent to a server to be transcribed. Also known as voice to text, audio to text, or audio transcription.
How to use the Voz a texto
- Choose Record (live microphone) or Upload Audio/Video (a file)
- Pick a language, then speak or press Transcribe
- Edit the transcript directly, then copy or download it as .txt
Example
Speaking "add milk, eggs, and bread to the shopping list" fills the transcript with that exact text as you speak; uploading a short voice-memo file transcribes the whole recording the same way, into an editable transcript you can correct, copy, or download.
Why use Voz a texto
Live voice typing
Speak into your microphone and watch the transcript fill in as you talk — no typing required.
Real file transcription
Upload an MP3, WAV, M4A, WEBM, or MP4/WEBM video and get a real transcript back — no server upload.
Editable transcript
Fix a misheard word or reformat a sentence directly in the output — it's a normal text box, not read-only.
Multiple languages
Pick a language for either mode instead of being locked to English.
Copy or download
Copy the transcript or save it as a .txt file.
Nothing uploaded, either way
Record mode's audio goes only to your browser's recognition engine; Upload mode's model runs and stays entirely on your device.
What can you create?
Meeting & interview notes
Upload a recording and get a working transcript to clean up and share.
Voice typing for drafts
Draft an email, note, or document by speaking instead of typing.
Content repurposing
Turn a podcast or video's audio track into a blog post or captions draft.
Accessibility
Get a text version of spoken content for anyone who can't easily listen to it.
Voice memos to text
Transcribe a phone voice memo without re-typing it by hand.
Frequently asked questions
Where does my speech get processed?
In Record mode, your browser's own speech recognition engine handles it — in Chrome and Edge, that means your microphone audio is sent to that browser vendor's speech-recognition service to be transcribed (not to Docomint's servers, which never see it). If you need transcription that never leaves your device, use Upload Audio/Video mode instead — its offline model runs and stays entirely on your device.
Which browsers support this?
Chrome and Edge have the most reliable support for the underlying Web Speech API; Firefox and Safari support varies by version, so if the microphone prompt never appears, try Chrome.
Does this work well in a noisy room?
Accuracy drops noticeably with background noise, overlapping speech, or a distant microphone — like any speech recognition system, it works best with a clear, close, single voice.
What audio and video formats can I upload to transcribe?
MP3, WAV, M4A, and WEBM audio, plus MP4 and WEBM video (the audio track is transcribed). Anything your browser itself can decode will generally work; formats outside this list aren't tested and may not.
Is my uploaded file sent to a server to be transcribed?
No — Upload Audio/Video mode runs an offline speech-recognition model entirely in your browser via WebAssembly. The file never leaves your device to produce the transcript.
Can I edit the transcript after it's generated?
Yes — the transcript is a normal, editable text box, whether it came from live dictation or an uploaded file. Fix a misheard word, reformat it, or add to it directly, then copy or download it as .txt.
How large an audio or video file can I upload to transcribe?
Up to 50MB on the free plan, or 200MB on Pro. That covers most voice memos, meeting recordings, and short videos; a longer file will need trimming or a Pro upgrade first.
Does this translate speech into another language, or only transcribe it?
It transcribes only — the output is text in the same language that was spoken, not a translation. Pick the matching language before you record or upload for the most accurate transcript.