Speech to Text
The AI Speech to Text feature converts spoken audio into accurate, editable text using advanced AI transcription models. Simply upload an audio file, let the AI process it, and export the generated transcript.
Step 1: Open AI Speech to Text
From the left sidebar, click AI Speech to Text under the Studio section.
This opens the transcription workspace.
Step 2: Upload Your Audio File
On the right side of the page, click the Select Audio File upload area.
Supported audio formats include:
- MP3
- WAV
- Other supported audio formats
Maximum file size: 50 MB
Once uploaded, your audio is prepared for transcription.
Step 3: Select the Language Model
The Auto-Detection model automatically identifies the spoken language and applies the appropriate transcription model, making it ideal for multilingual audio without requiring manual language selection.
Step 4: Start Transcription
After uploading your audio, click the Transcribe button at the top of the page.
The AI begins processing your audio and converting speech into text.
Processing time depends on:
- Audio duration
- Audio quality
- Number of speakers
- Server workload
Step 5: View the Transcription Output
Once processing is complete, the generated text appears in the Transcription Output section.
The transcript is fully readable and can be reviewed for accuracy.
The page also displays the total number of characters generated.
Step 6: Review and Edit
Carefully review the generated transcript.
For the best results:
- Check names and technical terms.
- Correct any punctuation if necessary.
- Verify timestamps or formatting if your workflow requires them.
Step 7: Export the Transcript
After reviewing the transcription, click Export Text.
The transcript can be downloaded for use in:
- Documents
- Subtitles
- Blog articles
- Meeting notes
- Research
- Video captions
- Content creation workflows
Best Practices
- Upload clear, high-quality audio for the best transcription accuracy.
- Minimize background noise whenever possible.
- Use recordings with a single speaker for optimal results.
- Ensure the audio is complete before starting transcription.
- Keep uploaded files within the supported size limit (50 MB).
Note: Transcription accuracy depends on audio quality, speaker clarity, background noise, accents, and recording conditions. Clear recordings generally produce the most accurate results.