Definition
Speech-to-text — Speech-to-text is technology that converts spoken audio into written text. Real-time systems transcribe as someone speaks, with a short delay, while other systems process a full recording afterward and return a transcript once it's complete. Accuracy depends on several factors: audio quality, background noise, accents, and how well the system handles natural speech patterns like filler words, interruptions, and overlapping talk. In an interview setting, speech-to-text is the step that makes everything downstream possible; scoring, keyword search, and transcript review all depend on an accurate written record of what was actually said, not a rough approximation of it.
Speech-to-text underlies both voice AI systems, which need it to understand a spoken response before generating one back, and interview transcription generally, where an accurate written record lets a recruiter or an automated scoring system review exactly what a candidate said without re-listening to the full recording.
Real-time transcription during a live AI interview also makes a genuine follow-up question possible, since the system needs the words on the page before it can reason about what was actually answered.
Frequently asked questions
How accurate is speech-to-text in a real interview?
Accuracy varies by system and conditions, and depends heavily on audio quality, background noise, and how clearly someone speaks. Purpose-built systems tuned for conversational speech tend to perform noticeably better than general dictation tools on natural, unscripted talk.
Does speech-to-text work in real time during a live interview?
Modern systems can transcribe with only a short delay, which is what allows a voice AI interview to understand an answer and generate a relevant follow-up question while the conversation is still happening.
What happens to filler words and pauses in a transcript?
Most speech-to-text systems capture them as spoken, including filler words like "um" and false starts, since an accurate transcript is meant to reflect what was actually said rather than a cleaned-up version of it.
Related pages
Get a transcript built from real speech
Intervieux transcribes every AI interview in real time, so the written record and the scoring behind it reflect exactly what was said.