Glossary

Speech-to-text

Speech-to-text is technology that converts spoken audio into written text, either as someone is speaking or afterward from a recording.

In an interview, it's what turns a spoken conversation into a transcript that can be reviewed, searched, or scored.

Definition

Speech-to-text — Speech-to-text is technology that converts spoken audio into written text. Real-time systems transcribe as someone speaks, with a short delay, while other systems process a full recording afterward and return a transcript once it's complete. Accuracy depends on several factors: audio quality, background noise, accents, and how well the system handles natural speech patterns like filler words, interruptions, and overlapping talk. In an interview setting, speech-to-text is the step that makes everything downstream possible; scoring, keyword search, and transcript review all depend on an accurate written record of what was actually said, not a rough approximation of it.

Speech-to-text underlies both voice AI systems, which need it to understand a spoken response before generating one back, and interview transcription generally, where an accurate written record lets a recruiter or an automated scoring system review exactly what a candidate said without re-listening to the full recording.

Real-time transcription during a live AI interview also makes a genuine follow-up question possible, since the system needs the words on the page before it can reason about what was actually answered.

Frequently asked questions

How accurate is speech-to-text in a real interview?

Accuracy varies by system and conditions, and depends heavily on audio quality, background noise, and how clearly someone speaks. Purpose-built systems tuned for conversational speech tend to perform noticeably better than general dictation tools on natural, unscripted talk.

Does speech-to-text work in real time during a live interview?

Modern systems can transcribe with only a short delay, which is what allows a voice AI interview to understand an answer and generate a relevant follow-up question while the conversation is still happening.

What happens to filler words and pauses in a transcript?

Most speech-to-text systems capture them as spoken, including filler words like "um" and false starts, since an accurate transcript is meant to reflect what was actually said rather than a cleaned-up version of it.

Related pages

Get a transcript built from real speech

Intervieux transcribes every AI interview in real time, so the written record and the scoring behind it reflect exactly what was said.