# Speech-to-text

Speech-to-text is technology that converts spoken audio into written text, either as someone is speaking or afterward from a recording.

In an interview, it's what turns a spoken conversation into a transcript that can be reviewed, searched, or scored.

**Speech-to-text** — Speech-to-text is technology that converts spoken audio into written text. Real-time systems transcribe as someone speaks, with a short delay, while other systems process a full recording afterward and return a transcript once it's complete. Accuracy depends on several factors: audio quality, background noise, accents, and how well the system handles natural speech patterns like filler words, interruptions, and overlapping talk. In an interview setting, speech-to-text is the step that makes everything downstream possible; scoring, keyword search, and transcript review all depend on an accurate written record of what was actually said, not a rough approximation of it.

Speech-to-text underlies both voice AI systems, which need it to understand a spoken response before generating one back, and interview transcription generally, where an accurate written record lets a recruiter or an automated scoring system review exactly what a candidate said without re-listening to the full recording. Real-time transcription during a live AI interview also makes a genuine follow-up question possible, since the system needs the words on the page before it can reason about what was actually answered.

## Frequently asked questions

### How accurate is speech-to-text in a real interview?

Accuracy varies by system and conditions, and depends heavily on audio quality, background noise, and how clearly someone speaks. Purpose-built systems tuned for conversational speech tend to perform noticeably better than general dictation tools on natural, unscripted talk.

### Does speech-to-text work in real time during a live interview?

Modern systems can transcribe with only a short delay, which is what allows a voice AI interview to understand an answer and generate a relevant follow-up question while the conversation is still happening.

### What happens to filler words and pauses in a transcript?

Most speech-to-text systems capture them as spoken, including filler words like "um" and false starts, since an accurate transcript is meant to reflect what was actually said rather than a cleaned-up version of it.

## Related pages

- [Voice AI](/glossary/voice-ai)
- [Interview transcript](/glossary/interview-transcript)
- [Browse open roles](/jobs)

## Get a transcript built from real speech

Intervieux transcribes every AI interview in real time, so the written record and the scoring behind it reflect exactly what was said.

Start practicing free: https://www.intervieux.ai/register · Hire with Intervieux: https://www.intervieux.ai/employers/signup
