A laptop screen showing a video call interface
Career & Applications9 min readAugust 2026

A Machine Is Interviewing You.
It Is Not Listening. It Is Reading a Transcript.

Read before your next recorded interview

Quick answer

Most automated interviews convert your speech to text and score the text. Published research finds speech recognition makes significantly more errors on non-native English, and a controlled experiment found that transcription errors themselves lower how people rate a speaker. So the fix is not losing your accent. It is giving the transcriber less to get wrong: slower delivery, full sentences, key terms said clearly and once more in plain words.

There is no face on the screen. A question appears, a countdown starts, and you talk into your own laptop camera for two minutes while nobody responds. No nod, no follow-up, no chance to read the room, because there is no room. Then it ends and you wait. Almost every candidate walks away from that experience with the same two questions. What was it actually looking for, and did any of that count?

This is now the normal way to be interviewed

Greenhouse surveyed 2,950 job seekers across the United States, the United Kingdom, Ireland, Germany and Australia. The results describe something that has become mainstream faster than the advice around it has kept up.

What the survey foundShare of candidates
Have already been through an AI interview (US respondents)63 percent
Say the use of AI was not clearly disclosed beforehand70 percent
Trust AI to assess them fairly26 percent
Have abandoned a hiring process because it used an AI interview38 percent

Greenhouse candidate experience research, published April 2026, based on 2,950 job seekers surveyed across the US, UK, Ireland, Germany and Australia.

The step nobody explains

Here is the part that changes how you should prepare. These systems do not listen to you in any meaningful sense. Your recording goes through automatic speech recognition, which turns your voice into a block of text. Almost everything that happens after that happens to the text. Content, structure, relevance, keywords, sometimes tone, are read off the transcript. The audio has already done its job by the time you are scored.

Which means the transcript is your real interview answer. Not the thing you said. The thing the software wrote down.

Where that goes wrong for a non-native speaker

Speech recognition is not equally accurate for everyone, and this is measured rather than anecdotal. A study published in npj Digital Medicine in March 2026 tested Whisper and WhisperX, two widely used speech recognition models, on native and non-native English speech and found error rates significantly higher for the non-native speakers. Peer-reviewed analysis in JASA Express Letters evaluating Whisper across a range of accents and speaker traits points the same way. If English is your second language, the machine writing your transcript is measurably more likely to get your words wrong.

The finding that matters most, and it is not the one you expect

A preregistered experiment by Kadoma, Shrivastava and Naaman in 2026 showed 207 participants talks from speakers with various accents, with either accurate or error-filled subtitles. Error-filled subtitles reduced how people rated both the speaker and the content. Here is the crucial detail: once subtitle quality was held constant, the gap between accent groups disappeared. The penalty was not coming from the accent. It was coming from the errors. People blame the speaker for the software's mistakes. So the harm is not that a system dislikes how you sound, it is that it garbles you and then everyone marks the garbled version.

You are not being judged on your accent. You are being judged on a transcript that your accent made harder to produce.

What actually helps

This reframes the whole task. You are not trying to sound British or American, which does not work and is not the problem anyway. You are trying to give the transcriber as little as possible to get wrong.

1

Slow down more than feels natural

Speed is the single biggest cause of recognition errors, and nerves plus a countdown timer push everyone faster. Deliberately slower delivery costs you nothing on content and removes a large share of the mistakes.

2

Say the important nouns twice, in different words

If your answer turns on the phrase supply chain reconciliation, the transcript surviving with that phrase intact is what the score depends on. Say it, then restate it plainly: reconciling what we ordered against what actually arrived. If one version is mangled, the other carries the meaning.

3

Speak in complete sentences

Fragments, restarts and trailing off give the model less context to correct itself with. A full sentence is easier for the system to transcribe and easier to score.

4

Front-load the answer

Give the point in the first sentence, then the example. If your recording is cut short, or the last part transcribes badly, the substance is already safely down on the page.

5

Fix the audio before you fix the answer

A wired headset with a microphone beats a laptop microphone. A quiet room with soft furnishings beats a bare kitchen. Background noise and echo raise the error rate for everyone and raise it further for accented speech.

6

Avoid filler in your own language

Switching briefly into a first-language filler word or connective is common under pressure and reliably produces nonsense in the transcript.

7

Practise out loud with a transcriber running

Record an answer, run it through any free speech-to-text tool, and read what comes out. It is the closest you can get to seeing your own interview through the system's eyes, and it shows you exactly which of your words it keeps mangling.

The scenario

The first 15 seconds of an answer about a difficult project

Avoid

So yeah, basically, I mean there was this, like, situation with the, um, the vendor onboarding thing, and it was quite complex actually, so we had to sort of, you know, work around it and, yeah, eventually it got sorted out in the end.

Better

I will describe a vendor onboarding project. Vendor onboarding means getting a new supplier approved and set up in our systems. The problem was that approvals were taking six weeks. I redesigned the checklist, and we brought it down to nine days.

The second version is slower, uses complete sentences, explains its key term in plain words, and puts the result up front. It also transcribes far more reliably, because there is almost nothing in it for the software to guess at.

What you are entitled to ask

Seventy percent of candidates in the Greenhouse survey were not told AI was being used, which is why so few people think to ask anything at all. You are allowed to. Asking the recruiter whether the interview is automated, whether a human reviews the recording, and whether adjustments are available is a reasonable question, not a difficult one. If you have a disability, a speech condition or any other reason an automated assessment disadvantages you, ask about reasonable adjustments, since employers in many countries have obligations here. The worst realistic outcome is that you learn what you are walking into.

QShould I try to hide or soften my accent?

No, and the research suggests it is aimed at the wrong target anyway. The penalty in the controlled study came from transcription errors, not from accent itself. Clarity, pace and sentence structure reduce those errors. Trying to imitate an accent that is not yours usually makes you slower, more self-conscious and less clear, which makes the transcript worse rather than better.

QDoes anyone actually watch the recording?

Sometimes, and often only for the shortlist the system produced. That is the practical problem: a human may well review you eventually, but only if the automated pass ranked you high enough to reach them. The transcript is the gate, so treat it as the thing you are optimising for.

The 10 minutes before you record

  • Plug in a headset with a microphone, and test it by recording 20 seconds and playing it back.
  • Close the window, turn off the fan, and move away from hard echoing surfaces.
  • Say your three or four key technical terms out loud once, slowly, so the first time you use them is not on the clock.
  • Decide your opening sentence for the obvious question, and make it a full sentence with the answer in it.
  • Drink water and start slower than feels right. You will speed up on your own once the timer starts.

There is a real unfairness sitting underneath this, and it is worth naming plainly. A candidate who has done the harder thing, learning to work in a second language, is more likely to be misheard by the software and then marked on the mishearing. That is not your failure and it is not something you should have to manage. Until the tools improve, though, the practical response is the same: slow, complete, plainly worded answers, said once and then said again in simpler words. That is a speaking skill, it is trainable, and it happens to be the same skill every speaking exam is testing.

Go deeper

Half of Rejected Candidates Blame a Robot.

Before the interview comes the application. The famous claim that software auto-rejects most CVs turns out to have no source at all, and the truth changes what you should fix.

Read it →

Did this help?

About the author

M

Muskan

Co-founder, WizardTeacher · Content and curriculum lead

Muskan is co-founder of WizardTeacher and leads all content and curriculum development. She researches exam formats, marking criteria, and preparation strategies to ensure every resource on WizardTeacher is grounded in what the examiners actually reward.