How to · Speech and audio

Turning speech into text on your own machine

Speech-to-text is the one job where a small model on your own machine genuinely competes with anything you could pay for. The models are a fraction of the size of the ones that write, they run happily without a graphics card, and an hour of audio takes minutes rather than an afternoon.

Updated 24 Sept 2026

Most people meet transcription through a website that wants an upload, and conclude that getting words out of audio is something you rent. It is not. This is the one job where a model small enough to sit on an ordinary laptop does work you would otherwise pay for, and it does it without the recording ever leaving the room.

That last part is why it is worth the afternoon. A meeting is mostly other people's words. Uploading it makes a decision on their behalf that they were not asked about, and doing the same work locally removes the question entirely.

The steps below are the short version of a first run. The stance is the part worth arguing with: start with the app rather than the model, and take the middle-sized one. Nearly everyone who tells you local transcription is too slow started at the top of the ladder on a machine that could not carry it.

  1. 01

    Pick the file you actually care about

    Not a test clip. Use a real recording — a meeting, an interview, a voice note — because the thing you want to know is whether this works on your audio, with your accent and your background noise. Ten minutes is plenty for a first run.

    checkpoint · You have one real audio file on the machine you are going to use.

  2. 02

    Install an app rather than a model

    On a Mac, MacWhisper. On Windows or Linux, Whisper-based desktop apps and the whisper.cpp tools do the same job. You are installing software that fetches the model for you, which is why this step is minutes rather than an afternoon.

    checkpoint · The app opens and offers you a model to download.

  3. 03

    Take the middle-sized model, not the biggest

    Every one of these apps offers a ladder of sizes. The largest is slower by a wide margin and better by a narrow one, and starting there is the single most common reason people decide local transcription is too slow. Start one or two rungs down and move up only if the transcript disappoints you.

    checkpoint · The model has downloaded and the app is ready to run.

  4. 04

    Transcribe it and read the first page

    Run the file and read the opening properly rather than skimming. You are looking for whether names and jargon survived, whether punctuation is usable, and whether speaker changes are guessable from context. That tells you more than any score.

    checkpoint · You have a transcript, and you know what it got wrong.

  5. 05

    Decide what to fix before doing it again

    If the words are right and the punctuation is ragged, you are done — that is normal and quick to tidy. If names are consistently wrong, most tools let you supply a list of terms to expect. If it was simply too slow, that is the moment to look at a faster model rather than a bigger one.

    checkpoint · You know whether to change the model, change the settings, or just carry on.

What to run it with

Questions people actually ask

Do I need a graphics card?+

No, and this is the job where that answer is most solidly true. Transcription models are small, so an ordinary laptop runs them. A graphics card turns twenty minutes of waiting into two, which is a convenience rather than an entry requirement.

How accurate is it, really?+

On clear audio with one speaker, good enough that you will correct punctuation rather than words. Accuracy drops on crosstalk, background noise and strong accents, in that order. The scores on our model pages are word error rates measured on public test sets, so treat them as a ranking rather than a promise about your recordings.

Will it know who was speaking?+

Not on its own. Separating speakers is a second step, called diarisation, and some tools bundle it while most do not. If you need "who said what" rather than "what was said", check that before you pick a tool.

What about languages other than English?+

Whisper handles a long list of them and is the safe default. Several of the faster models are English-only, which is the main thing to check before switching to one. Accuracy in any language is usually lower than the English figures suggest.

Is this genuinely private?+

If the model runs on your machine, the audio does not leave it. That is the whole reason to do it this way for recordings of other people. Some transcription apps offer a cloud option and a local one in the same interface, so confirm which one you are using before the first real file.

One app, one model, one audio file. If the first transcript reads well, the rest is just volume.