How to · Speech and audio

Build a second brain from the things you say out loud

Catching what was said is the easy half, now that a good transcription model runs on a laptop. Two things go wrong anyway: people choose by reputation, when what decides the transcript is how a model copes with a room full of voices, and almost nobody gets as far as asking a question of the folder afterwards. This guide does both.

Updated 24 Sept 2026

The phrase "second brain" covers two different ambitions. The modest one is a searchable pile of everything you said and heard: meetings, voice memos, the idea you dictated at a bus stop. The ambitious one links those notes together so that next March, when a client asks what was agreed, the answer surfaces itself. You can reach the modest version in an afternoon. The ambitious one is a habit you keep up for months, and you build it yourself: no model on this site will do it for you.

What models genuinely solve is the first three stages: turning speech into text, turning long text into short text, and answering a question across a folder of notes. All three run on your own machine, which matters more here than usual, because a meeting recording is mostly other people's words and transcribing it at home means the audio never leaves the room it was recorded in.

Start with the transcription model, because everything downstream inherits its mistakes. A meeting is the hard case for this kind of model: several voices, some of them talking across each other, and a microphone in the middle of a table instead of in front of a mouth. Models that read a clean studio recording almost perfectly lose several times as many words in that room, so the accuracy figure you see quoted in a headline is measured on the wrong thing. The board to judge on is the recorded-meetings board, which scores every speech model we track on real recordings made in a room, lowest error at the top.

Read it before you pick, because the ordering surprises people. The famous name sits a long way down it. The rows at the very top are closed models, marked so in the board's own column: there is nothing to download, and for the one in front we list no host to rent from either. The two best-placed models you can download are Granite Speech from IBM and Cohere Transcribe, both about the size an ordinary laptop swallows without complaint, both free to download, and both a long way above Whisper on this particular audio. The cards below name each model and link to its page, where the download size and the fit verdict for your machine stay live.

What the famous name still buys is the software around it. Whisper has free apps on every platform, it covers far more languages than either model above it, and you can have a transcript from it in twenty minutes without opening a terminal. The two named above are recent releases, so for now they are scripts and command lines. If a first transcript tonight matters more than the last word of accuracy, start with Whisper and move when the errors begin to annoy you. If this is going to be a nightly habit, start with the best-placed downloadable model on that board.

Speed is the other axis, and it stops mattering quickly. The speed board ranks the same models on how much faster than real time they run, and every model named here clears an hour of audio in far less than an hour. That gap only decides anything if you are transcribing live or clearing a backlog of years. The audio tab of the catalogue gathers every transcription model we list with its speed, languages and licence in one place. Read its accuracy column as a separate question, though: it averages all the English recordings the leaderboard tests, of which a meeting room is only one, so models that separate in a room sit almost level there. For the room itself, stay on the meetings board.

The last stage is the one everybody skips, and it is the one that makes the other four worth doing, so here is how it actually goes. You do not point a model at the folder. You search the folder, the way you search anything, for the word you half remember. Then you paste the two or three notes that came back into the model you already run and ask your question against those. That is retrieval done by hand, it needs nothing you have not already installed, and it beats the automatic versions for a long time because you can see precisely what the model was given before it answered.

Arithmetic is what makes it work this way. A month of meetings runs well past any context window you have, and handing a model more than it can hold makes it find less. Every tool that claims to answer questions across your documents is doing this same search on your behalf and hiding the result, which is fine when it picks the right note and confusing when it picks the wrong week.

When the folder outgrows hand-searching, and it will somewhere in the low hundreds of notes, the thing you want is a local tool that keeps an index of the folder for you. AnythingLLM is built for exactly this and runs entirely on your machine; Open WebUI, if you already use it as a front end for Ollama, will take a document set too. Insist on one thing from whichever you choose: it must show you which note an answer came from. These tools find notes by matching the shape of your words against the shape of theirs, and that is confidently wrong often enough that an answer about your own meetings is only worth having when you can open the note behind it.

We have no guide of our own for that indexing step yet. The gap is ours, and it is on our list.

Here is the whole pipeline, in five stages, each closing on a checkpoint you can verify before moving on.

  1. 01

    Capture

    Record everything you might want back — your phone's voice memos, a meeting recorder, call audio. If other people are in the recording, tell them.

    checkpoint · You can play the file back locally.

  2. 02

    Transcribe

    This is the choice the whole thing rests on. A meeting is the hard case: several voices, some of them talking across each other, and a microphone sitting in the middle of a table instead of in front of a mouth. Pick off the recorded-meetings board, which scores models on exactly that audio, and take the best-placed one you can download.

    checkpoint · The transcript reads correctly at a hard moment — two people talking at once.

  3. 03

    Distil

    A raw transcript is unreadable next week. A small text model turns it into a summary with decisions and open questions pulled out. This runs fine locally, on the same machine that transcribed.

    checkpoint · Yesterday's meeting fits on one screen and names its decisions.

  4. 04

    File it where search works

    Plain text or Markdown in one folder beats every app. The folder is the second brain; the tools on top of it come and go.

    checkpoint · A text search for a phrase you said last week finds the note.

  5. 05

    Ask it questions

    Search first, then ask. Find the two or three notes that mention pricing with an ordinary text search, paste those notes into the model from stage three, and put your question to it against them. That is the whole method, it works from the first week, and you can see exactly what the model was given. Once the folder outgrows hand-searching, a local tool that indexes it for you does the same search automatically, and the one thing to demand of it is that every answer names the note it came from.

    checkpoint · A question about an old meeting comes back with the right note named beside the answer.

What to run it with

Questions people actually ask

Everyone uses Whisper. Is it the wrong choice?+

It is a fair choice, and the board holds more accurate ones for this particular job. What Whisper buys you is the software: free apps on every platform, no terminal, and a longer list of languages than anything above it. What it costs you shows up on exactly the audio this guide is about, a room with several people in it, where the recorded-meetings board puts it well down the field. If you already have a Whisper app and the transcripts read well enough, keep it and spend the effort on the folder instead.

Do I need a GPU for this?+

No. The smaller transcription models run on ordinary laptops and Apple silicon, and a graphics card shortens the overnight batch without being the entry ticket.

Is it legal to record my meetings?+

The rules depend on where you are and who is in the recording, and in much of Europe telling people is not optional. This guide cannot answer it for you, so check what applies where you work before you hit record.

Why not just use a hosted transcription API?+

You can, and for public material it is the convenient route. A meeting is mostly other people's words, though, so sending it away makes a retention decision on everyone's behalf. If you go hosted, read who can see your prompts first.

Can I just paste a whole month of notes into a model and ask?+

Not usefully. A month of meetings runs past any context window you have, and a model handed too much finds less. Search the folder first, hand it the two or three notes that matched, and ask the question against those.

One week of meetings, one folder of Markdown files, and a transcription model chosen off the meetings board. Search the folder before you ask a model anything.