How to caption a Zoom recording

The files least suited to an upload, captioned without one.

Meeting recordings are near the top of the list of things an organisation should not be handing to a third-party service, and near the top of the list of things people need transcribed. Zoom's own cloud transcription exists but is tied to plan tiers and account settings, and it does not help at all with a local recording somebody saved to their laptop two months ago.

Open the subtitle editor → Reads the recording from your disk and transcribes it locally. Nothing is uploaded.

Find the file

A local Zoom recording lands in a Zoom folder in your documents or home directory, in a dated subfolder, usually as an MP4 with a separate M4A of the audio. Either works here — if you only want a transcript, the audio file is smaller and faster.

Drop it into the editor. If it is a long meeting, the editor will tell you roughly how much memory it expects to use before it starts. Recognition runs on your own processor or graphics card; the file is never sent anywhere, which is the entire point for this kind of material.

What meeting audio does to accuracy

Conference audio is harder than it sounds. Compression artefacts from variable connections, one person on a laptop microphone across a room, background noise, and several accents in the same call all cost accuracy. Use the Accurate model and set the language explicitly rather than leaving it on auto — both help noticeably on this kind of recording.

Expect names and internal vocabulary to be the weak point: project codenames, acronyms, product names, people's surnames. Find and replace fixes the repeated ones in a single pass, which is usually five minutes of work for an hour-long call.

Speakers are not labelled automatically. In a meeting with more than three people this is the real limitation, and the honest answer is that adding labels by hand across an hour is slow. If who-said-what matters more than exact wording, it is worth reviewing with the recording playing rather than reading cold.

Subtitles or transcript?

Decide which you actually need, because they are different documents. If the recording is going to be watched — shared with people who missed the call, posted internally, used for training — export SRT and keep the timings. If it is going to be read or searched, export plain text: sentences grouped into paragraphs with the timestamps stripped out.

For a document people will read, it is worth a cleanup pass that you would not bother with for subtitles: remove filler, fix the punctuation on long sentences, add speaker markers at the points where the thread changes hands. The result is the thing people actually wanted when they asked for the recording.

A note on consent

Transcribing a recording does not change who is allowed to have it. If the call was recorded with everyone's knowledge and you are allowed to hold the recording, a transcript is the same material in another shape. If it was not, a transcript makes it more portable and more quotable, not less sensitive.

The advantage of a tool that runs locally is that it does not add a party to that question — there is no service holding a copy and no retention policy to reason about. What it does not do is answer the question for you.

Frequently asked questions

Does this work with a Teams or Meet recording?

Yes. Any MP4, MOV, WebM, MKV or common audio file works. The platform it came from is irrelevant.

Can it label who is speaking?

No. Speaker separation is not implemented, so labels are manual. In a large meeting that is the main limitation.

Is the recording uploaded?

No. It is read from your disk by the page and processed on your own machine. That is the reason this approach suits confidential material.

Audio file or video file?

Either. If you only want the text, the M4A is smaller and gets to the same place faster.