7 min read · updated 13 Aug 2026

Local vs cloud transcription.

Two ways to turn speech into text, and the differences that actually change your week. The rows that do not favour running it yourself are in here too.

08

NameSize

2026-08-03-1030-launch.m4a38 MB

2026-08-03-1030-launch.md24 KB

2026-08-04-0930-review.m4a31 MB

2026-08-04-0930-review.md19 KB

Vaultmeetings202608

example.com/transcribe

Upload audio

2026-08-03-1030-launch.m4a

Uploading, 62%38 MB

The same recording, on your Mac and in the cloud.

Two machines, and most pages answer for one.

Where the audio becomes text and where the text becomes a note are two separate steps, and a tool can put them in two separate places. Most of the ones that call themselves local are local for the first one only.

Offline is not a degraded mode.

A flight, a basement, a conference wifi that has given up. On your own machine the transcript happens anyway, which is the difference between a tool that works and a tool that usually works.

Nothing is counted.

A four hour workshop costs what a ten minute standup costs, so the only meter in the room is the disk. That quietly changes what you bother recording at all.

And this is the room where it loses.

Three people at once, a room mic across the table, words that exist only inside your company. The larger model in a datacenter is worth the most exactly here, and our own run says so.

Row by row.

Eleven rows, and they do not all go one way: five of them favour the machine on your desk, four favour the datacenter, and two are honestly a draw.

Marks are our reading of each cell: a check is a plus for you, a cross is a limit, a dash is neither. We build one of these two, so read this as argued rather than neutral. Every row is about the transcription step alone, not about what writes the summary afterwards, which is a separate question further down this page.
What is being compared On your own machine Sent to a service
Where the audio goesPlus. Nowhere. It is read on the disk it was recorded toLimit. Uploaded, and kept as long as their policy says
OfflinePlus. Works on a plane, in a basement, on hotel wifiLimit. Every minute needs a connection
What it costsPlus. The machine you already own, and nothing per minuteLimit. Per minute, per seat, or a plan with a cap in it
SpeedNeither. Bounded by your Mac. Apple silicon does an hour in minutesNeither. Bounded by the upload and their queue, usually fast
Size of the modelLimit. Whatever fits in your memory and your patiencePlus. As large as their hardware allows
Difficult audioLimit. Strong accents, crosstalk and jargon cost you more herePlus. The larger model is worth the most exactly here
Getting better over timeNeither. When you update the app or the modelPlus. They swap the model and you download nothing
Long or frequent meetingsPlus. Nothing metered. Disk is the only limitLimit. The thing every pricing page counts
SetupPlus. A download, onceNeither. An account, and often a key and a bill
When it goes wrongNeither. Slow, or a warm laptopNeither. An outage or a rate limit, on their schedule
Where you can use itLimit. That one computer, awake, with the app runningPlus. A phone, a browser, Windows, a server

Where the cloud is genuinely ahead.

A datacenter runs a model that does not fit on a laptop, on hardware that is not on your desk, and it can replace that model with a better one overnight without anyone downloading anything. That advantage is not evenly spread. On one person speaking clearly into a decent microphone in a common language, the two approaches land close enough that you would have to look for the difference. The gap opens on the hard parts: heavy accents, three people talking at once, a room mic six feet away, and vocabulary that only exists inside your company.

The cloud is also the only answer if the meeting is not on the machine doing the work. A phone, a Windows laptop, a call you are not in at all. No amount of local processing reaches those.

Because this is the row where our own product is not flattered, the run is written up rather than summarised: the benchmark page has the six meetings, the hand corrected reference, the three hosted services that were compared, and a section on what the run does not cover. Read the method rather than a headline number, including ours.

Where running it yourself wins.

Offline and the meter are the two scenes at the top of the page. These three only show themselves over months.

  • The audio does not travel to become text. There is no upload, so there is no retention window, no sub-processor list and no question about what a model was trained on. That is a shorter conversation with a security team than any policy document.
  • It still works when the vendor is having a bad day. An outage upstream is somebody else’s incident, not your missing meeting.
  • It keeps working after the invoice stops. A local model on a machine you own has no renewal date attached to it.

The split almost everyone glosses over.

Transcription and note writing are two different machines. Speech to text is one model, and reading that text to produce a summary, decisions and action items is another. A product can be entirely local for the first and entirely hosted for the second, and most tools that describe themselves as local, private or on-device are exactly that, because a summary is where a large language model earns its keep.

So the useful question is two questions. Where does the audio become text, and where does the text become a summary? A page that answers only the first one is answering the easier half.

The fully local path for both halves exists. It means pointing the summary step at a model running on your own machine, which is slower and weaker than a hosted one, and it is an honest choice rather than an equal one.

How to choose.

Local, if

You record on one Mac, you deal with material you would rather not upload, you meet often enough for a meter to matter, or you want the thing to work with the wifi off. Also if the answer to where does this audio go has to be a short one.

Cloud, if

Your meetings happen on phones and Windows machines as well, your audio is genuinely hard, you need somebody else to run it while your laptop is shut, or you need a team of people covered by one administrator.

Both, honestly

Transcribe locally because it is free, offline and private, then send the transcript text to a good hosted model to write the note, and keep the choice of which one in your own hands.

Where Aside sits in this.

Aside records and transcribes on your Mac. The speech model ships inside the download, so the recording and the transcript happen offline from the first second, in around 99 languages, with nothing metered and no account involved in that step.

The second machine is the hosted one, and the front page says so rather than burying it: the note is written by Claude Haiku through Aside cloud, which means the transcript text leaves your Mac for that one call. The same goes for asking questions and for chat. Bring your own API key and it goes to your account instead. Point it at Ollama and nothing goes out at all. The full list, including the uncomfortable part, is on the front page.

Whichever way you set it up, your meeting notes are files on your Mac, which is a separate argument worth making.

Last updated 13 August 2026.

Ten meetings free. Then $8 a month, or $79 once.

Records and transcribes on your Mac, and every note is a Markdown file you keep.

Requires macOS 26 or later. Apple silicon, or Intel with much slower transcription.