How to search your own podcast transcripts

You know you said it. You know roughly when. You have no idea which episode, and scrubbing through six hours of audio to find forty seconds is not a plan.

This is the most common request hosts have about their own back catalogue, and it gets answered badly. The usual advice is "get transcripts," which solves a different problem. A transcript makes a single episode readable. It doesn't make two hundred episodes findable, and those aren't the same job.

Getting the transcripts is the easy half

Most hosting platforms now generate transcripts automatically, and if yours doesn't, a batch transcription service will do your whole back catalogue for a few dollars an hour of audio. Accuracy on clean interview audio is good enough that you'll rarely be misled.

Check your feed before you pay for anything. A lot of shows already publish transcripts through the podcast:transcript tag without the host knowing, because the host turned it on once and forgot. If yours are there, they're free and they're already speaker-labelled.

Then you have two hundred text files and the actual problem starts.

Why Cmd+F fails on spoken language

Keyword search assumes you remember the words. You almost never do.

What you remember is the shape of the thing. A guest talked about their factory burning down. Except they said "the fire," and before that they said "the incident," and the word "factory" never appears in that stretch at all because it was already obvious from context. You search "factory fire" across your archive and get nothing, so you conclude it wasn't in your show, and you don't use it.

Spoken language does this constantly. People use pronouns for whole concepts. They refer back to things they mentioned ten minutes ago with a "that." They interrupt themselves. A transcript is faithful to how people talk, which is exactly what makes it hostile to keyword search.

The second problem is that a hit isn't an answer. Search "burnout" across your archive and you'll get sixty matches, most of them somebody saying the word in passing. What you wanted was the one conversation that was about it. Ranking by relevance is the entire job, and Cmd+F has no concept of relevance.

What the options actually do

Search inside your hosting platform. Some hosts search episode titles and descriptions, which is metadata, not content. A few search transcript text. Worth checking, free if you have it, and it's keyword matching, so it inherits both problems above.

Grep a folder of text files. Free, fast, precise, and you have to know the words. Genuinely useful when you're looking for a specific proper noun. Useless when you're looking for a topic.

A notes app with everything pasted in. Notion, Obsidian, Apple Notes. Better than nothing and it degrades over time. The failure mode is that keeping it current is manual work with no deadline, so it stops happening around episode forty and you can't trust it afterwards.

Semantic search. Instead of matching words, this matches meaning. You describe the moment and it finds passages that are about that, whether or not they use your words. "The guest whose factory burned down" finds the fire conversation even though nobody said "factory" during it.

The tradeoff is real. Semantic search can be vague where keyword search is exact. Search for a specific person's name and keyword wins, because you want that string and nothing near it. The good setups run both and blend the results, which is why you'll see "hybrid search" in tooling descriptions.

Make it findable by more than text

Text search over transcripts is one layer. Two things make an archive substantially more useful, and neither requires new recording.

Timestamps that go back to the audio. A search result that says "episode 112" leaves you scrubbing. A result that says "episode 112, 47:20" means you can hear it in ten seconds, which is the difference between using a callback on air and deciding it isn't worth the risk.

A list of what your show has actually covered. Not tags you wrote by hand, which nobody maintains past the first month. The set of people, companies, works, and topics that came up, pulled from the transcripts themselves, with the episodes attached to each. This answers the question you have most often, which isn't "where did I say this phrase" but "have I covered this, and with whom."

That second one changes how you plan, too. When you can see that four different guests circled the same subject across two years and disagreed with each other, that's an episode. You'd never find it by searching, because you didn't know to look.

A workflow that survives contact with a real week

Do this once, then keep it cheap.

Get every episode transcribed in one pass, including the old ones you're slightly embarrassed by. Those are usually where the good callbacks live, because they're far enough back that your audience has forgotten them and you have too.

Search by description, not by keyword. Type the sentence you'd say to a producer. "The one where somebody argued that remote work was over." Save the searches that worked as a starting point for the next time.

Before you record, spend ten minutes searching your own archive for the topics on your rundown. You'll find two or three prior moments worth referencing, and having them in front of you is what makes a callback land instead of half-surface as "I think we talked about this once."

And accept that the archive is only useful if it's complete. A search tool that covers your last twenty episodes answers almost nothing. The value curve is steep and it's at the back.

Where this stops being a search problem

Everything above works between recordings, when you have both hands and time to look. The harder version is when you need the same answer with a microphone open, mid-sentence, and no way to go hunting.

That's what we build. Tally indexes your whole published archive, and searching it by description is one part of that. The other part is live intelligence, which means the right information appearing on screen while you're still talking. During a session the desktop app listens and surfaces the line you needed: the episode where you covered something before, the guest's bio when their name goes, a flag on a number that just went past. It makes no sound and it stays out of your recording chain.

It runs on macOS today and it's invite only while we onboard shows one at a time. Here's how live intelligence works, and if the thing you keep losing is a specific name or date, that's a different problem than dead air and worth reading separately.


Tally is invite only right now

We're onboarding shows in small waves so we can set each one up properly. Leave your email and we'll reach out when the next wave opens.