How to Summarize a Long Recording Without Losing the Middle
Summarize a 1- to 3-hour recording without losing the middle: transcribe it, split it into sections, summarize each one, then merge them with a time index.
To summarize a long recording, turn it into text first, split the transcript into sections wherever the topic changes, summarize each section on its own with the same template, and then write the final summary from those section notes rather than from the raw transcript. Working in sections stops the middle of a two-hour session from being squeezed out, and each section's start time becomes an index for jumping back to any moment that needs checking.
This guide is for recordings of an hour or more: a two-hour strategy call, a half-day workshop, a lecture, a long research interview.
Why a long recording is harder than three short ones
You cannot skim audio. Scrubbing a waveform tells you nothing about a section until you listen to it. Even at 2x speed, a two-hour recording takes an hour, and by the end the detail from the start has usually faded. A transcript turns an audio problem into a text problem: text can be searched, skimmed and handed to a language model. Audio cannot.
One-pass AI summaries lose the middle. The research paper Lost in the Middle: How Language Models Use Long Contexts (Liu et al., 2023) found that performance was often highest when the relevant information sat at the beginning or end of a long input, and dropped significantly when it sat in the middle. That study tested question answering and key-value retrieval. A later study of summarization itself, On Positional Bias of Faithfulness for Long-form Summarization (Wan et al., NAACL 2025), found the same U shape: models summarized the beginning and end of a document faithfully and neglected the middle. Paste a two-hour transcript in one go and most of the meeting is in the middle.
A fixed-length summary has to compress harder. Ask for one page from three hours and a decision made in ninety seconds competes for space with a fifteen-minute tangent. The tangent usually wins, because a summarizer tends to weight topics by how much was said about them.
Long sessions change their mind. A decision taken in the first half hour can be reopened ninety minutes later. A single-pass summary may report the first version, the second, or both as separate decisions. Sections make the reversal visible: the two moments land in different notes with different start times.
Step 1: Get the file into a shape you can work with
Check the recording is complete first: play a few seconds at the start, middle and end. A file that cut out partway is better found now than after you have summarized it.
Then think about length. A tool that transcribes inside your browser has to hold the audio in your computer's memory, and the browser's standard audio API decodes it into uncompressed 32-bit samples (MDN: AudioBuffer), resampled to the output device's rate, most commonly 44,100 samples a second (MDN: AudioContext). Two hours of stereo is therefore 44,100 × 2 channels × 4 bytes × 7,200 seconds, about 2.5 GB, before transcription even starts.
Splitting solves that, and it gives you chapters for free. Cut the file into parts of 30 to 60 minutes and number them so each part's start time is known. With the free command-line tool ffmpeg, one line does it:
ffmpeg -i meeting.m4a -f segment -segment_time 1800 -segment_start_number 1 -reset_timestamps 1 -c copy part%02d.m4a
That writes 30-minute parts (1,800 seconds each), numbered from 01 and each starting at 0:00, without re-encoding. The ffmpeg documentation for the segment option notes that split points may not be exact; for an index, close is enough. Any audio editor that cuts at set points works too. If the recording lives inside a meeting platform, export it first; for Zoom, see turning a Zoom recording into text on the free Basic plan.
Step 2: Transcribe, then map the structure
Transcribe each part. Before you run anything through a summarizer, skim the transcript and mark where things change. Three kinds of phrase make the structure obvious:
- Agenda transitions: "okay, let's move on to," "the next item is." These are your section boundaries.
- Decision language: "we agreed," "the plan is," "going forward we will." This is where conclusions land.
- Commitment language: "I'll take that," "can you send," "by end of week." These are your action items.
The scan also shows whether the transcript is good enough to summarize. If a stretch is garbled by crosstalk, a dropped connection or a noisy room, a summary will silently miss whatever was said there, so note the gap rather than let the summary paper over it.
Step 3: Build a time index
A long recording's summary needs an index as much as bullet points: the summary is where people start, the recording is where they settle disagreements. A simple table is enough:
| Start | Section | What happened |
|---|---|---|
| 0:00 | Introductions and agenda | Scope of the review agreed |
| 0:14 | Last quarter's results | One figure questioned; owner to re-check |
| 0:47 | Hiring plan | Two roles approved, one deferred |
| 1:32 | Vendor choice (reopened) | Earlier decision paused pending a second quote |
If your transcription tool prints timestamps, copy the time at each topic change. If it does not, the parts from step 1 are your anchors: with the command above, part03 starts at 1:00:00. Within a part, estimate where a topic began from how far into that part's transcript it starts: a third of the way through a 30-minute part is about 10 minutes in. Play from there, listen for the transition phrase you marked in step 2, and add the player's position to the part's start time. Within a minute is close enough.
What counts as a section depends on what was recorded:
| Recording | Natural section boundary | What each section note must capture |
|---|---|---|
| Meeting or workshop | Agenda item or exercise | Decisions, commitments with owner and date, open questions |
| Lecture or training | Topic or slide change | Key ideas, definitions, worked examples, anything flagged as important |
| Interview or research session | Each question | The answer in brief, plus one verbatim quote worth keeping |
| Panel or webinar | Each speaker or audience question | Each speaker's main claim and the evidence given for it |
Step 4: Summarize each section with the same template
Use one fixed template for every section, so the notes can be merged without rewriting:
- Section number, title and start time
- Decisions, and whether any of them changes something decided earlier
- Commitments as owner, task and date, written as "Unassigned" or "No date" when none was stated, never guessed
- Open questions
- Names, numbers and dates that were mentioned
- One line on anything that sounded reluctant or unresolved
Keep each section small enough for your summarizer to read all of it. If the tool has an input limit, that limit sets the maximum size, not the topic: split a long topic into as many runs as it needs and label them 4a, 4b, 4c.
Step 5: Merge the section notes into one summary
Now write the final summary from the section notes, not from the transcript. Section notes are short enough to read in one sitting, and nothing in them is buried in a middle. If your summarizer caps its input and the combined notes are over that cap, merge in rounds: summarize the notes for each part, then summarize those.
- Resolve reversals. If section 2 says "go with vendor A" and section 6 says "vendor choice paused," record the later position and say it changed.
- Merge repeated commitments. Tasks often get restated at the end of a long meeting. Keep one line with the clearest owner and date.
- Rank by consequence, not airtime. The decision that changes what people do on Monday goes first, even if it took ninety seconds.
- Attach the index. Put it under the summary so anyone can jump to the source.
The finished document opens with a short paragraph on what the session was for and what came of it, then decisions, action items, open questions and the index. For where the action items should live next, see keeping action items from slipping after a meeting.
Speaker attribution over a long session
Attribution gets harder as a recording gets longer. People join late, leave early and change seats, so a voice that was close to the microphone in the first hour may be across the room by the last. Some transcription tools, including the one in Meetings Brief, produce no speaker labels at all. Three habits fix most of it:
- Keep a roster with rough join and leave times. An action item cannot belong to someone who had already left.
- Ask whoever is chairing to restate the owner and date at the end of each item. "So that's Priya, by Friday?" makes the commitment easy to find and attribute.
- Check every attributed action item against the recording. Jump to the moment from the index and listen for thirty seconds.
What the summary will still miss
Sectioning fixes the length problem. It does not fix judgment. "That's an interesting idea," said three times about the same proposal and followed by "we can park that for now," is usually a polite rejection, and a transcript-based summary will often log it as "the team considered the proposal." Where the real outcome was what nobody said, a human has to write that line. For what to verify before sending any AI-written summary, see how AI summarization works and when to trust its output.
Some recordings need the transcript itself, with the summary as a way to navigate it: anything with legal or regulatory weight, calls where the exact words of a commitment or scope change will decide what happens next, and research interviews where the detail of what someone said is the output. Keep both. If the session needs formal minutes rather than a working summary, our guide to writing minutes fast without losing accuracy picks up from here.
Doing this with Meetings Brief
Meetings Brief is free and handles the transcribing and the summarizing. Splitting the file and building the time index are yours to do, and these are the limits worth knowing.
Transcribe an existing recording at meetingsbrief.com/transcribe. The speech model, a version of OpenAI's Whisper, runs in your browser, so the audio file is not uploaded. It reads mp3, m4a, wav, ogg, webm and the audio track of an mp4. Like Whisper itself, it works through audio in a sliding 30-second window, here with neighboring windows overlapping by 10 seconds so words are not cut at the seams. We set no length limit; your computer's memory is the limit, so split any file that will not load. The output is plain text with no timestamps and no speaker labels, which is why steps 1 and 3 matter. English is the default language; if the recording is in another language, choose it from the list before you start.
Summarize the sections at the meeting summary generator. One run accepts about 7,800 characters of notes, roughly 1,300 words, and the page tells you before sending if you are over. Paste one section at a time; dividing a transcript's word count by 1,300 gives a rough count of the runs it needs. A two-hour transcript takes many runs, and the notes from that many runs are usually too long for one run as well, so merge in two rounds: combine the notes for each part from step 1, summarize each group, then paste those group summaries for the final version. The text you paste is sent to Anthropic's API to write the summary and is not used to train models.
For long calls that have not happened yet, use live meeting notes. In Chrome or Edge on a desktop, you share the Google Meet, Zoom or Microsoft Teams tab with its audio; no bot joins. The transcript builds every 30 seconds and is saved on your device. When you stop, the brief you can email uses the same section-then-merge idea: a long transcript is cut into sections of about 15,000 characters, each is summarized, and the brief (summary, topics, action items with owners, open questions) is written from those notes, with the transcript attached. The sections follow length rather than topic, and the brief has no time index. That step sends the transcript text, not the audio, to Anthropic's API. It does not work on iPhone or Android.
If what you have is a short voice memo rather than a long session, skip the sectioning and see turning a voice note into a meeting summary.
Frequently asked questions
How do you summarize a two-hour recording?
Transcribe it, then split the transcript into sections at each change of topic and note each section's start time. Summarize every section with the same template: decisions, commitments with an owner and a date, and open questions. Write the final summary from those section notes rather than the raw transcript, and attach the time index.
Why do AI summaries miss things in the middle of a long recording?
Language models tend to use information at the beginning and end of a long input better than information in the middle, a pattern documented in the 2023 research paper "Lost in the Middle" by Liu and colleagues. A 2025 study of summarization by Wan and colleagues found the same U shape: models summarized the beginning and end of a document faithfully and neglected the middle. A two-hour transcript pasted in one go puts most of the meeting in that middle, and a fixed-length summary must compress harder, so short but important decisions lose out to long discussions.
Should I split a long audio file before transcribing it?
For recordings over an hour, usually yes, especially with a tool that runs in your browser. The browser's standard audio API decodes it into uncompressed 32-bit samples held in memory, so two hours of stereo at 44.1 kHz occupies about 2.5 GB before transcription starts. Parts of 30 to 60 minutes avoid that, let you rerun one part instead of the whole file, and give each part a known start time.
How do I add timestamps or chapters to a summary of a long recording?
If your transcription tool prints timestamps, copy the time at each topic change into an index. If it does not, split the recording into parts and use each part's start time as an anchor. Estimate where each topic begins from how far into that part's transcript it appears, play from there, and add the player's position to the part's start time. A table with start time, section title and one line on what happened is enough.
How many summary-generator runs does a two-hour transcript need?
Divide the transcript's word count by about 1,300, roughly what one run accepts. Speaking speed varies a lot between meetings, so count the words rather than guessing from the length of the recording. A 15,600-word transcript, for example, needs 12 runs for its sections, then at least one more to merge the section notes, or two rounds of merging if the combined notes are also over the limit.
Can Meetings Brief summarize a long recording in one go?
For a live call, yes: live meeting notes, in Chrome or Edge on a desktop, splits a long transcript into sections, summarizes each, and emails one brief with the transcript attached. For a recording you already have, transcribe it free in your browser, then summarize it section by section in the summary generator, which accepts about 7,800 characters of notes per run, and finish by summarizing the combined notes, in two rounds if they are over the limit.