Summarize Spoken Content in Videos
Upload a video that contains speech and get a transcript-based summary with key points you can check against the recording. Cloud processing accepts files up to 1GB and 60 minutes and uses account credits. Summary generation reuses the existing Pro or Max summary — this page does not run a separate video model.
- MP4, MOV, WebM, MKV, M4V
- Upload only — no YouTube link
- Summary on Pro and Max
- Transcript review beside summary
- TXT, DOCX, PDF export
Sign in to create a saved task. If your file is unavailable after signing in, select it again.
Dashboard
Estimate appears after audio is ready. This creates one saved task in My records.
Requires an active Pro or Max plan.
The two stages are charged separately. Summary is an explicit Pro or Max action and uses the saved transcript, so a summary retry never retranscribes the audio. The server verifies final usage and plan access.
How do you want to transcribe?
Free minutes are included. Upload a file or record audio to start.
Choose your file and settings before signing in; sign in to submit. Processing uploads media to cloud services and uses account credits; summary generation is a separate Pro or Max action. Sign in to see your current plan and available credits. Review the estimate first — see plans and usage limits and our privacy policy.
WhisperWeb vs Notta vs Otter vs Happy Scribe
A video summarizer is only as useful as the transcript under it. Meeting-first tools optimize for live calls; social summarizers optimize for a link. This page is built for a recording you already own and a summary you can check line by line. The table is a dated reading of public pricing and feature pages, not a scorecard. No accuracy percentages, star ratings, or traffic figures are claimed, and no competitor outbound links are placed on this page.
| What the job needs | WhisperWeb | Notta | Otter | Happy Scribe |
|---|---|---|---|---|
| Start from a video file | WhisperWebAvailable: Upload here and stay hereMP4, MOV, WebM, MKV, or M4V up to 1GB / 60 minutes. The transcript and summary finish on this page, with no email handoff or expiry link. | NottaLimited: Upload, then a result linkThe public converter asks for an email and returns a link reported to expire in about 72 hours. Free files are capped at 3 minutes. | OtterLimited: Meetings first, files secondLive capture is the core product. Basic accounts have been limited to a small number of lifetime file imports. | Happy ScribeAvailable: Upload or importAccepts uploads plus YouTube, Vimeo, and Drive imports on a subtitle-oriented workspace. |
| Is the summary tied to a transcript you can open? | WhisperWebAvailable: Same task, same recordThe overview, key points, and chapters come from the saved transcript on this task. Edit the wording and the existing summary is marked out of date until you regenerate it. | NottaLimited: AI summary beside the notesSummaries and templates are advertised across plans, but the compressed text and its source are not guaranteed to sit in one review view. | OtterLimited: Summary on the meeting recordSummaries and chat attach to the captured conversation rather than a deliberately reviewed upload. | Happy ScribeLimited: AI summaries on paid plansMeeting summaries and AI chat are paid features; custom summary templates sit higher up the tiers. |
| Check a bullet against the exact moment | WhisperWebAvailable: Playback plus timestampsReplay the file, keep the original text beside the summary, correct names and numbers, then export the version you actually reviewed. | NottaLimited: Editor with timestampsReview tools exist on paid plans; free access to the full recording after a downgrade has been limited. | OtterLimited: Transcript in the Otter appThe live transcript is readable in Otter; independent verification of a summary claim still falls to the reader. | Happy ScribeAvailable: Online editor plus exportsBrowser editing with a wide export set, though human proofreading is sold separately from about $2.00 per minute. |
| Free allowance vs a 40-minute recording | WhisperWebLimited: 5 free minutes, then creditsOne credit per started minute after 30 seconds. A 40-minute video therefore needs a plan or credit pack rather than the free start. | NottaLimited: 120 minutes, 3-minute capThe free month is generous by total minutes but each file stops at 3 minutes, so a single long video needs Pro. | OtterLimited: 300 minutes, 30-minute capThe free plan is built around half-hour conversations and a low lifetime import limit, not long uploads. | Happy ScribeLimited: 10-minute AI trialFree AI transcription is a short trial; the entry plan then includes about 120 uploaded AI minutes per month. |
| Where the summary and transcript can go | WhisperWebAvailable: TXT, DOCX, and PDF hereDownload the summary as its own TXT and the transcript you reviewed as TXT, DOCX, or PDF. No caption export is offered because this page is not a subtitle tool. | NottaLimited: Downloads after upgradePublic help pages describe downloading TXT, DOCX, Excel, PDF, or SRT as a Pro capability. | OtterLimited: TXT and MP3 on BasicPDF, DOCX, and SRT are listed as higher-tier exports. | Happy ScribeLimited: TXT and SRT on FreeMore formats appear as the plan rises; subtitle delivery and human review are separate products. |
| Starting paid price (per account unless noted) | WhisperWebAvailable: From $9.90/mo, or $4.90/mo yearlyStarter $9.90/mo ($58.80/yr), Pro $29.90/mo or $14.90/mo yearly, Max $49.90/mo or $24.90/mo yearly. Summary generation stays a Pro or Max action. | NottaLimited: Pro $8.17/mo billed yearlyListed at $97.99 per year on the public Pro page. Monthly billing exists; confirm the live toggle before you buy. | OtterLimited: From $16.99/user/mo, or $8.33 yearlyPro and Business are priced per user, which matters once a team shares the account. | Happy ScribeLimited: From $17/mo, or $8.50 yearlyThe Basic plan lists about 120 AI minutes per month with an overage rate near $0.20 per minute. |
Figures checked on 2026-09-19 against the vendors' public pricing and tool pages. Brand names are trademarks of their owners. Plans change; confirm details with each provider before you buy. For our own column, the current plans and usage limits page is the source of truth rather than this table.
Choose WhisperWeb when
- You already have the recording and want to finish the whole job on one page, with no bot in the call.
- You need the summary to point back at a transcript you can replay, edit, and export rather than trusting compressed bullets.
- The video is a one-off file in a container another tool will not accept, or the speech is outside the few languages a meeting tool lists.
- You need a written summary for a report or an article, not just a highlight reel of a social clip.
Another product may fit better when
- You want a bot to join Zoom, Meet, or Teams and type during a live meeting — that is Otter, Happy Scribe, or Notta, not WhisperWeb.
- You need a human-proofread record or a broadcast caption package. Happy Scribe sells human review separately from about $2.00 per minute.
- You want a free, open-ended summary. WhisperWeb summary generation is a Pro or Max action; free accounts can still transcribe within the free allowance.
- You need slide OCR, on-screen text recognition, or visual scene understanding. None of those run on this page.
How It Works
The loop is upload, review, then summarize — all on this URL. No handoff to the general workspace, and no waiting on an email link.
Upload a video with speech
Choose an MP4, MOV, WebM, MKV, or M4V you have permission to process. This page takes a file, not a YouTube or social link. Set the spoken language or leave detection on, and read the duration and credit estimate before you sign in and submit. Media with no audio track cannot become a summary, so the run is meant to fail rather than invent words.
Review the saved transcript
Progress, errors, and the finished text remain on this address. Play the video while you correct names, figures, and product terms. The summary step reads this saved transcript, so a clean transcript produces a more trustworthy summary. If the recognizer finishes but the file held no speech, the task is rejected instead of saved as empty output.
Generate and check the summary
On Pro or Max, request the overview and key points from the transcript view. Compare each bullet with the wording behind it before you reuse it. If you later edit the transcript, the existing summary is marked out of date. Download the summary as TXT, or the transcript as TXT, DOCX, or PDF. Everything stays under one task in My records.
Summarize speech in uploaded videos
A summarizer that reads speech does not watch the picture. It transcribes the audio track, then writes an overview, key points, action items, and chapters from that text. Two stages keep it honest: the first saves a durable transcript, the second reads the saved words and compresses them. If the summary step times out, the transcript is untouched and only the summary needs a retry.
That ordering matters for a video because the interesting facts are often said, not shown. A product demo may name a price aloud, a lecture may read a definition, and an interview may state a date. Those survive transcription and can be summarized. A slide that is never spoken does not. Because the transcript is kept on the same task, you can open the original wording whenever a bullet looks too confident.
For audio-only recordings, the AI Audio Summarizer runs the same summary path from an audio upload. When the deliverable is a full written record rather than a condensed view, use the Video to Text converter.
Review key points against the transcript
A summary is an interpretation, not a quotation. Treat every bullet as a claim to verify. Open the transcript view, expand the wording behind a point, and replay the video at that timestamp. Names, numbers, product versions, and dates are the usual places an automatic pass drifts — a spoken “fifteen” can land as “fifty,” and a two-word brand name can split in half.
- Check proper nouns first. People, companies, and products rarely appear in the summary correctly if they were misheard in the transcript.
- Check every figure. Currency, percentages, and quantities carry more weight in a recap than in a full transcript.
- Edit, then regenerate. Fixing the transcript marks an older summary out of date so the two versions cannot be mixed silently.
- Keep the source. Export the reviewed transcript beside the summary when the result will be cited or published.
What this tool does not analyze
This is a speech tool, and the limits follow from that. Being clear about them protects the reviewer from trusting a summary for something it never saw.
- Silent scenes and visual context. A quiet montage, a reaction shot, or a physical demonstration produces nothing to summarize. The summary describes what was said, not what happened on screen.
- On-screen text and slides. Titles, lower thirds, code editors, and slide copy stay in the picture. OCR is a separate workflow that this page does not run.
- Music without words. A score or ambient track has no speech to transcribe. On this route a file with no usable speech is rejected rather than returned as an empty summary.
- YouTube, TikTok, and social pages. This page takes an uploaded file. It does not download a watch page or a share link, and it does not strip captions from a platform.
See an Example
The clip is an original synthetic seminar recording created for the WhisperWeb video cluster. It is not a customer recording and not a claimed model result. Use it to rehearse the upload, review, and export path on a short file. The points below match what the narrator actually says, so you can trace every one back to a spoken sentence.


Reference spoken script
Welcome to the Harbor Lab seminar recording. Today we convert a lecture capture into a working transcript. Pause at each chapter marker, check every speaker name against the slate, and export captions only after you have reviewed the text. Do not treat an automatic transcript as a finished quotation.
This is an authored script, not a transcription-provider output, and matches the sample above. Recognition, punctuation, and noise robustness vary by recording.
Authored expected summary
- The recording is a Harbor Lab seminar about converting a lecture capture into a working transcript.
- Pause at each chapter marker and check every speaker name against the slate.
- Export captions only after the text behind them has been reviewed.
- Treat an automatic transcript as a draft, never as a finished quotation.
These points were authored from the script, not produced by the summarizer. A live summary can drop the slate check or merge two instructions; that is exactly why the transcript stays open beside it.
- Container
- MP4 (H.264 + AAC)
- Duration
- 19.36 seconds
- File size
- 318,097 bytes
- Credits
- 0.5 (under 30 seconds)
Supported Files and Plan Limits
Cloud processing accepts one video file up to 1GB and 60 minutes, whichever bound you reach first. Billing follows the clock: half a credit through the first 30 seconds, then a full credit for each minute you start. A forty-minute recording is about forty credits, so check the estimate and your balance before you submit.
- Base transcription uses transcription credits on every plan, including the free start. The server re-checks file type, size, duration, and balance on submit.
- Summary generation is a separate Pro or Max action on the saved transcript and is charged only when you request it. Retrying a failed summary does not re-run or re-charge the base transcription.
- Media is uploaded to cloud services, so do not treat this page as on-device Whisper. Review the privacy policy before sending footage that contains other people's voices.
Need a different starting point? The Lecture to Notes page is tuned for class recordings, and the general transcription workspace adds link sources and subtitle exports.
Frequently Asked Questions
Can I paste a YouTube URL?
No. This page accepts an uploaded video file. It does not download a YouTube watch page, a TikTok share link, or a Drive preview, and it does not copy captions from a platform. Use a file you have the rights to process, or a direct media URL that ends in a supported extension on the general workspace.
Does it understand silent scenes?
No. The summary is built from speech in the audio track. A silent scene, a music-only montage, or a visual demonstration with no narration leaves nothing to transcribe. If a run finds no usable speech, this page rejects the task instead of returning an empty or invented summary.
Can I inspect the original transcript?
Yes. The saved transcript stays on the same task as the summary. Open the original or edited text, replay the video at a timestamp, and correct names or numbers before you regenerate the summary. Download the reviewed transcript as TXT, DOCX, or PDF, and the summary as its own TXT file.
Does summarization cost extra?
Yes, and it is a separate stage. Transcription uses credits on any plan. Generating the summary is a Pro or Max action charged only when you request it. If the summary fails, the transcript remains saved and only the summary step is retried, so the transcription is not charged twice.
Where does the finished work live?
The task is filed in My records with source tool “video-summarizer” and reopens from this page with its task link. Only the owner can read or download it. Choose the file again if a refresh drops a local selection before submit, because a browser cannot rebuild file bytes from a filename.
Related Tools
Stay on published WhisperWeb routes for the full transcript, an audio-only summary, or a class-recording workflow.
- Video to Text ConverterStay with the full timed transcript when a written record matters more than a condensed overview.
- AI Audio SummarizerStart from an audio file or a browser recording when there is no picture to keep.
- Lecture to Notes with AIUse the lecture-tuned route when the recording is a class or seminar.
- General transcription workspaceOpen the general workspace when you also need link sources, subtitles, or translation.