Chinese speech to text

Chinese Transcription for Mandarin Recordings

Upload a Mandarin recording and get the transcript in Chinese characters. Play it back, correct names and numbers, and export TXT, DOCX, SRT, VTT, or JSON without leaving this page.

  • Chinese preselected
  • Written in Chinese characters
  • Play, edit, and export here
Chinese is preselected for Mandarin speech. For a Cantonese recording, pick Cantonese in the language list and check a short clip first.
My recordsCloud processing · up to 1 GB / 60 minutes. Final usage is verified by the server.

Sign in to create a saved task. If your file is unavailable after signing in, select it again.

Dashboard

New Transcription
0 min

How do you want to transcribe?

Upload audio or video
Estimated cost: 0 min

Free minutes are included. Upload a file or record audio to start.

Choose your file and settings before signing in; sign in to submit. Processing uploads audio to cloud services and uses account credits. Review the estimated credits and your balance first. See plans and usage limits and our privacy policy.

Three steps

How It Works

  1. 1

    Upload the Chinese Recording

    Choose an audio or video file, or record with your microphone. Chinese is already selected, so the model writes Chinese characters instead of guessing the language from a noisy intro. Check the duration and credit estimate, then sign in to submit. Upload progress and any error stay in the tool above.

  2. 2

    Play Back and Correct

    When processing finishes, the transcript opens here. Play the recording while you check personal names, company names, figures, and English terms spoken inside Chinese sentences. Save your edits. If processing fails or finds no speech, read the message and submit again after fixing the file.

  3. 3

    Export in Chinese

    Download TXT or DOCX for reading and editing, SRT or VTT for captions, or JSON with timestamps. Every text export is UTF-8, so hanzi open correctly in current editors. The task also appears in Recordings; reopen its link later to download again without paying for another run.

What you get back

What Chinese Transcription Returns

Chinese transcription on this page runs speech recognition with the language fixed to Chinese. You get a transcript written in Chinese characters, split into timed segments that play back against the audio and become caption cues. Segments are joined without the stray spaces that appear when Chinese text is stitched together like English words.

Spoken numbers usually come back as digits, dates as 2026年10月8日, and percentages with a percent sign. English product names and acronyms spoken inside a Chinese sentence stay in Latin letters, but they are also a common source of mistakes. Punctuation may be half-width, so tidy commas and question marks before you publish.

The transcript is not translated, summarized, or converted between simplified and traditional characters, and the model does not always pick the same script: in our test, adding background noise switched a Mandarin file to traditional characters. This is Chinese speech to text for people who need the original-language record: a researcher checking a Mandarin interview, an editor preparing captions, or a team member who wants a colleague fluent in Chinese to review a call.

Real sample, real output

See an Example

Mandarin interview transcript with a misheard name and English acronym highlighted for review
Illustration built from the recorded output below; not a screenshot of the editor.

This practice interview is synthetic: we wrote a six-line Mandarin script with two fictional speakers, then generated both voices with macOS text-to-speech. A second copy adds steady background noise. The script is the human reference.

We transcribed both files with Whisper large-v3, the model behind this tool, with the language set to Chinese and the same greedy decoding settings. The run used the same model weights on our own machine, so the cloud service may differ slightly. The table shows every difference that matters for review.

Human reference vs model output

Human referenceClean fileNoisy fileWhat to check
林晓雨林小雨林曉宇The interviewer’s given name came back with same-sounding characters.
陈思远陈思远陳思源The guest’s name was right on the clean file and one character off with noise.
导出导出老出An everyday word was misheard once background noise was added.
用API批量处理用RPP量处理用OPP量處理An English acronym inside a Chinese sentence was misheard in both files.
12个人 · 35% · 2026年10月8日12个人 · 35% · 2026年10月8日12個人 · 35% · 2026年10月8日Spoken numbers came back as digits and matched.
我们团队我们团队我們團隊Same Mandarin voice, but the noisy file came back in traditional characters.

Ignoring punctuation, the clean output differed from the reference by 4 of 164 characters. The noisy output differed by 6 after converting its script for comparison, and every comma and question mark in both files was half-width. Upload the clean MP3 above to try the full loop: fix 林小雨 and the acronym, save, and export TXT. Your own run may not match ours exactly.

Know the limits

Dialects, Names, and Silence

Mandarin is the tested case. Cantonese is a separate choice in the language list; our only check was one synthetic sentence, where the Cantonese setting wrote colloquial Cantonese characters and still misheard one word. We have not tested Hokkien, Taiwanese, Shanghainese and other Wu varieties, Sichuanese, Hakka, or heavy regional accents. Try a short excerpt before relying on those recordings.

Names are the weakest point. Many Chinese given names share a sound with other characters, so the model picks a plausible spelling rather than the person’s real one, as with 林晓雨 in the sample. Add important names and terms in the settings before submitting, keep a list of the correct characters, and confirm company names, places, and English acronyms against your notes.

Silence and pure noise are a known trap for this model: it can write a subtitle credit or a sign-off that nobody said. The tool removes those lines and fails the task with a clear no-speech message when no spoken words remain, rather than saving a full fake transcript. A short invented phrase inside a long, noisy recording can still slip through, so listen to quiet passages.

Choose for your workflow

WhisperWeb vs Sonix vs Notta vs Trint for Chinese

Chinese workflowWhisperWebSonixNottaTrint
Chinese varietiesWhisperWebLimited: Mandarin tested; Cantonese optionMandarin sample shown above. Cantonese appears in the language list but was only spot-checked.SonixLimited: Mandarin pageIts Mandarin page says Cantonese, Shanghainese, and Hokkien are separate languages that a Mandarin model will not handle accurately.NottaLimited: Mandarin pageIts Chinese converter page lists Mandarin with simplified and traditional output.TrintLimited: Mandarin pageIts Chinese page describes Mandarin transcription.
Simplified or traditionalWhisperWebNot available: No conversionKeeps the script the model writes; convert separately if needed.SonixLimited: Guidance onlyIts page advises choosing a script for your audience.NottaAvailable: ListedIts page names both simplified and traditional.TrintAvailable: Via translationIts page says it translates into simplified and traditional Chinese.
English translationWhisperWebLimited: Separate toolUse the Audio Translation Studio after transcription.SonixAvailable: ListedAI translation is described on its Mandarin page.NottaAvailable: ListedTranslation of Chinese transcripts is described.TrintAvailable: ListedTranslation is described alongside transcription.
ExportsWhisperWebAvailable: TXT, DOCX, SRT, VTT, JSONPDF is left out on this page because it cannot render Chinese characters.SonixAvailable: Word, PDF, text, SRT, VTTFormats named on its Mandarin page.NottaAvailable: TXT, DOCX, SRT, PDF, XLSXFormats named on its Chinese converter page.TrintAvailable: Multiple file typesIts page mentions several export types.

Checked September 13, 2026, from each vendor’s own Chinese page: Sonix, Notta, Trint. We did not run their products; accuracy claims on those pages are not repeated or compared here. Plans and features change.

Pick WhisperWeb when you want to upload a Chinese file, check it against playback, and export text or captions from one page using the same account as your other recordings. Consider Notta or Trint if built-in script choice or translation in the same editor matters more, and Sonix if you already use its editor. Run a representative clip through whichever tool you are considering before committing to a plan.

Before you upload

Supported Files and Plan Limits

Upload MP3, WAV, M4A, AAC, FLAC, OGG, OPUS, WebM, WMA, MP4, or MOV, or record in the browser. Each file can be up to 1GB and 60 minutes. Split a longer recording into parts; each part becomes its own task and transcript.

Processing runs in the cloud and uses account credits based on duration. The estimate appears before you submit, and the server checks your balance. Read plans and usage limits for current terms. Only the task owner can open the saved transcript and its exports.

Common questions

Frequently Asked Questions

Does this transcribe Mandarin or Cantonese?

The preselected Chinese setting is aimed at Mandarin, and our sample interview is Mandarin. The language list also offers Cantonese. In one short synthetic Cantonese test, the Chinese setting rewrote the speech as standard written Chinese, while the Cantonese setting kept colloquial characters such as 我哋 and 喺. Test a short clip of your own speaker before a long file.

Will I get simplified or traditional characters?

You get whatever script the model writes, and it can change between files. The clean copy of our Mandarin sample came back in simplified characters, while the noisy copy of the same voice came back in traditional characters. This page does not convert between simplified and traditional Chinese, so convert the exported text in a separate tool if your readers need one script.

Can it translate the Chinese transcript into English?

Not on this page; it keeps the transcript in Chinese so you can check it against the audio. For an English version, use the Audio Translation Studio, which transcribes first and then translates the saved transcript. Have a fluent reader check any translation you publish.

What happens if my file is silent or only background noise?

In our tests the model wrote a channel sign-off asking viewers to like and subscribe for 20 seconds of silence, and a “thanks for watching” line for pink noise. Those lines are removed, and when nothing spoken remains the task fails with a no-speech message instead of saving a made-up transcript. The credits reserved for that failed run are refunded.

How accurate is Chinese speech to text here?

It depends on the recording. On our short synthetic interview the transcript differed from the human reference by 4 of 164 characters. The noisy copy differed by 6 once its traditional characters were converted to simplified for comparison. That is one sample, not a benchmark. Real interviews with crosstalk, accents, dialect, or rare names can do much worse, so review before quoting anyone.