Hindi–English Auto Captions — Speech to SRT & VTT

Hindi–English Auto Captions — Speech to SRT & VTT

Turn a short recording into editable captions

A spoken explanation is easier to review when you can see its words and timing. This tool uses a speech model selected for your chosen language to turn a short Hindi or English recording into caption segments. You can correct the recognized text, adjust the segment times and download subtitle files. The output is an editable draft that needs a human review, especially for names, numbers and unfamiliar terms.

The tool accepts an audio file or a video file whose sound your browser can decode. It extracts sound for recognition and provides an audio player with a current-caption preview. It does not burn text into a new video, identify speakers, translate between languages or record a live microphone. Use a separate subtitle-burn tool after you have checked the downloaded SRT.

Select a suitable source

Choose a recording lasting at least half a second and no more than 60 seconds, with a file size of 50 MB or less. Supported extensions are WAV, MP3, M4A, AAC, OGG, FLAC, MP4, WebM and MOV. The extension alone cannot guarantee compatibility: the codec inside the file must also be readable by the current browser. A short WAV or MP3 is a useful alternative when video sound fails to decode.

Pick the spoken language, Hindi or English. This setting guides transcription; it is independent of the website interface language. For a mostly Hindi explanation containing a few English technical terms, start with Hindi and inspect those terms closely. The model may change spelling, punctuation or script when speech mixes languages. Neither option promises reliable transcription of every accent or mixed-language sentence.

Trim long recordings before selecting them. The tool rejects an overlength file rather than silently transcribing only its opening minute. Clear speech with one main speaker is more practical than music, overlapping conversations or a noisy group recording. A silent or extremely quiet clip is rejected before the model is loaded, but the sound check cannot guarantee that a non-silent file contains speech.

First use and local processing

Click Load model & create captions only when you are ready for the first-use download. The browser loads the pinned Transformers.js runtime from jsDelivr and the selected model files from Hugging Face: Whisper Base for English and the multilingual Whisper Small model for Hindi. Runtime and model downloads can use around 150 MB or more for English, and around 400 MB or more for Hindi, depending on the required files and browser cache. A stable connection, available memory and time to wait are necessary. Changing the spoken language loads that language’s model; the previous worker is released. Hindi recognition is intended to produce Devanagari text, and a result with no Devanagari characters is rejected rather than presented as a successful Hindi transcript.

After the downloads, recognition runs in a browser worker using WebAssembly. The custom implementation decodes the selected source locally, combines its sound into mono and resamples it to 16 kHz for the model. It does not upload the selected audio to a transcription endpoint or store a transcript on the website server. The download providers still receive normal requests for their assets, including connection information such as an IP address.

Some model files may be cached by the browser to reduce repeat downloads. This is not a promise that later sessions work fully offline: cache eviction, a different browser, storage restrictions or missing runtime assets can require the network again. Closing the page removes the current input and caption interface; it does not necessarily remove the downloaded model cache. Website advertising or analytics can also use the network separately.

Review the generated segments

When recognition finishes, an audio player and numbered caption segments appear. Each segment has a start time, an end time and an editable text field. Play the source and compare the current-caption display with what you hear. Recognition timestamps are estimates, so check both the first and last word of a segment rather than assuming that the boundaries are exact.

Correct personal names, place names, dates, amounts and technical vocabulary manually. A plausible-looking phrase can still be wrong. Whisper may produce unrelated words over music or pauses; remove or replace incorrect text by editing the segment, and adjust its timing where necessary. This interface does not add, delete or reorder segments, so substantial restructuring is better done in a subtitle editor after export.

Times are entered in seconds, with millisecond precision. Keep the start at zero or later, the end after the start and no later than the source duration, and avoid overlapping segments. Caption text must contain 1–1000 characters, with no empty paragraph or timing-arrow sequence. After editing, click Update preview & downloads. Old download links disappear when a segment changes so you do not unknowingly save the earlier wording.

Export formats and a practical workflow

SRT contains numbered segments with hour-minute-second-millisecond timestamps. VTT starts with WEBVTT and uses the timing syntax expected by WebVTT-compatible players. TXT contains the segment text in order, without timecodes. All three are UTF-8 text exports, not audio files, MP4 videos or a rendered picture of a caption. A plain text editor can open them for inspection.

For a short tutorial, transcribe the source, correct an unfamiliar product name, align the segment boundaries and export SRT for your editing app. Keep TXT as a readable transcript. For a web player that accepts a separate subtitle track, try VTT and check its rendering in that player. A service may require a specific caption workflow, so confirm that it accepts the chosen format before relying on the file.

Exported text is wrapped at approximately 42 Unicode characters per line for readability. This is character-based wrapping, not a guarantee of a particular pixel width, two-line layout or reading speed. Hindi marks, emoji and different fonts can occupy different space. Check the subtitles in their final player and split or shorten dense segments there if your layout needs tighter control.

Browser support and troubleshooting

The implementation needs audio decoding, an OfflineAudioContext, module workers and WebAssembly. It uses a single WebAssembly thread and does not require camera or microphone permissions. A desktop browser can be more practical than a phone with limited memory. Mobile processing may be slow, warm the device or fail when the browser restricts background work; keep the tab available while recognition runs.

A model-loading error can come from a blocked external asset, lost connection, browser restrictions or insufficient memory. A decode error concerns the source format rather than spelling quality. Cancel processing terminates the worker and discards partial captions. A new run may reload model assets or reuse cached files. Work that exceeds five minutes stops with an explanatory message; try a shorter clip or a stronger device.

Related tools and reading

Caption exports can reveal phone numbers, addresses or other private details spoken in the recording. Understand the harm of exposed personal details

Frequently asked questions

Does it create captions automatically from sound?

Yes, when the browser can decode the source and load the speech model. It recognizes spoken Hindi or English and returns estimated timed segments. Review the result; it is not a verified transcript or guaranteed word-for-word caption file.

Can I download a captioned MP4 video?

No. This tool exports SRT, VTT and TXT files. The player is a source-audio preview. To permanently put captions into a video, first check the SRT and then use a compatible subtitle-burn tool or editor.

Is my recording uploaded for transcription?

The custom tool does not send the selected recording to an online transcription service. Recognition is local, but the runtime and model are fetched from external providers. Ordinary site resources can also make network requests.

Why are Hindi names or English terms wrong?

The recognition model can confuse pronunciation, spelling, script and unfamiliar vocabulary. Mixed-language recordings add difficulty. Choose the main spoken language and edit every important name or number before exporting.

Can it work on mobile or without internet?

A capable mobile browser may work, but memory and processing limits matter. First use requires asset downloads, and later offline availability is not guaranteed. Browser caching can reduce traffic without ensuring a complete offline environment.

Why did a subtitle file stop matching my changes?

Editing clears the existing links. Use Update preview & downloads to validate the revised times and create new files. A file already saved to your device remains the older snapshot and must be replaced with the revised export.