I work on OpenSubs, a free, open source (AGPL-3.0) subtitle tool that runs entirely in the browser tab.

You drop in a video file, and Whisper transcribes it on your own machine, using transformers.js with WebGPU where available and WebAssembly otherwise. The model (40–250 MB) downloads once and is cached. There is no upload endpoint in the product, so the video has nowhere to go.

After that you can:

  • fix lines by typing over them (click a timestamp to jump to that moment)
  • translate into 20 languages with Chrome's built-in on-device translator
  • pick one of 12 caption styles, including word-by-word highlighting
  • export SRT / VTT / ASS, or burn the subtitles into an MP4 (libass compiled to WebAssembly, encoded with WebCodecs)

A few things I learned building it:

  • Whisper hallucinates on silence and music ("Thanks for watching!", or the Japanese equivalent). A Silero VAD pass runs before Whisper, and a cleanup step drops the known stock phrases.
  • Singing doesn't count as speech for the VAD, so a music video gets a "no speech found" warning. You can still force it.

Honest limits: it only takes video files, not audio-only files. Cue timings can't be edited yet. Builds are release candidates. Everything that runs locally is free with no account; the only paid part is optional cloud translation on our backend (US$5 for 1000 credits), and you can bring your own Claude / OpenAI / DeepL key instead.

Site: https://opensubs.app/ Code: https://github.com/open-subs/opensubs

Feedback welcome, especially on languages where the transcription goes wrong.

top 50 comments

sorted by: hot top controversial new old
[–] 44 points 1 week ago (1 child)

Whisper hallucinates on silence and music (“Thanks for watching!”, or the Japanese equivalent).

I've had Whisper ask me to donate to anime fansubbing groups. I wonder what it was trained on?

  • source
  • hideshow 1 child comment
  • [–] 32 points 1 week ago (12 children)

    Nice, another blindly generated app by Claude Code. 37 commits and already announcing it tells it all.

  • source
  • hideshow 12 child comments
  • [–] 6 points 1 week ago (11 children)

    yep, I would rather generate one with Claude to my liking and make it more efficient (target the CLI for example)

  • source
  • parent
  • hideshow 11 child comments
  • [–] 9 points 1 week ago (10 children)

    Which is absolutely legit. I only criticize the people releasing a “production ready” product, which is nothing but a PoC and then also claim that it is AGPL-3, while abandoning it after 2 months. This hurts the open source community imho.

  • source
  • parent
  • hideshow 10 child comments
  • [–] 2 points 1 week ago (7 children)

    This hurts the open source community imho.

    Nah. At worst the open source community is indifferent. If someone can pick up where this left off then it benefits the community.

  • source
  • parent
  • hideshow 7 child comments
  • load more comments (2 replies)
  • [–] 9 points 1 week ago (2 children)

    It sounds like a pretty cool project, thanks for sharing! As a browser-based project designed to run offline, you might want to consider shipping it as an electron app.

  • source
  • hideshow 2 child comments
  • [–] [S] 2 points 2 days ago

    There is a desktop app already — macOS, Windows and Linux. It's Tauri rather than Electron, so it's a much smaller download, and it's meant for the cases a page is bad at: long files and batches, with large-v3-turbo instead of the small models the browser is limited to.

    Worth knowing before you try it: the builds are unsigned, and burning subtitles into the picture needs an ffmpeg built with the ass filter (Homebrew's plain ffmpeg formula doesn't have it; the Windows build offers to install one for you).

  • source
  • parent
  • [–] 6 points 1 week ago* (last edited 1 week ago) (9 children)

    Did whisper improve since last year?

    Last time I tried to generate subtitles for a German TV show, the subtitles were more inacurrate than what I would have written down (German is my third language and I struggle to understand anything that isn’t voiceovered) and timing was so bad I didn’t even bother to generate subtitles for the second episode.

  • source
  • hideshow 9 child comments
  • [–] [S] 1 point 2 days ago

    That depends much more on the model size than on the year. In the browser we're limited to small models — tiny (English only), base and small, 40 to 250 MB — and on a German TV show those are exactly where Whisper gets shaky.

    The desktop app and the CLI can run large-v3-turbo (1.6 GB, 99 languages), which is a different experience on non-English speech. If your last try was a small model, that's the part worth changing before judging it.

  • source
  • parent
  • [–] 16 points 1 week ago* (6 children)

    No but other models have. parakeet-unified is currently the best small model.

    You can tell OP vibe coded the whole thing from their outdated model choice.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 10 points 1 week ago (5 children)

    You can tell OP vibe coded the whole thing from their outdated model choice.

    I hate this landscape we're in now. I'm always suspicious of whether something was created by a thieving computer program or a real human. Just like I'm always suspicious of crazy videos these days being AI or not. Neither can I watch an ad anymore without suspicion. It's so much extra mental overhead.

    I'm tired, boss.

  • source
  • parent
  • hideshow 5 child comments
  • [–] 10 points 1 week ago* (last edited 1 week ago) (4 children)

    You can also just click the repository and see that Claude is a contributor which usually gives it away.

    My point is more so that OP didn't even put in the effort to ask claude to look up the current best ASR model to use but instead one-shotted the thing

  • source
  • parent
  • hideshow 4 child comments
  • [–] 2 points 1 week ago (2 children)

    Which one is the best model now? Last I checked (well over a year ago) it was still whisper.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 4 points 1 week ago* (1 child)

    IBM Granite Speech 4.1 or Cohere Transcribe according to current benchmarks.

    IBN Granite Speech 5 appears to be a great fast option as well for fast transcription but it's English only. Parakeet Unified is a good fast modern option with multi language.

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 1 point 2 days ago

    Fair question about the model. The reason it's Whisper is the runtime rather than the benchmark: the web version and the extension run on transformers.js, on WebGPU where it exists and WebAssembly where it doesn't, and Whisper is what runs there today. The desktop and CLI builds use whisper.cpp with large-v3-turbo.

    Granite Speech and Cohere Transcribe aren't something I can run in a browser tab right now, so I can't claim to have compared them on equal footing. If you've got benchmark numbers for them on long-form, multi-language audio, I'd read them.

  • source
  • parent
  • [–] 6 points 1 week ago (1 child)

    Why not just use the whisper plugin for Bazarr?

  • source
  • hideshow 1 child comment
  • [–] 3 points 1 week ago

    i get ya, but also a lot of people don't have a dedicated self hosted setup. recently my hdd died and all my films were there, so now i'm downloading and deleting, which just feels like a drag to start bazarr for one film. this is handy for some people.

  • source
  • parent
  • [–] 4 points 1 week ago

    I just made an open-shoe! This is the first FOSS shoe ever. Everything is free but if you want both shoes, you'll need a subscription. OpenShoe is right foot only. Any foot, any size, durable and incredibly versatile. Left shoe only available in sizes M5, M7, M11.5, and M16.

  • source
  • [–] 1 point 1 week ago (1 child)

    I'm a lazy reader so don't know if this is implemented but it would be cool if it could be done live with the current audio out of the system. I know this was done for streamers etc using OBS/localvocal but it's a pain to get working in my experience - having to fiddle around with OBS filter settings.

  • source
  • hideshow 1 child comment
  • [–] [S] 1 point 5 days ago

    Not system-wide audio, but per-tab: there's a browser extension that puts subtitles over whatever is playing in the tab, transcribed as it plays. It's in the Chrome and Firefox stores now, free and AGPL like the rest. It won't replace OBS + localvocal for anything outside the browser, but there's no filter chain to set up — install it, open the tab, hit Start.

    Two things to set expectations. It runs behind the audio: it transcribes in windows, and on integrated graphics that's tens of seconds, so the overlay prints how far behind it currently is rather than pretending it's live. And it only works where the page itself plays the media — a site embedding someone else's player in an iframe, or a file served without CORS headers, can't be captured at all, and it tells you so instead of sitting there.

    On permissions: it asks for no site access when you install it. When you hit Start it requests that one site, and nothing else.

  • source
  • parent
  • [–] 1 point 1 week ago (1 child)

    What’s the max file size?

    I’ve been looking for the fan edit of marvels Infinity saga.

    It’s about 50 hours long divided into many many files ranging from 3-10 gb each. Would it be able to handle this? Or what’s the best way to split them up

  • source
  • hideshow 1 child comment
  • [–] [S] 1 point 2 days ago

    Honest answer: for files that size, use the desktop app rather than the web page. The browser build holds the decoded audio in memory, so multi-GB files are where it starts to struggle, and I haven't tested anything close to 10 GB in a tab.

    The desktop app (macOS, Windows, Linux) exists for exactly this — long files, no browser in the way — and it can use large-v3-turbo (1.6 GB, 99 languages) instead of the small models the web version is limited to. Two things worth knowing before you download: the builds are unsigned, and burning subtitles into the picture needs an ffmpeg built with the ass filter.

    Files are processed one at a time, so a 50-hour fan edit split across many files is a batch job rather than one run.

  • source
  • parent
  • load more comments
    view more: next ›