[–] [S] 1 point 1 day ago

Fair question about the model. The reason it's Whisper is the runtime rather than the benchmark: the web version and the extension run on transformers.js, on WebGPU where it exists and WebAssembly where it doesn't, and Whisper is what runs there today. The desktop and CLI builds use whisper.cpp with large-v3-turbo.

Granite Speech and Cohere Transcribe aren't something I can run in a browser tab right now, so I can't claim to have compared them on equal footing. If you've got benchmark numbers for them on long-form, multi-language audio, I'd read them.

  • source
  • parent
  • context
  • [–] [S] 1 point 1 day ago

    That depends much more on the model size than on the year. In the browser we're limited to small models — tiny (English only), base and small, 40 to 250 MB — and on a German TV show those are exactly where Whisper gets shaky.

    The desktop app and the CLI can run large-v3-turbo (1.6 GB, 99 languages), which is a different experience on non-English speech. If your last try was a small model, that's the part worth changing before judging it.

  • source
  • parent
  • context
  • [–] [S] 2 points 1 day ago

    There is a desktop app already — macOS, Windows and Linux. It's Tauri rather than Electron, so it's a much smaller download, and it's meant for the cases a page is bad at: long files and batches, with large-v3-turbo instead of the small models the browser is limited to.

    Worth knowing before you try it: the builds are unsigned, and burning subtitles into the picture needs an ffmpeg built with the ass filter (Homebrew's plain ffmpeg formula doesn't have it; the Windows build offers to install one for you).

  • source
  • parent
  • context
  • [–] [S] 1 point 1 day ago

    Honest answer: for files that size, use the desktop app rather than the web page. The browser build holds the decoded audio in memory, so multi-GB files are where it starts to struggle, and I haven't tested anything close to 10 GB in a tab.

    The desktop app (macOS, Windows, Linux) exists for exactly this — long files, no browser in the way — and it can use large-v3-turbo (1.6 GB, 99 languages) instead of the small models the web version is limited to. Two things worth knowing before you download: the builds are unsigned, and burning subtitles into the picture needs an ffmpeg built with the ass filter.

    Files are processed one at a time, so a 50-hour fan edit split across many files is a batch job rather than one run.

  • source
  • parent
  • context
  • [–] 1 point 3 days ago

    For the whole workflow you described—cuts, face blur, and subtitles—I would still use Kdenlive or Shotcut. A subtitle-only tool will not replace tracking a blur or arranging clips. If subtitles are the part making the editor feel heavy, you can split that step out: generate and proofread an SRT first, then bring it into the editor for the final render.

    OpenSubs has Linux desktop builds for that subtitle step: it can generate and edit subtitles, export SRT/VTT/ASS, or burn them into an MP4. The builds are unsigned, and burn-in requires an ffmpeg build with the ass filter. (Disclosure: I’m one of the developers.)

  • source
  • [–] [S] 1 point 4 days ago

    Not system-wide audio, but per-tab: there's a browser extension that puts subtitles over whatever is playing in the tab, transcribed as it plays. It's in the Chrome and Firefox stores now, free and AGPL like the rest. It won't replace OBS + localvocal for anything outside the browser, but there's no filter chain to set up — install it, open the tab, hit Start.

    Two things to set expectations. It runs behind the audio: it transcribes in windows, and on integrated graphics that's tens of seconds, so the overlay prints how far behind it currently is rather than pretending it's live. And it only works where the page itself plays the media — a site embedding someone else's player in an iframe, or a file served without CORS headers, can't be captured at all, and it tells you so instead of sitting there.

    On permissions: it asks for no site access when you install it. When you hit Start it requests that one site, and nothing else.

  • source
  • parent
  • context
  •  

    I work on OpenSubs, a free, open source (AGPL-3.0) subtitle tool that runs entirely in the browser tab.

    You drop in a video file, and Whisper transcribes it on your own machine, using transformers.js with WebGPU where available and WebAssembly otherwise. The model (40–250 MB) downloads once and is cached. There is no upload endpoint in the product, so the video has nowhere to go.

    After that you can:

    • fix lines by typing over them (click a timestamp to jump to that moment)
    • translate into 20 languages with Chrome's built-in on-device translator
    • pick one of 12 caption styles, including word-by-word highlighting
    • export SRT / VTT / ASS, or burn the subtitles into an MP4 (libass compiled to WebAssembly, encoded with WebCodecs)

    A few things I learned building it:

    • Whisper hallucinates on silence and music ("Thanks for watching!", or the Japanese equivalent). A Silero VAD pass runs before Whisper, and a cleanup step drops the known stock phrases.
    • Singing doesn't count as speech for the VAD, so a music video gets a "no speech found" warning. You can still force it.

    Honest limits: it only takes video files, not audio-only files. Cue timings can't be edited yet. Builds are release candidates. Everything that runs locally is free with no account; the only paid part is optional cloud translation on our backend (US$5 for 1000 credits), and you can bring your own Claude / OpenAI / DeepL key instead.

    Site: https://opensubs.app/ Code: https://github.com/open-subs/opensubs

    Feedback welcome, especially on languages where the transcription goes wrong.