Skip to main content

whisper

Transcribes and translates audio with OpenAI’s Whisper (openai-whisper Python package and whisper CLI), covering model sizes from tiny to large plus turbo, language specification, initial prompts, word and segment timestamps, temperature fallback, batch processing, and subtitle generation. Use when transcribing speech, podcasts, meetings, or video audio to text. Use when translating non-English speech to English. Use when working with noisy or multilingual audio in any of 99 languages. Use when choosing a Whisper model size for available VRAM. Use when generating subtitles from audio. Not for speaker diarization or live captioning; use AssemblyAI or Deepgram instead.
Category: multimodal-and-emerging · License: MIT · Version: 1.0.0

Install

When to use it

Transcribes and translates audio with OpenAI’s Whisper (openai-whisper Python package and whisper CLI), covering model sizes from tiny to large plus turbo, language specification, initial prompts, word and segment timestamps, temperature fallback, batch processing, and subtitle generation. Use when transcribing speech, podcasts, meetings, or video audio to text. Use when translating non-English speech to English. Use when working with noisy or multilingual audio in any of 99 languages. Use when choosing a Whisper model size for available VRAM. Use when generating subtitles from audio. Not for speaker diarization or live captioning; use AssemblyAI or Deepgram instead.

Full playbook

Read SKILL.md for the complete workflow, references and any scripts. The agent installer copies the full skill folder.