transcription-skill
by kajisho5
Transcribe speech in audio or video files (mp4, mov, mkv, wav, m4a, mp3) into a structured, timestamped Transcript with a local ASR engine (faster-whisper), Japanese and English first-class; export plain SRT/VTT; derive speech intervals. Use this skill whenever the user asks what is said in a recording, wants a transcript, timestamps of speech, word timings, or timed text to hand to a subtitle or editing tool. It does not edit video, style subtitles, detect speakers, or decide anything about a production.