ClawKit Logo
ClawKitReliability Toolkit
Back to Registry
Official Verified media Safety 4/5

speech-to-text

Transcribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large. Capabilities: transcription, translation, multi-language, timestamps. Use for: meeting transcription, subtitles, podcast transcripts, voice notes. Triggers: speech to text, transcription, whisper, audio to text, transcribe audio, voice to text, stt, automatic transcription, subtitles generation, transcribe meeting, audio transcription, whisper ai

Why use this skill?

Transcribe audio to text with OpenClaw using Whisper models. High accuracy, 99+ languages, timestamping, and translation support via inference.sh integration.

skill-install — Terminal

Install via CLI (Recommended)

clawhub install openclaw/skills/skills/okaris/speech-to-text
Or

What This Skill Does

The speech-to-text skill for OpenClaw provides high-fidelity audio transcription using state-of-the-art models like Fast Whisper Large V3 and Whisper V3 Large. By integrating the inference.sh CLI, the skill enables users to seamlessly convert spoken language into written text. It supports over 99 languages and offers advanced capabilities such as timestamp generation, language detection, and even automatic translation to English. This skill is engineered to handle diverse audio sources, including podcasts, voice notes, research interviews, and meeting recordings, providing accurate, structured data that is easy to search, index, or archive. By leveraging the inference.sh infrastructure, it delivers performance-optimized transcription directly within your workflow.

Installation

To integrate this skill into your environment, use the OpenClaw management CLI. Run the following command in your terminal:

clawhub install openclaw/skills/skills/okaris/speech-to-text

Ensure that you have the inference.sh CLI configured beforehand by running 'infsh login'. This ensures that the skill can authenticate against the inference engine to pull the necessary models and process your audio files efficiently.

Use Cases

  • Corporate Meetings: Automatically generate searchable records of board meetings or team syncs.
  • Content Creation: Effortlessly transcribe podcast episodes for show notes or blog posts.
  • Accessibility: Provide accurate subtitles for video content to make audio media inclusive.
  • Voice Productivity: Quickly convert dictated voice notes into organized, actionable text documentation.
  • Multilingual Research: Translate and transcribe interviews conducted in foreign languages to streamline analysis.

Example Prompts

  1. "Transcribe this meeting recording located at https://audio-source.mp3 and provide me with a full transcript including timestamps."
  2. "Can you take this French interview from https://interview.mp3 and translate it into English for my report?"
  3. "Generate a subtitle file for this video https://video.mp4 using the Whisper V3 Large model for maximum accuracy."

Tips & Limitations

  • Model Selection: Choose 'Fast Whisper' for high-speed batch processing, or 'Whisper V3 Large' if your priority is the highest possible word-error-rate accuracy.
  • Audio Quality: Ensure source audio is clear and free of extreme background noise; while Whisper is robust, high-quality input significantly improves punctuation and technical term accuracy.
  • Network Dependence: Because this skill utilizes the inference.sh API, you will need a stable internet connection to upload your audio files and retrieve the generated text results.
  • Formatting: The output is returned as structured JSON, making it ideal for piping into other OpenClaw workflows or automated database entries.

Metadata

Author@okaris
Stars1287
Views2
Updated2026-02-22
View Author Profile
AI Skill Finder

Not sure this is the right skill?

Describe what you want to build — we'll match you to the best skill from 16,000+ options.

Find the right skill
Add to Configuration

Paste this into your clawhub.json to enable this plugin.

{
  "plugins": {
    "official-okaris-speech-to-text": {
      "enabled": true,
      "auto_update": true
    }
  }
}

Tags(AI)

#transcription#whisper#speech-to-text#audio-analysis#ai-models
Safety Score: 4/5

Flags: network-access, external-api

Related Skills

product-changelog

Product changelog and release notes that users actually read. Covers categorization, user-facing language, visuals, and distribution. Use for: release notes, changelogs, product updates, feature announcements, versioning. Triggers: changelog, release notes, product update, version notes, what's new, feature announcement, product changelog, update log, release announcement, version release, product release, ship notes

okaris 1287

logo-design-guide

Logo design principles and AI image generation best practices for creating logos. Covers logo types, prompting techniques, scalability rules, and iteration workflows. Use for: brand identity, startup logos, app icons, favicons, logo concepts. Triggers: logo design, create logo, brand logo, logo generation, ai logo, logo maker, icon design, brand mark, logo concept, startup logo, app icon logo

okaris 1287

ai-content-pipeline

Build multi-step AI content creation pipelines combining image, video, audio, and text. Workflow examples: generate image -> animate -> add voiceover -> merge with music. Tools: FLUX, Veo, Kokoro TTS, OmniHuman, media merger, upscaling. Use for: YouTube videos, social media content, marketing materials, automated content. Triggers: content pipeline, ai workflow, content creation, multi-step ai, content automation, ai video workflow, generate and edit, ai content factory, automated content creation, ai production pipeline, media pipeline, content at scale

okaris 1287

explainer-video-guide

Explainer video production guide: scripting, voiceover, visuals, and assembly. Covers script formulas, pacing rules, scene planning, and multi-tool pipelines. Use for: product demos, how-it-works videos, onboarding videos, social explainers. Triggers: explainer video, how to make explainer, product video, demo video, video production, video script, animated explainer, product demo video, tutorial video, onboarding video, walkthrough video, video pipeline

okaris 1287

newsletter-curation

Newsletter curation with content sourcing, editorial structure, and subscriber growth strategies. Covers issue formatting, link roundups, commentary style, and sending cadence. Use for: email newsletters, link roundups, weekly digests, curated content, creator newsletters. Triggers: newsletter, email newsletter, newsletter curation, weekly digest, link roundup, curated newsletter, newsletter writing, newsletter format, subscriber growth, newsletter strategy, content curation, newsletter template

okaris 1287