Hugging Face model card
Fields: tags, gated, license, summary, languages, library name, pipeline tag
Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.
Provider endpoint: openai/whisper-large-v3
Fields: tags, gated, license, summary, languages, library name, pipeline tag
Fields: model identity
Fields: identity, description, modalities, context window, maximum output, pricing, supported parameters
Answers are generated from the same sourced model and provider facts shown above.
Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.
Model specifications and prices may vary by provider and change over time. AIToolly displays sources and verification dates so users can confirm critical details before production use.