マルチモーダル向けAIモデル

マルチモーダル向けの選定AIモデル269件を、追跡可能なプロバイダー情報で比較します。

269件の収録モデル

モデル一覧

この表示に 269 モデル

比較
利用可能

A flagship GPT-5.6 model listed for complex reasoning, coding and multi-step agent workflows.

コンテキスト
1.05M
入力
$2.5
出力
$15
ファイル画像テキスト
モデルを見る
利用可能

A lower-cost GPT-5.6 model for high-volume chat, classification and lightweight agent workflows.

コンテキスト
1.05M
入力
$0.20
出力
$1.2
ファイル画像テキスト
モデルを見る
利用可能

A Sonnet-class model listed for coding, agents and professional knowledge work with adaptive reasoning.

コンテキスト
1M
入力
$2
出力
$10
テキスト画像ファイル
モデルを見る
利用可能

A fast multimodal Gemini model listed for responsive agent workflows, coding and multi-step reasoning.

コンテキスト
1.05M
入力
$0.375
出力
$1.88
テキスト画像動画ファイル
モデルを見る
利用可能

A Grok model listed for coding, knowledge work and STEM tasks with text, image and file input.

コンテキスト
500K
入力
$2
出力
$6
テキスト画像ファイル
モデルを見る

Alibaba / Qwen

Qwen3.8 27B

利用可能

An open-weight vision-language model listed for coding, research, multimodal interaction and agent tasks.

コンテキスト
262K
入力
$0.45
出力
$3.2
テキスト画像動画
モデルを見る
利用可能

A dense instruction model listed for agent workflows, coding and complex professional tasks.

コンテキスト
262K
入力
$1.5
出力
$7.5
テキスト画像ファイル
モデルを見る
利用可能

A multimodal endpoint listed for workflows that combine reasoning, text and image generation.

コンテキスト
272K
入力
$8
出力
$15
画像テキストファイル
モデルを見る

Runway

Aleph 2.0

利用可能

Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change.…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画
モデルを見る

Sentence Transformers

all-MiniLM-L12-v2

利用可能

Sentence Transformers: all-MiniLM-L12-v2 is included in the AIToolly model catalog.

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る

Sentence Transformers

all-MiniLM-L6-v2

利用可能

Sentence Transformers: all-MiniLM-L6-v2 is included in the AIToolly model catalog.

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る

Sentence Transformers

all-mpnet-base-v2

利用可能

Sentence Transformers: all-mpnet-base-v2 is included in the AIToolly model catalog.

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る
利用可能

This model always redirects to the latest model in the Anthropic Claude Haiku family.

コンテキスト
200K
入力
$1
出力
$5
テキスト画像ファイル
モデルを見る

Deepgram

Aura-2

利用可能

Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages.

コンテキスト
0
入力
$30
出力
$0.00
テキスト
モデルを見る
利用可能

Auto Router (Beta) is a task-aware router from the provider catalog. It classifies each request, then routes it the most popular model for that task based on aggregate spend, filtered by your…

コンテキスト
2M
入力
不明
出力
不明
テキスト画像音声ファイル
モデルを見る
利用可能

If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3.

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る
利用可能

If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3.

コンテキスト
512
入力
$0.01
出力
$0.00
テキスト
モデルを見る

BAAI

bge-m3

利用可能

In this project, we introduce BGE-M3, which is distinguished for its versatility in Multi-Functionality, Multi-Linguality, and Multi-Granularity.

コンテキスト
8K
入力
$0.01
出力
$0.00
テキスト
モデルを見る

Google

Chirp 3

利用可能

Chirp 3 is Google's latest multilingual speech-to-text model. It offers enhanced transcription accuracy across 24 GA languages and 77+ preview languages, with support for automatic language detection, automatic punctuation, and…

コンテキスト
0
入力
$16,000
出力
$0.00
音声
モデルを見る
利用可能

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and…

コンテキスト
1M
入力
$10
出力
$50
テキスト画像ファイル
モデルを見る
利用可能

This model always redirects to the latest model in the Claude Fable family.

コンテキスト
1M
入力
$10
出力
$50
テキスト画像ファイル
モデルを見る
利用可能

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance…

コンテキスト
200K
入力
$1
出力
$5
テキスト画像ファイル
モデルを見る
利用可能

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and…

コンテキスト
200K
入力
$5
出力
$25
ファイル画像テキスト
モデルを見る
利用可能

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective…

コンテキスト
1M
入力
$5
出力
$25
テキスト画像ファイル
モデルを見る
利用可能

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on…

コンテキスト
1M
入力
$5
出力
$25
テキスト画像ファイル
モデルを見る
利用可能

Fast-mode variant of Opus 4.7 - identical capabilities with higher output speed at premium 6x pricing.

コンテキスト
1M
入力
$30
出力
$150
テキスト画像ファイル
モデルを見る
利用可能

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token…

コンテキスト
1M
入力
$5
出力
$25
テキスト画像ファイル
モデルを見る
利用可能

Fast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8.

コンテキスト
1M
入力
$10
出力
$50
テキスト画像ファイル
モデルを見る
利用可能

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis…

コンテキスト
1M
入力
$5
出力
$25
テキスト画像ファイル
モデルを見る
利用可能

Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5.

コンテキスト
1M
入力
$10
出力
$50
テキスト画像ファイル
モデルを見る
利用可能

This model always redirects to the latest model in the Claude Opus family.

コンテキスト
1M
入力
$5
出力
$25
テキスト画像ファイル
モデルを見る
利用可能

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with…

コンテキスト
1M
入力
$3
出力
$15
テキスト画像ファイル
モデルを見る
利用可能

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with…

コンテキスト
1M
入力
$3
出力
$15
テキスト画像ファイル
モデルを見る
利用可能

Mistral Codestral Embed is specially designed for code, perfect for embedding code databases, repositories, and powering coding assistants with state-of-the-art retrieval.

コンテキスト
8K
入力
$0.15
出力
$0.00
テキスト
モデルを見る

Sesame

CSM 1B

利用可能

CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversational and read-speech styles. At 1B parameters, it…

コンテキスト
4K
入力
$7
出力
$0.00
テキスト
モデルを見る
利用可能

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is…

コンテキスト
512K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Intfloat

E5-Base-v2

利用可能

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る

Intfloat

E5-Large-v2

利用可能

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022

コンテキスト
512
入力
$0.01
出力
$0.00
テキスト
モデルを見る

Perplexity

Embed V1 0.6B

利用可能

pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 0.6B parameter model targeting lightweight, low-latency…

コンテキスト
32K
入力
$0.004
出力
$0.00
テキスト
モデルを見る

Perplexity

Embed V1 4B

利用可能

pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 4B parameter model maximizing retrieval…

コンテキスト
32K
入力
$0.03
出力
$0.00
テキスト
モデルを見る
利用可能

Flux TTS is a text-to-speech model from Deepgram. It is suited for natural, expressive English speech synthesis across Deepgram's Flux voice catalog.

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト
モデルを見る

Black Forest Labs

FLUX.2 Flex

利用可能

FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing in the same unified architecture. Pricing is as follows, per the docs: We charge $0.06…

コンテキスト
67K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Black Forest Labs

FLUX.2 Klein 4B

利用可能

The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, **delivering state-of-the-art quality with end-to-end inference in as low as under a second**. Built for applications that require real-time image generation without sacrificing quality, and runs on consumer hardware, with as little as 13GB VRAM.

コンテキスト
41K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Black Forest Labs

FLUX.2 Max

利用可能

FLUX.2 [max] is the new top-tier image model from Black Forest Labs, pushing image quality, prompt understanding, and editing consistency to the highest level yet. Pricing is as follows, [per…

コンテキスト
47K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Black Forest Labs

FLUX.2 Pro

利用可能

A high-end image generation and editing model focused on frontier-level visual quality and reliability. It delivers strong prompt adherence, stable lighting, sharp textures, and consistent character/style reproduction across multi-reference inputs.…

コンテキスト
47K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Black Forest Labs

FLUX.3 Video

利用可能

FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画
モデルを見る
利用可能

The simplest way to get free inference. the provider catalog/free is a router that selects free models at random from the models available on the provider catalog. The router smartly filters for models that…

コンテキスト
200K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route…

コンテキスト
1M
入力
$5
出力
$30
テキスト画像
モデルを見る
利用可能

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool…

コンテキスト
1.05M
入力
$0.50
出力
$3
テキスト画像ファイル音声
モデルを見る
利用可能

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic…

コンテキスト
1.05M
入力
$0.25
出力
$1.5
テキスト画像動画ファイル
モデルを見る
利用可能

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across…

コンテキスト
1.05M
入力
$0.25
出力
$1.5
テキスト画像動画ファイル
モデルを見る
利用可能

Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output…

コンテキスト
33K
入力
$1
出力
$20
テキスト
モデルを見る
利用可能

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation…

コンテキスト
1.05M
入力
$2
出力
$12
音声ファイル画像テキスト
モデルを見る

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party…

コンテキスト
1.05M
入力
$2
出力
$12
テキスト音声画像動画
モデルを見る
利用可能

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution…

コンテキスト
1.05M
入力
$1.5
出力
$9
テキスト画像動画ファイル
モデルを見る
利用可能

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

コンテキスト
1.05M
入力
$0.30
出力
$2.5
テキスト画像動画ファイル
モデルを見る
利用可能

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and…

コンテキスト
1.05M
入力
$0.75
出力
$3.75
テキスト画像動画ファイル
モデルを見る
利用可能

gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark…

コンテキスト
20K
入力
$0.15
出力
$0.00
テキスト
モデルを見る
利用可能

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports…

コンテキスト
8K
入力
$0.20
出力
$0.00
テキスト画像ファイル音声
モデルを見る
利用可能

Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It…

コンテキスト
8K
入力
$0.20
出力
$0.00
テキスト画像ファイル音声
モデルを見る
利用可能

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

コンテキスト
262K
入力
$0.07
出力
$0.34
画像テキスト動画
モデルを見る
利用可能

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

コンテキスト
262K
入力
$0.10
出力
$0.34
画像テキスト動画
モデルを見る

Runway

Gen-4.5

利用可能

Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence.…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

コンテキスト
131K
入力
$0.30
出力
$0.90
画像テキスト動画
モデルを見る
利用可能

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,…

コンテキスト
203K
入力
$1.2
出力
$4
画像テキスト動画
モデルを見る
利用可能

This model always redirects to the latest model in the Google Gemini Flash family.

コンテキスト
1.05M
入力
$0.375
出力
$1.88
テキスト画像動画ファイル
モデルを見る
利用可能

This model always redirects to the latest model in the Google Gemini Pro family.

コンテキスト
1.05M
入力
$2
出力
$12
音声ファイル画像テキスト
モデルを見る

OpenAI

GPT Audio

利用可能

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced…

コンテキスト
128K
入力
$2.5
出力
$10
テキスト音声
モデルを見る
利用可能

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million…

コンテキスト
128K
入力
$0.60
出力
$2.4
テキスト音声
モデルを見る
利用可能

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates…

コンテキスト
400K
入力
$5
出力
$30
テキスト画像ファイル
モデルを見る
利用可能

OpenAI's GPT Image 1 generates and edits images via the dedicated Images API. Features accurate text rendering, transparent backgrounds, and up to 16 reference images for edits.

コンテキスト
400K
入力
$10
出力
$10
テキスト画像
モデルを見る
利用可能

A cost-efficient variant of GPT Image 1 for high-quality image generation at reduced latency and cost via OpenAI's dedicated Images API.

コンテキスト
400K
入力
$2.5
出力
$2.5
テキスト画像
モデルを見る
利用可能

OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API.

コンテキスト
400K
入力
$8
出力
$8
テキスト画像
モデルを見る
利用可能

GPT Transcribe is a high-accuracy speech-to-text model from OpenAI. It is suited for recorded audio, streamed file transcription, and committed Realtime turns, with free-form context, keyword hints, and multiple language…

コンテキスト
0
入力
$4,500
出力
$0.00
音声
モデルを見る
利用可能

GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities. It's priced per token (input and output), making it suitable for high-volume transcription workflows that…

コンテキスト
128K
入力
$1.25
出力
$5
音声
モデルを見る
利用可能

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities. It's priced per token (input and output), making it suitable for workflows that benefit from token-level billing transparency.

コンテキスト
128K
入力
$2.5
出力
$10
音声
モデルを見る
利用可能

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.…

コンテキスト
400K
入力
$0.625
出力
$5
テキスト画像
モデルを見る
利用可能

GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,…

コンテキスト
400K
入力
$10
出力
$10
画像テキストファイル
モデルを見る
利用可能

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text…

コンテキスト
400K
入力
$2.5
出力
$2
ファイル画像テキスト
モデルを見る

OpenAI

GPT-5 Pro

利用可能

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and…

コンテキスト
400K
入力
$15
出力
$120
画像テキストファイル
モデルを見る

OpenAI

GPT-5.1

利用可能

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning…

コンテキスト
400K
入力
$1.25
出力
$10
画像テキストファイル
モデルを見る
利用可能

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.…

コンテキスト
400K
入力
$1.25
出力
$10
テキスト画像
モデルを見る
利用可能

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic…

コンテキスト
400K
入力
$1.25
出力
$10
テキスト画像
モデルを見る
利用可能

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

コンテキスト
400K
入力
$0.25
出力
$2
画像テキスト
モデルを見る

OpenAI

GPT-5.2

利用可能

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly…

コンテキスト
400K
入力
$1.75
出力
$14
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on…

コンテキスト
128K
入力
$1.75
出力
$14
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning,…

コンテキスト
400K
入力
$21
出力
$168
画像テキストファイル
モデルを見る
利用可能

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.…

コンテキスト
400K
入力
$1.75
出力
$14
テキスト画像
モデルを見る
利用可能

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results…

コンテキスト
400K
入力
$1.75
出力
$14
テキスト画像ファイル
モデルを見る

OpenAI

GPT-5.4

利用可能

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for…

コンテキスト
1.05M
入力
$2.5
出力
$15
テキスト画像ファイル
モデルを見る
利用可能

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,…

コンテキスト
400K
入力
$0.75
出力
$4.5
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency…

コンテキスト
400K
入力
$0.20
出力
$1.25
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K…

コンテキスト
1.05M
入力
$30
出力
$180
テキスト画像ファイル
モデルを見る

OpenAI

GPT-5.5

利用可能

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token…

コンテキスト
1.05M
入力
$5
出力
$30
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for…

コンテキスト
1.05M
入力
$30
出力
$180
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

コンテキスト
1.05M
入力
$0.20
出力
$1.2
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

コンテキスト
1.05M
入力
$2.5
出力
$15
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic…

コンテキスト
1.05M
入力
$2
出力
$12
ファイル画像テキスト
モデルを見る
利用可能

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex tasks.

コンテキスト
1.05M
入力
$2
出力
$12
ファイル画像テキスト
モデルを見る

SpaceXAI

Grok 4.20

利用可能

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering…

コンテキスト
2M
入力
$1.25
出力
$2.5
テキスト画像ファイル
モデルを見る
利用可能

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information…

コンテキスト
2M
入力
$1.25
出力
$2.5
テキスト画像ファイル
モデルを見る

SpaceXAI

Grok 4.3

利用可能

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual…

コンテキスト
1M
入力
$1.25
出力
$2.5
テキスト画像ファイル
モデルを見る

SpaceXAI

Grok 4.5

利用可能

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

コンテキスト
500K
入力
$2
出力
$6
テキスト画像ファイル
モデルを見る
利用可能

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding…

コンテキスト
256K
入力
$1
出力
$2
テキスト画像ファイル
モデルを見る
利用可能

Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

This model always redirects to the latest Grok model from xAI.

コンテキスト
500K
入力
$2
出力
$6
テキスト画像ファイル
モデルを見る

SpaceXAI

Grok STT 1.0

利用可能

Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

コンテキスト
0
入力
$100,000
出力
$0.00
音声
モデルを見る
利用可能

Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara,…

コンテキスト
15K
入力
$15
出力
$0.00
テキスト
モデルを見る

Thenlper

GTE-Base

利用可能

General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る

Thenlper

GTE-Large

利用可能

General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning

コンテキスト
512
入力
$0.01
出力
$0.00
テキスト
モデルを見る

MiniMax

H3

利用可能

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画音声
モデルを見る

MiniMax

Hailuo 2.3

利用可能

Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Thinking Machines

Inkling

利用可能

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

コンテキスト
524K
入力
$0.95
出力
$4.05
テキスト画像音声
モデルを見る

Thinking Machines

Inkling Small

利用可能

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

コンテキスト
524K
入力
$0.45
出力
$1.2
テキスト画像音声
モデルを見る

MoonshotAI

Kimi K2.5

利用可能

📰   Tech Blog     |     📄   Paper

コンテキスト
262K
入力
$0.57
出力
$2.85
テキスト画像
モデルを見る

MoonshotAI

Kimi K2.6

利用可能

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

コンテキスト
262K
入力
$0.95
出力
$4
テキスト画像
モデルを見る

MoonshotAI

Kimi K2.7 Code

利用可能

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

コンテキスト
262K
入力
$0.71
出力
$3.5
テキスト画像
モデルを見る

MoonshotAI

Kimi K3

利用可能

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

コンテキスト
1.05M
入力
$3
出力
$15
テキスト画像動画
モデルを見る

hexgrad

Kokoro 82M

利用可能

Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese)…

コンテキスト
4K
入力
$0.62
出力
$0.00
テキスト
モデルを見る
利用可能

Krea 2 Large is Krea's high-capability image generation model, more than twice the size of Krea 2 Medium. Its lighter post-training gives images a rawer, more textured, and flexible character,…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Krea 2 Medium is Krea's balanced, cost-efficient image generation model and a practical starting point for a broad range of use cases. Its extensive post-training supports stable, consistent generations, with…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Krea 2 Medium Turbo is a distilled, speed-focused variant of Krea 2 Medium from Krea. It is designed for rapid iteration and graphic design exploration where fast generation is the…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG…

コンテキスト
10K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate…

コンテキスト
1.05M
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz…

コンテキスト
1.05M
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Microsoft

MAI-Image-2.5

利用可能

Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

コンテキスト
4K
入力
$5
出力
$0.00
テキスト画像
モデルを見る
利用可能

Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

コンテキスト
4K
入力
$5
出力
$0.00
テキスト画像
モデルを見る
利用可能

MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, with reliable transcription across 43 languages, diverse…

コンテキスト
0
入力
$360,000
出力
$0.00
音声
モデルを見る

Microsoft

MAI-Voice-2

利用可能

MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,…

コンテキスト
0
入力
$22
出力
$0.00
テキスト
モデルを見る
利用可能

MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages…

コンテキスト
0
入力
$15
出力
$0.00
テキスト
モデルを見る

Xiaomi

MiMo-V2.5

利用可能

🤗 HuggingFace  | 📰 Blog  | 🎨 Xiaomi MiMo API Platform  | 🗨️ Xiaomi MiMo Studio  |

コンテキスト
1.05M
入力
$0.14
出力
$0.28
テキスト音声画像動画
モデルを見る

MiniMax

MiniMax M3

利用可能

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

コンテキスト
524K
入力
$0.30
出力
$1.2
テキスト画像動画
モデルを見る
利用可能

The largest model in the Ministral 3 family, **Ministral 3 14B** offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language model with vision capabilities.

コンテキスト
262K
入力
$0.20
出力
$0.20
テキスト画像
モデルを見る
利用可能

The smallest model in the Ministral 3 family, **Ministral 3 3B** is a powerful, efficient tiny language model with vision capabilities.

コンテキスト
131K
入力
$0.10
出力
$0.10
テキスト画像
モデルを見る
利用可能

A balanced model in the Ministral 3 family, **Ministral 3 8B** is a powerful, efficient tiny language model with vision capabilities.

コンテキスト
262K
入力
$0.15
出力
$0.15
テキスト画像
モデルを見る
利用可能

Mistral Embed is a specialized embedding model for text data, optimized for semantic search and RAG applications. Developed by Mistral AI in late 2023, it produces 1024-dimensional vectors that effectively…

コンテキスト
8K
入力
$0.10
出力
$0.00
テキスト
モデルを見る
利用可能

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

コンテキスト
262K
入力
$0.50
出力
$1.5
テキスト画像ファイル
モデルを見る
利用可能

Mistral Small 4 is a powerful hybrid model capable of acting as both a general instruction model and a reasoning model. It unifies the capabilities of three different model families—**Instruct**, **Reasoning** (previously called Magistral), and **Devstral**—into a single, unified model.

コンテキスト
262K
入力
$0.15
出力
$0.60
テキスト画像
モデルを見る
利用可能

This model always redirects to the latest model in the MoonshotAI Kimi family.

コンテキスト
975K
入力
$2.6
出力
$13
テキスト画像動画
モデルを見る

Sentence Transformers

multi-qa-mpnet-base-dot-v1

利用可能

This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and was designed for **semantic search**. It has been trained on 215M (question, answer) pairs from diverse sources. For an introduction to semantic search, have a look at: SBERT.net - Semantic Search

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る
利用可能

Multilingual E5 Text Embeddings: A Technical Report. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, arXiv 2024

コンテキスト
512
入力
$0.01
出力
$0.00
テキスト
モデルを見る
利用可能

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon…

コンテキスト
131K
入力
$0.35
出力
$1.5
テキスト画像
モデルを見る
利用可能

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context…

コンテキスト
1.05M
入力
$1.25
出力
$4.25
テキスト画像動画ファイル
モデルを見る
利用可能

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context…

コンテキスト
1.05M
入力
$1.25
出力
$4.25
テキスト画像動画ファイル
モデルを見る

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,…

コンテキスト
33K
入力
$0.30
出力
$2.5
画像テキスト
モデルを見る

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines…

コンテキスト
66K
入力
$0.50
出力
$3
画像テキスト
モデルを見る

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation…

コンテキスト
66K
入力
$0.25
出力
$1.5
画像テキスト
モデルを見る

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and…

コンテキスト
66K
入力
$2
出力
$12
画像テキスト
モデルを見る

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and…

コンテキスト
66K
入力
$2
出力
$12
画像テキスト
モデルを見る
利用可能

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval…

コンテキスト
33K
入力
$0.00
出力
$0.00
テキスト
モデルを見る
利用可能

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family.

コンテキスト
256K
入力
$0.00
出力
$0.00
テキスト音声画像動画
モデルを見る

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice…

コンテキスト
0
入力
$3.33
出力
$0.00
音声
モデルを見る
利用可能

🤗 Model &nbsp&nbsp | &nbsp&nbsp 🔀 the provider catalog (Enjoy two weeks free starting June 9!) &nbsp&nbsp | &nbsp&nbsp 💻 Github &nbsp&nbsp | &nbsp&nbsp 🧭 ModelScope &nbsp&nbsp | &nbsp&nbsp 🚀 Nex-AGI

コンテキスト
262K
入力
$0.025
出力
$0.10
テキスト画像
モデルを見る

Nex AGI

Nex-N2-Pro

利用可能

🤗 Model &nbsp&nbsp | &nbsp&nbsp 🔀 the provider catalog (Enjoy two weeks free starting June 9!) &nbsp&nbsp | &nbsp&nbsp 💻 Github &nbsp&nbsp | &nbsp&nbsp 🧭 ModelScope &nbsp&nbsp | &nbsp&nbsp 🚀 Nex-AGI

コンテキスト
262K
入力
$0.25
出力
$1
テキスト画像
モデルを見る
利用可能

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing…

コンテキスト
1M
入力
$0.30
出力
$2.5
テキスト画像動画ファイル
モデルを見る
利用可能

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

コンテキスト
1M
入力
$2.5
出力
$12.5
テキスト画像
モデルを見る

Deepgram

Nova-3

利用可能

Deepgram Nova-3 general-purpose speech-to-text model with monolingual and multilingual transcription support.

コンテキスト
0
入力
$4,300
出力
$0.00
音声
モデルを見る
利用可能

This model always redirects to the latest model in the OpenAI GPT family.

コンテキスト
1.05M
入力
$2.5
出力
$15
ファイル画像テキスト
モデルを見る
利用可能

This model always redirects to the latest model in the OpenAI GPT Mini family.

コンテキスト
400K
入力
$0.75
出力
$4.5
ファイル画像テキスト
モデルを見る

Canopy Labs

Orpheus 3B

利用可能

Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and is suited for narration, voice assistants, and…

コンテキスト
4K
入力
$7
出力
$0.00
テキスト
モデルを見る
利用可能

Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across…

コンテキスト
0
入力
$1,500
出力
$0.00
音声
モデルを見る

Sentence Transformers

paraphrase-MiniLM-L6-v2

利用可能

Sentence Transformers: paraphrase-MiniLM-L6-v2 is included in the AIToolly model catalog.

コンテキスト
512
入力
$0.005
出力
$0.00
テキスト
モデルを見る

Perceptron

Perceptron Mk1

利用可能

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding…

コンテキスト
33K
入力
$0.15
出力
$1.5
テキスト画像動画
モデルを見る
利用可能

Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

コンテキスト
0
入力
$15
出力
$0.00
テキスト
モデルを見る
利用可能

Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.

コンテキスト
0
入力
$20
出力
$0.00
テキスト
モデルを見る
利用可能

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

コンテキスト
0
入力
$3.33
出力
$0.00
音声
モデルを見る
利用可能

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

コンテキスト
0
入力
$7.5
出力
$0.00
音声
モデルを見る
利用可能

Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages —…

コンテキスト
0
入力
$35
出力
$0.00
音声
モデルを見る
利用可能

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

コンテキスト
33K
入力
$0.02
出力
$0.00
テキスト
モデルを見る
利用可能

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

コンテキスト
33K
入力
$0.01
出力
$0.00
テキスト
モデルを見る
利用可能

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

コンテキスト
131K
入力
$0.13
出力
$0.52
テキスト画像
モデルを見る
利用可能

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

コンテキスト
131K
入力
$0.20
出力
$2.4
テキスト画像
モデルを見る
利用可能

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

コンテキスト
131K
入力
$0.104
出力
$0.416
テキスト画像
モデルを見る
利用可能

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

コンテキスト
131K
入力
$0.117
出力
$0.455
画像テキスト
モデルを見る
利用可能

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

コンテキスト
131K
入力
$0.18
出力
$2.1
画像テキスト
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.39
出力
$2.34
テキスト画像動画
モデルを見る
利用可能

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of…

コンテキスト
1M
入力
$0.26
出力
$1.56
テキスト画像動画
モデルを見る
利用可能

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This…

コンテキスト
1M
入力
$0.30
出力
$1.8
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.26
出力
$2.08
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.195
出力
$1.56
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.225
出力
$1.8
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.10
出力
$0.15
テキスト画像動画
モデルを見る
利用可能

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the…

コンテキスト
1M
入力
$0.065
出力
$0.26
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.60
出力
$3.6
テキスト画像動画
モデルを見る
利用可能

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

コンテキスト
262K
入力
$0.14
出力
$1
テキスト画像動画
モデルを見る
利用可能

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in…

コンテキスト
1M
入力
$0.1875
出力
$1.13
テキスト画像動画
モデルを見る
利用可能

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers…

コンテキスト
1M
入力
$0.325
出力
$1.95
テキスト画像動画
モデルを見る
利用可能

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world…

コンテキスト
1M
入力
$0.03
出力
$0.13
テキスト画像動画
モデルを見る
利用可能

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its…

コンテキスト
1M
入力
$0.32
出力
$1.28
テキスト画像
モデルを見る
利用可能

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,…

コンテキスト
1M
入力
$2
出力
$6
テキスト画像動画
モデルを見る

Recraft

Recraft V3

利用可能

Recraft V3 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. Supports the following image_config parameters:…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Recraft

Recraft V4

利用可能

Recraft V4 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. It delivers stronger compositional judgment,…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4 Pro is an image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios, double the resolution of…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios — double…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,…

コンテキスト
66K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

**Reka Edge** is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding, video analysis, object detection, and agentic tool-use.

コンテキスト
16K
入力
$0.10
出力
$0.10
画像テキスト動画
モデルを見る
利用可能

Riverflow V2 Fast is the fastest variant of Sourceful's Riverflow 2.0 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.0 series represents SOTA performance on image generation and…

コンテキスト
8K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Riverflow V2 Pro is the most powerful variant of Sourceful's Riverflow 2.0 lineup, best for top-tier control and perfect text rendering. The Riverflow 2.0 series represents SOTA performance on image…

コンテキスト
8K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family…

コンテキスト
33K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image…

コンテキスト
33K
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Fish Audio

S1

利用可能

S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported…

コンテキスト
0
入力
$15
出力
$0.00
テキスト
モデルを見る

Fish Audio

S2 Pro

利用可能

S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.

コンテキスト
0
入力
$15
出力
$0.00
テキスト
モデルを見る

Fish Audio

S2.1 Pro

利用可能

S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and…

コンテキスト
0
入力
$15
出力
$0.00
テキスト
モデルを見る
利用可能

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト
モデルを見る
利用可能

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,…

コンテキスト
262K
入力
$0.95
出力
$4
テキスト画像ファイル
モデルを見る

ByteDance Seed

Seed 1.6

利用可能

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

コンテキスト
262K
入力
$0.25
出力
$2
画像テキスト動画
モデルを見る

ByteDance Seed

Seed 1.6 Flash

利用可能

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of…

コンテキスト
262K
入力
$0.075
出力
$0.30
画像テキスト動画
モデルを見る

ByteDance Seed

Seed 2.1 Turbo

利用可能

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and…

コンテキスト
262K
入力
$0.50
出力
$2.5
テキスト画像動画
モデルを見る

ByteDance Seed

Seed-2.0-Code

利用可能

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude…

コンテキスト
262K
入力
$0.50
出力
$3
テキスト画像動画
モデルを見る

ByteDance Seed

Seed-2.0-Lite

利用可能

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across…

コンテキスト
262K
入力
$0.25
出力
$2
テキスト画像動画
モデルを見る

ByteDance Seed

Seed-2.0-Mini

利用可能

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,…

コンテキスト
262K
入力
$0.10
出力
$0.40
テキスト画像動画
モデルを見る
利用可能

ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

ByteDance

Seedance 2.0

利用可能

Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画音声
モデルを見る
利用可能

Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画音声
モデルを見る
利用可能

Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画音声
モデルを見る

ByteDance

Seedance 2.5

利用可能

Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像動画音声
モデルを見る

ByteDance Seed

Seedream 4.5

利用可能

Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements, especially in editing consistency, including better preservation of subject details,…

コンテキスト
4K
入力
$0.00
出力
$0.00
画像テキスト
モデルを見る

ByteDance Seed

Seedream 5.0 Lite

利用可能

Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

ByteDance Seed

Seedream 5.0 Pro

利用可能

Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Exclusively available on the the provider catalog API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based…

コンテキスト
200K
入力
$3
出力
$15
テキスト画像
モデルを見る
利用可能

OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

コンテキスト
0
入力
$100
出力
$0.00
テキスト
モデルを見る
利用可能

MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.

コンテキスト
0
入力
$60
出力
$0.00
テキスト
モデルを見る
利用可能

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters…

コンテキスト
256K
入力
$0.20
出力
$1.15
テキスト画像動画
モデルを見る
利用可能

text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two…

コンテキスト
8K
入力
$0.13
出力
$0.00
テキスト
モデルを見る
利用可能

text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces…

コンテキスト
8K
入力
$0.02
出力
$0.00
テキスト
モデルを見る

Fish Audio

Transcribe 1

利用可能

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

コンテキスト
0
入力
$100
出力
$0.00
音声
モデルを見る

Google

Veo 3.1

利用可能

Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る
利用可能

Voxtral Mini is an enhancement of Ministral 3B, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

コンテキスト
0
入力
$16.67
出力
$0.00
音声
モデルを見る
利用可能

Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings,…

コンテキスト
0
入力
$3,000
出力
$0.00
音声
モデルを見る
利用可能

Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sounding audio output.

コンテキスト
4K
入力
$16
出力
$0.00
テキスト
モデルを見る
利用可能

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

コンテキスト
32K
入力
$0.10
出力
$0.30
テキスト音声ファイル
モデルを見る
利用可能

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

コンテキスト
0
入力
$50
出力
$0.00
音声
モデルを見る

VoyageAI by MongoDB

voyage-4

利用可能

voyage-4 is a general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4 supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more…

コンテキスト
32K
入力
$0.06
出力
$0.00
テキスト
モデルを見る

VoyageAI by MongoDB

voyage-4-large

利用可能

voyage-4-large is a state-of-the-art general-purpose and multilingual embedding optimized for retrieval quality. Enabled by Matryoshka learning and quantization-aware training, voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with…

コンテキスト
32K
入力
$0.12
出力
$0.00
テキスト
モデルを見る

VoyageAI by MongoDB

voyage-4-lite

利用可能

voyage-4-lite is a lightweight, general-purpose embedding model optimized for low latency and cost. Enabled by Matryoshka learning and quantization-aware training, voyage-4-lite supports embeddings in 2048, 1024, 512, and 256 dimensions,…

コンテキスト
32K
入力
$0.02
出力
$0.00
テキスト
モデルを見る

VoyageAI by MongoDB

voyage-code-4

利用可能

voyage-code-4 is a code embedding model from Voyage AI, a MongoDB company. It is designed for coding agents and code retrieval, with Matryoshka embeddings at 2048, 1024, 512, and 256…

コンテキスト
32K
入力
$0.12
出力
$0.00
テキスト
モデルを見る

VoyageAI by MongoDB

voyage-multimodal-3.5

利用可能

voyage-multimodal-3.5 is a state-of-the-art multimodal embedding model capable of vectorizing not only text, images, and video individually, but also content that interleaves all three modalities. It delivers excellent performance for…

コンテキスト
32K
入力
$0.12
出力
$0.00
テキスト画像
モデルを見る

Alibaba

Wan 2.6

利用可能

Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio,…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

Alibaba

Wan 2.7

利用可能

Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content…

コンテキスト
0
入力
$0.00
出力
$0.00
テキスト画像
モデルを見る

OpenAI

Whisper 1

利用可能

Whisper is OpenAI's open-source automatic speech recognition model, available via API as whisper-1. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats…

コンテキスト
0
入力
$6,000
出力
$0.00
音声
モデルを見る
利用可能

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.

コンテキスト
0
入力
$7.5
出力
$0.00
音声
モデルを見る
利用可能

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.

コンテキスト
0
入力
$3.33
出力
$0.00
音声
モデルを見る
OUR METHOD

AIToollyのモデルデータについて

モデル固有の事実とプロバイダーエンドポイントの事実は分離して管理します。第三者カタログは検索とスナップショットに使い、モデル情報には公式文書と検証済みモデルカードを優先します。

01出典と確認
02確認日
03機能