가성비용 AI 모델

가성비에 적합한 선별 AI 모델 154개를 추적 가능한 제공사 정보로 비교합니다.

154개의 수집 모델

모델 디렉토리

현재 154개 모델

비교
사용 가능

A lower-cost GPT-5.6 model for high-volume chat, classification and lightweight agent workflows.

컨텍스트
1.05M
입력
US$0.20
출력
US$1.2
파일이미지텍스트
모델 보기
사용 가능

A fast multimodal Gemini model listed for responsive agent workflows, coding and multi-step reasoning.

컨텍스트
1.05M
입력
US$0.375
출력
US$1.88
텍스트이미지비디오파일
모델 보기

Alibaba / Qwen

Qwen3.8 27B

사용 가능

An open-weight vision-language model listed for coding, research, multimodal interaction and agent tasks.

컨텍스트
262K
입력
US$0.45
출력
US$3.2
텍스트이미지비디오
모델 보기

Alibaba / Qwen

Qwen3 Coder Next

사용 가능

An open-weight coding model listed for coding agents and local development workflows.

컨텍스트
262K
입력
US$0.12
출력
US$0.80
텍스트
모델 보기

Runway

Aleph 2.0

사용 가능

Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change.…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오
모델 보기

Sentence Transformers

all-MiniLM-L12-v2

사용 가능

Sentence Transformers: all-MiniLM-L12-v2 is included in the AIToolly model catalog.

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기

Sentence Transformers

all-MiniLM-L6-v2

사용 가능

Sentence Transformers: all-MiniLM-L6-v2 is included in the AIToolly model catalog.

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기

Sentence Transformers

all-mpnet-base-v2

사용 가능

Sentence Transformers: all-mpnet-base-v2 is included in the AIToolly model catalog.

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기
사용 가능

If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3.

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기
사용 가능

If you are looking for a model that supports more languages, longer texts, and other retrieval methods, you can try using bge-m3.

컨텍스트
512
입력
US$0.01
출력
US$0.00
텍스트
모델 보기

BAAI

bge-m3

사용 가능

In this project, we introduce BGE-M3, which is distinguished for its versatility in Multi-Functionality, Multi-Linguality, and Multi-Granularity.

컨텍스트
8K
입력
US$0.01
출력
US$0.00
텍스트
모델 보기
사용 가능

Mistral Codestral Embed is specially designed for code, perfect for embedding code databases, repositories, and powering coding assistants with state-of-the-art retrieval.

컨텍스트
8K
입력
US$0.15
출력
US$0.00
텍스트
모델 보기
사용 가능

> I have to praise this model for good focus. I said earlier that it still remembers it at 12K. I think my personal evaluation of it has already beaten the rest.

컨텍스트
131K
입력
US$0.30
출력
US$0.50
텍스트
모델 보기
사용 가능

We introduce **DeepSeek-V3.2**, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

컨텍스트
164K
입력
US$0.269
출력
US$0.40
텍스트
모델 보기
사용 가능

We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.

컨텍스트
164K
입력
US$0.27
출력
US$0.41
텍스트
모델 보기
사용 가능

We present a preview version of **DeepSeek-V4** series, including two strong Mixture-of-Experts (MoE) language models — **DeepSeek-V4-Pro** with 1.6T parameters (49B activated) and **DeepSeek-V4-Flash** with 284B parameters (13B activated) — both supporting a context length of **one million tokens**.

컨텍스트
1.02M
입력
US$0.0826
출력
US$0.1652
텍스트
모델 보기
사용 가능

**DeepSeek-V4-Flash-0731** is the official release of **DeepSeek-V4-Flash**, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

컨텍스트
1.05M
입력
US$0.14
출력
US$0.28
텍스트
모델 보기
사용 가능

This model always redirects to the latest model in the DeepSeek V4 Flash family.

컨텍스트
1.02M
입력
US$0.0786
출력
US$0.1572
텍스트
모델 보기
사용 가능

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is…

컨텍스트
512K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Intfloat

E5-Base-v2

사용 가능

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기

Intfloat

E5-Large-v2

사용 가능

Text Embeddings by Weakly-Supervised Contrastive Pre-training. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei, arXiv 2022

컨텍스트
512
입력
US$0.01
출력
US$0.00
텍스트
모델 보기

Perplexity

Embed V1 0.6B

사용 가능

pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 0.6B parameter model targeting lightweight, low-latency…

컨텍스트
32K
입력
US$0.004
출력
US$0.00
텍스트
모델 보기

Perplexity

Embed V1 4B

사용 가능

pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized for standard dense text retrieval with the 4B parameter model maximizing retrieval…

컨텍스트
32K
입력
US$0.03
출력
US$0.00
텍스트
모델 보기
사용 가능

Flux TTS is a text-to-speech model from Deepgram. It is suited for natural, expressive English speech synthesis across Deepgram's Flux voice catalog.

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

Black Forest Labs

FLUX.2 Flex

사용 가능

FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing in the same unified architecture. Pricing is as follows, per the docs: We charge $0.06…

컨텍스트
67K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Black Forest Labs

FLUX.2 Klein 4B

사용 가능

The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, **delivering state-of-the-art quality with end-to-end inference in as low as under a second**. Built for applications that require real-time image generation without sacrificing quality, and runs on consumer hardware, with as little as 13GB VRAM.

컨텍스트
41K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Black Forest Labs

FLUX.2 Max

사용 가능

FLUX.2 [max] is the new top-tier image model from Black Forest Labs, pushing image quality, prompt understanding, and editing consistency to the highest level yet. Pricing is as follows, [per…

컨텍스트
47K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Black Forest Labs

FLUX.2 Pro

사용 가능

A high-end image generation and editing model focused on frontier-level visual quality and reliability. It delivers strong prompt adherence, stable lighting, sharp textures, and consistent character/style reproduction across multi-reference inputs.…

컨텍스트
47K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Black Forest Labs

FLUX.3 Video

사용 가능

FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오
모델 보기
사용 가능

The simplest way to get free inference. the provider catalog/free is a router that selects free models at random from the models available on the provider catalog. The router smartly filters for models that…

컨텍스트
200K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model has consistently held a top spot on the Massive Text Embedding Benchmark…

컨텍스트
20K
입력
US$0.15
출력
US$0.00
텍스트
모델 보기
사용 가능

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports…

컨텍스트
8K
입력
US$0.20
출력
US$0.00
텍스트이미지파일오디오
모델 보기
사용 가능

Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It…

컨텍스트
8K
입력
US$0.20
출력
US$0.00
텍스트이미지파일오디오
모델 보기
사용 가능

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

컨텍스트
262K
입력
US$0.07
출력
US$0.34
이미지텍스트비디오
모델 보기
사용 가능

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

컨텍스트
262K
입력
US$0.10
출력
US$0.34
이미지텍스트비디오
모델 보기

Runway

Gen-4.5

사용 가능

Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence.…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

컨텍스트
203K
입력
US$0.06
출력
US$0.40
텍스트
모델 보기
사용 가능

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.

컨텍스트
131K
입력
US$0.075
출력
US$0.30
텍스트
모델 보기
사용 가능

📣 **Update [10-07-2025]:** Added a *default system prompt* to the chat template to guide the model towards more *professional, accurate, and safe* responses.

컨텍스트
131K
입력
US$0.017
출력
US$0.112
텍스트
모델 보기
사용 가능

**Model Summary:** Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from *Granite-4.1-8B-Base* using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.

컨텍스트
131K
입력
US$0.05
출력
US$0.10
텍스트
모델 보기
사용 가능

Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Thenlper

GTE-Base

사용 가능

General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기

Thenlper

GTE-Large

사용 가능

General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning

컨텍스트
512
입력
US$0.01
출력
US$0.00
텍스트
모델 보기

MiniMax

H3

사용 가능

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기

MiniMax

Hailuo 2.3

사용 가능

Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Hermes 4 70B is a frontier, hybrid-mode **reasoning** model based on Llama-3.1-70B by Nous Research that is aligned to **you**.

컨텍스트
131K
입력
US$0.13
출력
US$0.40
텍스트
모델 보기

Tencent

Hy3

사용 가능

**Hy3** is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.

컨텍스트
262K
입력
US$0.132
출력
US$0.528
텍스트
모델 보기
사용 가능

**Hy3 preview** is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Hy3 preview is the first model trained on our rebuilt infrastructure, and the strongest we've shipped so far. It improves significantly on complex reasoning, instruction following, context learning, coding, and agent tasks.

컨텍스트
262K
입력
US$0.18
출력
US$0.60
텍스트
모델 보기
사용 가능

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make…

컨텍스트
256K
입력
US$0.15
출력
US$0.60
텍스트
모델 보기

hexgrad

Kokoro 82M

사용 가능

Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese)…

컨텍스트
4K
입력
US$0.62
출력
US$0.00
텍스트
모델 보기
사용 가능

Krea 2 Large is Krea's high-capability image generation model, more than twice the size of Krea 2 Medium. Its lighter post-training gives images a rawer, more textured, and flexible character,…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Krea 2 Medium is Krea's balanced, cost-efficient image generation model and a practical starting point for a broad range of use cases. Its extensive post-training supports stable, consistent generations, with…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Krea 2 Medium Turbo is a distilled, speed-focused variant of Krea 2 Medium from Krea. It is designed for rapid iteration and graphic design exploration where fast generation is the…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Poolside

Laguna S 2.1

사용 가능

Laguna S 2.1 is a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work. It sits between Laguna XS 2.1 (33B-A3B) and Laguna M.1 (225B-A23B) in the Laguna series and shares the family recipe: a token-choice router with softplus gating over 256 routed experts plus one shared expert, grouped-query attention, and interleaved full/sliding-window attention.

컨텍스트
1.05M
입력
US$0.09
출력
US$0.18
텍스트
모델 보기
사용 가능

Laguna XS 2.1 is a 33B total parameter Mixture-of-Experts model with 3B activated parameters per token designed for agentic coding and long-horizon work on a local machine. This model is an upgraded version of our Laguna XS.2 model with a +5.4% jump on SWE-bench Multilingual as well as stronger performance on terminal-style tasks.

컨텍스트
262K
입력
US$0.06
출력
US$0.12
텍스트
모델 보기
사용 가능

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or…

컨텍스트
128K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

inclusionAI

Ling-2.6-1T

사용 가능

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast…

컨텍스트
262K
입력
US$0.075
출력
US$0.625
텍스트
모델 보기

inclusionAI

Ling-2.6-flash

사용 가능

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency.…

컨텍스트
262K
입력
US$0.01
출력
US$0.03
텍스트
모델 보기
사용 가능

🤗 Hugging Face    |   🤖 ModelScope    |   🐙 the provider catalog   

컨텍스트
262K
입력
US$0.021
출력
US$0.063
텍스트
모델 보기

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG…

컨텍스트
10K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate…

컨텍스트
1.05M
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz…

컨텍스트
1.05M
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Inception

Mercury 2

사용 가능

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving…

컨텍스트
128K
입력
US$0.25
출력
US$0.75
텍스트
모델 보기

Xiaomi

MiMo-V2.5

사용 가능

🤗 HuggingFace  | 📰 Blog  | 🎨 Xiaomi MiMo API Platform  | 🗨️ Xiaomi MiMo Studio  |

컨텍스트
1.05M
입력
US$0.14
출력
US$0.28
텍스트오디오이미지비디오
모델 보기
사용 가능

The largest model in the Ministral 3 family, **Ministral 3 14B** offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language model with vision capabilities.

컨텍스트
262K
입력
US$0.20
출력
US$0.20
텍스트이미지
모델 보기
사용 가능

The smallest model in the Ministral 3 family, **Ministral 3 3B** is a powerful, efficient tiny language model with vision capabilities.

컨텍스트
131K
입력
US$0.10
출력
US$0.10
텍스트이미지
모델 보기
사용 가능

A balanced model in the Ministral 3 family, **Ministral 3 8B** is a powerful, efficient tiny language model with vision capabilities.

컨텍스트
262K
입력
US$0.15
출력
US$0.15
텍스트이미지
모델 보기
사용 가능

Mistral Embed is a specialized embedding model for text data, optimized for semantic search and RAG applications. Developed by Mistral AI in late 2023, it produces 1024-dimensional vectors that effectively…

컨텍스트
8K
입력
US$0.10
출력
US$0.00
텍스트
모델 보기
사용 가능

Mistral Small 4 is a powerful hybrid model capable of acting as both a general instruction model and a reasoning model. It unifies the capabilities of three different model families—**Instruct**, **Reasoning** (previously called Magistral), and **Devstral**—into a single, unified model.

컨텍스트
262K
입력
US$0.15
출력
US$0.60
텍스트이미지
모델 보기

Sentence Transformers

multi-qa-mpnet-base-dot-v1

사용 가능

This is a sentence-transformers model: It maps sentences & paragraphs to a 768 dimensional dense vector space and was designed for **semantic search**. It has been trained on 215M (question, answer) pairs from diverse sources. For an introduction to semantic search, have a look at: SBERT.net - Semantic Search

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기
사용 가능

Multilingual E5 Text Embeddings: A Technical Report. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, arXiv 2024

컨텍스트
512
입력
US$0.01
출력
US$0.00
텍스트
모델 보기
사용 가능

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval…

컨텍스트
33K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

컨텍스트
262K
입력
US$0.05
출력
US$0.20
텍스트
모델 보기
사용 가능

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family.

컨텍스트
256K
입력
US$0.00
출력
US$0.00
텍스트오디오이미지비디오
모델 보기
사용 가능

> Use temperature=1.0 and top_p=0.95 across **all tasks and serving backends** — reasoning, tool calling, and general chat alike.

컨텍스트
262K
입력
US$0.085
출력
US$0.40
텍스트
모델 보기
사용 가능

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

컨텍스트
262K
입력
US$0.08
출력
US$0.20
텍스트
모델 보기
사용 가능

NVIDIA: Nemotron Nano 12B 2 VL (free) is included in the AIToolly model catalog.

컨텍스트
128K
입력
US$0.00
출력
US$0.00
이미지텍스트비디오
모델 보기
사용 가능

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

컨텍스트
128K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

🤗 Model &nbsp&nbsp | &nbsp&nbsp 🔀 the provider catalog (Enjoy two weeks free starting June 9!) &nbsp&nbsp | &nbsp&nbsp 💻 Github &nbsp&nbsp | &nbsp&nbsp 🧭 ModelScope &nbsp&nbsp | &nbsp&nbsp 🚀 Nex-AGI

컨텍스트
262K
입력
US$0.025
출력
US$0.10
텍스트이미지
모델 보기
사용 가능

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized…

컨텍스트
256K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

We introduce Olmo 3, a new family of 7B and 32B models both Instruct and Think variants. Long chain-of-thought thinking improves reasoning tasks like math and coding.

컨텍스트
66K
입력
US$0.15
출력
US$0.50
텍스트
모델 보기

Sentence Transformers

paraphrase-MiniLM-L6-v2

사용 가능

Sentence Transformers: paraphrase-MiniLM-L6-v2 is included in the AIToolly model catalog.

컨텍스트
512
입력
US$0.005
출력
US$0.00
텍스트
모델 보기
사용 가능

Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

컨텍스트
33K
입력
US$0.02
출력
US$0.00
텍스트
모델 보기
사용 가능

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

컨텍스트
33K
입력
US$0.01
출력
US$0.00
텍스트
모델 보기
사용 가능

Qwen3 Reranker 8B is a text reranking model from Alibaba Cloud built on the Qwen3 architecture. It evaluates query-document pairs to produce relevance scores for use in retrieval and RAG…

컨텍스트
41K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

컨텍스트
131K
입력
US$0.13
출력
US$0.52
텍스트이미지
모델 보기
사용 가능

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

컨텍스트
131K
입력
US$0.104
출력
US$0.416
텍스트이미지
모델 보기
사용 가능

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

컨텍스트
131K
입력
US$0.117
출력
US$0.455
이미지텍스트
모델 보기
사용 가능

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

컨텍스트
262K
입력
US$0.10
출력
US$0.15
텍스트이미지비디오
모델 보기
사용 가능

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the…

컨텍스트
1M
입력
US$0.065
출력
US$0.26
텍스트이미지비디오
모델 보기
사용 가능

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world…

컨텍스트
1M
입력
US$0.03
출력
US$0.13
텍스트이미지비디오
모델 보기

Recraft

Recraft V3

사용 가능

Recraft V3 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. Supports the following image_config parameters:…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Recraft

Recraft V4

사용 가능

Recraft V4 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. It delivers stronger compositional judgment,…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4 Pro is an image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios, double the resolution of…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios — double…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,…

컨텍스트
66K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

**Reka Edge** is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding, video analysis, object detection, and agentic tool-use.

컨텍스트
16K
입력
US$0.10
출력
US$0.10
이미지텍스트비디오
모델 보기
사용 가능

Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing…

컨텍스트
33K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing…

컨텍스트
33K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기
사용 가능

Rerank v3.5 is designed to reorder search results for improved relevance. It supports multi-aspect and semi-structured data reranking over 100+ languages. Ideal for refining results from semantic or keyword search…

컨텍스트
4K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

rerank-2.5

사용 가능

rerank-2.5 is a cutting-edge reranker optimized for quality, delivering a 7.94% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 12.70%…

컨텍스트
32K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

rerank-2.5-lite

사용 가능

rerank-2.5-lite is a reranker optimized for both latency and quality, delivering a 7.16% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5…

컨텍스트
32K
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

inclusionAI

Ring-2.6-1T

사용 가능

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool…

컨텍스트
262K
입력
US$0.075
출력
US$0.625
텍스트
모델 보기
사용 가능

Riverflow V2 Fast is the fastest variant of Sourceful's Riverflow 2.0 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.0 series represents SOTA performance on image generation and…

컨텍스트
8K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Riverflow V2 Pro is the most powerful variant of Sourceful's Riverflow 2.0 lineup, best for top-tier control and perfect text rendering. The Riverflow 2.0 series represents SOTA performance on image…

컨텍스트
8K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family…

컨텍스트
33K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image…

컨텍스트
33K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트
모델 보기

ByteDance Seed

Seed 1.6 Flash

사용 가능

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of…

컨텍스트
262K
입력
US$0.075
출력
US$0.30
이미지텍스트비디오
모델 보기

ByteDance Seed

Seed-2.0-Mini

사용 가능

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,…

컨텍스트
262K
입력
US$0.10
출력
US$0.40
텍스트이미지비디오
모델 보기
사용 가능

ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

ByteDance

Seedance 2.0

사용 가능

Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기
사용 가능

Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기
사용 가능

Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기

ByteDance

Seedance 2.5

사용 가능

Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기

ByteDance Seed

Seedream 4.5

사용 가능

Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements, especially in editing consistency, including better preservation of subject details,…

컨텍스트
4K
입력
US$0.00
출력
US$0.00
이미지텍스트
모델 보기

ByteDance Seed

Seedream 5.0 Lite

사용 가능

Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

ByteDance Seed

Seedream 5.0 Pro

사용 가능

Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized…

컨텍스트
131K
입력
US$0.15
출력
US$0.60
텍스트
모델 보기
사용 가능

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive…

컨텍스트
524K
입력
US$0.03
출력
US$0.12
텍스트
모델 보기
사용 가능

OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.…

컨텍스트
262K
입력
US$0.10
출력
US$0.30
텍스트
모델 보기
사용 가능

text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of text that can be used to measure the relatedness between two…

컨텍스트
8K
입력
US$0.13
출력
US$0.00
텍스트
모델 보기
사용 가능

text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces…

컨텍스트
8K
입력
US$0.02
출력
US$0.00
텍스트
모델 보기
사용 가능

text-embedding-ada-002 is OpenAI's legacy text embedding model.

컨텍스트
8K
입력
US$0.10
출력
US$0.00
텍스트
모델 보기

Google

Veo 3.1

사용 가능

Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
사용 가능

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding.

컨텍스트
32K
입력
US$0.10
출력
US$0.30
텍스트오디오파일
모델 보기

VoyageAI by MongoDB

voyage-4

사용 가능

voyage-4 is a general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4 supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more…

컨텍스트
32K
입력
US$0.06
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

voyage-4-large

사용 가능

voyage-4-large is a state-of-the-art general-purpose and multilingual embedding optimized for retrieval quality. Enabled by Matryoshka learning and quantization-aware training, voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with…

컨텍스트
32K
입력
US$0.12
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

voyage-4-lite

사용 가능

voyage-4-lite is a lightweight, general-purpose embedding model optimized for low latency and cost. Enabled by Matryoshka learning and quantization-aware training, voyage-4-lite supports embeddings in 2048, 1024, 512, and 256 dimensions,…

컨텍스트
32K
입력
US$0.02
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

voyage-code-4

사용 가능

voyage-code-4 is a code embedding model from Voyage AI, a MongoDB company. It is designed for coding agents and code retrieval, with Matryoshka embeddings at 2048, 1024, 512, and 256…

컨텍스트
32K
입력
US$0.12
출력
US$0.00
텍스트
모델 보기

VoyageAI by MongoDB

voyage-multimodal-3.5

사용 가능

voyage-multimodal-3.5 is a state-of-the-art multimodal embedding model capable of vectorizing not only text, images, and video individually, but also content that interleaves all three modalities. It delivers excellent performance for…

컨텍스트
32K
입력
US$0.12
출력
US$0.00
텍스트이미지
모델 보기

Alibaba

Wan 2.6

사용 가능

Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio,…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Alibaba

Wan 2.7

사용 가능

Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기
OUR METHOD

AIToolly의 모델 데이터 처리 방식

모델 고유 정보와 제공사 엔드포인트 정보를 분리합니다. 서드파티 카탈로그는 탐색과 스냅샷에 사용하며 모델 정보는 공식 문서와 검증된 모델 카드를 우선합니다.

01출처 및 검증
02검증일
03기능