로컬 배포용 AI 모델

로컬 배포에 적합한 선별 AI 모델 20개를 추적 가능한 제공사 정보로 비교합니다.

20개의 수집 모델

모델 디렉토리

현재 20개 모델

비교

Alibaba / Qwen

Qwen3.8 27B

사용 가능

An open-weight vision-language model listed for coding, research, multimodal interaction and agent tasks.

컨텍스트
262K
입력
US$0.45
출력
US$3.2
텍스트이미지비디오
모델 보기

Alibaba / Qwen

Qwen3 Coder Next

사용 가능

An open-weight coding model listed for coding agents and local development workflows.

컨텍스트
262K
입력
US$0.12
출력
US$0.80
텍스트
모델 보기
사용 가능

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is…

컨텍스트
512K
입력
US$0.00
출력
US$0.00
텍스트이미지
모델 보기

Z.ai

GLM 5

사용 가능

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

컨텍스트
198K
입력
US$0.60
출력
US$1.92
텍스트
모델 보기
사용 가능

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.

컨텍스트
131K
입력
US$0.075
출력
US$0.30
텍스트
모델 보기

MiniMax

H3

사용 가능

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and…

컨텍스트
0
입력
US$0.00
출력
US$0.00
텍스트이미지비디오오디오
모델 보기

Thinking Machines

Inkling

사용 가능

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

컨텍스트
524K
입력
US$0.95
출력
US$4.05
텍스트이미지오디오
모델 보기

Thinking Machines

Inkling Small

사용 가능

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

컨텍스트
524K
입력
US$0.45
출력
US$1.2
텍스트이미지오디오
모델 보기

MoonshotAI

Kimi K3

사용 가능

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

컨텍스트
1.05M
입력
US$3
출력
US$15
텍스트이미지비디오
모델 보기

hexgrad

Kokoro 82M

사용 가능

Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese)…

컨텍스트
4K
입력
US$0.62
출력
US$0.00
텍스트
모델 보기
사용 가능

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon…

컨텍스트
131K
입력
US$0.35
출력
US$1.5
텍스트이미지
모델 보기
사용 가능

🤗 Model &nbsp&nbsp | &nbsp&nbsp 🔀 the provider catalog (Enjoy two weeks free starting June 9!) &nbsp&nbsp | &nbsp&nbsp 💻 Github &nbsp&nbsp | &nbsp&nbsp 🧭 ModelScope &nbsp&nbsp | &nbsp&nbsp 🚀 Nex-AGI

컨텍스트
262K
입력
US$0.025
출력
US$0.10
텍스트이미지
모델 보기
사용 가능

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and…

컨텍스트
1M
입력
US$0.65
출력
US$3.25
텍스트
모델 보기
사용 가능

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

컨텍스트
131K
입력
US$0.21
출력
US$1.9
텍스트이미지
모델 보기
사용 가능

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

컨텍스트
262K
입력
US$0.14
출력
US$1
텍스트이미지비디오
모델 보기
사용 가능

> [!Note] > This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > > These artifacts are compatible with vLLM, SGLang, TokenSpeed, etc.

컨텍스트
1M
입력
US$2
출력
US$6
텍스트
모델 보기
사용 가능

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.…

컨텍스트
262K
입력
US$0.10
출력
US$0.30
텍스트
모델 보기
사용 가능

Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. Built on Trinity-Large-Base and post-trained with extended chain-of-thought reasoning and agentic RL, Trinity-Large-Thinking delivers state-of-the-art performance on agentic benchmarks while maintaining strong general capabilities.

컨텍스트
262K
입력
US$0.22
출력
US$0.85
텍스트
모델 보기

OpenAI

Whisper 1

사용 가능

Whisper is OpenAI's open-source automatic speech recognition model, available via API as whisper-1. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats…

컨텍스트
0
입력
US$6,000
출력
US$0.00
오디오
모델 보기
사용 가능

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.

컨텍스트
0
입력
US$7.5
출력
US$0.00
오디오
모델 보기
OUR METHOD

AIToolly의 모델 데이터 처리 방식

모델 고유 정보와 제공사 엔드포인트 정보를 분리합니다. 서드파티 카탈로그는 탐색과 스냅샷에 사용하며 모델 정보는 공식 문서와 검증된 모델 카드를 우선합니다.

01출처 및 검증
02검증일
03기능