Other Model Family

Explore 13 curated models in the Other family and compare context, API pricing, modalities and capabilities.

13models tracked

Model directory

13 models in this view

Compare

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG…

Context
10K
Input
$0.00
Output
$0.00
TextImage
View model

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval…

Context
33K
Input
$0.00
Output
$0.00
Text
View model

Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

Context
262K
Input
$0.05
Output
$0.20
Text
View model

NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows. It extends the Nemotron Nano family with integrated video+speech comprehension, Graphical User Interface (GUI), Optical Character Recognition (OCR), and speech transcription capabilities, enabling end-to-end processing of rich enterprise content such as meeting recordings, M&E assets, training videos, and complex business documents. NVIDIA Nemotron 3 Nano Omni was developed by NVIDIA as part of the Nemotron model family.

Context
256K
Input
$0.00
Output
$0.00
TextAudioImageVideo
View model
Active

> Use temperature=1.0 and top_p=0.95 across **all tasks and serving backends** — reasoning, tool calling, and general chat alike.

Context
262K
Input
$0.085
Output
$0.40
Text
View model
Active

For more details on how to deploy and use the model - see the Quick Start Guide below!

Context
512K
Input
$0.60
Output
$3.6
Text
View model

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice…

Context
0
Input
$3.33
Output
$0.00
Audio
View model

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

Context
262K
Input
$0.08
Output
$0.20
Text
View model

NVIDIA: Nemotron Nano 12B 2 VL (free) is included in the AIToolly model catalog.

Context
128K
Input
$0.00
Output
$0.00
ImageTextVideo
View model

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

Context
128K
Input
$0.00
Output
$0.00
Text
View model
Active

Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across…

Context
0
Input
$1,500
Output
$0.00
Audio
View model
OUR METHOD

How AIToolly handles model data

Model-native facts and provider endpoint facts stay separate. Third-party catalogs support discovery and provider snapshots, while official documentation and verified model cards take priority for model facts.

01Sources and verification
02Verified
03Capability