ngrok AI Gateway
ngrok AI Gateway: Unified AI Gateway
Use ngrok AI Gateway to route public and self-hosted models. Features include unified API access, LLM observability, key management, and failover.
2026-08-06
--K
Structured product information
ngrok AI Gateway: At a glance
Last checked: · Official source
Key features
Unified AI Gateway
A hosted gateway that brings public providers, custom endpoints, and self-hosted models into a single interface with built-in observability and access control.
Private Local LLM Connectivity
Connects to self-hosted models without requiring public IPs or inbound ports, maintaining private connectivity.
Key Management (BYOK)
Allows users to bring their own API keys from providers like OpenAI and Anthropic to manage them in one place.
Scoped Access Control
Provides separate access keys for different applications or developers, with permissions to restrict access to specific providers and models.
Automated Failover and Retries
Automatically reroutes requests to healthy alternatives if a model or key fails and performs instant retries without code changes.
Unified Observability
Tracks tokens, latency, and errors across all routed calls, providing visibility into usage by app, developer, or model.
Best for
- Developers using multiple AI model providers
- Teams running self-hosted or local LLMs
- Organizations requiring centralized access control and observability for AI infrastructure
Supported platforms
- OpenAI SDK
- Anthropic SDK
- Vercel AI SDK
- Terraform
- CLI
Privacy & security
- Private connectivity for self-hosted models without public IPs or inbound ports
- Scoped access keys for granular permission management
ngrok AI Gateway Product Information
The ngrok AI Gateway is a hosted platform designed to centralize the management of public AI providers, custom endpoints, and self-hosted models. By providing a single interface, it allows developers to consolidate infrastructure that typically requires multiple integrations. The service operates on the principle of "One gateway, every model," enabling teams to interact with various large language models (LLMs) through a unified AI API. This approach simplifies the technical overhead of managing disparate provider requirements and authentication methods.
Managing AI Model Routing with an AI Gateway
Centralized AI model routing allows developers to switch between different providers or models without significant code modifications. The ngrok AI Gateway uses a "One URL, any SDK" approach, where users change their baseURL to a single ngrok endpoint and update their API key to begin routing traffic. This setup is compatible with the OpenAI SDK, Anthropic SDK, and Vercel AI SDK.
Through a feature known as Bring Your Own Keys (BYOK), organizations can use their existing credentials from providers like OpenAI and Anthropic. This ensures that users maintain their direct relationships and established rates with those providers while managing all keys in one location. The gateway acts as a programmable layer that can be configured via APIs, Terraform, the ngrok CLI, or other custom tooling, making it adaptable to existing DevOps workflows.
Self-Hosted LLM Connectivity and Private Access
For organizations running their own infrastructure, the gateway provides specialized self-hosted LLM connectivity. This feature allows users to connect to local LLMs that are reachable from the gateway without the need to configure public IP addresses or open inbound ports. By maintaining private connectivity, the gateway treats self-hosted models with the same routing logic as public providers.
This capability is particularly relevant for teams that need to connect local LLMs to public applications while adhering to strict networking requirements. The gateway handles the underlying connection, ensuring that internal models remain shielded from the public internet while still being accessible to authorized applications through the unified interface.
LLM Observability and Usage Tracking
Effective management of AI infrastructure requires detailed visibility into how models are being utilized. The ngrok AI Gateway includes built-in LLM observability to track performance metrics that are often fragmented across different provider dashboards. It monitors token usage, latency, and error rates across every routed call.
These analytics are rolled up to provide a clear view of usage patterns by specific applications, individual developers, or specific models. Instead of relying on separate provider-specific dashboards that may not distinguish between different internal projects, the gateway provides a consolidated view of the entire AI stack's performance and consumption.
Resilience Features and Access Control
To maintain application uptime, the gateway includes automated failover and retry mechanisms. If a specific model or provider key slows down or fails, the system can reroute before users notice by directing requests to a pre-defined healthy alternative. Additionally, the gateway performs instant retries for failed requests, allowing applications to continue running without requiring manual error-handling logic within the application code.
Security is managed through scoped access control. Administrators can issue separate access keys to different developers or applications and set granular permissions. These permissions define exactly which providers and models a specific key is authorized to call, preventing the need to share master API keys across an entire organization. This structure ensures that access is restricted to the specific resources required for a given task or environment.
Frequently Asked Questions
How do I integrate the ngrok AI Gateway with my existing AI applications?
You can integrate by changing your baseURL to https://gateway.ngrok.ai and swapping your API key. It is compatible with the OpenAI SDK, Anthropic SDK, and Vercel AI SDK.
Can I use the ngrok AI Gateway with models I run myself?
Yes, you can route to any local LLM reachable from the AI Gateway. It uses private connectivity, so you do not need to manage public IPs or inbound ports.
How does the gateway handle model failures or errors?
The gateway automatically reroutes requests to a defined healthy alternative if a model or key fails. It also performs instant retries for failed requests so that applications can continue running without manual error handling in the code.
How can I control which developers or apps have access to specific models?
You can issue separate access keys to each developer or application and configure which specific providers and models each key is permitted to call.
What does the ngrok AI Gateway do?
Unified AI Gateway: A hosted gateway that brings public providers, custom endpoints, and self-hosted models into a single interface with built-in observability and access control. Private Local LLM Connectivity: Connects to self-hosted models without requiring public IPs or inbound ports, maintaining private connectivity. Key Management (BYOK): Allows users to bring their own API keys from providers like OpenAI and Anthropic to manage them in one place. Scoped Access Control: Provides separate access keys for different applications or developers, with permissions to restrict access to specific providers and models. Automated Failover and Retries: Automatically reroutes requests to healthy alternatives if a model or key fails and performs instant retries without code changes. Unified Observability: Tracks tokens, latency, and errors across all routed calls, providing visibility into usage by app, developer, or model.








