Move from prototype to production agents in hours, not months. Start by connecting your environment of choice—public cloud, your VPC, or on‑prem clusters—and create a workspace for each team. Add credentials and secrets once, then route traffic through the unified control plane to whichever model or backend you prefer. Set organization policies up front: define data residency by region, put guardrails on rate limits and token spend, and choose default fallbacks if a provider degrades. With these basics in place, your developers get a single endpoint for consistent auth, observability, and cost governance across every project.
Next, assemble your agent. Register tools and internal APIs with clear input/output contracts so the system can validate calls before they run. Plug in memory: connect vector or relational stores, pick retention rules, and set which facts an agent may write or read. Configure planning behavior—single step for simple tasks or multi‑hop for research, support, or workflow automation. Manage prompts like code: branch, version, tag, and test them against a replay set. Run A/B evaluations on safety, latency, and task success before promoting a new prompt or toolchain to production, and enable automatic rollback if performance drifts.
When you’re ready to ship, choose the runtime that fits: GPUs for heavy LLMs, CPUs for lightweight APIs, or a mix that autos-scales by load and budget. Use staged rollouts—canary or blue‑green—to release safely. Monitor live metrics in one place: latency, tokens, cost per request, tool error rates, and end‑to‑end traces showing each step the agent took. Turn on streaming for responsive UIs, caching for repeated queries, and circuit breakers with retries to handle flaky dependencies. Set latency targets and configure model fallbacks by policy so user experience remains stable even under provider incidents.
Operate at enterprise scale without losing control. Assign role‑based permissions for teams and service accounts, integrate SSO, and capture immutable audit trails for every prompt change, tool invocation, and data access. Enforce privacy rules like field‑level redaction, data retention windows, and geographic storage constraints to satisfy regulatory needs such as SOC 2, HIPAA, or GDPR. Typical workflows include: a support copilot that pulls answers from your knowledge base while respecting entitlements; a coding assistant wired to internal repos and issue trackers; a content studio that auto‑drafts assets with human approval; or an analytics bot that plans structured queries and explains the results. Real‑time cost dashboards, usage quotas, and 24/7 SLA support keep your operations predictable as you scale.
Developer
Free
Requests per month - 50K
Users - 3
Al Gateway
<ul>
<li>Universal AΡΙ
RBAC on models
Virtual models
Self-hosted models
Playground
Observability
<ul>
<li>Logs
Traces
Custom Metadata
Custom pricing / model
Cost per team/user/model/application
Metadata filtering
MCP Gateway
<ul>
<li>MCP Servers - Register up to 5
Tool calls per month - 50K
RBAC on MCPS
Metrics
Logs
Support for advanced authentication
Self-hosted MCPs
Prompt Management
<ul>
<li>Number of saved prompts - Up to 10
Versioning & Variables
Security & Authentication
<ul>
<li>SOC2
GDPR, HIPAA Compliance Certificates
Deployment modes
<ul>
<li>SaaS
Deployment customization
<ul>
<li>Gitops (Infrastructure as code)
Support
<ul>
<li>Support - Community Support
Pro
$499.00 per month
Requests per month - 1M
Users - 10
Al Gateway
<ul>
<li>Universal AΡΙ
RBAC on models
Self-hosted models
Playground
Simple caching
Semantic Caching
Control Center
<ul>
<li>Weight-based Routing
Latency-based Routing
Priority-based Routing
Fallbacks - With advanced features
Budget limiting
Rate limiting
Observability
<ul>
<li>Logs
Traces
Custom Metadata
Custom pricing / model
Cost per team/user/model/application
Metadata filtering
MCP Gateway
<ul>
<li>MCP Servers - Register up to 25
Tool calls per month - 1M
RBAC on MCPS
Metrics
Logs
Support for advanced authentication
Self-hosted MCPs
Prompt Management
<ul>
<li>Number of saved prompts - Unlimited
Versioning & Variables
Guardrails
<ul>
<li>Partner Guardrails integration
Security & Authentication
<ul>
<li>Role based access control
SOC2
Deployment modes
<ul>
<li>SaaS
Deployment customization
<ul>
<li>Gitops (Infrastructure as code)
Support
<ul>
<li>Support - Production
SLA - Standard SLA
Enterprise
Custom
Requests per month - Custom 10M Plus per month
Users - Custom
Al Gateway
<ul>
<li>Universal AΡΙ
RBAC on models
Virtual models
Self-hosted models
Playground
Multiple gateway endpoints
Simple caching
Semantic Caching
Control Center
<ul>
<li>Weight-based Routing
Latency-based Routing
Priority-based Routing
Fallbacks - With advanced features
Budget limiting
Rate limiting
Observability
<ul>
<li>Logs - With custom retention
Traces - Export to custom storage buckets
Feedback on traces
Custom Metadata
Custom pricing / model
Cost per team/user/model/application
Metadata filtering
Alerts
Export to other monitoring platforms
MCP Gateway
<ul>
<li>MCP Servers - Custom
Tool calls per month - Custom
RBAC on MCPS
Virtual MCP Servers
Metrics - Comprehensive
Logs
Support for advanced authentication
Self-hosted MCPs
Prompt Management
<ul>
<li>Number of saved prompts - Unlimited
Versioning & Variables
Guardrails
<ul>
<li>Partner Guardrails integration
Custom Guardrail Hooks
Security & Authentication
<ul>
<li>Role based access control
SSO
SOC2
GDPR, HIPAA Compliance Certificates
Org management
Audit logs
Deployment modes
<ul>
<li>SaaS
VPC / On-prem
Air-gapped deployment
Deployment customization
<ul>
<li>Data Lake Export
Connect multiple storage bucket
Multiple gateway planes
Gitops (Infrastructure as code)
Support
<ul>
<li>Support - Priority Support,Dedicated Onboarding
SLA - Enterprise-Grade SLA
Comments