Pay-As-You-Go, Simple & Flexible
No monthly subscription, no minimum spend, pay only for what you use
OpenRely operates on a pure pay-as-you-go model. Deposited credits never expire, token usage is settled in real-time per request, with native Prompt Caching discounts to maximize compute efficiency.
Developer Free Trial
Ideal for developers to verify API gateway streaming and build early prototypes with instant credits and referral rewards.
- New User Bonus: Get $5 free test credit upon registration
- Referral Rewards: Invite new users to register, both inviter and invitee receive $5 credit
- Access to basic GPT and DeepSeek model series
- Up to 60 TPS request rate limit
- Community support and online developer docs
Pay-As-You-Go
Billed on exact token consumption with zero base fee. Enjoy automatic discounts as volume scales.
- 100% compatible with OpenAI / Anthropic / Gemini SDKs
- Full aggregation across DeepSeek, Claude, GPT, and Gemini
- Sub-second intelligent routing and automatic failover
- 500 TPS concurrency ceiling and granular API key ACLs
- Standard ticketing and dedicated technical expert groups
Enterprise Custom
Dedicated multi-node gateway cluster hosting with 99.99% SLA guarantee and 24/7 architect support.
- Dedicated gateway nodes with proximity edge acceleration
- 99.99% production SLA high availability guarantee
- Private model deployment and custom upstream channel routing
- Uncapped custom concurrency and throughput allocation
- 24/7 dedicated solutions architect priority support
Zero Monthly Fee
No forced monthly lock-in or minimum spend. Top up as you go, pay only for what you actually use.
Real-time Settlement
Tokens are audited and billed instantly per request with complete transparency down to single tokens.
Non-expiring Credits
Your deposited balance never expires. Eliminate the anxiety of end-of-month credit resets.
Prompt Cache Discounts
Native support for LLM Prompt Caching, slashing input token costs by up to 90% on repeated contexts.
Check Real-Time Model Rates in Console
OpenRely continuously aggregates 100+ leading LLMs globally (DeepSeek, Claude, GPT, Gemini, etc.). Because upstream providers may adjust rates dynamically, real-time rates, context windows, and availability status are centrally published in the Console Model Plaza.
View Complete Live Model PricingPricing & Billing FAQ
Common questions regarding billing mechanics, balance validity, and top-ups
QHow is the cost calculated for each request?
Costs are calculated based on actual Prompt Tokens (input), Completion Tokens (output), and whether Prompt Caching is hit. Formula: Cost = (Input Tokens × Input Rate + Output Tokens × Output Rate) / 1,000,000. Streaming and non-streaming requests share the exact same billing formula.
QDoes my account balance expire?
Never. Deposited funds remain valid indefinitely without any time limits as long as your account is in good standing.
QHow does Prompt Caching discount work?
When calling models that support Prompt Caching (e.g. Claude, DeepSeek) with repeated prefix prompts or long document context, the gateway automatically identifies cached tokens and discounts input rates by 80%~90%.
QWhat happens when my account balance is depleted? Will I go into negative debt?
You will never incur negative debt. OpenRely operates on strict prepaid balance deduction. When your balance cannot cover a request, the gateway returns a clean 402/403 status without surprise overage charges.
Ready to get started?
Sign up today to get free credits and connect seamlessly with Cursor, Claude Code, and over 100 developer tools.