Quick Access
Pay-As-You-Go System
Neosantara AI operates on a flexible Pay-As-You-Go billing system. You only pay for what you use—no mandatory subscriptions, no restrictive volume limits.No Subscriptions
Flexible Scaling
Transparent Pricing
Legacy Users (Transition)
Quota Usage
Automatic Migration
Welcome Credit
Pricing Structure
Batch Processing Discount
Save 50% with Batch API
/v1/batches endpoint. No special configuration needed!Chat Completion Models
Pricing per 1 million tokens (1M tokens):- Indonesian Models
- Open Source Models
- Specialized Models
Understanding Caching
Understanding Caching
How Token Pricing Works
How Token Pricing Works
- 1 token ≈ 4 characters in English
- 1 token ≈ 2-3 characters in Indonesian
- 1M tokens ≈ 750,000 words in English
- Model:
nusantara-base - Input: 1,000 tokens (Rp 300 / 1M × 1,000 = Rp 0.3)
- Output: 500 tokens (Rp 1,500 / 1M × 500 = Rp 0.75)
- Total: Rp 1.05 per request
Image Generation
High Resolution
Medium Resolution
Low Resolution
OCR (Optical Character Recognition)
Extract text from images at an affordable rate:Tool Usage & Function Calling
- Tool definitions count as input tokens
- Tool calls count as output tokens
- Tool results count as input tokens
nusantara-base:
- Tool definition: 200 tokens input = Rp 0.06
- Model’s tool call: 50 tokens output = Rp 0.075
- Tool result: 150 tokens input = Rp 0.045
- Total tool overhead: Rp 0.18
Throughput Tiers
Automatic Tier Upgrades
Rate Limits by Tier
What do RPM, ITPM, and OTPM mean?
What do RPM, ITPM, and OTPM mean?
- RPM (Requests Per Minute): Maximum API requests you can make per minute
- ITPM (Input Tokens Per Minute): Maximum input tokens processed per minute across all requests
- OTPM (Output Tokens Per Minute): Maximum output tokens generated per minute across all requests
How tier upgrades work
How tier upgrades work
- You deposit Rp 100,000 → Unlock Basic tier
- You spend Rp 50,000 → Still Basic tier (deposit history counts)
- You deposit another Rp 600,000 → Unlock Standard tier
Need higher limits?
Need higher limits?
- Custom concurrent request limits
- Dedicated infrastructure
- Volume discounts
- Priority support
Balance Management
Checking Your Balance
Usage Dashboard
Billing Page
Adding Funds
Navigate to Billing
Select Top Up Amount
Complete Payment
Supported Payment Methods
QRIS
E-Wallets
Bank Transfer
Credit Card
Cost Optimization Tips
Use Batch API for 50% Savings
Use Batch API for 50% Savings
- Data labeling and classification
- Bulk translations
- Content generation pipelines
- Embeddings for large datasets
Choose the Right Model
Choose the Right Model
- Simple tasks: Use
llama-3.2-11b(Rp 100/1M) ornusantara-base(Rp 300/1M) - Medium tasks: Use
gpt-oss-20b(Rp 1,600/1M) orgemma2-9b-it(Rp 200/1M) - Complex tasks: Use
claude-sonnet-4.5or specialized models only when needed
Leverage Prompt Caching
Leverage Prompt Caching
- Put static instructions/examples at the start
- Cache reads cost 90% less than fresh inputs
- Ideal for repeated queries with consistent context
Optimize Token Usage
Optimize Token Usage
- Set appropriate
max_tokenslimits - Use concise system prompts
- Implement streaming to stop generation early if needed
- Remove unnecessary whitespace and formatting
Frequently Asked Questions
What happens if my balance runs out?
What happens if my balance runs out?
Can I get a refund?
Can I get a refund?
Do credits expire?
Do credits expire?
How accurate is the usage tracking?
How accurate is the usage tracking?
Can I set spending limits?
Can I set spending limits?