Rate limiting controls how frequently a user or application can call an AI feature.
Why is it especially important for AI?
Every request may consume:
- provider quota
- tokens
- money
- server connections
- tool calls
- database and vector-search resources
Without limits, one user, a bug, or an attacker could create a large bill or make the feature unavailable.
Simple example
Maximum 20 AI requests per user every minute.
After the limit:
HTTP 429 Too Many Requests
Retry-After: 30
What should be limited?
Do not rely only on request count.
Possible limits include:
- requests per minute
- input tokens per request
- output tokens per request
- total tokens per day
- concurrent generations
- tool actions per hour
- spending per user or organization
One request containing 100,000 tokens is not equivalent to a small request.
Where to apply limits
Use multiple scopes when needed:
| Scope | Purpose |
| User ID | Fair usage for signed-in users |
| Organization | Shared team budget |
| IP address | Basic protection for anonymous traffic |
| Endpoint | Stricter limits for expensive features |
| Global | Protect provider quota and system capacity |
IP limits alone are unreliable because many users can share an IP and attackers can rotate addresses.
Common algorithms
- Fixed window: count requests during a fixed minute; simple but can allow bursts at the boundary
- Sliding window: count over the most recent time period; smoother but more work
- Token bucket: users collect tokens over time and spend one per request; allows controlled bursts
You only need to understand the behaviour. A rate-limit library or gateway can implement the algorithm.
Simplified backend example
const allowed = await limiter.consume({
key: `ai:user:${user.id}`,
cost: 1,
limit: 20,
windowSeconds: 60
});
if (!allowed) {
return response.status(429).json({
error: "AI request limit reached. Try again shortly."
});
}
For multiple backend instances, store counters in a shared system such as Redis rather than local memory.
Rate limiting vs budget limiting
- Rate limit: controls how fast requests arrive
- Budget limit: controls total usage or cost over a longer period
A production AI feature may need both.
Common mistakes
- limiting only at the frontend
- trusting a user-provided ID as the key
- keeping counters separately on each server
- using the same limit for cheap and expensive endpoints
- retrying provider rate-limit errors immediately
- returning no information about when to retry
Rate limiting protects availability and speed. Token and spending budgets protect cost.