Rate limiting controls how frequently a user or application can call an AI feature.

Why is it especially important for AI?

Every request may consume:

Without limits, one user, a bug, or an attacker could create a large bill or make the feature unavailable.

Simple example

Maximum 20 AI requests per user every minute.

After the limit:

HTTP 429 Too Many Requests
Retry-After: 30

What should be limited?

Do not rely only on request count.

Possible limits include:

One request containing 100,000 tokens is not equivalent to a small request.

Where to apply limits

Use multiple scopes when needed:

Scope Purpose
User ID Fair usage for signed-in users
Organization Shared team budget
IP address Basic protection for anonymous traffic
Endpoint Stricter limits for expensive features
Global Protect provider quota and system capacity

IP limits alone are unreliable because many users can share an IP and attackers can rotate addresses.

Common algorithms

You only need to understand the behaviour. A rate-limit library or gateway can implement the algorithm.

Simplified backend example

const allowed = await limiter.consume({
  key: `ai:user:${user.id}`,
  cost: 1,
  limit: 20,
  windowSeconds: 60
});

if (!allowed) {
  return response.status(429).json({
    error: "AI request limit reached. Try again shortly."
  });
}

For multiple backend instances, store counters in a shared system such as Redis rather than local memory.

Rate limiting vs budget limiting

A production AI feature may need both.

Common mistakes

Rate limiting protects availability and speed. Token and spending budgets protect cost.