AI API cost is commonly calculated from the number of input and output tokens used.

Input and output tokens may have different prices.

What counts as input tokens?

What counts as output tokens?

Basic cost formula

Input cost  = input tokens  × input price per token
Output cost = output tokens × output price per token
Total cost  = input cost + output cost

Prices are often listed per one million tokens:

Input cost =
(input tokens ÷ 1,000,000) × input price per million

Output cost =
(output tokens ÷ 1,000,000) × output price per million

Example calculation

Assume these illustrative prices:

Input:  ₹100 per 1,000,000 tokens
Output: ₹400 per 1,000,000 tokens

One request uses:

5,000 input tokens
1,000 output tokens

Calculation:

Input  = (5,000 ÷ 1,000,000) × ₹100 = ₹0.50
Output = (1,000 ÷ 1,000,000) × ₹400 = ₹0.40

Total = ₹0.90

These prices are examples only. Always use the current pricing for the exact model.

Reading usage from a response

A provider may return something like:

{
  "input_tokens": 5000,
  "output_tokens": 1000,
  "total_tokens": 6000
}

Store usage by user, feature, model, and request.

Why input can become expensive

In a long conversation, your application may resend the conversation history with every request.

Request 1: short history
Request 20: large history + new message

RAG documents, tool definitions, and repeated instructions also add input tokens.

How to control cost

Cost should be controlled on the server

The frontend can request a long answer, but the backend should enforce:

Cost vs context window

These are related but different:

A request may fit inside the context window and still be unnecessarily expensive.

Track cost per request, per user, and per feature. A small cost multiplied by thousands of requests becomes a large bill.