The context window is the maximum amount of information an AI model can consider at one time. It is measured in tokens.

Think of it as the model's working memory for the current request.

What goes inside the context window?

The window may contain:

flowchart TB
    A["Context-window token budget"] --> B["Instructions"]
    A --> C["Conversation history"]
    A --> D["Current prompt and documents"]
    A --> E["Generated answer"]

All of these compete for space inside the same token budget.

Simple example

Assume a model has a 10,000-token context window.

Content Tokens
Instructions 1,000
Conversation history 4,000
Your new prompt and documents 2,000
Space remaining for the answer 3,000
10,000 - 1,000 - 4,000 - 2,000 = 3,000 tokens remaining

This is a simplified example.

Exact input and output limits depend on the model and API.

What happens when a conversation becomes too long?

The application must usually do one or more of these:

If an important old detail is removed or summarized badly, the model may answer without it.

Example

At the beginning, you say:

My project uses PostgreSQL. Do not suggest MongoDB.

After a very long conversation, that message may no longer be present in the active context.

The AI might then suggest MongoDB because it cannot see the earlier instruction.

Why long context can still cause mistakes

Even when information technically fits inside the context window, a very long input can make the relevant detail harder to identify.

The model may focus on a more recent or strongly worded instruction.

A larger context window provides more space, but it does not guarantee perfect memory or reasoning.

How to handle long conversations

Context window is temporary working memory, not permanent memory.

The model only knows the information currently placed inside its context.

Context window vs model training

Model training Context window
Knowledge learned before the conversation Information supplied for the current request
Not changed by an ordinary chat Changes with every message
Long-term model parameters Temporary working information

Interview answer

The context window is the maximum number of tokens a model can consider during one request. It contains instructions, conversation history, user input, documents, and generated output. When the limit is reached, older or less relevant information may need to be removed, summarized, or retrieved separately.