Temperature is a generation setting that controls how strongly the model prefers the most likely next token over other possible tokens.
It does not add knowledge to the model. It changes how the model chooses from the knowledge and language patterns it already has.
First understand how text is generated
An AI does not write the complete answer at once.
It repeatedly performs this process:
- Read all tokens generated so far.
- Calculate a score for every possible next token.
- Convert those scores into probabilities.
- Select one token.
- Repeat the process for the next token.
flowchart LR
A["Existing text"] --> B["Score possible next tokens"]
B --> C["Temperature adjusts probabilities"]
C --> D["Select one token"]
D --> E["Add it to the text"]
E --> A
Temperature affects step 3, just before a token is selected.
Why is temperature needed?
There is often more than one sensible continuation for a sentence.
Consider:
After finishing work, I like to ___
Possible continuations include:
- relax
- read
- exercise
- ride
- cook
If the model always selected only the highest-probability token:
- repeated prompts would produce nearly the same answer
- stories and marketing text could become repetitive
- brainstorming would provide fewer unusual ideas
- the model would avoid less common but interesting word choices
If the model freely selected any token:
- answers could become random
- sentences could lose meaning
- factual responses could become unreliable
Temperature provides control between these two extremes.
Temperature is needed because different tasks require different behavior: consistency for extraction and coding, but variety for brainstorming and creative writing.
What exactly does temperature change?
Before temperature is applied, imagine the model has these possible next-token probabilities:
| Possible next token | Original probability |
| relax | 50% |
| read | 25% |
| exercise | 15% |
| ride | 10% |
Temperature reshapes this probability distribution before selection.
Lower temperature: sharp distribution
The most likely token becomes even more dominant.
| Token | Illustrative probability after low temperature |
| relax | 80% |
| read | 14% |
| exercise | 5% |
| ride | 1% |
The model will probably select relax, so the result is more predictable.
Higher temperature: flatter distribution
The probabilities move closer together.
| Token | Illustrative probability after high temperature |
| relax | 35% |
| read | 27% |
| exercise | 21% |
| ride | 17% |
Now ride or exercise has a greater chance of being selected. The answer becomes more varied.
The numbers above are simplified examples. Temperature does not manually add or subtract a fixed probability. It mathematically reshapes all token scores.
A little technical detail
The model first produces raw token scores called logits.
Temperature is applied before converting logits into probabilities:
P(token_i) = e^(z_i / T) / sum(e^(z_j / T))
Where:
z_iis the model's raw score for a tokenTis temperatureP(token_i)is the adjusted probability
You do not need to calculate this manually. Understand its effect:
- T below 1: differences between token scores become stronger
- T around 1: the original distribution is roughly preserved
- T above 1: differences become smaller
- T approaching 0: selection approaches choosing the highest-scoring token
The exact allowed range and behavior depend on the model and API.
Does low temperature produce the correct result?
Not necessarily. It produces the most likely and consistent result, not automatically the correct result.
Imagine the model incorrectly believes:
The capital of Australia is Sydney.
If Sydney has the highest score:
- low temperature may select Sydney very consistently
- high temperature might select another answer
- neither setting verifies the real fact
The correct answer is Canberra, but temperature itself does not check facts.
Low temperature = more predictable
Low temperature ≠ guaranteed correct
Correctness mainly depends on:
- what the model learned
- whether the prompt is clear
- whether correct context was provided
- whether trusted documents or tools are used
- whether the output is verified
Lower temperature is recommended for factual tasks because it reduces unnecessary variation, not because it turns the model into a fact checker.
Same prompt at different temperatures
Prompt:
Write a short description of a motorcycle.
Very low temperature
A motorcycle is a two-wheeled motor vehicle used for transportation.
Direct, safe, and predictable.
Medium temperature
A motorcycle is a compact two-wheeled machine that makes everyday travel quick and engaging.
More natural and expressive.
High temperature
A motorcycle is freedom balanced on two wheels, turning an ordinary road into an open invitation.
More creative, but less suitable for a technical definition.
When should we use each level?
| Task | Preferred behavior | Reason |
| Extract invoice fields | Low | The output should follow the same structure every time. |
| Generate SQL from a schema | Low | Creativity is less important than consistency. |
| Customer-support answer | Low to medium | The answer should be stable but still natural. |
| Explain a concept | Low to medium | Some variation can improve clarity without becoming unfocused. |
| Brainstorm product names | Medium to high | Different and unusual suggestions are useful. |
| Write a fictional story | Higher | Variety and surprising choices are desirable. |
These are general guidelines, not universal numeric rules. Test the actual model with your application.
Backend API example
A simplified request might look like:
const response = await ai.generate({
prompt: "Extract the order number from this invoice.",
temperature: 0.2
});
For brainstorming:
const response = await ai.generate({
prompt: "Suggest ten unusual names for a motorcycle application.",
temperature: 1.0
});
The exact SDK fields and supported values depend on the provider and model.
Temperature and hallucination
A higher temperature can increase the chance of unusual or unsupported output because less likely tokens receive more opportunity.
However:
- high temperature does not always cause hallucinations
- low temperature does not eliminate hallucinations
- a low-temperature model can repeat the same false claim consistently
Grounding the model with reliable data and verifying its output are more important for factual accuracy.
Common misunderstandings
Temperature is not model intelligence
Changing temperature does not make the model more knowledgeable or better at reasoning.
Temperature does not directly control answer length
Length is primarily controlled through instructions and output-token limits.
Temperature does not rewrite the prompt
It changes token sampling during output generation.
Temperature zero may not guarantee identical output
Some models, APIs, and computing systems can still produce small differences. Also, some modern models do not expose a temperature setting at all.
Final mental model
| Setting | Mental model |
| Low temperature | "Strongly prefer the safest, most likely continuation." |
| Medium temperature | "Allow reasonable variation while staying focused." |
| High temperature | "Give less-common continuations a greater chance." |
Interview answer
Temperature is a generation parameter applied to a model's token scores before sampling. It controls how strongly the model favors high-probability tokens. Lower temperature creates a sharper probability distribution and more consistent output; higher temperature creates a flatter distribution and more varied output. It is useful because factual extraction and creative brainstorming need different generation behavior. Temperature affects randomness, not knowledge, so a low value does not guarantee correctness.