Temperature is a generation setting that controls how strongly the model prefers the most likely next token over other possible tokens.

It does not add knowledge to the model. It changes how the model chooses from the knowledge and language patterns it already has.

First understand how text is generated

An AI does not write the complete answer at once.

It repeatedly performs this process:

  1. Read all tokens generated so far.
  2. Calculate a score for every possible next token.
  3. Convert those scores into probabilities.
  4. Select one token.
  5. Repeat the process for the next token.
flowchart LR
    A["Existing text"] --> B["Score possible next tokens"]
    B --> C["Temperature adjusts probabilities"]
    C --> D["Select one token"]
    D --> E["Add it to the text"]
    E --> A

Temperature affects step 3, just before a token is selected.

Why is temperature needed?

There is often more than one sensible continuation for a sentence.

Consider:

After finishing work, I like to ___

Possible continuations include:

If the model always selected only the highest-probability token:

If the model freely selected any token:

Temperature provides control between these two extremes.

Temperature is needed because different tasks require different behavior: consistency for extraction and coding, but variety for brainstorming and creative writing.

What exactly does temperature change?

Before temperature is applied, imagine the model has these possible next-token probabilities:

Possible next token Original probability
relax 50%
read 25%
exercise 15%
ride 10%

Temperature reshapes this probability distribution before selection.

Lower temperature: sharp distribution

The most likely token becomes even more dominant.

Token Illustrative probability after low temperature
relax 80%
read 14%
exercise 5%
ride 1%

The model will probably select relax, so the result is more predictable.

Higher temperature: flatter distribution

The probabilities move closer together.

Token Illustrative probability after high temperature
relax 35%
read 27%
exercise 21%
ride 17%

Now ride or exercise has a greater chance of being selected. The answer becomes more varied.

The numbers above are simplified examples. Temperature does not manually add or subtract a fixed probability. It mathematically reshapes all token scores.

A little technical detail

The model first produces raw token scores called logits.

Temperature is applied before converting logits into probabilities:

P(token_i) = e^(z_i / T) / sum(e^(z_j / T))

Where:

You do not need to calculate this manually. Understand its effect:

The exact allowed range and behavior depend on the model and API.

Does low temperature produce the correct result?

Not necessarily. It produces the most likely and consistent result, not automatically the correct result.

Imagine the model incorrectly believes:

The capital of Australia is Sydney.

If Sydney has the highest score:

The correct answer is Canberra, but temperature itself does not check facts.

Low temperature = more predictable
Low temperature ≠ guaranteed correct

Correctness mainly depends on:

Lower temperature is recommended for factual tasks because it reduces unnecessary variation, not because it turns the model into a fact checker.

Same prompt at different temperatures

Prompt:

Write a short description of a motorcycle.

Very low temperature

A motorcycle is a two-wheeled motor vehicle used for transportation.

Direct, safe, and predictable.

Medium temperature

A motorcycle is a compact two-wheeled machine that makes everyday travel quick and engaging.

More natural and expressive.

High temperature

A motorcycle is freedom balanced on two wheels, turning an ordinary road into an open invitation.

More creative, but less suitable for a technical definition.

When should we use each level?

Task Preferred behavior Reason
Extract invoice fields Low The output should follow the same structure every time.
Generate SQL from a schema Low Creativity is less important than consistency.
Customer-support answer Low to medium The answer should be stable but still natural.
Explain a concept Low to medium Some variation can improve clarity without becoming unfocused.
Brainstorm product names Medium to high Different and unusual suggestions are useful.
Write a fictional story Higher Variety and surprising choices are desirable.

These are general guidelines, not universal numeric rules. Test the actual model with your application.

Backend API example

A simplified request might look like:

const response = await ai.generate({
  prompt: "Extract the order number from this invoice.",
  temperature: 0.2
});

For brainstorming:

const response = await ai.generate({
  prompt: "Suggest ten unusual names for a motorcycle application.",
  temperature: 1.0
});

The exact SDK fields and supported values depend on the provider and model.

Temperature and hallucination

A higher temperature can increase the chance of unusual or unsupported output because less likely tokens receive more opportunity.

However:

Grounding the model with reliable data and verifying its output are more important for factual accuracy.

Common misunderstandings

Temperature is not model intelligence

Changing temperature does not make the model more knowledgeable or better at reasoning.

Temperature does not directly control answer length

Length is primarily controlled through instructions and output-token limits.

Temperature does not rewrite the prompt

It changes token sampling during output generation.

Temperature zero may not guarantee identical output

Some models, APIs, and computing systems can still produce small differences. Also, some modern models do not expose a temperature setting at all.

Final mental model

Setting Mental model
Low temperature "Strongly prefer the safest, most likely continuation."
Medium temperature "Allow reasonable variation while staying focused."
High temperature "Give less-common continuations a greater chance."

Interview answer

Temperature is a generation parameter applied to a model's token scores before sampling. It controls how strongly the model favors high-probability tokens. Lower temperature creates a sharper probability distribution and more consistent output; higher temperature creates a flatter distribution and more varied output. It is useful because factual extraction and creative brainstorming need different generation behavior. Temperature affects randomness, not knowledge, so a low value does not guarantee correctness.