Term explained

What is inference in AI?

Inference is the stage where a trained AI model is actually used — taking your input and computing an answer — as opposed to the earlier, one-off training stage.

Added 9 Oct 2026

In plain English

AI systems have two distinct phases. Training is the expensive, one-off process of learning from data, adjusting billions of internal numbers. Inference is everything afterwards: the model, now fixed, takes an input and computes an output. Every time you send a chatbot a message, ask for an image, or let your phone transcribe a voice note, you are running inference.

It helps to think of training as writing a reference book and inference as looking something up in it. The book does not change as people read it — which is also why a model does not learn from your conversation unless the provider deliberately uses it in a later round of training.

Why it matters

Inference is where AI meets its costs and its limits in practice. Training grabs the headlines, but a popular service runs inference billions of times, so in total it can consume more computing power, electricity and money than training ever did. Inference cost drives what providers charge, how fast replies arrive and whether a feature is viable at all. It is also why there is so much effort to shrink models so they can run on a phone or laptop rather than in a data centre.

An example

You take a photo of a restaurant menu in another language and your phone overlays a translation. No learning is happening — a trained model is being run on your image, right then, to produce text.

What to watch out for

People often assume a chatbot “remembers” what they told it and improves accordingly. Usually it does not: within a conversation it simply re-reads the recent exchange each turn, and once that window is exceeded, earlier details drop away. Permanent changes to a model require retraining, not conversation.

← All terms