Inference

In plain English

If training is studying, inference is sitting the exam. Every time you ask a chatbot a question and it answers, that's inference.

In practice

Inference is where ongoing AI costs sit: per-token API pricing, GPU servers and response times all relate to inference. For high-volume tasks, a smaller, faster model often beats a large one on cost without losing much quality.

Under the hood

Inference runs a forward pass of the model on new inputs with fixed weights. For language models it is autoregressive, generating one token at a time, so latency grows with output length. Quantisation, batching and caching reduce its cost.

Example

"Inference costs dropped after we moved routine questions to a smaller model."

•