inference means Running a trained AI model to produce output. It's mostly used in casual online conversation. It's casual and generally safe with friends and online. It's standard online vocabulary; nothing inherently concerning about hearing it.
Running a trained AI model to produce output.
Patrick Aschermayr, Konstantinos Kalogeropoulos, Nikolaos Demiris: Semi-Markov Models with Particle-Based Bayesian Inference for Epidemics https://arxiv.org/abs/2605.02706 https://arxiv.org/pdf/2605.02706 https://arxi…
CellxPert: Inference-Time MCMC Steering of a Multi-Omics Single-Cell Foundation Model for In-Silico Perturbation #SingleCell 🧪🧬🖥️ https://arxiv.org/abs/2605.00930
Running a trained AI model to produce output. Training teaches the model; inference is every time you use it. Inference is what costs money per token.
"Our inference bill doubled after we shipped the AI search."
inference means Running a trained AI model to produce output. It's mostly used in casual online conversation. It's casual and generally safe with friends and online. It's standard online vocabulary; nothing inherently concerning about hearing it.
Slang is read by different audiences for very different reasons. The For parents view re-explains the term in plain, non-judgemental language so a caregiver can understand what their child means without needing the surrounding subculture. The For English learners view focuses on register (when it's safe to use), formal alternatives, and common mistakes — what a textbook won't teach you.
Add your own interpretation of "inference".
The vocabulary of software engineers, AI researchers, and anyone living in a terminal or on GitHub — from LLM to MCP, CORS to vibe coding, agentic to enshittification.
See all Tech, Dev & AI slang on Slangora.
Browse all .
Using more compute at inference time — via longer chain-of-thought, sampling, or verification — to boost model quality without retraining. The OpenAI o1 + Claude extended-thinking paradigm. You pay seconds to gain accuracy.
Training a smaller, cheaper model to mimic the output of a larger one. "The 1B distillation of our 70B model runs fine on a laptop." Critical for shipping AI on devices, edge inference, and cost reduction. Trade-off: distilled models lose nuanced reasoning the parent could do.
An AI model trained on comprehensive data, enabling its application across various use cases such as chatbots and generative AI.
The amount of computational power or processing time required for an artificial intelligence system to generate a result or perform a task after it has already been trained.
Luca M. Possati: How Light Reshapes the Mind. An Active Inference Framework for the Cognitive and Emotional Effects of Indoor Lighting https://arxiv.org/abs/2605.01290 https://arxiv.org/pdf/2605.01290 https://arxiv.or…
inference means Running a trained AI model to produce output. It's mostly used in casual online conversation. It's casual and generally safe with friends and online. It's standard online vocabulary; nothing inherently concerning about hearing it.
"inference" is slang. It means: Running a trained AI model to produce output. Register: casual. Use it in: casual conversation and online posts. Avoid in: formal writing and professional emails.
“Patrick Aschermayr, Konstantinos Kalogeropoulos, Nikolaos Demiris: Semi-Markov Models with Particle-Based Bayesian Inference for Epidemics https://arxiv.org/abs/2605.02706 https://arxiv.org/pdf/2605.02706 https://arxiv.org/html/2605.02706”
“CellxPert: Inference-Time MCMC Steering of a Multi-Omics Single-Cell Foundation Model for In-Silico Perturbation #SingleCell 🧪🧬🖥️ https://arxiv.org/abs/2605.00930”
“Luca M. Possati: How Light Reshapes the Mind. An Active Inference Framework for the Cognitive and Emotional Effects of Indoor Lighting https://arxiv.org/abs/2605.01290 https://arxiv.org/pdf/2605.01290 https://arxiv.org/html/2605.01290”
“A third type of inference, abduction, has been proposed, notably by Charles Sanders Peirce.”
“Inference is traditionally divided into deduction and induction, a distinction that dates at least to Aristotle (300s BC).”
No comments yet — say something.
Specifically in AI/ML: running a trained model to produce output, as opposed to training. "Inference cost" = the cost per API call.
"Inference is most of the GPU spend."
No comments yet — say something.