Skip to content

Repository files navigation

How does AI inference work?

InferenceClear

How does AI inference work?
Training a model happens once; using it happens billions of times. Watch a frozen network answer questions, light up a grid of multiply-adds, see why a chatbot's speed is set by memory, shrink a model to 4 bits and break it at 2, and follow one question into a 120 kW rack and back.

▶ Play with it  ·  Read the 60-second explainer  ·  Watch the 40-second video

Glassbox No. 077 AI & Data Code: MIT Content: CC BY 4.0 Privacy: explained

In 60 seconds

  1. Train once, answer forever. Training writes a model's weights, slowly and once. Inference freezes them and runs the inputs forward to answer a question. Our spiral network cost 192 million multiply-adds to train and 304 per answer, so after about 630,000 questions answering has cost more than learning.
  2. It is all multiply-add. A forward pass is mostly one step: multiply an input by a weight and add it to a total. A layer is a matrix times a vector. An LLM needs about 2 × parameters FLOPs per token, and GPUs, TPUs and phone NPUs do thousands of these sums at once.
  3. Memory is the real limit. To write each token, a chatbot must read every weight from memory. So tokens per second ≈ memory bandwidth ÷ model size in bytes: about 200 for an 8B model in 16-bit on an H100. A growing KV cache adds more to read; batching many users shares the reads.
  4. Fewer bits, smaller models. Quantisation rounds weights onto a few levels: FP32 → FP16 → INT8 → INT4. Size and read time fall; accuracy holds until too few bits are left, and our net breaks at 2 bits. Pruning and distillation shrink models too, which is how phones run AI.
  5. The journey of one question. Your prompt travels to a data centre, is tokenised, then prefilled in one big pass. The answer is decoded one token at a time and streamed back. Racks draw around 120 kW and need liquid cooling; a median chatbot text prompt was reported at about 0.24 Wh.
  6. Phone or cloud. On-device AI keeps data private, works offline and has no network wait, but the model must be small. The cloud runs far bigger models but needs a connection and servers. Each answer uses little energy, yet data centres used about 1.5% of the world's electricity in 2024.

Words worth knowing

Term Meaning
Inference Using a trained model, with its weights frozen, to answer a new question.
Multiply-add (MAC) Multiply one input by one weight and add it to a running total; two FLOPs.
FLOP One floating-point operation, such as a single multiply or add.
Memory bandwidth How many bytes per second can move from memory to the chip; it sets LLM decoding speed.
KV cache Saved keys and values for every earlier token, so they need not be recomputed.
Batching Serving many users' next tokens with one read of the weights.
Quantisation Storing weights with fewer bits by rounding them onto a few allowed levels.
Prefill and decode Processing the whole prompt at once, then writing the answer one token at a time.
On-device AI Running a model on your own phone or laptop instead of in a data centre.

A short history

From motor-driven knobs in 1960 to data centres answering billions of questions a day: how we learned to run trained models fast, small and cheap.

  • 1980 · XCON: expert rules at work in a factory (John McDermott (Carnegie Mellon University) for Digital Equipment Corporation, Salem, New Hampshire, United States)
  • 2009 · GPUs make deep learning practical (Rajat Raina, Anand Madhavan and Andrew Ng, Stanford University, United States)
  • 2012 · AlexNet wins on two gaming GPUs (Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton, University of Toronto, Canada)
  • 2016 · TPU: a chip built for inference (Norman Jouppi and a Google team, Google, Mountain View, United States)
  • 2017 · A neural engine in a phone (Apple (A11 Bionic); Huawei had announced its Kirin 970 with an NPU days earlier, Cupertino, United States; Shenzhen, China)
  • 2017 · INT8 inference goes into production tools (Benoit Jacob and colleagues, Google, United States)
  • 2023 · A chatbot model on a laptop (Georgi Gerganov, Sofia, Bulgaria)
  • 2023 · vLLM and PagedAttention (Woosuk Kwon, Zhuohan Li, Ion Stoica and colleagues, UC Berkeley, United States)

The full story, with 30 moments, charts, people and 56 sources: glassbox.how/e/inferenceclear/history. The data lives in history.json.

Video and slides

Made with the Glassbox studio from this box's storyboard (window.glassbox.director). Free to reuse under CC BY 4.0.

Video: How does AI inference work?

Carousel slide-1 Carousel slide-2 Carousel slide-3 Carousel slide-4

File What Size
glassbox/reel.mp4 Reel / Short, with captions and soundtrack 1080×1920
glassbox/video.mp4 YouTube video, with captions and soundtrack 1920×1080
glassbox/slide-1…10.jpg Instagram carousel 1080×1350
glassbox/thumb.jpg YouTube thumbnail 1280×720
glassbox/cover.jpg Share card and repo social preview 1200×630
glassbox/history-reel.mp4 “History in 10 moments” Reel / Short 1080×1920
glassbox/history-slide-*.jpg History carousel 1080×1350
glassbox/post.json Post copy and schedule used by the publish kit

Privacy

This box has no accounts and no ads, and it ships its own fonts and libraries. When you run it yourself it sends nothing anywhere. On glassbox.how, the site's /bar.js also loads Glassbox's analytics: Google Analytics to count visits (it asks first in the EU, UK and Switzerland, and stays off when your browser sends Global Privacy Control or Do Not Track) and ClickTrust to detect bots.

It remembers a few things in your own browser only, and never sends them anywhere:

Browser storage key What it holds
inferenceclear.v1 Which chapters you have opened, your best quiz scores, and sound on or off.

Exactly what each one sees is at glassbox.how/privacy.

Licences

  • Code: MIT. Use it, change it, ship it.
  • Explanations, text, images and videos (glassbox.json, glassbox/): CC BY 4.0. Credit “Glassbox, glassbox.how/e/inferenceclear”.
  • Third-party parts keep their own licences: three.js (MIT), Geist, Instrument Serif (SIL OFL 1.1).
  • The Glassbox name and logo aren't covered by either licence. See the terms.

Found a mistake? Open an issue. Corrections happen in public.

Run it

It's plain HTML, CSS and JavaScript. No build step and no dependencies. Run locally, it contacts no other website.

python3 -m http.server 8000

Three.js and the fonts ship in vendor/ and fonts/, so it also works offline.

Then open http://localhost:8000.

How it's built

File What
index.html, css/app.css The page and its styles
js/app.js, js/stage.js, js/ui.js, js/kit.js The shared Glassbox 3D engine: chapters, 3D stage, controls, quiz, video director
js/infer.js The maths, from scratch: a small network trained once in your browser then frozen, exact multiply-add counts, quantisation and pruning, and the LLM memory-bandwidth speed model
js/view.js Drawing helpers: canvas boards, text tiles, the 3D network, a GPU-style chip and framing for phones and the tall video
js/chapters/*.js One file per chapter: the 3D model, controls, text, key terms, quiz and video scenes
glassbox.json Title, question, explainer beats, key terms, browser storage and credits shown on glassbox.how
reel in each chapter The storyboard the Glassbox studio records into short videos
glassbox/ The published video, slides, thumbnail and post copy
fonts/, vendor/three/ Self-hosted Geist and Instrument Serif (SIL OFL 1.1) and three.js (MIT)

About

How does AI inference work? An interactive, open-source explainer. Glassbox No. 077.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages