Run large language models in the browser with WebGPU. three-llm implements transformer inference with Three.js and its TSL compute shader system, so model execution stays on the user's GPU without a server-side inference runtime.
Try the live demo: threekit-llm.ben3d.ca · Read the technical write-up
See packages/three-llm/README.md for full documentation, including features, requirements, install, and usage.
To work on the chat app or test against the hosted model bucket:
corepack enable
pnpm install
pnpm devOpen http://localhost:3000. The demo loads checkpoints through the website's /api/models/ proxy backed by the public gs://three-llm bucket, and falls back to Hugging Face if a file is missing.
This monorepo uses pnpm workspaces:
packages/three-llm: the inference library, model loaders, tokenizers, and TSL kernelspackages/website: the React chat demo
Requirements: Node.js 20 or newer, pnpm 11.
pnpm dev # watch the library and run the demo
pnpm build # build every workspace package
pnpm test # run type checks and unit tests
pnpm test:e2e # run Playwright tests
pnpm lint # check source with Oxlint
pnpm format # format the repository with Oxfmt
pnpm size # check the compiled bundle against the size budgetCoverage and bundle-size gates, Codecov setup, and the release process live in RELEASING.md.
MIT © 2026 Ben Houston
See CONTRIBUTING.md for the issue, branch, PR, and release workflow, RELEASING.md for release-specific gates and maintainer setup, and SECURITY.md for private vulnerability reporting.
