AI & Development

Running AI Inference in the Browser with WebAssembly (2026)

How to run small AI models directly in the browser using WebAssembly and ONNX Runtime Web - no API calls, no latency, no cost per inference.

Every production AI feature that calls a cloud API has three dependencies: a network connection, an API provider that is up, and a budget. For features that run frequently - real-time classification, instant suggestions, local content filtering - those dependencies compound into latency, cost, and fragility that degrades the user experience. Browser-based inference via WebAssembly removes all three. The model runs in the user's browser, responses are instantaneous, and the inference cost is zero. In 2026, for models up to a few hundred megabytes, this is not an experimental technique - it is a production-viable architecture that several large-scale web apps use in production.

What WebAssembly inference actually means

WebAssembly (WASM) is a binary instruction format that runs in all modern browsers at near-native speed. ONNX Runtime Web is a JavaScript library that runs ONNX-format AI models using WebAssembly as the execution backend. The practical result: you export a model to ONNX format, load it in the browser via ONNX Runtime Web, and run inference locally using the user's CPU or GPU (via WebGPU, where supported).

The key constraint is model size. A model that fits in memory and runs in under 200ms on a mid-range phone is viable for real-time inference. A model that needs 4GB of VRAM is not. The practical floor in 2026 is approximately:

  • Classifiers and embedders (10 to 50 MB) - Sentiment analysis, topic classification, intent detection, semantic embedding generation. These run in 10 to 50ms even on low-end hardware and are fully viable for real-time use.
  • Small generative models (100 to 500 MB) - Phi-3 Mini, Gemma 2B, TinyLlama, Qwen 1.5B. These run in 1 to 3 seconds per response on modern hardware, suitable for user-initiated tasks where a brief wait is acceptable.
  • Medium generative models (500 MB, 2 GB) - Viable only with WebGPU on capable hardware. Currently not suitable for targeting all users.

Setting up ONNX Runtime Web

The minimal setup to run a model:

import * as ort from 'onnxruntime-web';

// Load the model (can be a local file or a URL to a CDN/your server)
const session = await ort.InferenceSession.create('/models/sentiment-classifier.onnx', {
  executionProviders: ['wasm'], // or ['webgpu', 'wasm'] to prefer GPU
});

// Run inference
const feeds = {
  input_ids: new ort.Tensor('int64', tokenIds, [1, tokenIds.length]),
  attention_mask: new ort.Tensor('int64', attentionMask, [1, attentionMask.length]),
};
const results = await session.run(feeds);
const logits = results.logits.data; // Float32Array

The executionProviders array specifies fallback order. Specifying ['webgpu', 'wasm'] uses WebGPU acceleration where available and falls back to WASM - this is the recommended default for production since WebGPU is now supported in Chrome, Edge, and Safari on modern hardware.

Model loading strategy: the download problem

The practical obstacle in browser inference is not the inference itself - it is the model download. A 50 MB model takes 2 to 10 seconds to download on a typical mobile connection. A 500 MB model is a non-starter as a synchronous user experience. The strategies that work in production:

  • Cache aggressively with Cache API - Store the model in the browser's Cache API after the first download. Subsequent visits use the cached model with zero download time. The cache persists across sessions until the user clears site data.
  • Progressive enhancement - Default to the cloud API for all users. Download the model in the background after the first interaction. Switch to local inference once the download is complete. Users who return get local inference; first-time users get cloud inference.
  • Quantised models - INT8 and INT4 quantisation reduces model size 2 to 4× with minimal quality loss for most classification and embedding tasks. A 50 MB model quantised to INT8 becomes 12 MB - a 3-second download becomes under 1 second.
  • Split loading - For generative models, load the model in chunks and begin inference on the first chunk while the rest downloads. ONNX Runtime Web supports this with the sessionOptions.externalData API.

Tokenisation in the browser

Most transformer-based models require tokenisation before inference - converting text to token IDs using the model's vocabulary. The recommended approach in 2026 is the @huggingface/transformers library (the JavaScript port of Transformers), which handles tokenisation and provides a higher-level pipeline API built on ONNX Runtime Web:

import { pipeline } from '@huggingface/transformers';

// Creates a sentiment analysis pipeline backed by ONNX Runtime Web
const classifier = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english');

// Single inference call - handles tokenisation internally
const result = await classifier('This is a great product');
// [{ label: 'POSITIVE', score: 0.9997 }]

The pipeline API is significantly simpler for standard NLP tasks. Use ONNX Runtime directly only when you need fine-grained control over the inference session or are running a custom model architecture.

WebGPU acceleration: when it matters

WebGPU is the new browser GPU API, supported in Chrome 113+, Edge 113+, and Safari 18+. For models above ~50 MB, WebGPU inference is typically 3 to 8× faster than WASM. For small classifiers and embedders, the overhead of GPU data transfer negates the speedup - WASM is often faster for these.

Feature-detecting WebGPU correctly:

const hasWebGPU = 'gpu' in navigator && (await navigator.gpu?.requestAdapter()) !== null;
const provider = hasWebGPU ? 'webgpu' : 'wasm';

Always specify ['webgpu', 'wasm'] as the provider array rather than checking manually and branching - ONNX Runtime handles the fallback automatically and logs which provider was actually used.

Use cases where browser inference wins

The architecture makes the most sense for specific feature categories:

  • Real-time input classification - Spam detection, profanity filtering, intent classification on every keystroke. At cloud API latency (200 to 500ms), these features feel laggy. At WASM latency (5 to 30ms), they feel instant.
  • Privacy-sensitive features - Content that should never leave the device: personal journal analysis, sensitive document summarisation, private conversation suggestions. Browser inference means no data is transmitted to any server.
  • Offline features - AI features in PWAs or apps that need to work without connectivity. The model is cached locally; inference runs regardless of network state.
  • High-frequency, low-stakes inference - UI suggestions, autocomplete, tag recommendations. The per-inference cost of cloud APIs makes high-frequency features prohibitively expensive; local inference eliminates the cost entirely.

Browser inference is not a replacement for cloud LLMs - it is a complement. For complex reasoning, long-form generation, and tasks that require the full capacity of frontier models, cloud APIs remain the right choice. But for the growing category of features that need speed, privacy, or offline capability, WebAssembly inference has crossed the viability threshold and is worth serious consideration in any AI-capable web application.