AI & Development
How to run small AI models directly in the browser using WebAssembly and ONNX Runtime Web - no API calls, no latency, no cost per inference.
Every production AI feature that calls a cloud API has three dependencies: a network connection, an API provider that is up, and a budget. For features that run frequently - real-time classification, instant suggestions, local content filtering - those dependencies compound into latency, cost, and fragility that degrades the user experience. Browser-based inference via WebAssembly removes all three. The model runs in the user's browser, responses are instantaneous, and the inference cost is zero. In 2026, for models up to a few hundred megabytes, this is not an experimental technique - it is a production-viable architecture that several large-scale web apps use in production.
WebAssembly (WASM) is a binary instruction format that runs in all modern browsers at near-native speed. ONNX Runtime Web is a JavaScript library that runs ONNX-format AI models using WebAssembly as the execution backend. The practical result: you export a model to ONNX format, load it in the browser via ONNX Runtime Web, and run inference locally using the user's CPU or GPU (via WebGPU, where supported).
The key constraint is model size. A model that fits in memory and runs in under 200ms on a mid-range phone is viable for real-time inference. A model that needs 4GB of VRAM is not. The practical floor in 2026 is approximately:
The minimal setup to run a model:
import * as ort from 'onnxruntime-web';
// Load the model (can be a local file or a URL to a CDN/your server)
const session = await ort.InferenceSession.create('/models/sentiment-classifier.onnx', {
executionProviders: ['wasm'], // or ['webgpu', 'wasm'] to prefer GPU
});
// Run inference
const feeds = {
input_ids: new ort.Tensor('int64', tokenIds, [1, tokenIds.length]),
attention_mask: new ort.Tensor('int64', attentionMask, [1, attentionMask.length]),
};
const results = await session.run(feeds);
const logits = results.logits.data; // Float32Array
The executionProviders array specifies fallback order. Specifying ['webgpu', 'wasm'] uses WebGPU acceleration where available and falls back to WASM - this is the recommended default for production since WebGPU is now supported in Chrome, Edge, and Safari on modern hardware.
The practical obstacle in browser inference is not the inference itself - it is the model download. A 50 MB model takes 2 to 10 seconds to download on a typical mobile connection. A 500 MB model is a non-starter as a synchronous user experience. The strategies that work in production:
sessionOptions.externalData API.Most transformer-based models require tokenisation before inference - converting text to token IDs using the model's vocabulary. The recommended approach in 2026 is the @huggingface/transformers library (the JavaScript port of Transformers), which handles tokenisation and provides a higher-level pipeline API built on ONNX Runtime Web:
import { pipeline } from '@huggingface/transformers';
// Creates a sentiment analysis pipeline backed by ONNX Runtime Web
const classifier = await pipeline('sentiment-analysis', 'Xenova/distilbert-base-uncased-finetuned-sst-2-english');
// Single inference call - handles tokenisation internally
const result = await classifier('This is a great product');
// [{ label: 'POSITIVE', score: 0.9997 }]
The pipeline API is significantly simpler for standard NLP tasks. Use ONNX Runtime directly only when you need fine-grained control over the inference session or are running a custom model architecture.
WebGPU is the new browser GPU API, supported in Chrome 113+, Edge 113+, and Safari 18+. For models above ~50 MB, WebGPU inference is typically 3 to 8× faster than WASM. For small classifiers and embedders, the overhead of GPU data transfer negates the speedup - WASM is often faster for these.
Feature-detecting WebGPU correctly:
const hasWebGPU = 'gpu' in navigator && (await navigator.gpu?.requestAdapter()) !== null;
const provider = hasWebGPU ? 'webgpu' : 'wasm';
Always specify ['webgpu', 'wasm'] as the provider array rather than checking manually and branching - ONNX Runtime handles the fallback automatically and logs which provider was actually used.
The architecture makes the most sense for specific feature categories:
Browser inference is not a replacement for cloud LLMs - it is a complement. For complex reasoning, long-form generation, and tasks that require the full capacity of frontier models, cloud APIs remain the right choice. But for the growing category of features that need speed, privacy, or offline capability, WebAssembly inference has crossed the viability threshold and is worth serious consideration in any AI-capable web application.