API β every way to call it
The complete call surface of everything-webgpu, one page. README.md is the pitch and
the migration story; this is the catalogue. Asserted against the code by
test/api-doc.test.mjs β every engine method named here exists, every
package export appears, and the error table equals ERROR, so this page cannot drift from the
source without failing npm test.
The page is in three tiers. Start here β load, ask, conversation,
environment β is the dead-simple path, and most apps need nothing else. Native
passthrough is chat.completions.create(), the WebLLM/OpenAI compatibility
layer, which never changes (see Stability). When you need
more is the rest of the surface: the general calls and their scheduling
fields, embeddings, residency, device inspection, configuration, lifecycle. Everything outside the
passthrough is pre-1.0 and may move; the ergonomic verbs are being consolidated in
ROADMAP.md.
The four lines
import { CreateScheduledEngine } from "everything-webgpu"; // 1. import
const engine = await CreateScheduledEngine("Llama-3.2-1B-Instruct-q4f16_1-MLC"); // 2. load a model
const reply = await engine.ask("Name three primary colours."); // 3. ask
console.log(reply); // 4. the answer
reply is a plain string. Line 2 downloads ~0.8 GB the first time and prints throttled progress to
the console unless you pass initProgressCallback (render it yourself) or initProgressCallback: null (silence). After that first run it is a cache read and needs no network.
On Vite, add one plugin β see Bundlers.
Start here
Four calls. Get an engine with CreateScheduledEngine (below, under Getting an
engine), then load a model, ask it things or hold a conversation, and
environment tells you whether the machine is up to it.
Load a model β engine.load(src, opts?)
One call, four source shapes. engine.load(src, opts?) works out what you handed it, registers
whatever needs registering, and brings the model up.
src |
route |
|---|---|
"Llama-3.2-1B-Instruct-q4f16_1-MLC" |
a prebuilt id, or anything you registered earlier. A typo is answered with near matches. |
"https://huggingface.co/mlc-ai/Foo-MLC" |
an HF repo. /resolve/main/ is not derived β WebLLM appends it. |
"https://cdn.example/models/foo/" + { modelLib } |
any base URL you host. modelLib is required and never guessed (0 of 163 prebuilt models have a derivable lib name or same-origin lib). |
{ model, modelLib } |
the explicit remote spec. |
input.files | dropEvent.dataTransfer | { files } |
a folder off disk. No network at any point. |
opts β keepResident, signal, modelType, contextWindow, vramRequiredMB, id,
onProgress, defer.
idoverrides the id derived from the URL's last segment.defer: trueregisters the source without building a pool β the drop-now-load-later flow. It returns the registry record.deferon a bare prebuilt id is an error, not a silent load.keepResident: trueholds this model in VRAM alongside whatever is already up. The default unloads everything else first β the safe choice on a 16 GB machine.
await engine.load("Llama-3.2-1B-Instruct-q4f16_1-MLC");
await engine.load("https://cdn.example/models/my-model/", { modelLib: "https://cdn.example/models/my-model/lib.wasm" });
await engine.load(dropEvent.dataTransfer);
await engine.load(input.files, { defer: true }); // register now, build the pool on first use
load() composes the lower-level registerModel, ingestModelFolder and the download primitives;
prefetch() warms the cache with no GPU. All of that is under More on
loading.
Ask one question β engine.ask(input, opts?)
One question, its own task, no session β two ask()s never supersede each other. Returns the
reply string. opts.onDelta to stream. Goes through the same scheduler as everything else: priority
bands, one-task-one-engine, opt-in preemption.
Hold a conversation β engine.conversation(opts)
A multi-turn chat that keeps its own history. engine.conversation({ system?, keep?, ...defaults })
β one stable task for every turn, turns serialised, history bounded at keep: 12 exchanges
(Infinity opts out).
const chat = engine.conversation({ system: "You are terse." });
await chat.say("capital of France?"); // β { text, finishReason }
await chat.say("and its population?", onDelta); // remembers
chat.messages; chat.length; chat.reset(); chat.restore(messages);
Inspect the machine β engine.environment(opts?)
| call | answers |
|---|---|
await engine.environment(opts?) |
the preflight. A report; every line has severity Β· affects Β· cause Β· fix Β· operable. { scope: "local" } never touches the model layer (poll freely); { scope: "device" } is hardware only. |
await engine.environment.measure() |
one calibration generation β measured tok/s for the current model |
environment() only reports. Writes go through configure(); passing a setting to
environment() is an error that names configure(). Per-model "will it run" is
canRun(modelId); the fuller device surface is there too.
Native passthrough
The one call that never changes. @mlc-ai/web-llm is OpenAI-shaped, and so is this: the
migration off it is a one-line import swap, and chat.completions.create() then takes and returns
exactly the same shapes β same streamed chunk objects, same finish reasons, same non-streaming
envelope. See Stability.
| call | shape |
|---|---|
engine.chat.completions.create(params) |
the WebLLM/OpenAI shape, unchanged. Streams the same chunks, same finish reasons. session/priority/task/preemptible are additive. |
params = the OpenAI generation fields WebLLM already speaks (messages, temperature,
max_tokens, response_format, extra_body, β¦) plus the scheduling fields that are the only
thing this adds over calling WebLLM directly:
| field | meaning |
|---|---|
modelId |
load/route to this model instead of the current one |
id |
job id; also what cancel(id) takes |
task |
the unit that owns an engine; a whole batch shares one |
session |
a later job with this key supersedes the earlier one |
priority |
"interactive" | "normal" | "background" |
preemptible |
may be interrupted by an interactive job (set it on the work that can afford to lose) |
complete(), completeRaw() and batch() take the same scheduling fields and expose cancelled /
preempted as first-class outcomes the OpenAI shape has no room for β see The general
calls.
When you need more
Getting an engine
| call | when |
|---|---|
await CreateScheduledEngine(modelId?, opts?) |
the common case. Loads modelId before returning, like WebLLM's CreateMLCEngine. Omit it for an engine that loads later. |
new ScheduledEngine({ store, workerUrl?, loadWebLLM?, prebuilt? }) |
when you must pass a store explicitly β a worker, a test, an extension. Does not load anything. |
opts for CreateScheduledEngine β store, initProgressCallback, workerUrl, loadWebLLM,
prebuilt. Anything not store/initProgressCallback is forwarded to the constructor.
| constructor field | default | meaning |
|---|---|---|
store |
IndexedDB (CreateScheduledEngine only; the constructor requires it) |
a ModelStore, or a bare StorageAdapter it wraps |
prebuilt |
true |
expose WebLLM's 163 HuggingFace models. false = offline-only: load() resolves registered models and nothing else, and an unknown id fails before the WebLLM bundle is fetched |
workerUrl |
new URL("./engine-worker.js", import.meta.url) |
the decode worker's module URL |
loadWebLLM |
() => import("../../vendor/web-llm.js") |
override the bundle source (tests) |
Stores β import { indexedDBStorage } from "everything-webgpu/adapters/idb" (pages, plus
ensurePersistent()), everything-webgpu/adapters/memory (memoryStorage(), tests),
everything-webgpu/adapters/webext (webExtensionStorage() + attachWebExtensionTransport()).
import { ScheduledEngine, ModelStore } from "everything-webgpu";
import { memoryStorage } from "everything-webgpu/adapters/memory";
const engine = new ScheduledEngine({ store: new ModelStore(memoryStorage()) });
More on loading
Warming the cache first β await engine.prefetch(modelId, { onProgress, signal }). Downloads
the weights without building an engine and without WebGPU, so an app can warm the cache before
it knows whether the machine can run the model. Interrupted downloads resume; a second call is free.
WebLLM cannot express this β reload() needs a GPU before it fetches a shard.
Low-level, still exported β load() composes these rather than replacing them:
| call | does |
|---|---|
engine.registerModel(spec) |
add a { modelId, model, modelLib } or { modelId, files } record, no pool |
ingestModelFolder(entries, { store }) |
folder β populated Cache Storage, returns the record |
filesFromInput(input.files) / filesFromDataTransfer(dt) |
either browser shape β flat { path, file }[] (the latter is async) |
prefetchModel({ modelId, record, ... }) / resolveModelUrl(...) |
the engine-free download primitives |
The general calls
Every call here goes through the same scheduler as ask / conversation / chat.completions:
priority bands, session supersession, one-task-one-engine, opt-in preemption.
| call | shape |
|---|---|
await engine.complete(payload, onChunk?) |
{ text, usage, finishReason, cancelled?, preempted? }. onChunk(delta) streams plain text. |
await engine.completeRaw(payload, onRawChunk?) |
same, but onRawChunk gets WebLLM's chunk object verbatim. |
await engine.batch({ requests, task?, ...sched }, onItem?) |
requests fanned across the pool as one task. Returns BatchItem[] β each with index, engineIndex, startedAt, finishedAt, and text/usage or error. |
payload is the OpenAI generation fields plus the scheduling fields β the same table as Native
passthrough (modelId, id, task, session, priority, preemptible).
Result flags: cancelled: true (superseded or cancel()ed), preempted: true (text is partial).
Ghost text β engine.ghostText(opts)
engine.ghostText({ prompt, debounceMs?, maxTokens?, session?, ...defaults }) β debounce + one
session key + interactive priority + resolves null when stale. prompt is required, no
default: prompts are model-specific and belong to whoever owns the feature.
const ghost = engine.ghostText({ prompt: (before) => `Continue:\n${before}` });
const hint = await ghost.suggest(editor.textBefore()); // string | null
ghost.cancel(); // on blur / accept
Embeddings
Needs an embedding model (snowflake-arctic-embed-*, from 239 MB), usually held resident
alongside a chat model.
| call | returns |
|---|---|
await engine.embed(input, opts?) |
number[][] β one vector per input, in order |
await engine.embedRaw(input, opts?) |
WebLLM's OpenAI envelope (data[].embedding) |
opts: modelId, task, session, priority, preemptible, id. A running embedding cannot
be interrupted β one forward pass has no decode loop to break out of; queued embeddings supersede
normally.
Residency and cache
A resident model is a full copy of its weights in VRAM, and nothing reports free VRAM to a page β so residency is explicit.
| call | frees | keeps |
|---|---|---|
await engine.unload() |
current model's VRAM | cache + registry |
await engine.unload(id) |
that model's VRAM | cache + registry |
await engine.unload(id, "cache") |
VRAM + cached bytes | registry entry |
await engine.unloadAll() |
every resident model's VRAM | cache + registry |
await engine.remove(id) |
bytes + registry entry | nothing β for an injected model, means re-supplying the folder |
engine.evict(id) |
low-level primitive unload(id, "cache") is built on |
registry entry |
Routing without loading β engine.use(id) points unaddressed requests at an already-resident
model (free and instant; load() is what costs). engine.resident lists model ids with a live
pool. await engine.cacheState(id) says what is on disk ("complete" / "partial" / absent).
Inspecting the machine
The device surface behind the Start here preflight.
| call | answers |
|---|---|
await engine.canRun(modelId) |
per-model: { ok, blockers, warnings }, before anything downloads |
await engine.recommendModels(opts?) |
which models this device should be asked to run, best first. opts: maxVramMB, needsVision, needsToolCalling, prefer |
await engine.estimateSpeed(modelId?) |
projected decode tok/s (uses the measured rate once one generation has happened) |
await engine.probe() |
raw device probe: WebGPU, adapter, shader-f16, the five limits, storage quota. Cached. |
await engine.features() |
what is switched on now, vs what the device could support. multiStepOff is non-null when decode fell back to one GPU sync per token β the silent halving environment() reports as degraded |
engine.hasWebGPU |
Boolean(navigator.gpu) |
await engine.listAvailableModels() |
registered + prebuilt, normalised. Costs one bundle fetch. |
engine.listModels() |
registered only β cheap, no bundle load |
Configuration
await engine.configure(patch) β applies a runtime knob and persists it as the default.
| knob | effect |
|---|---|
decodeSteps |
forward steps per GPU sync. Hot, no reload. 1β32 (DEFAULT_DECODE_STEPS = 15). |
engineCount |
pool size. Persisted; live pools keep the size they came up with. |
temperature, maxTokens, systemPrompt |
generation defaults (DEFAULT_SETTINGS) |
Not operable from JS, report-only via environment(): KV reuse (derived from the 9-storage-buffer
cap), compute-pass batching (build-time NO_PASS_MERGE), shader-f16, GPU, about:config flags.
Lifecycle and cancellation
| call | |
|---|---|
const stop = engine.subscribe(listener) |
listener(state) fires immediately, then on every change. Returns unsubscribe. |
engine.state |
snapshot: status, modelId, progress, error, pool {size,busy,queued,maxSize,growthBlocked}, resident, decode |
engine.store |
the ModelStore, so a host can drive the registry without a second handle |
engine.cancel(idOrSession) |
cancel by job id or by session key |
engine.load(id, { signal }) |
an AbortController signal tears down an in-flight download |
state.status is one of ENGINE_STATE: "idle" Β· "loading" Β· "ready" Β· "error".
Errors
Every failure is an EngineError with a .code, a human-readable .message (the thing you print),
and structured .detail. import { isEngineError, ERROR } from "everything-webgpu".
| code | what to do |
|---|---|
NO_WEBGPU |
tell the user to check flags/hardware; retrying is futile |
NO_MODEL |
nothing registered β send them to your setup flow |
UNKNOWN_MODEL |
that id is not resolvable; listAvailableModels() says what is |
CACHE_INCOMPLETE |
a locally-registered model was evicted; re-register the folder |
INVALID_MODEL_FOLDER |
not a compiled MLC model; detail says what is missing |
BAD_REQUEST |
the caller's arguments are wrong β a bug in the caller |
ABORTED |
the caller cancelled it. Not a failure; do not report it as one |
GENERATION_FAILED |
the model failed mid-generation |
PACKAGE_INCOMPLETE |
your build is wrong, not your code β missing vendor/ bundle, or a decode worker the bundler did not emit. message names the fix; detail.cause says which |
try { await engine.load(id); }
catch (err) {
if (isEngineError(err, ERROR.CACHE_INCOMPLETE)) return reRegisterFolder();
throw err;
}
Bundlers
The engine spawns its decode worker with new Worker(new URL("./engine-worker.js", import.meta.url), { type: "module" }). On Vite, its dependency pre-bundler rewrites that URL to
a path that 404s β in vite dev, on a real (non-linked) install only. Add the plugin:
import { everythingWebGPU } from "everything-webgpu/vite";
export default defineConfig({ plugins: [everythingWebGPU()] });
Equivalent by hand: optimizeDeps: { exclude: ["everything-webgpu"] }. Skip both and load()
throws PACKAGE_INCOMPLETE naming the fix rather than hanging. vite build is unaffected either
way. Other bundlers that honour new URL(..., import.meta.url) for workers (Webpack 5, Rollup,
Parcel 2) need nothing.
Every export
import { β¦ } from "everything-webgpu" β 43 names.
Engine & entry β ScheduledEngine, CreateScheduledEngine, EnginePool
Model sources β ModelStore, ingestModelFolder, filesFromInput, filesFromDataTransfer,
prefetchModel, resolveModelUrl, isInjected, baseUrlFor, groupKeysByScope, toAppConfig
Recipes (also methods on the engine) β ask, conversation, ghostText
Device β probeDevice, canRun, projectSpeed, rankModels, REFERENCE_DECODE_BYTES_PER_SECOND
Multi-step decoding β installMultiStepDecoding, burstSize, clampSteps,
DEFAULT_DECODE_STEPS, MAX_DECODE_STEPS
Errors β EngineError, ERROR, isEngineError, asEngineError
Formatting β formatBytes
Enums / constants β PRIORITY, PRIORITY_ORDER, ENGINE_STATE, UNLOAD_LEVEL, SEVERITY,
MODEL_TYPE, SOURCE, DEFAULT_SETTINGS, WORKER_CONFIGURE, CACHE_CONFIG, CACHE_MODEL,
CACHE_WASM
Enum values
| enum | values |
|---|---|
PRIORITY |
interactive Β· normal Β· background |
ENGINE_STATE |
idle Β· loading Β· ready Β· error |
UNLOAD_LEVEL |
vram Β· cache |
SEVERITY |
blocked Β· degraded Β· tune Β· info Β· ok |
SOURCE |
prebuilt Β· remote Β· injected |
MODEL_TYPE |
llm = 0 Β· embedding = 1 Β· vlm = 2 |
CACHE_* |
webllm/config Β· webllm/model Β· webllm/wasm |
DEFAULT_SETTINGS |
engineCount: 2, decodeSteps: 15, temperature: 0.6, maxTokens: 1024, systemPrompt: "" |
Subpath exports
| specifier | |
|---|---|
everything-webgpu |
everything above |
everything-webgpu/vite |
everythingWebGPU() Vite plugin |
everything-webgpu/worker |
the decode worker entry (for a custom workerUrl) |
everything-webgpu/adapters/idb |
indexedDBStorage() |
everything-webgpu/adapters/memory |
memoryStorage() |
everything-webgpu/adapters/webext |
webExtensionStorage(), attachWebExtensionTransport() |
everything-webgpu/adapters/protocol |
the wire-protocol constants |