File size: 36,149 Bytes
1944112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
# Archive β€” done, and why

Completed work, kept for its **reasoning** rather than its status. A decision recorded without its
cause gets re-litigated by the next session, or quietly reversed.

Open work is in [ROADMAP.md](ROADMAP.md). What is true and measured is in [AI.md](AI.md).

---

## The extraction β€” `main` became a library

`main` is a library for developers embedding a local model in their own app; `demo` keeps the Firefox
extension and is the library's first consumer.

**Phase 0 β€” froze the extension.** `demo` branch, `demo-baseline` tag.

**Phase 1 β€” decoupled the three platform dependencies.** `background.js` went from 390 lines to 22.
The engine core (`src/engine/`) references no WebExtension API, asserted by a test β€” that claim is
only ever broken in the host nobody ran.

| | before | after |
| --- | --- | --- |
| Registry | `browser.storage.local` Γ—4 | `ModelStore` over an injected `StorageAdapter` β€” deliberately `get(key)`/`set(obj)`, the exact shape of `browser.storage.local`, so the WebExtension adapter is a passthrough |
| Worker | `browser.runtime.getURL(…)` | `new URL("./engine-worker.js", import.meta.url)` β€” understood by Vite/webpack/esbuild *and* correct on `moz-extension://`, so it replaces `getURL` rather than sitting beside it |
| Transport | 6 `browser.runtime` listeners | `attachWebExtensionTransport(engine)`; wire format unchanged byte-for-byte |

The wire protocol was **demoted**, not removed: in a page there is no `sendMessage`, so `OP`/`PORT_OP`
became one adapter's vocabulary rather than the interface.

**Phase 2 β€” the developer-facing surface.**

- `CreateScheduledEngine()` + `chat.completions.create()`. Migration off `@mlc-ai/web-llm` costs one
  line; everything after it is unchanged. `complete()`/`batch()` keep the names so `chat` stays free
  for the facade.
- `EngineError { code, message, detail }` β€” 8 codes, each existing because a caller does something
  *different* about it. A test walks `src/engine/` for bare `throw new Error`, because one untyped
  throw forces every caller back to string matching.
- `.d.ts` generated from the JSDoc that was already there; only the request/result typedefs were
  added. Source stays plain ESM.
- `@mlc-ai/web-llm` pinned to exactly `0.2.84` β€” `build.mjs` patches the bundle by string anchor and
  throws on a miss, so a caret range would break every consumer's build.

## Zero-download stopped being a constraint

It existed because the *extension* could not reasonably download. A developer embedding a model
usually can, and often must. So a model now arrives by one of three routes, resolved by `load()`:

| route | how | validation |
| --- | --- | --- |
| `prebuilt` | one of WebLLM's 163 HuggingFace models | none needed |
| `remote` | `registerModel({ modelId, model, modelLib })`, any base URL | none β€” a bad URL is reported far better by WebLLM's loader than by a HEAD request |
| `injected` | `registerModel({ modelId, files })`, **no network at any point** | exhaustive, before the first byte is written |

Three things fell out that were not obvious going in:

1. **It cost almost no code.** `toAppConfig` already emitted `{model, model_id, model_lib}`; it had
   only ever been handed `.invalid` URLs.
2. **It retired the top risk.** Cache injection on an ordinary page origin was risk #1 because the
   whole design rested on it. It is now the offline route only.
3. **The routes cannot be confused, so downloads default on.** `.invalid` is reserved by RFC 6761
   and can never resolve, so an injected model whose cache was evicted fails with a DNS error rather
   than silently pulling a gigabyte. That is the *mechanism* of the offline guarantee, not a label
   for it β€” a test asserts every URL such a record carries is on a `.invalid` host.

**Why local upload keeps a synthetic origin.** WebLLM composes every artifact URL as
`new URL(relative, base)` and runs the base through `cleanModelUrl`, which calls `new URL(...)` β€” so
the base must be absolute and resolvable. A `blob:` URL cannot serve as one, and there is no hook to
hand the loader bytes directly. Pre-populating the cache under WebLLM's own scopes and keys *is* its
native path. Seeding the cache for a URL the developer hosts was rejected: it would unify the record
shapes but make eviction silently re-download a gigabyte, which is the failure the design prevents.

## Model lifecycle β€” four states, three operations

The old API collapsed them, and that was a real bug: `remove()` deleted the cache **and** the
registry record, so for a remote model it threw away the only URL the bytes could be fetched from.

| state | VRAM | cache | record | leave via |
| --- | --- | --- | --- | --- |
| resident | βœ… | βœ… | βœ… | `unload(id)` |
| cached | β€” | βœ… | βœ… | `evict(id)` |
| registered | β€” | β€” | βœ… | `remove(id)` |
| unknown | β€” | β€” | β€” | β€” |

**Multiple resident models.** `#pool` became `#pools: Map<modelId, EnginePool>` with a current
selection. `use(id)` switches for free; a request naming a resident model routes to it *without*
changing which is current. Additive residency is opt-in (`{ keepResident: true }`) because each
resident model is a full copy of its weights and nothing reports free VRAM to a page.

## Device and compatibility

`probeDevice()` / `canRun()` / `rankModels()` answer "will this run here" before a byte is fetched.
The rules are this project's platform scars as code: the blocklisted adapter, Firefox's
9-storage-buffer cap, `q4f16_1` on a device without `shader-f16`.

Two things learned while building it:

- **Blockers and warnings must stay separate.** The 9-buffer cap costs KV reuse but is a *warning* β€”
  blocking it would refuse the exact configuration this project ships on.
- **"Largest that fits" is bad default advice.** Decode is memory-bandwidth-bound, so the largest
  model that fits is also the slowest thing that fits. `prefer: "quality" | "speed"` makes it the
  caller's choice rather than an assumption.

`probeDevice()` never throws β€” an unusable device is a result to explain, not an exception.

## Stop reinventing WebLLM

An audit ([WEBLLM-SURFACE.md](WEBLLM-SURFACE.md)) found three functions reimplemented that WebLLM
exported all along. Cause: the bundle was treated as something to `grep` for narrow facts rather than
an API to survey once β€” `cleanModelUrl` was even *read on screen* and then rewritten, without asking
what else used it. This violated AI.md's own **Reuse First** principle.

| Removed | Replaced by |
| --- | --- |
| `cleanModelUrl()` reimplemented | nothing β€” it only propped up the two below |
| `ModelStore.cacheKeysFor()` | WebLLM derives the keys it fetched |
| `ModelStore.cacheState()` | `engine.cacheState()` β†’ `hasModelInCache` for remote |
| `ModelStore.evict()` (all sources) | `evictInjected()` + `engine.evict()` β†’ `deleteModelAllInfoInCache` |
| speed re-derived from the worker probe | `usage.extra.decode_tokens_per_s`, already on every response |

The routing rule is now explicit: **whoever wrote the bytes owns the keys.** Our path survives only
where it demonstrably does more β€” WebLLM's delete and cache-check both read `tensor-cache.json` to
enumerate shards, so once *that* file is evicted they can neither find nor clean the shards it
indexes. Injected records carry an explicit key list and have no such failure. There is a test for
exactly that case, and it is the only justification for keeping the code.

**Speed was worse than duplication:** the pool already set `include_usage` and already stored
`chunk.usage`, so the measurement was being *received and discarded* so the worker probe could
recompute it.

## Raw chunk pass-through, and the tool-calling bug

The pool stripped every chunk to `delta.content` and the facade rebuilt an envelope from scratch β€”
so `tool_calls` was dropped entirely (**tool calling returned nothing usable**), `logprobs` was
always `null`, and `created` was restamped per chunk. Chunks now pass through verbatim.

Nothing is synthesized on the normal path: WebLLM emits its own terminal `finish_reason` chunk and
its own usage chunk. The one exception is an interrupted generation, where the stream simply stops
and a consumer would otherwise never learn why.

**A correction to the plan that produced this.** It specified `mergeToolCallDeltas()` for
"standard OpenAI fragment accumulation". WebLLM does not stream fragments β€” it parses the whole
output message at the end and emits tool calls complete in one terminal chunk. Building the merge
would have been machinery for a wire shape that is never produced: the plan's own failure mode,
inside the plan meant to prevent it.

## Bugs found and fixed along the way

| | |
| --- | --- |
| `EnginePool.load()` leaked an engine | It awaited `createEngine` before assigning `#slots`, so an `unload()` in that window tore down an *empty* pool and the engine then installed itself into a pool nobody referenced β€” leaking a worker and a full copy of the weights. `#grow()` had always guarded this; `load()` never did. Fixed with a `#generation` counter. |
| `registerModel` accepted URLs that fail at load | WebLLM's `cleanModelUrl` ends in `new URL()` with no base, so a relative `/models/x/` throws deep in the loader. Now resolved at registration. |
| `state.modelId` / `resident` went stale | Views onto `#pools` that nothing re-synced after unload, and left pointing at a model that never came up after a failed load. |
| `probeDevice` threw on a partial `navigator.gpu` | A polyfill without `requestAdapter` produced a TypeError from a function documented never to throw. |
| The pool discarded `finish_reason` | A `max_tokens` truncation was indistinguishable from the model choosing to stop. |
| `store.remove()` stranded remote shards | It iterates `groupKeysByScope`, empty for remote records β€” and deleting the entry destroys the only URL those bytes could be derived from. Now `engine.remove()` evicts first. |
| `throw`-as-`goto` in `multistep.js` | Caught two lines below; replaced with the control flow it was emulating. |
| `features()` called a dead fast path healthy | It answered "is multi-step on?" with `decodeSteps > 1`, but the worker keeps posting the *configured* K β€” 15 β€” long after the contract check routed decoding to stock single-step. So the one call documented as "what is switched **on** now" reported `multiStepDecoding: true` for the exact fault it exists to surface, and never exposed `multiStepOff` at all. `environment()` escaped it only by reading `state.decode.multiStepOff` itself rather than trusting `features()`. Now `multiStepDecoding` is `decodeSteps > 1 && !multiStepOff?.length`, and `multiStepOff` is returned beside it. `decodeSteps` deliberately keeps reporting the knob's value: dropping it to 1 would make `environment()` advise `configure({ decodeSteps: 15 })` for a fault no setting can fix. |

## Corrections to the record

Kept because each was stated confidently and was wrong; a future reader should not re-derive them.

- **"A second engine measured 1.06x, so batched decode is the only route to concurrent throughput."**
  This framed two complementary mechanisms as substitutes. A second engine buys *task isolation* β€”
  a translation and a ghost-text completion running at once β€” and never claimed aggregate
  throughput; the GPU is already saturated by a 2B. Batched decode makes *one task's* many requests
  faster. Neither substitutes for the other.
- **"Multi-model residency via `#pools` was not necessary."** Wrong. `reload()` unconditionally calls
  `unload()` first, so `reload([a,b])` is all-or-nothing β€” adding a third model reloads the first two
  (~51 s each). Additive residency does not exist upstream.
- **"No load time is measured in the repo."** It is: 51 s, [AI.md](AI.md) line 77. A grep for the
  wrong phrasing missed the table row.


## `npm run e2e` verified the extraction on real hardware

First run against the post-extraction tree: real Firefox, real GPU (Apple Silicon, `shader-f16`),
real `Qwen3.5-0.8B-q4f16_1-MLC`, drag-and-drop ingestion through the production `src/engine/` and
`src/adapters/webext.js` paths. **`e2e PASS`.**

- Ingest: 443,129,354 bytes, 11 shards, 2531 ms. Load: 48.1 s (AI.md's 51 s figure is for the larger
  2B; the 0.8B here loading faster is consistent with that being memory-bandwidth-bound).
- Decode: 27.4 tok/s over 127 tokens β€” inside AI.md's measured 16.6–27.9 tok/s range for this model.
  Decode probe: 664 kernels/token (639 forward + 25 sampling), 16.1 flushes/token β€” the same shape
  the compute-pass-batching patch targets, and it is still applying (41.3 kernels/flush).
- KV-reuse path exercised and correct: paged prefill and forced ragged re-prefill produced identical
  output on a multi-round conversation. Re-prefill slope 2.29 ms/token, in the neighbourhood of
  AI.md's 5.27 ms/token figure (different model, different history length β€” not a direct comparison).
- The scheduler's own two-tasks-two-engines check: 3.8 s concurrent vs 4.0 s sequential = **1.05x**,
  consistent with AI.md's measured 1.06x. This is the number the "second engine buys isolation, not
  throughput" framing rests on, now reconfirmed after the pool moved to `#pools: Map<modelId,
  EnginePool>` β€” evidence the multi-model split did not regress the single-model scheduling behaviour
  it was built on top of.

**One number worth a second look, not treated as a finding here:** this run reported
`storageBuffersPerStage=9`. It did not block anything β€” the KV-reuse path was exercised in the same
run and passed β€” but it is the exact threshold `probeDevice()`'s `NO_KV_REUSE` warning keys off, so a
future run reporting the same value is worth cross-checking against `device.test.mjs`'s assumptions
rather than assumed benign a second time.

> **[resolved] A second run reported 9 again, and the cross-check says 9 is the baseline, not an
> anomaly.** AI.md has said so all along β€” the Firefox Metal backend caps
> `maxStorageBuffersPerShaderStage` at 9, which is the entire reason the `storage-buffer-limit`
> patch exists. What the cross-check *did* surface is sharper than the original worry:
> `device.test.mjs` defines a healthy device as `storageBuffers = 10`, so the "good" default in the
> test matrix describes hardware nobody here has. On the reference M4, `probe.kvReuse` is
> **always** `false`, `engine-worker.js` forces `resetChat()` on every prefill, and
> **`batch_prefill_paged_kv_kernel` has therefore never executed on real hardware** β€” it is
> mock-tested only. It would first run on a >=10-buffer device, i.e. Chrome (ROADMAP, Gate B).
>
> That also made the e2e's own multiround check misleading: with reuse forced off, its "with KV
> reuse" and "forced reprefill" branches both ran the ragged kernel, so `identical` was guaranteed
> and the line `paged prefill is fine` claimed something the run had not tested. The check now
> reports `UNVERIFIED for paged prefill` on a sub-10-buffer device and still fails if two ragged
> re-prefills of the same history disagree.

The manifest.json restore left a diff β€” `restore()` round-trips the file through
`JSON.parse`/`stringify`, which turns `\uXXXX`-escaped em-dashes back into literal UTF-8. Cosmetic,
not a behaviour change, reverted with `git checkout`. Worth knowing before the next e2e run leaves the
same diff and it looks like something broke.

> **[fixed]** It did leave the same diff on the next run. The snapshot now only round-trips through
> JSON when the tree is *actually* dirty (a run killed mid-flight leaves the patch behind); a clean
> file is restored byte-for-byte. The cost of the old behaviour was not untidiness β€” it was training
> the reader to ignore a dirty tree after an e2e, which is precisely when a real diff matters.


## Lossless WebLLM upgrade β€” a bump is minutes, not an afternoon

`@mlc-ai/web-llm` is pinned exactly because `build.mjs` rewrites the bundle by matching source text.
"Lossless" was never meant as "automatic" β€” it means a bump *fails at the right line* instead of
somewhere deep in a half-patched loader. The standing runbook is in
[WEBLLM-SURFACE.md](WEBLLM-SURFACE.md), "Upgrading"; this is why each piece exists.

Three guards, because there are three distinct ways an upgrade breaks us:

| drift | caught by | the failure it prevents |
| --- | --- | --- |
| **surface** β€” code moved or reformatted | `build/patches.mjs` verify-then-write | a patch anchor silently landing in the wrong place, or the build half-applying and reporting only the first miss |
| **semantic** β€” a symbol survives, its meaning changed | `test/webllm-contract.test.mjs` | an export deleted, a field renamed, an enum gaining a case β€” none of which throw |
| **behavioural** β€” every name and shape intact, output wrong | `npm run e2e` | a tvmjs refactor that changes numerics |

**Contract tests (2a).** Makes WEBLLM-SURFACE.md executable β€” every export and shape the project
depends on, asserted statically against the bundle, GPU-free, first in `npm test`. The
highest-value piece: it catches semantic drift, which the patches cannot see. Two guards beyond the
obvious list: the monkeypatch member list is *derived from `multistep.js`'s own source* so it cannot
go stale, and `model_lib` unguessability is asserted rather than assumed (if it became derivable,
the "do not guess" rule in the verb-consolidation section should be revisited). Each assertion class
was mutation-tested β€” which found a real bug: `bundle.includes(name)` still passes when
`processNextToken` becomes `processNextTokenV2`, since the old name stays a substring. Now
word-bounded.

**Patch self-check and fuzzy diagnostics (2b).** Patches moved to `build/patches.mjs` as data,
applied by a shared verifier. Every anchor is checked before anything is rewritten. `patch-manifest.json`
records the version the anchors last held against, so a bump announces `0.2.84 -> 0.2.85: verifying
4 anchor(s)` rather than silently succeeding. Vanished identifiers are matched against survivors by
trigram overlap β€” simulated upstream renaming `requiredMaxStorageBuffersPerShaderStage`, the
diagnostic found the replacement at 85% similarity with its line. Ambiguity is a hard stop too: an
anchor matching 1995 sites refuses rather than rewriting one at random. Two corrections the
simulation forced: rank candidate lines by summed *rarity* not hit count (raw count returned
`const msg = {` β€” true and useless), and a rename needs a human to *approve* the new anchor, not to
*find* it.

**Structured patches (2c).** The corrected expectation held: AST parsing survives *formatting*
drift, not renames β€” an AST search by name fails exactly as a string match does. So the gain is
narrower than "structured = durable", and the work matched the correction rather than the original
proposal. `in: { enclosing }` scopes an anchor to the function a *sibling anchor* matched in β€” a
matched anchor, not a function name, so it adds no identifier upstream could rename. This removes the
false-failure class around `compute.end();`, a string generic enough that any unrelated new compute
pass in tvmjs failed the build. Anchors also match modulo whitespace and are word-bounded β€” the
latter not in the plan and found the same way 2a's bug was: `compute.end();` is a substring of
`precompute.end();`. The rebuilt bundle is byte-identical to the string-replacing applier's output.
`acorn` is a devDependency, ~565 KB unpacked (the original estimate was off 10x), never shipped.

**Runtime monkeypatch guard (2d).** `multistep.js` drives ~30 undocumented tvmjs internals; a rename
turns the fast path off *silently* β€” ~18 -> ~10 tok/s with nothing in the log. 2a covers the static
half. The runtime half: `PIPELINE_CONTRACT` + `missingPipelineMembers()`, checked against the live
pipeline at first decode (there is no pipeline at install time β€” the engine gets one per `reload()`),
verdict cached. Three buckets, because presence is not the failure that hurts: `calls` must be
callable (a rename throws β€” loud), `numbers` must be numbers (`x += 1` on an absent member creates a
property and the KV accounting drifts β€” silent), `reads` need only exist. Failure is loud once per
pipeline and posts `multiStepOff` to the host, which is otherwise indistinguishable from an idle
engine since `onBurst` is the only thing that reports stats. Found while building it: 2a's derivation
matched `\bpipeline\.` and missed members reached across a line break β€” the softmax at the heart of
the burst had no rename guard for as long as that test existed. Now whitespace-tolerant, comments
stripped first, and `PIPELINE_CONTRACT` is asserted to equal what the source actually reaches for.

**The flow, documented (2e).** Moved to WEBLLM-SURFACE.md so the doc you must revise on a bump is
the doc that tells you how.


## Verb consolidation β€” the ergonomic layer

`chat.completions.create()` is the compatibility layer and never changes. Everything here is
*additional* β€” the verbs a developer reaches for when they are not porting WebLLM code.

**`load(src, opts)` β€” one polymorphic entry.** Absorbs `load` + `registerModel` +
`ingestModelFolder`. Dispatch is a pure, synchronous `classifySource()` in
[src/engine/sources.js](src/engine/sources.js), so all four shapes β€” prebuilt id, HF/hosted URL,
`{model, modelLib}`, folder off disk β€” are testable with no GPU and no store. `registerModel` and
`ingestModelFolder` stay exported unchanged; `load()` composes them.

Two dispatch rules were **dropped after measuring**, both because a wrong guess surfaces as a 404
deep inside the loader:

- **`modelLib` is never guessed.** `<base><id>-webgpu.wasm` matches **0 of 163** prebuilt models
  (real names carry a `_cs1k`-style suffix, drop `-MLC`) and **0 of 163** host the lib on the
  weights' origin (they live on `raw.githubusercontent.com`). A remote source without `modelLib`
  fails in the classifier with that sentence, before any fetch.
- **`/resolve/main/` is not derived for HF URLs.** WebLLM's `cleanModelUrl` already appends it; doing
  it ourselves double-applies. A test asserts the stored URL is byte-identical to what was passed.

The id *is* derived from the URL's last segment β€” safe where `modelLib` is not, because an id is a
key in our own registry, never a path anything fetches, so a wrong guess is visible immediately and
free. `{ id }` overrides. `defer: true` on a bare prebuilt id is an **error**, not a silent load β€”
that silent load is the trap the whole section exists to avoid. Unknown ids get near-match hints.

Six mutation tests. One false pass worth remembering: the near-match hint appears at **two** error
sites and `String.replace` mutated only the first, so a working guard looked untested β€” a mutation
that does not apply is indistinguishable from a guard that does not work. Two latent crashes fixed
on the way: `filesFromInput`/`filesFromDataTransfer` spread their argument, so an array-like-but-not-
iterable `FileList`/`DataTransferItemList` died with `fileList is not iterable` three frames from
the caller's drop handler. Both use `Array.from` now.

**`unload(id, level)` β€” two depths, not two verbs.** `UNLOAD_LEVEL` is `"vram"` (default: free
VRAM, keep cache + record) or `"cache"` (also delete bytes, keep record β€” the old `evict()`).
`remove()` keeps its own verb: it is the one that cannot be undone without re-supplying the source.
An unrecognised level is refused with an error pointing at `remove()`, because "forget this model"
is the reading someone will try to spell as a level and it is the destructive one.

**[settled] A bare `unload()` frees only the current model**, with `unloadAll()` explicit for the
rest. The plan had the bare call free *everything*; shipping that silently would trap anyone already
calling `unload()`. "Free everything" is the more destructive reading and should be asked for by
name. `#evictBytes()` was split out of `evict()` so `unload(id, "cache")` reaches the bytes without
re-entering the class for a pool just torn down.

**`environment()` β€” read-only report; `.measure()` on it.** Absorbs `probe` + `features` +
`estimateSpeed` for the *reporting* half. [src/engine/environment.js](src/engine/environment.js), a
callable `engine.environment` cached like `chat`. Three open questions were all resolved by one
decision β€” **split read from write**: `environment()` reports only, writes go through `configure()`,
and passing a setting to `environment()` is an *error naming `configure()`*, not a silent no-op.
Implicit read/write dispatch by argument shape is the opposite of foolproof β€” the "reject or
write-then-report?" question had no intuitive answer precisely because one function was doing two
jobs.

Every report line carries `severity` Β· `affects` Β· `cause` Β· `fix` Β· `operable`, with **`fix: null`
⟺ `operable: false`** asserted for every line β€” hardware, build-time flags and browser settings
report a consequence with no remedy, which is still the difference between a bug report and an
informed decision. A **blocked device short-circuits the report**: "K=15 forward steps per GPU sync"
next to "no model can load" is true and useless. `configure()` grew `engineCount` because the report
advertises it as operable and a report naming a call that throws is worse than no report β€” it is
persisted, not hot, and `environment()` reports that gap rather than pretending. The `multiStepOff`
guard (Β§2d) finally has a consumer: a `degraded` line naming the missing internal, where before it
was posted by the worker and read by nothing.

Seven mutation tests. One real hole found: "`local` never fetches" was tested with a fetch counter,
but `load()` caches the model's size so `estimateSpeed()` short-circuits and *neither* scope fetches
after a load. The guarantee is structural β€” `local` never consults the model layer β€” and is tested
that way now.


## Engine capability β€” prefetch, embeddings, recipes

**`prefetch(modelId)`** β€” [src/engine/prefetch.js](src/engine/prefetch.js). Downloads a model with
no engine and **no WebGPU at all**: an app can warm the cache before it knows whether the machine
can run the model. Resumes; a second call is free.

The hard part: fetching the artifacts ourselves means deriving their URLs β€” the `/resolve/main/`
rule the verb-consolidation work above refused to derive. That refusal still holds; it was about not
deriving a URL *WebLLM will derive again at load*, which double-applies. Here WebLLM is not in the
loop β€” we are the loader. What makes it safe is not trusting the derivation: a key off by one
character writes a cache the loader never reads, and prefetch would report success while the user
downloads the model twice. So every prefetch ends by asking **WebLLM's own `hasModelInCache`** β€”
which derives through the very function we mirror β€” and throws if it says no. The contract test also
pulls `cleanModelUrl` out of the bundle and *runs* it against ours on six URL shapes, so an upstream
scheme change fails a test, not a download. Seven mutation tests, all caught.

**Embeddings (`engine.embed()`)** β€” a `kind` field on the job and one branch in
[pool.js](src/engine/pool.js) `#start`. **One pool, not two**: priority, supersession, preemption
and one-task-one-engine are identical for both kinds; only the call at the far end differs. A second
pool would have duplicated the scheduler to change one line. `embed()` returns bare vectors,
`embedRaw()` keeps WebLLM's envelope. **Known limit:** a running embedding cannot be interrupted β€”
`interruptGenerate()` works by making a decode loop break out, and one forward pass has no loop, so
a cancel that lands after the job starts marks it cancelled without stopping it. Stated in the
JSDoc, the README and a `[known limit]` test rather than left to be discovered. Six mutations, five
caught; the sixth was *equivalent* β€” `#start` decides on an explicit `=== EMBEDDING` and any unknown
kind routes to chat either way.

**Recipes β€” `ask()`, `conversation()`, `ghostText()`** β€” [src/engine/recipes.js](src/engine/recipes.js),
also methods on the engine. Scope grew on request: one command for each of the three things apps
actually want. The scheduling shipped as specified β€” one stable `session`, `interactive`, short
`max_tokens`, debounce, `cancel()` on blur, stale contexts dropped β€” and **prompts stayed with the
caller**: `ghostText({ prompt })` is required with no default; `ask`/`conversation` carry the
caller's text through. The engine authors nothing.

The piece worth keeping: `suggest()` **resolves `null` when stale**. The engine already superseded
the work; what a caller still had to remember was not to *paint* the answer that came back anyway.
Returning `null` removes the choice β€” the difference between a policy and a wrapper.
`conversation()` bounds history at 12 exchanges, derived from AI.md's numbers: with no cross-turn KV
reuse every turn re-prefills at ~5.27 ms/token, so unbounded history is quadratic and a turn near
the limit waits ~22 s. `keep: Infinity` opts out.

Found and fixed: a **promise leak in the debounce**. A newer keystroke called `clearTimeout` on the
previous waiter, whose `await` then had nothing to resolve it β€” every superseded keystroke leaked a
promise that never settled, and `Promise.all` over a burst hung forever. A superseded waiter has to
be woken and told it lost, not merely disarmed. Twelve mutation tests; two initially passed β€” one
equivalent, one genuinely vacuous: `sent[0].session === sent[1].session` also holds when *neither*
has a session, which is exactly the regression it was meant to catch. Presence is asserted before
equality now.


## Shipping 0.1.0 β€” installable, documented, on npm

The library was extracted, tested and complete, and served the project's goal β€” "make WebLLM easier
to use, foolproof to build on" β€” **for nobody**, because it was unpublished, undocumented as a whole
surface, and un-installable. This section is the gap between "the code is done" and "a developer can
`npm i` it and run four lines."

**Four-line target, met without an engine change.** `import` / `CreateScheduledEngine(id)` /
`engine.ask(prompt)` / read the string. Probing the shape found three things that stopped it being
*usable*:

1. **Un-installable.** `vendor/web-llm.js` is a build product and is gitignored; there was only a
   `prepublishOnly`, and npm runs **`prepare`** for a git dependency. So `npm i` 404'd (`"private":
   true`) and a git dependency installed but could not run, failing with `GENERATION_FAILED: Cannot
   find module .../vendor/web-llm.js` β€” wrong twice, since nothing had begun generating and the path
   named is ours. Now `"prepare": "node build.mjs"`, `private` removed. Verified by deleting
   `vendor/` and running `npm install` β€” it comes back.
2. **Looked like a hung process.** No `initProgressCallback` meant zero output during a minutes-long
   ~0.8 GB download. `CreateScheduledEngine` now distinguishes three states: `undefined` β†’ a
   throttled console reporter (1 line/second; 58 shard callbacks β†’ 2 lines; the 100% report is never
   dropped), `null` β†’ explicit silence, a function β†’ unchanged. **`new ScheduledEngine()` stays
   silent** β€” a library core that logs is wrong in a worker, an extension background page, or a
   test. This is the getting-started facade only.
3. **The Vite worker-URL break.** Vite's dependency pre-bundler copies `new Worker(new
   URL("./engine-worker.js", import.meta.url))` verbatim into `node_modules/.vite/deps/`, where the
   sibling file does not exist β€” `vite dev` only, real (non-linked) install only, `vite build`
   unaffected. `everything-webgpu/vite` ships a plugin (`optimizeDeps.exclude`, the manual
   equivalent still documented). And if a consumer does neither, `load()` now fails with
   `PACKAGE_INCOMPLETE` naming the fix, because `new Worker()` does not throw on a 404 β€” it fires one
   `error` event and goes quiet, so the handshake is raced against it.

**`PACKAGE_INCOMPLETE` is one code for two causes** (`detail.cause` separates them). No caller
writes a different `catch` branch: both mean "your build is wrong, this app has not shipped," both
are fixed in config. A second code would grow the table a caller reads without giving them anything
to do.

**`verify-consumer` β€” the only test that can see the consumer's world.** Everything under `test/`
and every `examples/` project reaches the package through a *linked* path, and Vite never
pre-bundles a linked package β€” so none of them can exercise the one failure that reaches users. This
blind spot produced a **wrong claim in the docs**: that `optimizeDeps.exclude` was needed for `vite
build` and that the examples proved it. Measured on a real tarball install, neither holds β€” build
output is byte-identical with and without it. `npm run verify-consumer` packs the tarball, installs
it for real, and asserts three outcomes separately: `vite build` emits the worker chunk and keeps
WebLLM lazy; `vite dev` **without** the plugin still 404s the worker; `vite dev` with it resolves.
The middle one is asserted as a *failure* on purpose β€” a fix whose absence changes nothing is not a
fix, and if Vite ever stops pre-bundling this package that assertion says the plugin is dead weight.

**`API.md` β€” every call form on one page,** asserted by
[test/api-doc.test.mjs](test/api-doc.test.mjs), derived from the source the way `readme.test.mjs`
is: every `engine.x(` named resolves to a real member, **no public member is left undocumented**
(the reverse direction the README test lacks), the error table equals `ERROR`, every export
appears, enum-value rows match the real objects, the subpath table equals `package.json` `exports`.
Writing it found `engine.store` undocumented and a regex reading enum rows as error codes.

**`examples/` β€” bare, react, webext,** each a standalone project depending on the package as
`file:../..` so it resolves through the **`exports` map** β€” an example importing
`../../src/engine/index.js` would still run and would still leave the exports map, the `files` list
and every entry path untested. [test/examples.test.mjs](test/examples.test.mjs) derives its checks
from the example sources, so a fourth example is covered the moment its directory exists. Also
closed a silent `files`/`exports` gap: a new export path that `files` would not publish resolves in
the checkout and 404s in the tarball. Asserted from `package.json` now.

**Bundle-size story, measured not estimated:** **53 kB (~19 kB gzip)** entry chunk before a model
loads; the 6 MB WebLLM bundle is a **lazy chunk** fetched on the first `load()` or
`listAvailableModels()` and never by a visitor who does neither; the IndexedDB adapter is a further
0.8 kB lazy chunk that vanishes when a host brings its own store β€” the webext build emits no `idb`
chunk at all, which is that claim tested by construction.

**Licence compliance.** Publishing `vendor/web-llm.js` redistributes WebLLM (Apache-2.0) and its
dependency `loglevel` (MIT), and the esbuild bundle was built `legalComments: "none"` β€” no notice
survived, a violation. `THIRD-PARTY-NOTICES.md` now carries the full texts, generated from the
installed packages; [test/license.test.mjs](test/license.test.mjs) fails the build if a bundled
dependency ever lacks a notice, catching a future web-llm bump that inlines a new dep. `build.mjs`
uses `legalComments: "eof"` now β€” upstream has already stripped every `@license` banner (the bundle
is byte-identical either way today), but `"none"` would silently drop one a future dep adds. `LICENSE`
added (ISC). `files` scopes `vendor` to `web-llm.js` β€” the stale, unreferenced `vendor/web-llm.d.ts`
was shipping and made the tarball depend on disk state.

**Published:** `everything-webgpu@0.1.0`, 46 files, 2.3 MB packed; `dist.integrity` matched the
dry-run exactly. Still deferred to a later version: the `demo` extension rebuilding on the package
(the source-tree acceptance test for the extraction), and a `repository` field once the repo has a
remote.