webbrain-one's picture
Correct native WebGPU runtime dependency closure without changing the model
71866f0 verified
|
Raw History Blame Contribute Delete
5.08 kB
# FP32 Decision Adapter
The package adapter is a byte-preserving copy of the actual WebBrain extension's
new `src/chrome/src/providers/d1-runtime.js` and `d1-preprocess.js`, not a demo UI.
The latter differs from the preserved old baseline helper only by present-null
noul criteria lookup (`Object.hasOwn`, matching Python `dict.get`) and finite small
fractional Python-style JSON exponent formatting. JS Number cannot recover Python
int-versus-integral-float lexical types; pass preformatted state/criterion strings
when exact source lexical representation matters.
This file describes wiring, not a claim of packed-extension execution or real-site
quality. The source bidirectional decision model generates zero output tokens.
## Local Runtime
Use `runtime/vendor/ort.webgpu.bundle.min.mjs` and
`runtime/vendor/transformers.web.js`. The tokenizer module has only two documented
import-linkage changes to the local ORT bundle, reproducing the verified runner's
import map. ORT `1.31.0-dev.20260914-8d85527a0` and Transformers `4.3.1` are isolated
from the extension's existing chat runtime. JavaScript and native-WebGPU asyncify WASM are bundled
locally; no remote executable module loading is required. The preserved JSEP pair
is not selected: this native bundle needs the asyncify factory's webgpuInit.
Verify `package-manifest.json` against a trusted, separately pinned SHA-256 and
immutable HF revision before accepting its hashes. Verify every graph/external
data and tokenizer asset against its descriptor. For large data, contiguous
4 MiB `chunks` provide bounded per-chunk verification; do not allocate a second
1.5 GB buffer merely to run WebCrypto. Caching alone is not integrity verification.
The actual extension implements its own pinned download/cache/session orchestration.
```js
import * as ort from './vendor/ort.webgpu.bundle.min.mjs';
import { PreTrainedTokenizer } from './vendor/transformers.web.js';
import { createD1Runtime } from './d1-runtime.js';
ort.env.wasm.numThreads = 1;
ort.env.wasm.wasmPaths = {
mjs: new URL('./vendor/ort-wasm-simd-threaded.asyncify.mjs', import.meta.url).href,
wasm: new URL('./vendor/ort-wasm-simd-threaded.asyncify.wasm', import.meta.url).href
};
const tokenizer = new PreTrainedTokenizer(tokenizerJSON, tokenizerConfig);
const sessions = {};
for (const name of ['decision', 'vision', 'projector']) {
sessions[name] = await ort.InferenceSession.create(verifiedGraphs[name], {
executionProviders: [{ name: 'webgpu', device }],
graphOptimizationLevel: 'basic',
extra: { session: { disable_cpu_ep_fallback: '0' } },
externalData: [{ path: `${name}.data`, data: verifiedExternalData[name] }]
});
}
const judge = createD1Runtime({ ort, tokenizer, config, ratios, sessions, device, model });
const response = await judge.evaluate({ state, images: [inlinePngDataUrl], questions, signal });
```
The example assumes you already requested/checked the real WebGPU adapter/device
and fetched verified bytes/config/ratios. Ordinary Chrome may choose another GPU;
never label a software adapter or an unverified device as an RTX 5090 result.
CPU/WASM shape/control and floating mask/position construction remain possible.
Actual node placement requires a separate profile.
`questions` is an insertion-ordered object of named source-schema questions.
Choice criteria are an insertion-ordered name-to-description object; score
criteria are an ordinal array; noul follows source false/true marker scoring and
public yes/no answer semantics. Inline PNG/JPEG/WebP data URLs are accepted; the
adapter rejects page-chosen remote image URLs. A screenshot is encoded once and
each question becomes its own padded encoder row. The source 65,536 padded-token
subbatch budget is retained; named answers return in original order. No
generation/causal-mask path is involved.
Always use actual `float32` model/graph/feed precision. The unchanged root
`config.json` has a legacy `dtype: float16` label, which must not drive casting.
Do not alter temperatures, answer computation, option markers or token budget to
make a parity test pass. `dispose()` releases sessions/device owned by the adapter.
## Checkpoint
Root `model.safetensors` is the same clean 380-tensor FP32 artifact. Its file SHA
is distinct from the logical tensor-state SHA. No automatic Transformers Python
loader is advertised: legacy architecture metadata has no package `auto_map`.
`provenance/upstream-source/` is attribution/reference, not a native loading API;
the archived original audio.py is upstream source only, not audio runtime weights.
The original LFM model license and dependency notices remain mandatory. Synthetic
evaluation and finite runtime parity do not establish safe real-site automation,
completion reliability, or commercial eligibility.
The actual WebBrain GPL-3.0-or-later project notice is separately retained at
`../licenses/webbrain/LICENSE` for copied runtime source; `NOTICE` here describes
the component scopes. This package does not substitute MIT for that project notice
or offer a legal compatibility determination.