--- license: other license_name: lfm1.0 license_link: LICENSE base_model: LiquidAI/d1-omni-600M tags: - decision - webgpu - onnx - fp32 - experimental - image-text-to-decision --- # d1 Browser Decision FP32 An audio-free derivative of [LiquidAI/d1-omni-600M](https://huggingface.co/LiquidAI/d1-omni-600M/tree/02b55d7076f15129e59ab3f94783f32c4b088674). It directly scores named options for `choice`, `noul` and ordinal `score` questions using text/JSON state and optional screenshots. It does **not** generate text or tokens. Do not call `generate()` or treat it as a causal chat model. The reported quality data are authored synthetic browser pages, with AI-authored labels independently checked through AI DOM/pixel review, not human-expert annotations. Actual Chrome/WebGPU runtime parity was measured separately; that does not establish reliable completion detection on independently developed websites. ## Immutable Candidate The published bytes are the already validation-selected `projector_head_text_lora` epoch 06 from the balanced-V3 experiment. The lock was written before held-out model evaluation. Packaging performed no training, new selection, weight transfer, merge, quantization or ONNX re-export. - Source revision: `02b55d7076f15129e59ab3f94783f32c4b088674`. - 380 clean FP32 tensors: text encoder, decision head, vision tower and projector; no audio or LM head. - Prior adaptation updated 31 head tensors, four projector tensors and four encoder Q/V matrices through one rank-4, alpha-8 LoRA merge. The other 341 tensors, including all 197 vision-tower tensors, remain exact original-A values. - Physical `model.safetensors` SHA-256: `75ab6d7d0ec2966c969a95c548b91f4015b07cba82fa839fcfb5c19b40e9f940` (1,899,912,876 bytes). - Logical clean-tensor-state SHA-256: `76409dd958673e2028f1da23a909033876169f89603c16c7cbcdfbcc7404cdb5`. This is **not** the file SHA-256. - Selected delta SHA-256: `82b6844524adf1ff1f7add1c9ef57475af5fcfa07ada10b9edf70c4d24c1ba35`. - Selection-lock SHA-256: `4ab2441d9b3cf7742957a374988fc50fc400b29081b7c38b6920be88cd65bb19`. The three ONNX graphs and their external data total 1,900,356,386 bytes. They are full FP32, unquantized, and copied without binary changes. `package-manifest.json` and `SHA256SUMS` enumerate payload checksums; large files also have 4 MiB chunk checksums. **Config caveat:** the unchanged legacy `config.json` says `dtype: float16` and `architectures: [NoAudioModel]`. Actual checkpoint/graphs/feeds are **float32**, as explicitly required by the release manifest. Do not derive precision from the legacy label. This repository does not advertise an `AutoModel.from_pretrained()` loader or executable remote Python code. ## Synthetic Held-Out Results New final test: 72 screens, 144 questions (126 categorical choice/noul, 18 ordinal score). Old test: 60 screens, 120 questions (90 categorical, 30 score). All four historical stages are shown, not only the released one. | Historical stage | New categorical | Completion precision | Completion recall | Unknown correct | New score MAE | Old categorical | Old score MAE | |---|---:|---:|---:|---:|---:|---:|---:| | Original A | 48/126 | 21/59 | 21/24 | 6/28 | 0.736238 | 43/90 | 0.973761 | | Head only | 51/126 | 11/27 | 11/24 | 9/28 | 0.723945 | 42/90 | 0.913887 | | Projector + head | 93/126 | 22/23 | 22/24 | 23/28 | 0.574737 | 52/90 | 0.605605 | | Released projector + head + text LoRA | 92/126 | 22/23 | 22/24 | 23/28 | 0.574219 | 52/90 | 0.570224 | The released stage has one false completion among 48 non-complete/uncertain new-test cases: 0/24 known negatives and 1/24 uncertain cases. It has two missed new-test positives. **All four stages miss all seven positive completion examples in the older test.** No positives are predicted there, so old-test completion precision is undefined; zero false positives is not evidence of successful completion recognition. The LoRA stage did not add new-test categorical accuracy over projector+head. It remains the release candidate because selection was validation-only; held-out results did not reselect it. The original unlabelled 58-request/70-question regression suite is not an accuracy benchmark: the released stage changes 9/46 choice/noul decisions versus current original A (maximum probability drift 0.734239; maximum expected-score drift 1.496044). Limitations include synthetic layout/text/color regularities, finite family splits, overconfidence/train-versus-validation loss separation, residual Ready-identifier/completion association (0.622556 bits within sparse target groups), and joint family/option-position association (0.584963 bits). Zero conditional viewport MI in the specified QC groups is not proof of universal nuisance independence. No independently developed real-site dataset was collected for this release. ## Measured Runtime Scope The existing own-checkpoint FP32 source API was compared with desktop ONNX and actual Chrome 154 WebGPU on an NVIDIA RTX 5090, non-software adapter. Fixtures were 72 validation screenshots plus 20 neutral text requests: 92 requests/169 questions, not held-out quality examples. Each browser path had two warm-ups and ten hot repeats; all 145 categorical questions agreed on every hot repeat. Maximum absolute probability/expected-score errors stayed below the unchanged 0.001 gate. | Browser path | Max probability difference | Max expected-score difference | |---|---:|---:| | Same native media prefix | 0.0000563264 | 0.000109192 | | Actual PNG preprocessing + vision/projector | 0.0000483990 | 0.0000722781 | Tokens, pixels, masks and shapes matched their references exactly. Position-interpolation FP32 order differences reached 0.000000774860; the trained intermediate media-prefix difference reached 0.0565567. There is no claim of bit-exact intermediate activations or a 0.001 intermediate gate. Observed hot p50/p95 milliseconds on that machine: text decision 30.778/151.935; saved image-prefix decision 67.290/272.953; full image inference 219.943/751.239; PNG preprocessing plus inference 337.175/957.691. These finite request-balanced measurements are not throughput guarantees. Model/graph storage bytes are not VRAM consumption. Separate three-fixture profiling observed 1,465 WebGPU nodes and 407 CPU/WASM nodes, including four floating mask/position construction nodes; this is **not pure GPU execution**. Separate bounded memory sampling observed adapter-total memory, not isolated model VRAM. Existing runner parity is not, by itself, proof of a newly integrated packed WebBrain extension. ## Adapter Use The package contains the dependency-injected adapter copied from the actual [WebBrain](https://github.com/webbrain-one/webbrain) extension module `src/chrome/src/providers/d1-runtime.js`. Its preprocessing helper preserves the old verified bytes except two explicit source-semantic corrections: present-null noul criteria (`Object.hasOwn`, matching Python `dict.get`) and small fractional Python-style JSON exponent formatting. The old baseline helper is untouched. JavaScript numbers cannot recover Python int-versus-integral-float lexical types: use preformatted state/criterion strings when exact original lexical representation matters. The adapter also restores the source 65,536 padded-token subbatch budget while returning named answers in original order. Executable JavaScript/WASM must be bundled locally for extension CSP; do not load remote executable code. Immutable-revision model graph/tokenizer downloads are data and must be checksum-verified before caching/creating sessions. The exact API is `createD1Runtime({ort, tokenizer, config, ratios, sessions, device, model})`, then `evaluate({state, images, questions, signal})`. See [runtime/README.md](runtime/README.md) for session wiring. `noul` answers expose the source public yes-probability; score answers expose expected ordinal level, not argmax-class accuracy. Option insertion order, masks, media prefix, calibration and the image text limit of 896 tokens must remain unchanged. ```js import { createD1Runtime } from './runtime/d1-runtime.js'; // Supply verified package assets, the bundled ORT/tokenizer, and three FP32 sessions. const judge = createD1Runtime({ ort, tokenizer, config, ratios, sessions, device, model: 'd1-browser-decision-fp32' }); const result = await judge.evaluate({ state: { task: 'Check whether a visible receipt establishes completion.' }, images: [inlinePngDataUrl], questions: { completion: { type: 'choice', instructions: 'Judge only visible evidence.', criteria: { completed: 'An explicit receipt confirms the named task.', not_completed: 'Visible evidence establishes failure or an unfinished task.', unknown: 'The screenshot does not establish the outcome.' } } } }); // result.usage.output_tokens === 0; this is not text generation. ``` ## License And Notices The model is governed by the exact upstream [LFM Open License v1.0](LICENSE), not Apache/MIT. Its commercial-use provisions include annual-revenue threshold terms of US$10 million; determine eligibility and obtain any required separate license before commercial deployment. This card is not legal approval or an endorsement by Liquid AI. [NOTICE](NOTICE) and modified-binary sidecars retain attribution and identify the historical derivative changes without changing verified binary bytes. The underlying [LFM2.5-Encoder license reference](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M/blob/b886781f7c6f10ca9b7096e21b83e30a073c2f39/LICENSE) is separately retained; that license-reference revision is not a claim about the historical base-weight revision used by d1. Bundled ONNX Runtime 1.31.0-dev.20260914-8d85527a0 is MIT-licensed; Transformers.js 4.3.1 is Apache-2.0-licensed. The actual WebBrain GPL-3.0-or-later project notice is retained for the copied runtime source at [licenses/webbrain/LICENSE](licenses/webbrain/LICENSE). Separate component notices do not replace the model license or constitute a legal compatibility/commercial-eligibility opinion. ## Native WebGPU Runtime Dependency Correction This revision retains the same checkpoint, six ONNX graph/data files, tokenizer, config, decision adapter and preprocessing helper. It adds the exact same-version ORT asyncify MJS/WASM pair and corrects the local loader example. The native WebGPU bundle calls `webgpuInit`; explicitly forcing the JSEP factory (`jsepInit` only) failed during an actual packed-extension initialization before neural inference. The earlier successful numerical browser runner used the same ORT distribution's directory-prefix loader, which selected the native asyncify pair. Existing JSEP files remain preserved but must not be selected for this native WebGPU bundle. See `runtime/README.md`, `runtime/vendor/vendor-manifest.json` and the pinned `package-manifest.json` checksums. This is a dependency closure correction, not a new export, precision change or new quality result. It does not itself establish packed-extension inference success or real-site generalization. The existing synthetic evaluation, license restrictions and unverified real-site scope remain.