kleeedolinux commited on
Commit ·
111207e
1
Parent(s): 0949494
Readme update
Browse files
README.md
CHANGED
|
@@ -69,6 +69,10 @@ Julia 1 starts from [JHU CLSP's mmBERT-small](https://huggingface.co/jhu-clsp/mm
|
|
| 69 |
|
| 70 |
The parameter and context figures for mmBERT-small come from its [model card](https://huggingface.co/jhu-clsp/mmBERT-small). Julia's values describe this checkpoint and its evaluated runtime, not a claim that mmBERT-small itself is limited to Julia's interface. Julia is not a chat or text-generation model and is not interchangeable with `AutoModelForMaskedLM`.
|
| 71 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 72 |
## Start here
|
| 73 |
|
| 74 |
Python 3.11+ is required. CPU inference works with the standard PyTorch installation; no native router build is needed. Download the complete repository, including its 550.5 MiB checkpoint, then install its Python package:
|
|
|
|
| 69 |
|
| 70 |
The parameter and context figures for mmBERT-small come from its [model card](https://huggingface.co/jhu-clsp/mmBERT-small). Julia's values describe this checkpoint and its evaluated runtime, not a claim that mmBERT-small itself is limited to Julia's interface. Julia is not a chat or text-generation model and is not interchangeable with `AutoModelForMaskedLM`.
|
| 71 |
|
| 72 |
+
## WebGPU and ONNX
|
| 73 |
+
|
| 74 |
+
The separate [Julia-1-ONNX repository](https://huggingface.co/SupersonicLabs/Julia-1-ONNX) contains the full ONNX export and JavaScript WebGPU adapter. It loads the model once, warms it in memory, and runs inference through ONNX Runtime WebGPU. The WebGPU adapter source and Rust N-API/WebAssembly tokenizer live in that repository.
|
| 75 |
+
|
| 76 |
## Start here
|
| 77 |
|
| 78 |
Python 3.11+ is required. CPU inference works with the standard PyTorch installation; no native router build is needed. Download the complete repository, including its 550.5 MiB checkpoint, then install its Python package:
|