kleeedolinux commited on
Commit
111207e
·
1 Parent(s): 0949494

Readme update

Browse files
Files changed (1) hide show
  1. README.md +4 -0
README.md CHANGED
@@ -69,6 +69,10 @@ Julia 1 starts from [JHU CLSP's mmBERT-small](https://huggingface.co/jhu-clsp/mm
69
 
70
  The parameter and context figures for mmBERT-small come from its [model card](https://huggingface.co/jhu-clsp/mmBERT-small). Julia's values describe this checkpoint and its evaluated runtime, not a claim that mmBERT-small itself is limited to Julia's interface. Julia is not a chat or text-generation model and is not interchangeable with `AutoModelForMaskedLM`.
71
 
 
 
 
 
72
  ## Start here
73
 
74
  Python 3.11+ is required. CPU inference works with the standard PyTorch installation; no native router build is needed. Download the complete repository, including its 550.5 MiB checkpoint, then install its Python package:
 
69
 
70
  The parameter and context figures for mmBERT-small come from its [model card](https://huggingface.co/jhu-clsp/mmBERT-small). Julia's values describe this checkpoint and its evaluated runtime, not a claim that mmBERT-small itself is limited to Julia's interface. Julia is not a chat or text-generation model and is not interchangeable with `AutoModelForMaskedLM`.
71
 
72
+ ## WebGPU and ONNX
73
+
74
+ The separate [Julia-1-ONNX repository](https://huggingface.co/SupersonicLabs/Julia-1-ONNX) contains the full ONNX export and JavaScript WebGPU adapter. It loads the model once, warms it in memory, and runs inference through ONNX Runtime WebGPU. The WebGPU adapter source and Rust N-API/WebAssembly tokenizer live in that repository.
75
+
76
  ## Start here
77
 
78
  Python 3.11+ is required. CPU inference works with the standard PyTorch installation; no native router build is needed. Download the complete repository, including its 550.5 MiB checkpoint, then install its Python package: