Instructions to use mlx-community/Nemotron-Parse-2.0-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/Nemotron-Parse-2.0-8bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("mlx-community/Nemotron-Parse-2.0-8bit") config = load_config("mlx-community/Nemotron-Parse-2.0-8bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Nemotron-Parse-2.0-8bit
8-bit MLX quantization of
nvidia/NVIDIA-Nemotron-Parse-2.0, a 903M-parameter
encoder-decoder OCR model (C-RADIO vision encoder + mBART-style decoder).
Converted with mlx-vlm 0.6.12 (mlx 0.32.0) for inference on Apple Silicon.
This is the first MLX conversion of the Nemotron-Parse architecture. The model is image-to-text only: the prompt text is ignored, the image alone drives the transcription (this matches the source implementation).
Fidelity was measured against the source, not assumed. The numbers below and the method that produced them are in the tables that follow.
Loading requires the
nemotron_parsearchitecture port, new in mlx-vlm (PR #1866). Until it merges, install the port branch:pip install git+https://github.com/axiom-of-choice/mlx-vlm@feat/nemotron-parse
pip install mlx-vlm
python -m mlx_vlm generate \
--model mlx-community/Nemotron-Parse-2.0-8bit \
--image document.png \
--prompt "Extract the text" \
--max-tokens 512
Quantization
| Bits | 8 |
| Group size | 64 |
| Mode | affine |
| Effective bits/weight | 14.689 |
| Size on disk | 1.6 GB |
| Quantized tensors | 102 (decoder, group 64) |
| Unquantized tensors | 766 (vision tower + neck, bf16) |
Effective bits per weight exceeds the nominal 8 because mlx-vlm
quantizes only the language side and leaves the C-RADIO vision encoder in
bf16 by design. The vision tower is 72% of all parameters, which is why the
effective figure is much higher than on a decoder-only model. Verified, not
assumed: of 766 unquantized tensors, zero carry a .scales entry.
Measured fidelity
Weight-level metrics compare dequantized weights against the source tensor by tensor, streamed from disk (the source ships fp32; the 8-bit SNR band for a bf16-comparable baseline is 42.6–42.8 dB). No prompts, no sampling, hardware-independent.
| Variant | bpw | SNR | Relative L2 | Cosine | Decode tok/s | Peak RAM |
|---|---|---|---|---|---|---|
| 4-bit | 13.290 | 20.87 dB | 9.05% | 0.995908 | 236.2 | 14.34 GB |
| 8-bit | 14.689 | 42.76 dB | 0.73% | 0.999974 | 215.5 | 14.50 GB |
| bf16 (source) | 16 | n/a | n/a | n/a | 63.7 | 14.79 GB |
Measured on an M2 Pro / 32 GB, greedy decoding.
OCR accuracy vs ground truth
Because OCR has ground truth, transcription can be scored directly rather than inferred from token distributions. Six documents were rendered locally with exactly known content (invoice, clinical lab report, shipping label, receipt, motor-controller datasheet, rotated receipt).
| Variant | Field accuracy | Content accuracy | Numeric recall |
|---|---|---|---|
| 4-bit | 0.8026 | 0.8026 | 0.8333 |
| 8-bit | 0.9474 | 0.9474 | 0.8914 |
| bf16 (source) | 0.9474 | 0.9474 | 0.9353 |
- Field accuracy: required strings present, exact match (format-sensitive).
- Content accuracy: the same, ignoring markup and whitespace.
- Numeric recall: ground-truth numbers present, ignoring markup.
Per-document field accuracy:
| Doc | bf16 | 4-bit | 8-bit |
|---|---|---|---|
| invoice | 1.0000 | 1.0000 | 1.0000 |
| lab_report | 0.8571 | 1.0000 | 0.8571 |
| shipping_label | 1.0000 | 1.0000 | 1.0000 |
| receipt | 1.0000 | 1.0000 | 1.0000 |
| spec_table | 0.8667 | 0.0000 | 0.8667 |
| rotated_receipt | 1.0000 | 1.0000 | 1.0000 |
The 4-bit spec_table regression
At 4 bits the datasheet document transcribes as a degenerate & \ repetition
loop (field accuracy 0.00 vs 0.87 at 8 bits). The argmax flips at the second
decoding step and the loop is self-reinforcing; NVIDIA's own generation config
(repetition_penalty 1.1, applied in this harness) does not break it. This is a
real quantization trade at 4 bits on dense table layouts — the same document
class where the model's native LaTeX output is token-heavy. Every other
document scores 1.00 at 4 bits.
The greedy path is otherwise near-tie sensitive: on lab_report the 4-bit
variant transcribes all fields where bf16 and 8-bit miss two (a 0.00-vs-0.05
logit margin flips the path), so per-document ordering between 4/8/bf16 is
not meaningful; the aggregate tables are.
What was not measured
- No distributional metrics (perplexity, top-1 agreement, KL). The model is image-to-text only — text prompts are ignored — so a text-passage harness cannot score it.
- No standard OCR benchmark. No OmniDocBench, no DocVQA. The six documents are synthetic and rendered locally: exact ground truth, but not photographs, not handwriting, not skewed scans, and not non-Latin scripts.
- No judged quality, no layout-structure scoring, no throughput figures for any chip other than an M2 Pro.
- The source's own golden generation sequence is hardware-sensitive. The model's greedy path diverges between CUDA bf16 (NVIDIA's golden file) and CPU/Metal bf16 at the third token; the MLX port reproduces the CPU reference byte-for-byte. Verified against the Hugging Face implementation on the same inputs.
Sibling variants
| Variant | bpw | Field accuracy |
|---|---|---|
| Nemotron-Parse-2.0-4bit | 13.290 | 0.8026 |
| Nemotron-Parse-2.0-8bit | 14.689 | 0.9474 |
Credits
All credit for the underlying model belongs to
NVIDIA. This is a format conversion and
quantization; no training or fine-tuning was performed. The architecture port
(mlx_vlm/models/nemotron_parse/) is new in mlx-vlm.
- Downloads last month
- -
8-bit
Model tree for mlx-community/Nemotron-Parse-2.0-8bit
Base model
nvidia/NVIDIA-Nemotron-Parse-2.0