--- title: README emoji: ๐ŸŸง colorFrom: yellow colorTo: red sdk: static pinned: false ---

Primitive โ€” inference economics for production LLM systems

primitive.com Inference economics vLLM and ONNX Nine open-weight builds

--- ## Serving LLMs well is two problems **Which model should answer this request?** and **how cheaply can that model run?** Most teams solve neither โ€” they pin one frontier model to every route and pay frontier prices on the easy majority of their traffic. We build the two pieces that fix that.
### ๐Ÿ”€ Route A small, fast model that reads the prompt **before** you spend anything on it, then ranks which LLM is actually worth calling. Runs on CPU next to your API server โ€” no GPU, no PyTorch, no extra service to operate. ### ๐Ÿ—œ๏ธ Compress Mixed-precision quantization that shrinks a checkpoint and speeds up serving **without quietly trading away accuracy** โ€” and that still loads on the GPUs you already have, not just the newest ones.
--- ## Out now: nine open-weight builds

Every published build on one pinned protocol โ€” 1,370 items across fourteen public benchmarks, split into knowledge, tool calls and abstentions. Qwen3.8-Flash-Next leads overall at 90.3; Ornith-1.5 leads on knowledge at 91.7 but sits near the bottom on calls at 74.4; Laguna-S-2.1 is second from last overall yet best of all on abstentions at 81.0

Weights-only mixed precision, `compressed-tensors`, stock `vllm serve`. Every figure above is on the model's own card, next to the run that produced it. **One protocol for everything we publish.** 1,370 items across fourteen public benchmarks โ€” a 1,170-item knowledge suite over nine of them and a 200-item tool-calling suite over five more, scored as calls and abstentions separately and never blended. Same decode settings for every model, same 16k budget, thinking on. Tool-calling figures are means of repeated runs, because a single run of that suite can move three points on the same weights. * ๐Ÿš€ **[Qwen3.8-Flash-Next โ€” mixed NVFP4/FP8](https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-mixed-NVFP4-FP8)** ยท A **180 B model on one 96 GB GPU**. 360 GB in BF16, two data-center cards in FP8 โ€” this serves on one, with the 51 B n-gram table offloaded to host RAM and no patched runtime. Our highest score anywhere: **90.3 overall, 92.2 knowledge**. The [plain NVFP4 build](https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-NVFP4) is the same deployment; this one puts the attention and GDN projections at FP8 and is **13% faster single-stream** for an accuracy tie. * ๐Ÿงฎ **[Qwen3.8-27B โ€” mixed NVFP4/FP8](https://huggingface.co/primitive-ai/Qwen3.8-27B-mixed-NVFP4-FP8)** ยท 20.7 GiB, **2.5ร— smaller than BF16** and 28% smaller than Qwen's own FP8. The **shortest answers of any build we measured** (651 tokens) and the highest completion rate. **Loads on an A100**, not just Blackwell โ€” the BF16 `lm_head` is what keeps that path open โ€” and the MTP head survives, so speculative decoding still works. * ๐Ÿฆ… **[Ornith-1.5-35B-A3B โ€” mixed NVFP4/FP8](https://huggingface.co/primitive-ai/Ornith-1.5-35B-A3B-mixed-NVFP4-FP8)** ยท A 35 B MoE vision-language model at 22.6 GiB, **3.0ร— smaller than BF16**, and **the best knowledge score in the lineup at 91.7**. Tool calling is its weak axis, which is why it has a sibling โ†“ * ๐Ÿ› ๏ธ **[Ornith-1.5-35B-A3B โ€” agentic](https://huggingface.co/primitive-ai/Ornith-1.5-35B-A3B-agentic-NVFP4-FP8)** ยท Same model, same size, calibrated for tool use. Level with BF16 on tool calling and **2.8ร— steadier run to run** (sd 1.3 against 3.1). We do *not* claim it is more accurate at tool calling than the general build โ€” that gap does not survive the noise. What reproduces is the consistency. * โšก **[Nemotron-3.5-Lightning-30B-A3B โ€” mixed INT4/INT8](https://huggingface.co/primitive-ai/Nemotron-3.5-Lightning-30B-A3B-mixed-INT4-INT8)** ยท A MoE+Mamba hybrid at 19.2 GiB and 5.01 bits/weight, **3.3ร— smaller than BF16 at 2.7ร— its throughput**, and **8.6% smaller and ~7.7% faster than nvidia's own NVFP4**. The strongest tool-calling profile here โ€” 85.4 on calls and 70.5 on abstentions, the best balance of the two in the lineup. * ๐Ÿ‘๏ธ **[Muse-Glimmer-30B โ€” mixed NVFP4](https://huggingface.co/primitive-ai/Muse-Glimmer-30B-mixed-NVFP4)** ยท A vision-language model at 20.4 GiB, the smallest of its field, tied with BF16 on knowledge. **The vision tower stays BF16** โ€” a VLM loses its visual grounding before it loses its words. * ๐Ÿชถ **[Laguna-XS-2.1 โ€” mixed NVFP4/MXFP8](https://huggingface.co/primitive-ai/Laguna-XS-2.1-mixed-NVFP4-MXFP8)** ยท 19.3 GiB at 4.96 bits/weight, **3.2ร— smaller than BF16 and smaller than every official quant** โ€” and **2 points of pooled score ahead of the official NVFP4**, almost all of it in the tool-calling half. * ๐Ÿ˜ **[Laguna-S-2.1 โ€” mixed NVFP4/MXFP8](https://huggingface.co/primitive-ai/Laguna-S-2.1-mixed-NVFP4-MXFP8)** ยท 219 GB of source weights down to 64 GiB at 4.67 bits/weight โ€” **it serves from one 96 GiB GPU**, the smallest of the 4-bit field, and **the best abstention score of anything we publish (81.0)**: the one that knows when *not* to call a tool. * ๐Ÿงฐ **[Qwen3.8-Flash-Next โ€” quantized PLE tables](https://huggingface.co/primitive-ai/Qwen3.8-Flash-Next-PLE-quant)** ยท Not a checkpoint โ€” the 51.2 B-parameter n-gram tables that Flash-Next offloads, quantized from 95.4 GB to **28.8 GB** and served memory-mapped from disk. Validated end to end in a **48 GB container**, so a 64 GB-RAM host can serve the 180 B model. **Read the halves, not just the pooled score.** `overall` weights tool calling at 14.6% because that is what item count says โ€” which shrinks a tool-calling-only divergence to a seventh of its size. Ornith leads the lineup on knowledge and sits near the bottom on calls; Laguna-S is second from last overall and best of all on abstentions. Abstention is the weak axis on every model anyone has measured, ours included, so a build can look strong on one number and still over-call in an agent. ## Still in private beta The router is not published yet. We're working with a small number of teams running real production traffic โ€” the ones who can tell us where the sharp edges are before everyone else finds them. **A good fit if you're:** paying per-token to a frontier API on traffic you suspect doesn't need it, serving an open-weights model at meaningful volume, or squeezing a large checkpoint onto GPUs you already own.

Request access at primitive.com โ†’

--- ## How we publish * **Numbers or nothing.** Every figure on our model cards comes from a held-out suite and is reproducible from files in the repo. Where a metric is relative, we say what it's relative to. * **We document what breaks.** Our cards carry the failure modes, not just the wins โ€” the dtype choices that matter, the upstream quirks that bite, and the optimisations we tried and *rejected*. * **No silent quality loss.** A compression or routing win that changes answers isn't a win. When a cheaper option moved outputs too much, we don't ship it โ€” and we say so on the card. * **A tie is a tie.** Our accuracy deltas sit inside the run-to-run spread of the suite, and we label them as ties rather than dressing them up as wins. * **Runs on hardware you have.** Older-generation GPU fallbacks for the quants, CPU-only for the router. Nothing here assumes you just bought the newest rack. --- ## Repositories | repo | what it does | status | |---|---|---| | `Qwen3.8-Flash-Next-mixed-NVFP4-FP8` | 180 B MoE on one 96 GB GPU, FP8 attention + GDN projections. Our highest-scoring build. | **Open weights** | | `Qwen3.8-Flash-Next-NVFP4` | The same single-GPU deployment, plain NVFP4. | **Open weights** | | `Qwen3.8-Flash-Next-PLE-quant` | The 51.2 B n-gram tables quantized to 28.8 GB, mmapped from disk. Not a checkpoint. | **Open weights** | | `Qwen3.8-27B-mixed-NVFP4-FP8` | Mixed-precision quant of Qwen3.8-27B. vLLM-native, A100 fallback, MTP head preserved. | **Open weights** | | `Ornith-1.5-35B-A3B-mixed-NVFP4-FP8` | Mixed-precision quant of Ornith-1.5-35B-A3B, a 35 B MoE VLM. Best knowledge score here. | **Open weights** | | `Ornith-1.5-35B-A3B-agentic-NVFP4-FP8` | Same model calibrated for tool use โ€” 2.8ร— steadier run to run. | **Open weights** | | `Nemotron-3.5-Lightning-30B-A3B-mixed-INT4-INT8` | INT4/INT8 quant of nvidia's MoE+Mamba hybrid. Smaller and faster than the official NVFP4. | **Open weights** | | `Muse-Glimmer-30B-mixed-NVFP4` | NVFP4 quant of the Muse-Glimmer-30B VLM, GPTQ-calibrated. Vision tower left at BF16. | **Open weights** | | `Laguna-XS-2.1-mixed-NVFP4-MXFP8` | Mixed-precision quant of poolside's Laguna-XS-2.1. Smallest build of it we've measured. | **Open weights** | | `Laguna-S-2.1-mixed-NVFP4-MXFP8` | Mixed-precision quant of poolside's Laguna-S-2.1. 219 GB of weights onto a single 96 GiB card. | **Open weights** | | `llm-router-lite-80m-onnx` | CPU-only router: reads a chat `messages` array, ranks which LLM to call. ONNX, no torch or transformers. | Private beta | Published weights carry their upstream licence โ€” Apache-2.0 for the Qwen3.8-27B and Muse-Glimmer builds, Qwen Community 1.0 for the Flash-Next builds, OpenMDW-1.1 for the Laguna and Nemotron builds, MIT for both Ornith builds, and Apache-2.0 for the PLE tables. Derivative models carry upstream attribution and a `NOTICE` or `LICENSE` file. ---

Cutting your inference bill?

We'd like to hear what you're serving and where it hurts.
primitive.com


Primitive
primitive ยท inference economics for production LLM systems