--- license: mit base_model: microsoft/harrier-oss-v1-0.6b base_model_relation: quantized pipeline_tag: feature-extraction inference: false quantized_by: shadowrock-io library_name: vllm metrics: - cosine_similarity language: - en - multilingual model-index: - name: Harrier-OSS-v1-0.6B-Community-FP8 results: - task: type: STS dataset: name: MTEB STSBenchmark type: mteb/stsbenchmark-sts config: default split: test revision: b0fddb56ed78048fa8b90373c8a3cfc37b684831 metrics: - type: cosine_spearman name: STSBenchmark main score (vLLM-served FP8) value: 0.82532 source: name: ShadowRock eval (raw JSON) url: https://huggingface.co/shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8/tree/main/eval/evidence - task: type: STS dataset: name: MTEB STS17 (multilingual, 11 language-pair subsets, mean) type: mteb/sts17-crosslingual-sts config: default split: test revision: faeb762787bd10488a50c8b5be4a3b82e411949c metrics: - type: cosine_spearman name: STS17 mean main score (vLLM-served FP8) value: 0.81317 source: name: ShadowRock eval (raw JSON) url: https://huggingface.co/shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8/tree/main/eval/evidence - task: type: Retrieval dataset: name: MTEB SciFact type: mteb/scifact config: default split: test revision: 0228b52cf27578f30900b9e5271d331663a030d7 metrics: - type: ndcg_at_10 name: SciFact NDCG@10 (vLLM-served FP8) value: 0.73009 source: name: ShadowRock eval (raw JSON) url: https://huggingface.co/shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8/tree/main/eval/evidence - task: type: Retrieval dataset: name: MTEB NFCorpus type: mteb/nfcorpus config: default split: test revision: ec0fa4fe99da2ff19ca1214b7966684033a58814 metrics: - type: ndcg_at_10 name: NFCorpus NDCG@10 (vLLM-served FP8) value: 0.32504 source: name: ShadowRock eval (raw JSON) url: https://huggingface.co/shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8/tree/main/eval/evidence tags: - fp8 - e4m3 - w8a8 - llm-compressor - compressed-tensors - vllm - embeddings - retrieval - semantic-search - sentence-similarity - quantized - safetensors - qwen3 - multilingual --- ShadowRock # Harrier OSS v1 0.6B — Community FP8 **Unofficial community quantization — not a microsoft release.** FP8 (W8A8, `FP8_DYNAMIC`) quantization of [microsoft/harrier-oss-v1-0.6b](https://huggingface.co/microsoft/harrier-oss-v1-0.6b) (revision [`f9b9dc8`](https://huggingface.co/microsoft/harrier-oss-v1-0.6b/commit/f9b9dc8d367d443f2479d27aa5d8d2850c0774ee)), a 0.6B multilingual Qwen3Model embedding model using last-token pooling + L2 normalization (the same pooling/prompting family as the 270M sibling in this collection). All credit for the model belongs to Microsoft; this repo only changes the numeric precision of the transformer's `Linear` projection weights and activations. 1.19 GB BF16 → 0.70 GB FP8, and fixture cosine parity vs the BF16 golden is 0.9985 mean / 0.9957 min (see Benchmarks below) with MTEB deltas under 0.002 absolute on every task measured. Pick this variant for the best fidelity/size tradeoff on a GPU that supports FP8 (Ada/Hopper/ Blackwell). The companion NVFP4 variant trades additional fidelity for a further ~2x size reduction on Blackwell-class GPUs — see its own README for the pass/fail verdict. ## What is quantized / Quantization details - Method: `FP8_DYNAMIC` (compressed-tensors preset) via [llm-compressor](https://github.com/vllm-project/llm-compressor) 0.13.0 — static per-channel FP8 E4M3 weight scales computed directly from the weights, dynamic per-token FP8 E4M3 activation quantization at inference. - Scope: all 196 `nn.Linear` projections inside the 28 Qwen3Model decoder layers (`self_attn.{q,k,v,o}_proj`, `mlp.{gate,up,down}_proj`). Left at full precision: the embedding table (`embed_tokens`, tied to `lm_head`), every RMSNorm module, `lm_head`, and the sentence-transformers last-token pooling + L2-normalize heads (outside the transformer module tree, always full precision). - Calibration: 448-sample pinned manifest (`calibration/manifest.json`, seed 20260818973, own pool — disjoint from the 270m sibling's calibration set), bucketed by (input_type, token-length) — passages sampled from `mteb/sts17-crosslingual-sts` (disjoint from this collection's own eval fixture set), queries authored across the model's three documented instruct-prefix types (`web_search_query`/`sts_query`/`bitext_query`, confirmed byte-identical to the 270m sibling's prompts). Not strictly required for `FP8_DYNAMIC`'s data-free weight scales, but run for parity with the NVFP4 export path. - Full module inventory / provenance file: [`QUANT_PROVENANCE.json`](https://huggingface.co/shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8/blob/main/QUANT_PROVENANCE.json). - Quantize/eval scripts ship under `quantization/`, `eval/`, `calibration/`; raw eval JSON under `eval/evidence/`. ## Benchmarks / Validation Two complementary evals, both against this collection's own BF16 golden baseline computed on the same hardware (arcade, RTX 5070 Ti) with the same fixture/prompt pinning (see `eval/README.md` for the full protocol — own 190-fixture set, seed 20260818602, not reused from the 270m sibling): (1) fixture cosine parity via vLLM `llm.embed()`, gated by `fixtures/tolerances.json` (schema v1, `parity_harness.py`); (2) MTEB on the same 4-task subset used for the 270m sibling, run via a thin vLLM-backed MTEB encoder wrapper (`eval/run_mteb_vllm.py`) — **required** for this variant, since a `compressed-tensors` FP8 checkpoint cannot execute a raw `transformers`/`sentence-transformers` forward pass. | eval | metric | BF16 (vLLM) | FP8 | delta | gate | verdict | |---|---|---|---|---|---|---| | fixture parity (190 cases, non-stress) | cosine, p5 aggregate | 1.0 (self) | 0.9985 | — | ≥ 0.995 | PASS | | fixture parity (190 cases, non-stress) | cosine, per-case floor | 1.0 (self) | ≥ 0.9957 | — | ≥ 0.99 | PASS | | MTEB STSBenchmark | main score | 0.82529 | 0.82532 | +0.00003 | ≤ 0.01 | PASS | | MTEB STS17 (mean/11 subsets) | main score | 0.81352 | 0.81317 | -0.00035 | ≤ 0.01 | PASS | | MTEB SciFact | NDCG@10 | 0.72841 | 0.73009 | +0.00168 | ≤ 0.01 | PASS | | MTEB NFCorpus | NDCG@10 | 0.32494 | 0.32504 | +0.00010 | ≤ 0.01 | PASS | | MTEB macro delta (4 tasks) | mean absolute delta | — | — | 0.00054 | ≤ 0.005 | PASS | All gates pass with wide margin — comfortably ahead of both the FP8 fixture-parity and MTEB gates, and materially better than the 270m sibling's own (also-passing) FP8 numbers, consistent with a larger model having more redundancy to absorb 8-bit quantization noise. No case (incl. the long-input stress fixture) fell below its gate. Raw result JSON: `eval/evidence/fp8_fixture_parity_report.json`, `eval/evidence/fp8_mteb.json`, `eval/evidence/golden_bf16_mteb_vllm.json`. ### Protocol deviations from the paper / base card - MTEB numbers for this variant use a custom vLLM-backed MTEB encoder wrapper (`eval/run_mteb_vllm.py`), not the base card's documented `sentence-transformers` load path — required because compressed-tensors FP8 checkpoints only run through vLLM's fused kernels. The BF16-vs-FP8 deltas above compare **BF16-via-vLLM vs FP8-via-vLLM** (both in `eval/evidence/`) for an apples-to-apples backend comparison; the sentence-transformers BF16 numbers (also in `eval/evidence/golden_bf16_mteb.json`) differ only slightly (max 0.0013 per task) from the vLLM-served path, confirming vLLM serving reproduces the base card's documented behavior. - STS17/NFCorpus/SciFact dataset revisions above are the `mteb`-library-pinned snapshot hashes as resolved at eval time (2026-08-18). ## Serving / Usage ```python from vllm import LLM llm = LLM( model="shadowrock-io/Harrier-OSS-v1-0.6B-Community-FP8", runner="pooling", # task=embed trust_remote_code=True, max_model_len=8192, # base model supports up to 32768; raise if you need longer inputs gpu_memory_utilization=0.85, # lower if sharing the GPU with other processes dtype="bfloat16", # activation/compute dtype; FP8 weights dequant/fuse automatically ) # Documents: no prefix needed. doc_emb = llm.embed(["Paris is the capital and largest city of France."]) # Queries: apply the base model's documented instruct prefix (three task-specific variants exist # in the base repo's config_sentence_transformers.json `prompts` map — this is the general one). query = "Instruct: Given a web search query, retrieve relevant passages that answer the query\\nQuery: what is the capital of France" query_emb = llm.embed([query]) ``` No `PoolerConfig` override is needed: vLLM auto-resolves `pooling_type=LAST` and L2-normalization directly from this repo's `1_Pooling/config.json` + `config_sentence_transformers.json` (carried over unchanged from the BF16 base repo) — same as the BF16 model and the 270m sibling. FP8 requires no special vLLM flags beyond loading this repo; vLLM detects the `compressed-tensors` FP8 config in `config.json` automatically and uses fused FP8 kernels on GPUs that support them (Ada/Hopper/Blackwell — tested on an RTX 5070 Ti, Blackwell sm_120). Notes that matter for correct embedding output: - Pooling is **last-token**, not mean/CLS — confirmed via `pooling_mode_lasttoken: true` in this repo's `1_Pooling/config.json` and empirically (all fixture embeddings are unit-norm to within 6e-8). - The query instruct-prefix meaningfully changes the embedding vs. the same text posed as a bare passage (asymmetric by design) — apply it only on the query side, never on documents. Measured on: arcade (RTX 5070 Ti, 16 GB, Blackwell sm_120), vLLM 0.26.0, llm-compressor 0.13.0, compressed-tensors 0.18.0, torch 2.11.0+cu130, mteb (pinned per `eval/evidence/*.json`). ## Caveats - The 448-sample calibration set skews toward short-to-medium sentence-length inputs (STS-style sentences + short authored queries); very long documents were exercised only by this collection's single adversarial stress fixture, which showed no measurable fidelity drop beyond the general FP8 noise floor. - MTEB coverage is 4 small tasks (STSBenchmark, STS17, SciFact, NFCorpus); the base model claims broad multilingual and MTEB task-type coverage not evaluated here — only STS/retrieval-style tasks and 11 STS17 language pairs were measured. ## Intended use & limitations Intended uses are the base model's: general-purpose multilingual text embeddings for retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The base card's intended-use, safety, and coverage statements — [microsoft/harrier-oss-v1-0.6b](https://huggingface.co/microsoft/harrier-oss-v1-0.6b) — carry over unchanged; quantization alters none of the model's behavior boundaries, only its numeric precision. Our evaluation establishes parity on the benchmarks listed above and nothing beyond them: other languages, MTEB task types, and document-length regimes inherit the base model's behavior with FP8 quantization noise that we have not measured there. ## Attribution & citation Quantization, validation harness, and card by [Matt Busi](https://www.linkedin.com/in/matt-busi) ([@mattbusi](https://huggingface.co/mattbusi) on Hugging Face) at [ShadowRock](https://shadowrock.io). If you use this build, cite the microsoft base model — the underlying model architecture and training are theirs: ```bibtex @misc{harrier-oss-v1-0.6b, title = {Harrier OSS v1 0.6B}, author = {Microsoft}, year = {2026}, url = {https://huggingface.co/microsoft/harrier-oss-v1-0.6b} } ``` ## License MIT, inherited unchanged from the base model (see [LICENSE](LICENSE); the base repo declares `license: mit` in its card metadata but ships no LICENSE file, so the canonical MIT text is reproduced here). Community build by [ShadowRock](https://shadowrock.io); no Microsoft affiliation or endorsement. ## About ShadowRock [ShadowRock](https://shadowrock.io) is an AI-specialized systems integrator and Zendesk Premier Partner. We help businesses get real value from their go-to-market technology, from CRM and support platforms to applied AI like the models in this collection. Find us at [shadowrock.io](https://shadowrock.io) or on [LinkedIn](https://www.linkedin.com/company/shadowrock).