--- library_name: vllm base_model: google/gemma-4-26B-A4B-it base_model_relation: finetune license: apache-2.0 license_link: https://ai.google.dev/gemma/docs/gemma_4_license language: - en tags: - structured-extraction - json - html - schema-constrained - vllm - bf16 - merged-lora - moe pipeline_tag: text-generation --- # continuity-a4b-26b **Turn any web page into typed JSON against any schema you supply.** `continuity-a4b-26b` is a fine-tune of [`google/gemma-4-26B-A4B-it`](https://huggingface.co/google/gemma-4-26B-A4B-it) for schema-constrained extraction from web pages, merged and quantized to **NVFP4 (experts-only)**. It reads a cleaned DOM representation of a page plus a JSON Schema and emits one JSON instance matching that schema. **This repo is the bf16 build (49 GB).** For the 4-bit build see [`GestaltLabs/continuity-a4b-26b-nvfp4`](https://huggingface.co/GestaltLabs/continuity-a4b-26b-nvfp4) (17.5 GB, same weights, 21x the throughput). - **0.8932** strict success rate on IR / **0.9515** on HTML — best HTML success of any model we measured - **100% schema-valid output** (103/103), zero parse failures - **~3.8B active** of 25.2B language-model parameters (MoE); 26.3B including the vision tower --- ## Results Held-out set: **103 real pages**, one schema per page, greedy decoding, 2048-token cap, one scorer, identical prompt construction for every row below. `fna` = field-level accuracy after normalization; `success` = every required field present and correct. ### IR representation (cleaned-DOM IR — the representation we serve) | model | size | fna | success | schema-valid | pages/s | |---|---|---|---|---|---| | **continuity-a4b-26b (this repo, bf16)** | **49 GB** | **0.9210** | **0.8932** | **1.00** | **0.167** | | `continuity-a4b-26b-nvfp4` | 18 GB | 0.9178 | 0.9029 | 1.00 | 3.53 | | gemma-4-26B-A4B (zero-shot, no fine-tune) | 52 GB | 0.9137 | 0.9029 | 0.00 | 0.033 | | diffusiongemma-26B (zero-shot) | 52 GB | 0.9046 | 0.6796 | 0.00 | 0.213 | | NuExtract3 (4B dense) | 8 GB | 0.8905 | 0.6699 | 0.94 | 0.085 | | ReaderLM-v2 (1.5B) | 3 GB | 0.7573 | 0.3786 | 0.00 | 0.263 | ### HTML representation (raw markup) | model | fna | success | schema-valid | pages/s | |---|---|---|---|---| | **gemma-4-26B-A4B + LoRA** | 0.9203 | **0.9515** | 0.99 | 0.066 | | diffusiongemma-26B (zero-shot) | 0.9263 | 0.6408 | 0.00 | 0.200 | | gemma-4-26B-A4B (zero-shot) | 0.9684 | 0.6311 | 0.00 | 0.032 | | NuExtract3 | 0.8779 | 0.3010 | 1.00 | 0.085 | | ReaderLM-v2 | 0.8206 | 0.4175 | 0.00 | 0.192 | ### What the numbers mean - **Every accuracy column is apples-to-apples**: same 103 examples, same schema per page, same scorer, same 2048-token cap, same greedy decode. - **`pages/s` is not apples-to-apples.** Ours is vLLM on an RTX PRO 6000 Blackwell; the comparison rows were measured with HuggingFace `generate_batch` on other GPUs. Treat it as the deployment throughput of this artifact, not as a controlled engine-for-engine speed comparison. - **Quantizing to NVFP4 did not cost accuracy** (0.9178 vs 0.9210 field accuracy on the same weights before quantization) while shrinking the artifact **49 GB → 18 GB**. - **Schema validity is 1.00** because generation is post-checked against the schema; a row that cannot be parsed or repaired is reported separately rather than silently accepted. Every other zero-shot model here returns 0.00 on this column — their output is not schema-valid JSON. - Zero-shot `gemma-4-26B-A4B` ties on IR success (0.9029) but produces **no schema-valid output at all** (fenced/prose-wrapped JSON), and its HTML success is 0.6311 against 0.9515. The fine-tune is what makes the output usable. --- ## Usage Serve with vLLM. NVFP4 requires Blackwell (sm100/sm120) for the FP4 MoE kernels. ```bash pip install vllm ``` ```python import os # FlashInfer's JIT sampler does not build on sm120 (RTX PRO 6000). Disable it # unless you are on sm100, where it works. os.environ["VLLM_USE_FLASHINFER_SAMPLER"] = "0" from vllm import LLM, SamplingParams llm = LLM( model="GestaltLabs/continuity-a4b-26b", max_model_len=8192, gpu_memory_utilization=0.85, limit_mm_per_prompt={"image": 0}, # text-only use attention_backend="TRITON_ATTN", ) prompt = ( "Extract the fields defined by this JSON Schema from the page.\n" "Reply with a single JSON object and nothing else.\n\n" '\n{"type":"object","properties":{"name":{"type":"string"},' '"price":{"type":"number"}},"required":["name","price"]}\n\n\n' "\n...cleaned page text...\n\n" ) out = llm.generate([prompt], SamplingParams(max_tokens=2048, temperature=0.0)) print(out[0].outputs[0].text) ``` Notes: - Use `temperature=0.0`. The model was fine-tuned for greedy, single-shot extraction. - The page should be **cleaned** first. We feed a compact DOM-IR rendering of the page (script/style/nav stripped, text and structure kept), not raw HTML. - Validate the returned object against your schema before trusting it. The model is accurate on our set; it is not a formal guarantee. --- ## Training | | | |---|---| | Base | `google/gemma-4-26B-A4B-it` @ `4d7ae4984b7db7de8f8457170b3f1a419ee76d52` | | Method | LoRA, **attention + dense-MLP + router projections only** (fused expert weights frozen) | | LoRA | r=16, alpha=32, dropout=0.0, bf16 | | Trainable | 19,998,720 / 26,301,723,952 total checkpoint params (0.076%) | | Targets | `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj, proj` | | Steps | 514 (2 epochs), lr 1e-4 cosine, bsz 1 × grad-accum 4, seq 2048 | | Data | 1,027 fitted pages (572 longer pages dropped at the 2048 chat-token budget) | | Adapter sha256 | `98cee055d7f0ac6a4f0dc9f1aacc1112d34b76da508376b549b7fa32860e4fe9` | | Train data sha256 | `4f44b06d5c663b0485d0e819f3d30a1d30c2b77ca4363f8dfa12709405c40eee` | **Why attention-only LoRA.** The fused expert weights are 3-D `nn.Parameter`s (`experts.gate_up_proj` / `down_proj`, 128 experts), ~45.7 GB of the model. Adapting them via peft's `target_parameters` either hangs the first backward or OOMs on an 80 GB card. Freezing them and adapting the attention/MLP/router linears trains comfortably in 80 GB and is what this checkpoint is. Expert weights are unchanged from base. ### Quantization | | | |---|---| | Tool | NVIDIA ModelOpt 0.46.1 | | Recipe | `NVFP4_EXPERTS_ONLY` (4-bit, group_size 16, targets `Linear`) | | Quantized | the fused MoE expert projections | | Kept bf16 (93 modules) | all `self_attn`, `mlp`, `router`, `lm_head`, vision tower | | Output | 2 shards, 18 GB, `hf_quant_config.json` + `config.json.quantization_config` | Quantizing only the experts is deliberate: they are the bulk of the weights and the bulk of per-token decode traffic, while leaving attention in bf16 avoids the accuracy loss that aggressive whole-model 4-bit quantization causes. --- ## Intended use and limitations **Intended:** batch or interactive extraction of structured records from web pages where you supply the schema — product data, listings, articles, documents, directories. **Not intended:** as a general-purpose chat model, for safety-critical decisions, or on untrusted schemas without validating the returned object. **Known limitations** - Trained and evaluated on English pages; multilingual extraction is untested. - Pages longer than ~2,500 prompt tokens need truncation or chunking; the training data capped at 2048 chat tokens, so very long pages are out of distribution. - Multi-entity pages that require binding several entities into an array of objects are the hardest class for every model we tested, including this one. A separate 47-page multi-entity challenge set scores 0.627 field accuracy, and we publish that number rather than hiding it — if your workload is list-heavy, measure on your own data first. - Greedy decode only; sampling will degrade output validity. **This is a fine-tune of a Google Gemma model.** Gemma 4 terms and the Gemma 4 license apply. See https://ai.google.dev/gemma/docs/gemma_4_license. --- ## Provenance - Base model revision `4d7ae4984b7db7de8f8457170b3f1a419ee76d52`, weights unchanged except for the merged LoRA on attention/MLP/router linears. - NVFP4 quantization applied to the **merged** fine-tune (not the base), so the quantized artifact and the fine-tune are the same model. - Evaluation receipt (`merged-lora-nvfp4-experts_ir.json`) records the per-example scores for all 103 pages, the 2048 cap, the page representation, and `finish_reason` for every row (103/103 `stop`, no truncation). Every number in this card is measured, not estimated. Where a measurement is not apples-to-apples we say so.