Danny-Dasilva's picture
Upload folder using huggingface_hub
f252dd7 verified
|
Raw
History Blame Contribute Delete
4.75 kB
metadata
base_model: google/gemma-4-12B-it
license: gemma
library_name: transformers
pipeline_tag: text-generation
language:
  - en
tags:
  - gemma
  - gemma4
  - antidoom
  - ftpo
  - anti-repetition
  - int4
  - w4a16
  - gptq
  - llmcompressor
  - compressed-tensors
  - quantization
  - speculative-decoding
  - dspark
  - vllm
  - blackwell

gemma-4-12B-it — Antidoom + W4A16 GPTQ (group=32) · DSpark-verified

A Antidoom (FTPO anti-repetition) version of google/gemma-4-12B-it, quantized post-training to GPTQ int4 (W4A16, group=32, symmetric) and verified as a DSpark speculative-decoding target (draft head: deepseek-ai/dspark_gemma4_12b_block7).

TL;DR: 207 tok/s single-stream on one RTX 5090 (32 GB) — the fastest gemma-4-12B configuration we have measured on that card (beats our non-antidoom GPTQ at 198), at 7.8 GB of weights.

What was done

  1. Antidoom FTPO pass: 152 preference pairs mined at temperature 0.01 from the LiquidAI/antidoom-mix-v1.0 prompt mix; QLoRA (r=128, 4-bit training pass, max_seq 2048, lr 3e-5), early-stopped at chosen_win 0.575 (epoch ~1.9) — the strongest anti-loop convergence of our 4-model antidoom family; adapter merged into the bf16 base on CPU.
  2. Post-training GPTQ: identical recipe to gemma-4-12B-it-W4A16-GPTQ-g32-DSpark.

DSpark acceptance survived and improved: 35.6% (non-antidoom GPTQ) → 38.3% (this model). The FTPO patch suppresses exactly the degenerate repetition the draft also fails on, so speculation gets slightly better.

Measured speed — RTX 5090 (32 GB, Blackwell), single stream, greedy, 256-tok gens

config tok/s DSpark accept
this model + DSpark k=7 + CUDA graphs + autotune 207.5 (peaks 260) 38.3%
this model, native (no speculation) 119
reference: non-antidoom GPTQ + DSpark k=7 + autotune 198.2 35.6%
reference: official QAT w4a16-ct + DSpark slower than native 2.1%

Usage (vLLM + DSpark)

from vllm import LLM, SamplingParams
from transformers import AutoTokenizer

MODEL = "Danny-Dasilva/gemma-4-12B-it-antidoom-W4A16-GPTQ-g32-DSpark"
DRAFT = "deepseek-ai/dspark_gemma4_12b_block7"

llm = LLM(
    model=MODEL,
    max_model_len=8192,
    attention_backend="FLASHINFER",       # gemma-4 full-attn layers have head_dim 512
    enable_flashinfer_autotune=True,      # +2% in our runs
    speculative_config={
        "method": "dspark",
        "model": DRAFT,
        "num_speculative_tokens": 7,      # = draft block size; k<7 or k=14 are both slower
        "attention_backend": "TRITON_ATTN",
    },
)

tok = AutoTokenizer.from_pretrained(MODEL)
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain transformers step by step."}],
    add_generation_prompt=True, tokenize=False,
)
print(llm.generate([prompt], SamplingParams(temperature=0.0, max_tokens=256))[0].outputs[0].text)

Notes:

  • Always apply the chat template — gemma-4-it uses <|turn>-style tokens and add_bos_token=False; raw prompts produce garbage on any backend.
  • num_speculative_tokens must be ≤ 7 or a multiple of 7 (draft block size). 7 is optimal.
  • Works without DSpark too, as a normal compressed-tensors W4A16 checkpoint (kernel path identical to the official QAT release: group=32, symmetric, pack-quantized).
  • kv_cache_dtype="fp8" measured slower (−3%) at short context on this card.

Quantization recipe

llm-compressor 0.12 GPTQ, one-shot:

  • 256 calibration samples from HuggingFaceH4/ultrachat_200k (train_sft), chat template applied, max_seq_length 2048
  • scheme: int4, group_size=32, symmetric, weights-only (W4A16), pack-quantized format — mirrors the official QAT checkpoint's config so every vLLM kernel path is identical
  • sequential_targets=["Gemma4UnifiedTextDecoderLayer"]; lm_head and vision/audio embedder projections kept in bf16 (same 17-entry ignore list as the official QAT release, spelled so both HF and vLLM module names match)
  • processor_config.json included (vLLM's multimodal processor init requires it)

Provenance

  • Base: google/gemma-4-12B-it @ 5926caa4 → antidoom FTPO LoRA merge → GPTQ W4A16 g32
  • Built 2026-07-11 on a single RTX 5090; outputs verified coherent
  • Gemma is provided under and subject to the Gemma Terms of Use: https://ai.google.dev/gemma/terms