Google Gemma 3 27B-IT - TevunahAi Ultra-Hybrid GPTQ with EoRA

Model Details

Property Value
Base Model google/gemma-3-27b-it
Architecture Gemma 3 Dense Decoder + Vision Tower
Parameters 27B
Context Length 128,000 tokens
Languages 140+
Multimodal โœ… Text + Image input
Quantization TevunahAi Ultra-Hybrid GPTQ + EoRA
Original Size ~54 GB (BF16)
Quantized Size ~20-21 GB
Compression ~60% reduction
Active VRAM ~21.3 GB (with inference overhead)
License Gemma Terms of Use

Architecture Breakdown

Google Gemma 3 27B-IT is a state-of-the-art multimodal model trained on 14T tokens across 140+ languages:

Language Model (62 layers - QUANTIZED)

  • 62 Transformer Decoder Layers: Dense attention architecture
  • 32 Attention Heads: GQA with 16 KV heads (2:1 ratio)
  • Hidden Size: 5,376
  • Intermediate Size: 21,504
  • Head Dimension: 128
  • Vocab Size: 262,208
  • GELU (pytorch_tanh) Activation
  • RMSNorm + RoPE: theta=1M, factor=8 for extended context
  • Hybrid Attention: Sliding window (1024) with pattern 6

Vision Tower (27 layers - PRESERVED IN BF16)

  • 27 Vision Encoder Layers: Full precision for image understanding
  • Patch Embedding: Preserved for accurate image tokenization
  • All attention Q/K/V/O projections: BF16
  • All MLP fc1/fc2 layers: BF16

Why preserve the vision tower? Quantizing vision encoders significantly degrades image understanding. By keeping the vision tower in full precision, this quantization maintains excellent multimodal capabilities while compressing the language model.

Why This Matters

  • 128K context window: Process extremely long documents
  • 140+ languages: Comprehensive multilingual support
  • Multimodal: Text and image understanding (vision preserved!)
  • Hybrid attention: Efficient local/global attention patterns
  • 14T training tokens: Massive pretraining

Quantization Strategy

TevunahAi Ultra-Hybrid Mixed-Precision with EoRA-2048 Boundary Layers

This quantization uses EoRA (Error-optimized Low-Rank Adaptation) - NVIDIA's technique for recovering quantization error through learned low-rank adapters applied during the quantization process.

Language Model Quantization (62 layers)

Component Precision EoRA Rank Rationale
Layer 0 (all projections) INT8 2048 Maximum error correction at input
Layer 61 (all projections) INT8 2048 Maximum error correction at output
Attention Q/K/V/O (layers 1-60) INT8 64 Quality preservation for multilingual
MLP gate/up/down (layers 1-50) INT4 64 Maximum compression in middle layers
MLP gate/up/down (layers 51-60) INT8 64 Higher precision near output
Embeddings BF16 - Preserved for 262K vocab accuracy
LM Head BF16 - Preserved for output quality

Vision Tower (PRESERVED - NOT QUANTIZED)

Component Precision Rationale
vision_tower.vision_model.embeddings.patch_embedding BF16 Image tokenization accuracy
vision_tower.vision_model.encoder.layers.0-26 (all) BF16 Full image understanding

Why EoRA-2048 Boundaries?

  • Layer 0: First layer errors compound through all 61 subsequent layers
  • Layer 61: Final layer directly determines next token prediction
  • Rank 2048: Maximum error correction capacity (~8-10% recovery vs ~5% for rank 64)
  • 434 layer-specific rules: Not a blanket quantization - each projection optimized individually

Calibration

  • 1,700 samples (7x industry standard of 256)
  • 1,024 sequence length (optimized for 32GB VRAM)
  • Diverse datasets: UltraChat, SlimOrca, Code-Feedback, Orca-Math
  • Premium calibration for multilingual instruction following

Performance Benchmarks

Qualitative Tests (9/9 passed)

Test Result Speed Details
Basic Instruction โœ… PASS 9.4 tok/s Self-introduced as Gemma
Reasoning โœ… PASS 10.5 tok/s Apple math with edge cases
Code (Fibonacci) โœ… PASS 10.2 tok/s 5/5 elements + error handling
Japanese โœ… PASS 10.2 tok/s Four seasons - native quality
Chinese โœ… PASS 10.5 tok/s AI intro - comprehensive
Arabic โœ… PASS 10.7 tok/s Reading benefits - RTL correct
Korean โœ… PASS 11.3 tok/s Traditional food - detailed
Creative Writing โœ… PASS 12.0 tok/s Robot painting story
Summarization โœ… PASS 10.9 tok/s 6/6 key terms

Inference Performance

Metric Value
VRAM Usage 21.29 GB
Generation Speed 9-12 tok/s
Load Time ~113 seconds
Tests Passed 9/9

Multilingual Quality Examples

Japanese (ๆ—ฅๆœฌ่ชž) - Four seasons explanation:

ๆ—ฅๆœฌใฎๅ››ๅญฃใฏใ€ๆ˜ฅใ€ๅคใ€็ง‹ใ€ๅ†ฌใฎ4ใคใงใ€ใใ‚Œใžใ‚Œ็‰นๅพด็š„ใชๆฐ—ๅ€™ใ‚„้ขจๆ™ฏใŒใ‚ใ‚Šใพใ™ใ€‚
ๆ˜ฅ๏ผˆ3ๆœˆ๏ฝž5ๆœˆ๏ผ‰: ๆš–ใ‹ใใชใ‚Šใ€ๆกœใŒๅ’ฒใ่ช‡ใ‚‹็พŽใ—ใ„ๅญฃ็ฏ€ใงใ™...

Chinese (ไธญๆ–‡) - AI introduction:

ไบบๅทฅๆ™บ่ƒฝ๏ผŒ็ฎ€็งฐAI๏ผŒ็ฎ€ๅ•ๆฅ่ฏดๅฐฑๆ˜ฏ่ฎฉ่ฎก็ฎ—ๆœบๅƒไบบไธ€ๆ ทๆ€่€ƒๅ’Œ่กŒๅŠจ็š„ๆŠ€ๆœฏใ€‚
ๅฎƒไธๆ˜ฏไธ€ไธชๅ•ไธ€็š„ๆŠ€ๆœฏ๏ผŒ่€Œๆ˜ฏไธ€ๅ †ๆŠ€ๆœฏ็š„้›†ๅˆ...

Arabic (ุงู„ุนุฑุจูŠุฉ) - Benefits of reading:

ุงู„ู‚ุฑุงุกุฉ ู„ู‡ุง ููˆุงุฆุฏ ุฌู…ุฉุŒ ู„ุง ุชุนุฏ ูˆู„ุง ุชุญุตู‰ุŒ ุชู…ุณ ุฌูˆุงู†ุจ ู…ุฎุชู„ูุฉ ู…ู† ุญูŠุงุชู†ุง...

Korean (ํ•œ๊ตญ์–ด) - Traditional food:

ํ•œ๊ตญ ์Œ์‹์€ ์˜ค๋žœ ์—ญ์‚ฌ์™€ ๋…ํŠนํ•œ ๋ฌธํ™”๋ฅผ ๋‹ด๊ณ  ์žˆ์œผ๋ฉฐ, ๋ฐœํšจ ์Œ์‹๊ณผ ๋‹ค์–‘ํ•œ ์ฑ„์†Œ๋ฅผ 
๋งŽ์ด ์‚ฌ์šฉํ•˜๋Š” ๊ฒƒ์ด ํŠน์ง•์ž…๋‹ˆ๋‹ค...

Usage

GPTQModel (Recommended)

from gptqmodel import GPTQModel
from transformers import AutoTokenizer

model = GPTQModel.from_quantized(
    "TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
    device_map="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
    trust_remote_code=True
)

messages = [
    {"role": "user", "content": "Explain quantum computing in simple terms."},
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=512,
    temperature=0.7,
    top_p=0.9,
    do_sample=True,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
    device_map="auto",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
    "TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
    trust_remote_code=True
)

# Use same generation code as above

vLLM (Production)

pip install -U vllm

vllm serve TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ \
    --max-model-len 8192 \
    --trust-remote-code

Installation

pip install gptqmodel transformers>=4.48

Multilingual Capabilities

Gemma 3 27B supports 140+ languages with strong performance across language families:

Tested Languages (All Passed)

Language Script Quality
English Latin โญโญโญโญโญ
Japanese ๆ—ฅๆœฌ่ชž โญโญโญโญโญ
Chinese ไธญๆ–‡ โญโญโญโญโญ
Arabic ุงู„ุนุฑุจูŠุฉ โญโญโญโญโญ
Korean ํ•œ๊ตญ์–ด โญโญโญโญโญ

Multilingual Examples

# Japanese
messages = [{"role": "user", "content": "ๆ—ฅๆœฌใฎๅ››ๅญฃใซใคใ„ใฆ่ชฌๆ˜Žใ—ใฆใใ ใ•ใ„ใ€‚"}]

# Chinese
messages = [{"role": "user", "content": "็ฎ€ๅ•ไป‹็ปไธ€ไธ‹ไบบๅทฅๆ™บ่ƒฝใ€‚"}]

# Arabic
messages = [{"role": "user", "content": "ู…ุง ู‡ูŠ ููˆุงุฆุฏ ุงู„ู‚ุฑุงุกุฉุŸ"}]

# Korean
messages = [{"role": "user", "content": "ํ•œ๊ตญ์˜ ์ „ํ†ต ์Œ์‹์— ๋Œ€ํ•ด ๊ฐ„๋‹จํžˆ ์†Œ๊ฐœํ•ด ์ฃผ์„ธ์š”."}]

Multimodal Capabilities

The vision tower is preserved in full precision, maintaining excellent image understanding:

# Note: Multimodal inference requires additional setup
# See google/gemma-3-27b-it documentation for image input format

Known Issues

  • Tokenizer regex warning: Can be safely ignored or fixed with fix_mistral_regex=True when loading tokenizer

Memory Requirements

Inference (quantized model)

Context Length VRAM Required
Short (4K) 21-22 GB
Medium (16K) 24-28 GB
Long (32K) 32-36 GB
Extended (64K) 40-48 GB
Full (128K) 64+ GB

Tested on: RTX 5000 Ada (32GB) - 21.29 GB active VRAM during inference

Quantization (reproduction)

  • GPU: RTX 5000 Ada 32GB (with CPU offload)
  • RAM: 80GB+ recommended
  • Method: GPU Hessian + CPU offload

Quantization Details

Specification Value
Method GPTQ + Ultra-Hybrid + EoRA
Quantizer GPTQModel
EoRA Boundary Rank 2048 (layers 0 & 61)
EoRA Standard Rank 64 (layers 1-60)
Calibration Samples 1,700 (7x industry standard)
Sequence Length 1,024 tokens
Group Size 128
desc_act False
sym True (symmetric quantization)
Bits (default) 4
Language Model Rules 434 custom precision rules
Vision Tower Preserved in BF16 (27 layers)

Use Cases

Ideal for:

  • ๐ŸŒ Multilingual applications (140+ languages)
  • ๐Ÿ–ผ๏ธ Multimodal tasks (vision tower preserved)
  • ๐Ÿ“„ Long document processing (128K context)
  • ๐Ÿ’ฌ General chat and assistance
  • ๐Ÿ’ป Code generation and explanation
  • ๐Ÿ“ Content creation across languages
  • ๐Ÿ”ง 24GB GPU deployment (fits RTX 3090/4090)

Technical Specifications

Specification Value
Model Family Google Gemma 3
Variant 27B-IT (Instruction Tuned)
Total Parameters 27B
Language Model Layers 62
Vision Encoder Layers 27
Hidden Size 5,376
Intermediate Size 21,504
Attention Heads 32
KV Heads 16 (GQA)
Head Dimension 128
Activation GELU (pytorch_tanh)
Normalization RMSNorm
Position Encoding RoPE (theta=1M, factor=8)
Sliding Window 1024 (pattern: 6)
Context Length 128,000
Vocab Size 262,208
Training Tokens 14T
Languages 140+
Multimodal Text + Image

Acknowledgments

  • Google DeepMind for developing the Gemma 3 model family
  • NVIDIA for the EoRA (Error-optimized Low-Rank Adaptation) technique used in this quantization
  • GPTQModel team for the excellent quantization framework

License

Gemma Terms of Use - See Google's license terms for usage restrictions and requirements.

Citation

@software{gemma3_27b_gptq_2025,
  title = {Google Gemma 3 27B-IT - TevunahAi Ultra-Hybrid GPTQ with EoRA},
  author = {TevunahAi},
  year = {2025},
  note = {Ultra-Hybrid GPTQ with EoRA-2048 boundary layers, vision tower preserved},
  url = {https://huggingface.co/TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ}
}

@misc{gemma3_2025,
  title = {Gemma 3},
  author = {Google DeepMind},
  year = {2025},
  url = {https://huggingface.co/google/gemma-3-27b-it}
}

https://huggingface.co/TevunahAi

Downloads last month
32
Safetensors
Model size
36B params
Tensor type
BF16
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ

Quantized
(145)
this model

Collection including TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ