- Google Gemma 3 27B-IT - TevunahAi Ultra-Hybrid GPTQ with EoRA
Google Gemma 3 27B-IT - TevunahAi Ultra-Hybrid GPTQ with EoRA
Model Details
| Property | Value |
|---|---|
| Base Model | google/gemma-3-27b-it |
| Architecture | Gemma 3 Dense Decoder + Vision Tower |
| Parameters | 27B |
| Context Length | 128,000 tokens |
| Languages | 140+ |
| Multimodal | โ Text + Image input |
| Quantization | TevunahAi Ultra-Hybrid GPTQ + EoRA |
| Original Size | ~54 GB (BF16) |
| Quantized Size | ~20-21 GB |
| Compression | ~60% reduction |
| Active VRAM | ~21.3 GB (with inference overhead) |
| License | Gemma Terms of Use |
Architecture Breakdown
Google Gemma 3 27B-IT is a state-of-the-art multimodal model trained on 14T tokens across 140+ languages:
Language Model (62 layers - QUANTIZED)
- 62 Transformer Decoder Layers: Dense attention architecture
- 32 Attention Heads: GQA with 16 KV heads (2:1 ratio)
- Hidden Size: 5,376
- Intermediate Size: 21,504
- Head Dimension: 128
- Vocab Size: 262,208
- GELU (pytorch_tanh) Activation
- RMSNorm + RoPE: theta=1M, factor=8 for extended context
- Hybrid Attention: Sliding window (1024) with pattern 6
Vision Tower (27 layers - PRESERVED IN BF16)
- 27 Vision Encoder Layers: Full precision for image understanding
- Patch Embedding: Preserved for accurate image tokenization
- All attention Q/K/V/O projections: BF16
- All MLP fc1/fc2 layers: BF16
Why preserve the vision tower? Quantizing vision encoders significantly degrades image understanding. By keeping the vision tower in full precision, this quantization maintains excellent multimodal capabilities while compressing the language model.
Why This Matters
- 128K context window: Process extremely long documents
- 140+ languages: Comprehensive multilingual support
- Multimodal: Text and image understanding (vision preserved!)
- Hybrid attention: Efficient local/global attention patterns
- 14T training tokens: Massive pretraining
Quantization Strategy
TevunahAi Ultra-Hybrid Mixed-Precision with EoRA-2048 Boundary Layers
This quantization uses EoRA (Error-optimized Low-Rank Adaptation) - NVIDIA's technique for recovering quantization error through learned low-rank adapters applied during the quantization process.
Language Model Quantization (62 layers)
| Component | Precision | EoRA Rank | Rationale |
|---|---|---|---|
| Layer 0 (all projections) | INT8 | 2048 | Maximum error correction at input |
| Layer 61 (all projections) | INT8 | 2048 | Maximum error correction at output |
| Attention Q/K/V/O (layers 1-60) | INT8 | 64 | Quality preservation for multilingual |
| MLP gate/up/down (layers 1-50) | INT4 | 64 | Maximum compression in middle layers |
| MLP gate/up/down (layers 51-60) | INT8 | 64 | Higher precision near output |
| Embeddings | BF16 | - | Preserved for 262K vocab accuracy |
| LM Head | BF16 | - | Preserved for output quality |
Vision Tower (PRESERVED - NOT QUANTIZED)
| Component | Precision | Rationale |
|---|---|---|
| vision_tower.vision_model.embeddings.patch_embedding | BF16 | Image tokenization accuracy |
| vision_tower.vision_model.encoder.layers.0-26 (all) | BF16 | Full image understanding |
Why EoRA-2048 Boundaries?
- Layer 0: First layer errors compound through all 61 subsequent layers
- Layer 61: Final layer directly determines next token prediction
- Rank 2048: Maximum error correction capacity (~8-10% recovery vs ~5% for rank 64)
- 434 layer-specific rules: Not a blanket quantization - each projection optimized individually
Calibration
- 1,700 samples (7x industry standard of 256)
- 1,024 sequence length (optimized for 32GB VRAM)
- Diverse datasets: UltraChat, SlimOrca, Code-Feedback, Orca-Math
- Premium calibration for multilingual instruction following
Performance Benchmarks
Qualitative Tests (9/9 passed)
| Test | Result | Speed | Details |
|---|---|---|---|
| Basic Instruction | โ PASS | 9.4 tok/s | Self-introduced as Gemma |
| Reasoning | โ PASS | 10.5 tok/s | Apple math with edge cases |
| Code (Fibonacci) | โ PASS | 10.2 tok/s | 5/5 elements + error handling |
| Japanese | โ PASS | 10.2 tok/s | Four seasons - native quality |
| Chinese | โ PASS | 10.5 tok/s | AI intro - comprehensive |
| Arabic | โ PASS | 10.7 tok/s | Reading benefits - RTL correct |
| Korean | โ PASS | 11.3 tok/s | Traditional food - detailed |
| Creative Writing | โ PASS | 12.0 tok/s | Robot painting story |
| Summarization | โ PASS | 10.9 tok/s | 6/6 key terms |
Inference Performance
| Metric | Value |
|---|---|
| VRAM Usage | 21.29 GB |
| Generation Speed | 9-12 tok/s |
| Load Time | ~113 seconds |
| Tests Passed | 9/9 |
Multilingual Quality Examples
Japanese (ๆฅๆฌ่ช) - Four seasons explanation:
ๆฅๆฌใฎๅๅญฃใฏใๆฅใๅคใ็งใๅฌใฎ4ใคใงใใใใใ็นๅพด็ใชๆฐๅใ้ขจๆฏใใใใพใใ
ๆฅ๏ผ3ๆ๏ฝ5ๆ๏ผ: ๆใใใชใใๆกใๅฒใ่ชใ็พใใๅญฃ็ฏใงใ...
Chinese (ไธญๆ) - AI introduction:
ไบบๅทฅๆบ่ฝ๏ผ็ฎ็งฐAI๏ผ็ฎๅๆฅ่ฏดๅฐฑๆฏ่ฎฉ่ฎก็ฎๆบๅไบบไธๆ ทๆ่ๅ่กๅจ็ๆๆฏใ
ๅฎไธๆฏไธไธชๅไธ็ๆๆฏ๏ผ่ๆฏไธๅ ๆๆฏ็้ๅ...
Arabic (ุงูุนุฑุจูุฉ) - Benefits of reading:
ุงููุฑุงุกุฉ ููุง ููุงุฆุฏ ุฌู
ุฉุ ูุง ุชุนุฏ ููุง ุชุญุตูุ ุชู
ุณ ุฌูุงูุจ ู
ุฎุชููุฉ ู
ู ุญูุงุชูุง...
Korean (ํ๊ตญ์ด) - Traditional food:
ํ๊ตญ ์์์ ์ค๋ ์ญ์ฌ์ ๋
ํนํ ๋ฌธํ๋ฅผ ๋ด๊ณ ์์ผ๋ฉฐ, ๋ฐํจ ์์๊ณผ ๋ค์ํ ์ฑ์๋ฅผ
๋ง์ด ์ฌ์ฉํ๋ ๊ฒ์ด ํน์ง์
๋๋ค...
Usage
GPTQModel (Recommended)
from gptqmodel import GPTQModel
from transformers import AutoTokenizer
model = GPTQModel.from_quantized(
"TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
device_map="auto",
trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
"TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
trust_remote_code=True
)
messages = [
{"role": "user", "content": "Explain quantum computing in simple terms."},
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
"TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ",
trust_remote_code=True
)
# Use same generation code as above
vLLM (Production)
pip install -U vllm
vllm serve TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ \
--max-model-len 8192 \
--trust-remote-code
Installation
pip install gptqmodel transformers>=4.48
Multilingual Capabilities
Gemma 3 27B supports 140+ languages with strong performance across language families:
Tested Languages (All Passed)
| Language | Script | Quality |
|---|---|---|
| English | Latin | โญโญโญโญโญ |
| Japanese | ๆฅๆฌ่ช | โญโญโญโญโญ |
| Chinese | ไธญๆ | โญโญโญโญโญ |
| Arabic | ุงูุนุฑุจูุฉ | โญโญโญโญโญ |
| Korean | ํ๊ตญ์ด | โญโญโญโญโญ |
Multilingual Examples
# Japanese
messages = [{"role": "user", "content": "ๆฅๆฌใฎๅๅญฃใซใคใใฆ่ชฌๆใใฆใใ ใใใ"}]
# Chinese
messages = [{"role": "user", "content": "็ฎๅไป็ปไธไธไบบๅทฅๆบ่ฝใ"}]
# Arabic
messages = [{"role": "user", "content": "ู
ุง ูู ููุงุฆุฏ ุงููุฑุงุกุฉุ"}]
# Korean
messages = [{"role": "user", "content": "ํ๊ตญ์ ์ ํต ์์์ ๋ํด ๊ฐ๋จํ ์๊ฐํด ์ฃผ์ธ์."}]
Multimodal Capabilities
The vision tower is preserved in full precision, maintaining excellent image understanding:
# Note: Multimodal inference requires additional setup
# See google/gemma-3-27b-it documentation for image input format
Known Issues
- Tokenizer regex warning: Can be safely ignored or fixed with
fix_mistral_regex=Truewhen loading tokenizer
Memory Requirements
Inference (quantized model)
| Context Length | VRAM Required |
|---|---|
| Short (4K) | 21-22 GB |
| Medium (16K) | 24-28 GB |
| Long (32K) | 32-36 GB |
| Extended (64K) | 40-48 GB |
| Full (128K) | 64+ GB |
Tested on: RTX 5000 Ada (32GB) - 21.29 GB active VRAM during inference
Quantization (reproduction)
- GPU: RTX 5000 Ada 32GB (with CPU offload)
- RAM: 80GB+ recommended
- Method: GPU Hessian + CPU offload
Quantization Details
| Specification | Value |
|---|---|
| Method | GPTQ + Ultra-Hybrid + EoRA |
| Quantizer | GPTQModel |
| EoRA Boundary Rank | 2048 (layers 0 & 61) |
| EoRA Standard Rank | 64 (layers 1-60) |
| Calibration Samples | 1,700 (7x industry standard) |
| Sequence Length | 1,024 tokens |
| Group Size | 128 |
| desc_act | False |
| sym | True (symmetric quantization) |
| Bits (default) | 4 |
| Language Model Rules | 434 custom precision rules |
| Vision Tower | Preserved in BF16 (27 layers) |
Use Cases
Ideal for:
- ๐ Multilingual applications (140+ languages)
- ๐ผ๏ธ Multimodal tasks (vision tower preserved)
- ๐ Long document processing (128K context)
- ๐ฌ General chat and assistance
- ๐ป Code generation and explanation
- ๐ Content creation across languages
- ๐ง 24GB GPU deployment (fits RTX 3090/4090)
Technical Specifications
| Specification | Value |
|---|---|
| Model Family | Google Gemma 3 |
| Variant | 27B-IT (Instruction Tuned) |
| Total Parameters | 27B |
| Language Model Layers | 62 |
| Vision Encoder Layers | 27 |
| Hidden Size | 5,376 |
| Intermediate Size | 21,504 |
| Attention Heads | 32 |
| KV Heads | 16 (GQA) |
| Head Dimension | 128 |
| Activation | GELU (pytorch_tanh) |
| Normalization | RMSNorm |
| Position Encoding | RoPE (theta=1M, factor=8) |
| Sliding Window | 1024 (pattern: 6) |
| Context Length | 128,000 |
| Vocab Size | 262,208 |
| Training Tokens | 14T |
| Languages | 140+ |
| Multimodal | Text + Image |
Acknowledgments
- Google DeepMind for developing the Gemma 3 model family
- NVIDIA for the EoRA (Error-optimized Low-Rank Adaptation) technique used in this quantization
- GPTQModel team for the excellent quantization framework
License
Gemma Terms of Use - See Google's license terms for usage restrictions and requirements.
Citation
@software{gemma3_27b_gptq_2025,
title = {Google Gemma 3 27B-IT - TevunahAi Ultra-Hybrid GPTQ with EoRA},
author = {TevunahAi},
year = {2025},
note = {Ultra-Hybrid GPTQ with EoRA-2048 boundary layers, vision tower preserved},
url = {https://huggingface.co/TevunahAi/Gemma-3-27B-IT-TevunahAi-GPTQ}
}
@misc{gemma3_2025,
title = {Gemma 3},
author = {Google DeepMind},
year = {2025},
url = {https://huggingface.co/google/gemma-3-27b-it}
}
- Downloads last month
- 32