Instructions to use GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "question-answering" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("question-answering", model="GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit") model = AutoModelForCausalLM.from_pretrained("GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Gemma-3-270m-HuHotPotQA
This model is a fine-tuned version of unsloth/gemma-3-270m-it on the HuHotpotQA_8k dataset. It is designed for multi-step, context-based question answering in Hungarian.
Model Details
- Model Type: Causal Language Model (Decoder-only)
- Base Model: unsloth/gemma-3-270m-it
- Language: Hungarian (hu)
- Task: Context-based Question Answering (Extractive / Abstractive)
- Frameworks: Transformers, Unsloth, TRL
- License: Gemma Terms of Use
Dataset
The model was trained on HuHotpotQA, a Hungarian-language, multi-step question-and-answer dataset derived from Wikipedia articles. It follows the HotpotQA format, requiring the model to synthesize information from multiple provided context paragraphs to answer complex questions.
Prompt Format
The model uses the Gemma-3 chat template with a specific instruction for context-grounded generation:
<start_of_turn>system
Kizárólag a megadott kontextus alapján válaszolj a kérdésre!<end_of_turn>
<start_of_turn>user
Kérdés: {question}
Kontextus: {context}<end_of_turn>
<start_of_turn>model
{answer}<end_of_turn>
Evaluation
On a held-out test set:
| Metric | Value |
|---|---|
| n_samples | 113.0000 |
| exact_match | 61.9469 |
| token_f1 | 66.5318 |
| rouge_l | 66.0725 |
Training Details
The model was fine-tuned using Low-Rank Adaptation (LoRA) via the Unsloth library for efficient training and inference.
Training Hyperparameters
Training Regime: Mixed Precision (bf16), LoRA (r=64, alpha=64)
Target Modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Optimizer: adamw_8bit
Learning Rate: 2e-5
Learning Rate Scheduler: Cosine with 20 warmup steps
Epochs: 2
Batch Size: 1 (with gradient accumulation steps = 16 -> effective batch size 16)
Maximum Sequence Length: 9216 tokens
Weight Decay: 0.001
Loss Masking: Trained on assistant responses only (ignoring prompt/user loss)
Gradient Checkpointing: Unsloth optimized
How to Use
Inference
from unsloth import FastLanguageModel
from transformers import TextStreamer
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="GaborMadarasz/gemma_3_270_HuHotPotQA_4bit",
max_seq_length=9216,
load_in_4bit=False,
)
messages = [
{"role": "system", "content": "Kizárólag a megadott kontextus alapján válaszolj a kérdésre!"},
{"role": "user", "content": "Kérdés: Kik a főszereplői az 1956-os forradalomnak?\\nKontextus: Az 1956-os forradalom és szabadságharc..."},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True).removeprefix('<bos>')
inputs = tokenizer(text, return_tensors="pt").to("cuda")
_ = model.generate(
**inputs,
max_new_tokens=256,
temperature=1.0,
top_p=0.95,
top_k=64,
do_sample=True,
streamer=TextStreamer(tokenizer, skip_prompt=True),
)
Hardware
Trained in Google Colab, using Unsloth optimizations to maximize memory efficiency and speed.
Limitations and Biases
As a 270M parameter model fine-tuned on a specific QA dataset:
It may hallucinate facts not present in the provided context.
It might struggle with highly complex reasoning that requires more than 2-3 steps of deduction, despite being trained on multi-step QA.
Performance is highly dependent on the quality and relevance of the provided context.
Being based on Gemma-3, it inherits the base model's biases and safety guardrails.
Citation
@software{gemma_3_huhotpotqa,
title = {Gemma-3-270m Finetuned on HuHotpotQA},
author = {Gabor Madarasz},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/GaborMadarasz/gemma_3_270m_HuHotPotQA_16bit}
}
This gemma3_text model was trained 2x faster with Unsloth and Huggingface's TRL library.
- Downloads last month
- 51
