Instructions to use jokernifty/gemma-4-12B-it-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use jokernifty/gemma-4-12B-it-mlx-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("jokernifty/gemma-4-12B-it-mlx-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use jokernifty/gemma-4-12B-it-mlx-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "jokernifty/gemma-4-12B-it-mlx-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jokernifty/gemma-4-12B-it-mlx-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use jokernifty/gemma-4-12B-it-mlx-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "jokernifty/gemma-4-12B-it-mlx-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jokernifty/gemma-4-12B-it-mlx-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use jokernifty/gemma-4-12B-it-mlx-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "jokernifty/gemma-4-12B-it-mlx-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "jokernifty/gemma-4-12B-it-mlx-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jokernifty/gemma-4-12B-it-mlx-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use jokernifty/gemma-4-12B-it-mlx-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "jokernifty/gemma-4-12B-it-mlx-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jokernifty/gemma-4-12B-it-mlx-4bit
Run Hermes
hermes
gemma-4-12B-it (MLX, 4-bit quantized, text-only)
4-bit MLX conversion of google/gemma-4-12B-it.
Note: Gemma 4 12B is a unified encoder-free multimodal model (text + image + audio). This conversion is text-only — the vision and audio embedder weights are stripped because
mlx-lmdoesn't yet handle thegemma4_unifiedencoder-free multimodal path. The text backbone is intact.
Conversion details
- 4-bit, group size 64
- Token embedding kept in bf16 (~2 GB) — quantizing 262K × 3840 in one op exceeded the macOS Metal command-buffer watchdog
- Total size: ~8.1 GB across 5 shards
- Per-layer quantization to dodge GPU timeouts
Quick start
This model uses the gemma4_unified model_type, which isn't registered in mlx-lm yet. Until upstream adds it, drop this 4-line alias into your mlx_lm/models/ directory:
# mlx_lm/models/gemma4_unified.py
from . import gemma4
ModelArgs = gemma4.ModelArgs
_SKIP = ("vision_embedder.", "audio_embedder.", "audio_input.", "vision_input.",
"embed_vision_tokens.", "embed_audio_tokens.", "embed_video_tokens.")
class Model(gemma4.Model):
def sanitize(self, weights):
f = {k:v for k,v in weights.items()
if not any((k[len("model."):] if k.startswith("model.") else k).startswith(p) for p in _SKIP)}
return super().sanitize(f)
Then:
python -m mlx_lm chat --model jokernifty/gemma-4-12B-it-mlx-4bit
Recommended sampling (from Google's model card)
temperature = 1.0,top_p = 0.95,top_k = 64- Enable thinking mode by prepending
<|think|>to the system prompt.
License
Apache 2.0 (inherited). Use is also subject to the Gemma Terms and Prohibited Use Policy. All credit for the model goes to Google DeepMind.
- Downloads last month
- 8
4-bit