Gemma 3 270M Jev (MLX)

Gemma trained with Tunix to score supplied actions for Maze and ViZDoom. This is the complete FP32 checkpoint with its learned scoring head.

Run on an Apple Silicon Mac

Use the supplied MLXGameEngine. This export has a custom scalar scoring head; mlx_lm.generate is not the inference interface. The Transformers export contains the same trained parameters for CPU/CUDA use.

Create a virtual environment inside your project and install mlx==0.32.2, mlx-lm==0.29.1, transformers==4.57.6 and numpy there. Set HF_HOME to a directory inside the project before downloading.

import json
import sys
from pathlib import Path
from huggingface_hub import snapshot_download

folder = snapshot_download(
    "bzantium/gemma-3-270m-jev-mlx",
    local_dir="artifacts/gemma-3-270m-jev-mlx",
)
sys.path.insert(0, folder)
from gemmajev.mlx_backend import MLXGameEngine

engine = MLXGameEngine(folder, precision="float32")
request = json.loads((Path(folder) / "examples/navigation_request.json").read_text())
print(json.dumps(engine.predict(request), indent=2))

Use examples/doom_request.json for Doom. Saved responses are references only; the engine runs the weights and never reads those response files.

Conversion verification

On 40 frozen Tunix reference questions (32 Maze movement, 8 Doom), MLX FP32 changed no top-ranked answers. The maximum probability difference was 0.00000681. On an Apple M2 Max, one movement question took about 35.1 ms averaged over ten warm calls on one observation, including tokenization and response construction. This is a small runtime check, not a full-game or broad latency benchmark. Details are in validation.json.

Model and training

This is a game-specific candidate scorer based on google/gemma-3-270m-it (revision ac82b4e820549b854eebf28ce6dedaf9fdfa17b3). It was trained with Tunix 0.1.7. Both the Gemma backbone and a shared scalar scoring head were updated. The export contains all 268,098,816 parameters in FP32.

For each candidate, the model reads the state, question and candidate description through the official Gemma chat template. It scores the last valid token. Softmax over the candidate scores produces the response probabilities. Python builds the JSON structure; the model does not generate JSON tokens.

The released checkpoint is navigation-rehearsal, step 800. Its ancestry is:

Run Updates Learning rate Questions per batch
baseline 800 1e-5 4 Maze local-safety + 4 Doom
maze-warmup 240 1e-4 8 Maze, from the first 64 training questions
maze-expanded 1,600 1e-5 4 expanded Maze local-safety + 4 Doom
navigation-v1 1,200 3e-5 6 Maze movement + 2 Doom
navigation-rehearsal 800 1e-5 4 Maze movement + 4 Doom

All runs use seed 17, batch size 8, a 512-token capacity per candidate and one GPU. The movement pool contains 1,322 Maze examples and 2,174 original NanoJev Doom examples. The three recorded demo maps are included in training: 441 movement examples come from them; 881 come from 20 generated maps. An offline teacher uses the full map to label directions; model inputs contain only a 5ร—5 view, coordinates, goal offset and visit/attempt history. No RLCD or reinforcement learning was used.

Evaluation

Frozen validation task Questions Accuracy
Maze next direction 196, from four separate maps 70.4%
ViZDoom expert action 201 90.0%

The three fitted demo maps finish in 226, 131 and 200 moves. Code masks walls and explored branches and handles corridors and backtracking; Gemma is called only at junctions. These completions are not evidence of unseen-map generalization. On four separate development maps, both parent and final checkpoints finish all four; two routes improve and two worsen. Those maps were inspected during development and are not a fresh blind benchmark. The fixed Doom case succeeds with 13 decisions, 49 game ticks and one shot.

Intended use and limitations

Use for studying bounded game decisions and integrating a candidate scorer into code. This checkpoint has not been evaluated as a general chatbot, planner, calibrated confidence model, or a model that switches between System 1 and System 2. It does not reproduce Jev's proprietary architecture or training method. It is a text-only model; Doom observations contain structured text, not pixels. Do not interpret a high candidate probability as a validated probability of success. Candidate inputs above 512 tokens are rejected by the supplied adapter. Recordings are accelerated replays and do not measure inference latency.

Source and recipes: bzantium/gemmajev. Inspired by Jev and NanoJev.

License

Weights are a modified Gemma derivative subject to the Gemma Terms of Use, including the use restrictions in Section 3.2 and the incorporated Prohibited Use Policy. A copy is included in GEMMA_TERMS.txt. These files are modified: the backbone has been fine-tuned and a learned candidate head has been added; tensor layouts have been converted for inference. They are not original Google checkpoint files. The included inference code is Apache-2.0 (CODE_LICENSE). Data and game assets retain their original terms; game assets are not bundled with these weights.

Downloads last month
61
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bzantium/gemma-3-270m-jev-mlx

Finetuned
(1158)
this model