Does it have a Vision module?
No, it's Qwen 3 based, not the newer ones.
If you want uncensored (and actually trained) vision, you can try SicariusSicariiStuff/X-Ray_Alpha
Oh, I see. Since you were actually training it, can you give some insight about what parts of the model are actually censored or blind? I mean, encoder, projector, or the LLM itself?
You can read more info about it here:
https://arxiv.org/abs/2303.15343
https://arxiv.org/abs/2503.19786
And here's some hands on notebook for actual tuning:
https://colab.research.google.com/github/google-gemini/gemma-cookbook/blob/main/Demos/Emoji-Gemma-on-Web/resources/Fine_tune_Gemma_3_270M_for_emoji_generation.ipynb
You can read more info about it here:
https://arxiv.org/abs/2303.15343
https://arxiv.org/abs/2503.19786
what do these papers have to do with censorship problem?