Instructions to use jmkang212/qwen3vl-32b-vqa-lora-full-grpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jmkang212/qwen3vl-32b-vqa-lora-full-grpo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("visual-question-answering", model="jmkang212/qwen3vl-32b-vqa-lora-full-grpo")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("jmkang212/qwen3vl-32b-vqa-lora-full-grpo", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Recycling Classification VQA Model (QLoRA + GRPO Fine-tuned)
์ฌํ์ฉํ ์ด๋ฏธ์ง๋ฅผ ๋ถ๋ฅํ๋ ์๊ฐ์ง์์๋ต(VQA) ๋ชจ๋ธ์ ๋๋ค. ์ฌํ์ฉ ๊ฐ๋ฅ ์ฌ๋ถ ํ๋ณ, ๊ฐ์ ์นด์ดํ , ์ธ๋ถ ์ข ๋ฅ ๋ถ๋ฅ ๊ณผ์ ๋ฅผ ์ํํ๋๋ก ํ์ธํ๋๋์์ต๋๋ค.
A Vision-Language model fine-tuned for recycling-waste classification. Given an image, the model answers questions about (1) whether the item is recyclable, (2) how many items are present, and (3) the specific recycling category.
Model Details
Model Description
- Developed by: Kang Jaemin (๊ฐ์ฌ๋ฏผ)
- Model type: Vision-Language Model (Visual Question Answering)
- Language(s): Korean
- Finetuned from model:
- Fine-tuning method: QLoRA (4-bit quantization + LoRA) + GRPO
Task
๋ณธ ๋ชจ๋ธ์ ์ฌํ์ฉ ๋ถ๋ฅ ์ฑ๋ฆฐ์ง๋ฅผ ์ํด ๊ฐ๋ฐ๋์์ผ๋ฉฐ, ๋ค์ ๊ณผ์ ๋ฅผ ์ํํฉ๋๋ค:
- ์ฌํ์ฉ ๊ฐ๋ฅ ์ฌ๋ถ ํ๋ณ โ ์ด๋ฏธ์ง ์ ๋ฌผ์ฒด๊ฐ ์ฌํ์ฉ ๋์์ธ์ง ๋ถ๋ฅ
- ๊ฐ์ ์นด์ดํ โ ์ด๋ฏธ์ง ๋ด ์ฌํ์ฉํ ๊ฐ์ ์ธ์
- ์ธ๋ถ ์ข ๋ฅ ๋ถ๋ฅ โ ์ฌํ์ฉํ์ ๊ตฌ์ฒด์ ์นดํ ๊ณ ๋ฆฌ ๋ถ๋ฅ
Training Details
Training Procedure
์ฑ๋ฅ ๊ฐ์ ์ ๋จ๊ณ์ ์ผ๋ก ์งํํ์ต๋๋ค:
- QLoRA ๊ธฐ๋ฐ ํ์ธํ๋ โ ๋ํ ๋ชจ๋ธ์ 4-bit ์์ํํ์ฌ ์ ํ๋ GPU ํ๊ฒฝ์์ ๊ตฌ๋ํ๊ณ , LoRA๋ก ํจ์จ์ ์ผ๋ก ํ์ธํ๋
- ๋ชจ๋ธ ์ค์ผ์ผ์ โ ๋ ํฐ ํ๋ผ๋ฏธํฐ์ ๋ฒ ์ด์ค ๋ชจ๋ธ๋ก ๊ต์ฒดํ์ฌ ๊ธฐ๋ฐ ์ฑ๋ฅ ํฅ์
- ํ์ดํผํ๋ผ๋ฏธํฐ ํ๋ โ ํ์ต๋ฅ ๋ฑ ์ฃผ์ ํ์ดํผํ๋ผ๋ฏธํฐ ์ต์ ํ
- ์ด๋ฏธ์ง ์ ์ฒ๋ฆฌ ๊ฐ์ โ ๋์ ํฌ๊ธฐ ์ด๋ฏธ์ง ์ ๋ ฅ ์ ๋ฐ์ํ๋ ๋น์จ ์๊ณก ๋ฐ ๋ชจ์๋ฆฌ ์์ญ ์ธ์ ๋๋ฝ ๋ฌธ์ ๋ฅผ ๋ฐ๊ฒฌ, ์๋ณธ ๋น์จ์ ๋ณด์กดํ๋ ๋ฐฉ์์ผ๋ก ๊ฐ์
- GRPO ์ ์ฉ โ ์ค๋ต ์ฌ๋ก์ ๋ํ ํ์ต์ ๊ฐํํ๋ GRPO(Group Relative Policy Optimization)๋ฅผ ์ ์ฉํ์ฌ ์ถ๊ฐ ์ฑ๋ฅ ํฅ์
Training Hyperparameters
- Fine-tuning: QLoRA (4-bit)
Evaluation
Results
- Accuracy: 85% โ 93.3% (๋จ๊ณ์ ๊ฐ์ ํ)
- Challenge ranking: ์ ์ฒด 200ํ ์ค ์ต๊ณ 20์ ๊ธฐ๋ก, ์ต์ข 50์ (์์ 25%)
Summary
QLoRA ์์ํ๋ก ๋ํ ๋ชจ๋ธ์ ์ ํ๋ ํ๊ฒฝ์์ ๊ตฌ๋ํ๊ณ , ๋ชจ๋ธ ์ค์ผ์ผ์ ยทํ์ดํผํ๋ผ๋ฏธํฐ ํ๋ยท์ด๋ฏธ์ง ์ ์ฒ๋ฆฌ ๊ฐ์ ยทGRPO๋ฅผ ๋จ๊ณ์ ์ผ๋ก ์ ์ฉํ์ฌ ์ ํ๋๋ฅผ 85%์์ 93.3%๊น์ง ํฅ์์์ผฐ์ต๋๋ค.
Bias, Risks, and Limitations
- ๋ณธ ๋ชจ๋ธ์ ๋จ๊ธฐ ์ฑ๋ฆฐ์ง๋ฅผ ์ํด ํน์ ์ฌํ์ฉ ๋ถ๋ฅ ๋ฐ์ดํฐ์ ์ ํ์ธํ๋๋์์ผ๋ฉฐ, ํ์ต ๋ฐ์ดํฐ์ ๋ถํฌ๋ฅผ ๋ฒ์ด๋ ์ด๋ฏธ์ง(๋ค๋ฅธ ์กฐ๋ช , ๋ฐฐ๊ฒฝ, ๋ฌผ์ฒด ์ข ๋ฅ)์์๋ ์ฑ๋ฅ์ด ์ ํ๋ ์ ์์ต๋๋ค.
- ์ค์ ์ฌํ์ฉ ๋ถ๋ฆฌ๋ฐฐ์ถ ์์ฌ๊ฒฐ์ ์ ๋จ๋ ์ผ๋ก ์ฌ์ฉํ๊ธฐ์๋ ๊ฒ์ฆ์ด ๋ ํ์ํฉ๋๋ค.
How to Get Started with the Model
# TODO: ์ค์ ๋ก๋ฉ/์ถ๋ก ์ฝ๋ ์ถ๊ฐ from transformers import AutoModel, AutoProcessor # model = AutoModel.from_pretrained("...") # processor = AutoProcessor.from_pretrained("...")Model Card Authors
Kang Jaemin (๊ฐ์ฌ๋ฏผ)