Audio-Text-to-Text
Transformers
Safetensors
qwen2_5_omni
text-to-audio
audio
audio-question-answering
audio-classification
candidate-scoring
Instructions to use shlv/AudioJev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shlv/AudioJev with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shlv/AudioJev") model = AutoModelForMultimodalLM.from_pretrained("shlv/AudioJev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 1,139 Bytes
25bc587 d506a7b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 | Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) Alibaba Cloud. All Rights Reserved.
Built with Qwen.
AudioJev is a derivative of Qwen/Qwen2.5-Omni-3B. The saved Thinker weights were modified through three-stage fine-tuning of the audio and language decision path, using random-derangement paired candidate cross-entropy and aligned symmetric KL (lambda=0.5, seed=20261001, 4096 updates). The visual branch was excluded from training; the Talker is not included.
Modified weight files relative to the upstream pretrained model:
model-00001-of-00005.safetensors
model-00002-of-00005.safetensors
model-00003-of-00005.safetensors
model-00004-of-00005.safetensors
model-00005-of-00005.safetensors
These shards are byte-identical to the evaluated AudioJev checkpoint. The model/configuration and shard layout were saved by the training pipeline. The release corrects the optional total_parameters index metadata to the number of elements actually stored; tensor data and the weight map are unchanged. README.md documents this derivative. Inference software is maintained separately in AudioJev-Inference.
|