Audio-Text-to-Text
Transformers
Safetensors
qwen2_5_omni
text-to-audio
audio
audio-question-answering
audio-classification
candidate-scoring
Instructions to use shlv/AudioJev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shlv/AudioJev with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shlv/AudioJev") model = AutoModelForMultimodalLM.from_pretrained("shlv/AudioJev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Keep only standard model assets and documentation
Browse filesFollow the Qwen2.5-Omni-3B model repository layout: remove auxiliary training, evaluation, and release manifests. Retain checkpoint assets, model card, license, and the required derivative notice.
- Notice +1 -1
- README.md +3 -3
- evaluation_results.json +0 -37
- release_manifest.json +0 -111
- training_recipe.json +0 -61
Notice
CHANGED
|
@@ -11,4 +11,4 @@ Modified weight files relative to the upstream pretrained model:
|
|
| 11 |
model-00004-of-00005.safetensors
|
| 12 |
model-00005-of-00005.safetensors
|
| 13 |
|
| 14 |
-
These shards are byte-identical to the evaluated AudioJev checkpoint. The model/configuration and shard layout were saved by the training pipeline. The release corrects the optional total_parameters index metadata to the number of elements actually stored; tensor data and the weight map are unchanged. README.md
|
|
|
|
| 11 |
model-00004-of-00005.safetensors
|
| 12 |
model-00005-of-00005.safetensors
|
| 13 |
|
| 14 |
+
These shards are byte-identical to the evaluated AudioJev checkpoint. The model/configuration and shard layout were saved by the training pipeline. The release corrects the optional total_parameters index metadata to the number of elements actually stored; tensor data and the weight map are unchanged. README.md documents this derivative. Inference software is maintained separately in AudioJev-Inference.
|
README.md
CHANGED
|
@@ -38,13 +38,13 @@ AudioJev takes an audio waveform, a natural-language question, and a list of can
|
|
| 38 |
| Recommended inference dtype | BF16 |
|
| 39 |
| Candidate labels | `0`–`9`, then `A`–`Z` (2–36 candidates) |
|
| 40 |
|
| 41 |
-
The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint
|
| 42 |
|
| 43 |
For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.
|
| 44 |
|
| 45 |
## Inference
|
| 46 |
|
| 47 |
-
Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone **AudioJev-Inference** project. This model repository distributes the model weights, tokenizer/processor assets, model
|
| 48 |
|
| 49 |
After installing AudioJev-Inference, start the service with:
|
| 50 |
|
|
@@ -66,7 +66,7 @@ The following scores are for **this single checkpoint in the original candidate
|
|
| 66 |
| MMAR | 1,000 | 56.30% |
|
| 67 |
| MMSU evaluated short-audio subset | 3,931 | 62.45% |
|
| 68 |
|
| 69 |
-
|
| 70 |
|
| 71 |
## Scope
|
| 72 |
|
|
|
|
| 38 |
| Recommended inference dtype | BF16 |
|
| 39 |
| Candidate labels | `0`–`9`, then `A`–`Z` (2–36 candidates) |
|
| 40 |
|
| 41 |
+
The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint.
|
| 42 |
|
| 43 |
For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.
|
| 44 |
|
| 45 |
## Inference
|
| 46 |
|
| 47 |
+
Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone **AudioJev-Inference** project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files.
|
| 48 |
|
| 49 |
After installing AudioJev-Inference, start the service with:
|
| 50 |
|
|
|
|
| 66 |
| MMAR | 1,000 | 56.30% |
|
| 67 |
| MMSU evaluated short-audio subset | 3,931 | 62.45% |
|
| 68 |
|
| 69 |
+
The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind.
|
| 70 |
|
| 71 |
## Scope
|
| 72 |
|
evaluation_results.json
DELETED
|
@@ -1,37 +0,0 @@
|
|
| 1 |
-
{
|
| 2 |
-
"training_seed": 20261001,
|
| 3 |
-
"skl_weight": 0.5,
|
| 4 |
-
"candidate_order": "native",
|
| 5 |
-
"aggregation": "single checkpoint; not a seed average",
|
| 6 |
-
"metrics": {
|
| 7 |
-
"mmau_full": {
|
| 8 |
-
"questions": 10000,
|
| 9 |
-
"correct": 6941,
|
| 10 |
-
"accuracy": 0.6941
|
| 11 |
-
},
|
| 12 |
-
"mmau": {
|
| 13 |
-
"questions": 9000,
|
| 14 |
-
"correct": 6230,
|
| 15 |
-
"accuracy": 0.6922222222222222
|
| 16 |
-
},
|
| 17 |
-
"mmau_mini": {
|
| 18 |
-
"questions": 1000,
|
| 19 |
-
"correct": 711,
|
| 20 |
-
"accuracy": 0.711
|
| 21 |
-
},
|
| 22 |
-
"mmar": {
|
| 23 |
-
"questions": 1000,
|
| 24 |
-
"correct": 563,
|
| 25 |
-
"accuracy": 0.563
|
| 26 |
-
},
|
| 27 |
-
"mmsu": {
|
| 28 |
-
"questions": 3931,
|
| 29 |
-
"correct": 2455,
|
| 30 |
-
"accuracy": 0.6245230221317731
|
| 31 |
-
}
|
| 32 |
-
},
|
| 33 |
-
"mmau_scoring": "1000 public questions plus 9000 questions scored by the official service",
|
| 34 |
-
"mmsu_scope": "3931 questions in the evaluated short-audio subset",
|
| 35 |
-
"source_overlap_note": "The experiment reports 119 hidden MMAU questions whose audio sources overlap general-stage fitting/development data. Full scores retain them.",
|
| 36 |
-
"source_report_sha256": "dc53decb9db340a6b4b995b6ba475e23af09fcbe9e1cd091325fe1a6219ad5e3"
|
| 37 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
release_manifest.json
DELETED
|
@@ -1,111 +0,0 @@
|
|
| 1 |
-
{
|
| 2 |
-
"model_name": "AudioJev",
|
| 3 |
-
"repo_id": "shlv/AudioJev",
|
| 4 |
-
"created_utc": "2026-09-28T16:06:42.633063+00:00",
|
| 5 |
-
"training_seed": 20261001,
|
| 6 |
-
"skl_weight": 0.5,
|
| 7 |
-
"checkpoint": "general/checkpoint-2048",
|
| 8 |
-
"total_training_updates": 4096,
|
| 9 |
-
"weight_shards_unchanged": true,
|
| 10 |
-
"stored_tensor_parameters": 4703464448,
|
| 11 |
-
"stored_tensor_dtypes": {
|
| 12 |
-
"F32": 4703464448
|
| 13 |
-
},
|
| 14 |
-
"index_metadata_correction": {
|
| 15 |
-
"original_total_parameters": 1519790592,
|
| 16 |
-
"actual_total_parameters": 4703464448
|
| 17 |
-
},
|
| 18 |
-
"source_index_sha256": "7ab45af26aa38dd632422333408a6dfea16c180d9fc5d382cfd1a7999b58a502",
|
| 19 |
-
"files": {
|
| 20 |
-
"LICENSE": {
|
| 21 |
-
"size": 7387,
|
| 22 |
-
"sha256": "ae8ef9d3fb476f735d9b9abeaab73b7f778b121c77427cd918ec3da01eefbc44"
|
| 23 |
-
},
|
| 24 |
-
"Notice": {
|
| 25 |
-
"size": 1211,
|
| 26 |
-
"sha256": "8ad741a5ad2d7c46d7eba46772e2dead6c627ab383dacbb1ecd7ef152a56b3ac"
|
| 27 |
-
},
|
| 28 |
-
"README.md": {
|
| 29 |
-
"size": 4992,
|
| 30 |
-
"sha256": "527dd557d508e9a52d147a8bd0709c9ddd7a6f27ff8bdc65d3d382bd9616538d"
|
| 31 |
-
},
|
| 32 |
-
"added_tokens.json": {
|
| 33 |
-
"size": 579,
|
| 34 |
-
"sha256": "e81a2cc3bd867a1217019eed202d1d8a07e1063ece716f22060dda14f6cc07d8"
|
| 35 |
-
},
|
| 36 |
-
"chat_template.jinja": {
|
| 37 |
-
"size": 1281,
|
| 38 |
-
"sha256": "d027e14db2e64eb2c23597240d42ea747b4fd5efd6e6a7604a990949c610f79c"
|
| 39 |
-
},
|
| 40 |
-
"config.json": {
|
| 41 |
-
"size": 16911,
|
| 42 |
-
"sha256": "08755783271e6c0f425fa705bccb85ae6e0b2c3db4ce904b576fb385f6d7811f"
|
| 43 |
-
},
|
| 44 |
-
"evaluation_results.json": {
|
| 45 |
-
"size": 1090,
|
| 46 |
-
"sha256": "9ce7bd97ff76bfc0dd297e9ff75c5269528b501513ba4e44d2bd3684167301d6"
|
| 47 |
-
},
|
| 48 |
-
"generation_config.json": {
|
| 49 |
-
"size": 69,
|
| 50 |
-
"sha256": "3ee3cbe9acb55059b37a4e90bb843812eb86f70c5ebea404f137cd9c80127870"
|
| 51 |
-
},
|
| 52 |
-
"merges.txt": {
|
| 53 |
-
"size": 1671853,
|
| 54 |
-
"sha256": "8831e4f1a044471340f7c0a83d7bd71306a5b867e95fd870f74d0c5308a904d5"
|
| 55 |
-
},
|
| 56 |
-
"model-00001-of-00005.safetensors": {
|
| 57 |
-
"size": 3995066200,
|
| 58 |
-
"sha256": "435a6fbfe7c120119b86076831c2c48542f5018c227d3cfc0856a210d351bf3b"
|
| 59 |
-
},
|
| 60 |
-
"model-00002-of-00005.safetensors": {
|
| 61 |
-
"size": 3926509296,
|
| 62 |
-
"sha256": "3087437831e84d8430707b4a3be2dca239af4049b0f944f76f6062721b9c89c0"
|
| 63 |
-
},
|
| 64 |
-
"model-00003-of-00005.safetensors": {
|
| 65 |
-
"size": 3917844848,
|
| 66 |
-
"sha256": "35b3a36a2d092b19a0427c016c8d28cea4a4e0b5f3dcf59934b7d74eded521af"
|
| 67 |
-
},
|
| 68 |
-
"model-00004-of-00005.safetensors": {
|
| 69 |
-
"size": 3917844920,
|
| 70 |
-
"sha256": "dbf98c2f42746b10be25de4087fd65935415bacd1853c7a911ebac9e153cec3a"
|
| 71 |
-
},
|
| 72 |
-
"model-00005-of-00005.safetensors": {
|
| 73 |
-
"size": 3056764904,
|
| 74 |
-
"sha256": "ebd56db24da75a8b4fc022f2f2e6dbc75a4e7995d8c140be2d69289ce9c07390"
|
| 75 |
-
},
|
| 76 |
-
"model.safetensors.index.json": {
|
| 77 |
-
"size": 127717,
|
| 78 |
-
"sha256": "6ac3736c9a2d653b21c98a41e82e963eae26008d362b9328cd1efe4427896fdb"
|
| 79 |
-
},
|
| 80 |
-
"preprocessor_config.json": {
|
| 81 |
-
"size": 667,
|
| 82 |
-
"sha256": "b47055ce61463ce143e9aab741d55c0aa520801a0a5d63be73c5b17cecb6bc69"
|
| 83 |
-
},
|
| 84 |
-
"special_tokens_map.json": {
|
| 85 |
-
"size": 833,
|
| 86 |
-
"sha256": "db7fd39f5dc9ee37998c3ed04e3a3386989182a550064d2a2a9af16822cf22f4"
|
| 87 |
-
},
|
| 88 |
-
"spk_dict.pt": {
|
| 89 |
-
"size": 259544,
|
| 90 |
-
"sha256": "6a05609b28f5d42b7b748f0f07592545c8f1f6885b9ae8fff64baf56e86b2a18"
|
| 91 |
-
},
|
| 92 |
-
"tokenizer_config.json": {
|
| 93 |
-
"size": 5181,
|
| 94 |
-
"sha256": "d2705fa47e1aecad75e3ac83ca42ec983d47bd48e9c740791ebf2c005922704b"
|
| 95 |
-
},
|
| 96 |
-
"training_recipe.json": {
|
| 97 |
-
"size": 2551,
|
| 98 |
-
"sha256": "9ea8795d0ff10e5c94042d9fd22a4305101abbddc53968cefd1267242e0d5d52"
|
| 99 |
-
},
|
| 100 |
-
"video_preprocessor_config.json": {
|
| 101 |
-
"size": 1229,
|
| 102 |
-
"sha256": "3e7254028ff085aa9e0322c2db0f2b86781fb186d1d690ad4e990eb01fd466c2"
|
| 103 |
-
},
|
| 104 |
-
"vocab.json": {
|
| 105 |
-
"size": 3383407,
|
| 106 |
-
"sha256": "87a257b04b17642a0688c98cd1df89c398bda4fee532d6f88b38a659ecb4ac8d"
|
| 107 |
-
}
|
| 108 |
-
},
|
| 109 |
-
"layout_updated_utc": "2026-09-28T16:34:03.104869+00:00",
|
| 110 |
-
"inference_project": "AudioJev-Inference"
|
| 111 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
training_recipe.json
DELETED
|
@@ -1,61 +0,0 @@
|
|
| 1 |
-
{
|
| 2 |
-
"model_name": "AudioJev",
|
| 3 |
-
"base_model": "Qwen/Qwen2.5-Omni-3B",
|
| 4 |
-
"seed": 20261001,
|
| 5 |
-
"skl_weight": 0.5,
|
| 6 |
-
"method": "random_derangement_skl",
|
| 7 |
-
"loss": "paired_candidate_ce_plus_aligned_skl",
|
| 8 |
-
"pair_kind": "random",
|
| 9 |
-
"max_prompt_tokens": 4096,
|
| 10 |
-
"stages": [
|
| 11 |
-
{
|
| 12 |
-
"name": "semantic",
|
| 13 |
-
"learning_rate": 1e-06,
|
| 14 |
-
"warmup_steps": 32,
|
| 15 |
-
"eval_every": 256,
|
| 16 |
-
"examples": 4096,
|
| 17 |
-
"steps": 1024,
|
| 18 |
-
"development_examples": 256,
|
| 19 |
-
"ordered_train_sha256": "a9bd81622009a8b3ac0703eebdc9aeedd47c0c40e13403a3551fd83478210d50",
|
| 20 |
-
"source_train_sha256": "dcce30c0f53fd226f626ab9e55962671bbb5c82104f99de20e830c74c2fcdc74",
|
| 21 |
-
"development_sha256": "d6514e47cb62896de71c1418320106f32f997dc4e0ea68284b3784d861428368"
|
| 22 |
-
},
|
| 23 |
-
{
|
| 24 |
-
"name": "joint",
|
| 25 |
-
"learning_rate": 1e-06,
|
| 26 |
-
"warmup_steps": 32,
|
| 27 |
-
"eval_every": 256,
|
| 28 |
-
"examples": 4096,
|
| 29 |
-
"steps": 1024,
|
| 30 |
-
"development_examples": 320,
|
| 31 |
-
"ordered_train_sha256": "22d9fc5a9722b69363f1d614375d396af76e39e95c80ab8a0dec0af9bf949f9d",
|
| 32 |
-
"source_train_sha256": "7393eb4bb002ad2810d06c520f08e52eeca5feddb60becc1c58b9b826da06c19",
|
| 33 |
-
"development_sha256": "1af43e24f729f29be82b8229d341336a7a0a66549ffe0756dca0d7d172fe9172"
|
| 34 |
-
},
|
| 35 |
-
{
|
| 36 |
-
"name": "general",
|
| 37 |
-
"learning_rate": 1e-06,
|
| 38 |
-
"warmup_steps": 32,
|
| 39 |
-
"eval_every": 512,
|
| 40 |
-
"examples": 8192,
|
| 41 |
-
"steps": 2048,
|
| 42 |
-
"development_examples": 576,
|
| 43 |
-
"ordered_train_sha256": "5c2ae98a4a57bef65f77a43c8329ec9d9015d4b7b10094485d73380dc90c2842",
|
| 44 |
-
"source_train_sha256": "c6803170e0a5cee9f978f22e562855ca2ca37240475b823bf1efa22e2b13247d",
|
| 45 |
-
"development_sha256": "5bb427425dc64822e35cda72b270b7bfd3bb3d2fb1405c6edaa5a610a6a43c45"
|
| 46 |
-
}
|
| 47 |
-
],
|
| 48 |
-
"completed_training_updates": 4096,
|
| 49 |
-
"final_stage": "general",
|
| 50 |
-
"final_stage_step": 2048,
|
| 51 |
-
"checkpoint_choice": "fixed_final_step",
|
| 52 |
-
"checkpoint_selection_uses_development": false,
|
| 53 |
-
"checkpoint_selection_uses_test": false,
|
| 54 |
-
"hyperparameter_note": "Lambda 0.5 was chosen after the weight-comparison experiments; the fixed-step statement describes checkpoint selection only.",
|
| 55 |
-
"weights_dtype": "float32",
|
| 56 |
-
"evaluated_compute_dtype": "bfloat16",
|
| 57 |
-
"candidate_labels": "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZ",
|
| 58 |
-
"probability_readout": "softmax over supplied candidate-label logits at the final input position",
|
| 59 |
-
"evaluation_plan_sha256": "c00a67288c7f49b59891ce6bbcc69352f8f8895be0ea49f27627b207cee053b1",
|
| 60 |
-
"checkpoint_metadata_sha256": "81c2e6c0ac8d1a033f37c74e143ec1a704623891eb6ac5b83a41950f7677e094"
|
| 61 |
-
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|