blazeofchi commited on
Commit
a63c508
·
verified ·
1 Parent(s): 54e3507

Release v0.1.0 preview checkpoint and evaluation

Browse files
EVALUATION.md ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Evaluation: Aural One E2B Preview
2
+
3
+ This page reports the **selected Stage-B step-1,000 checkpoint**, SHA-256 `ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96`. Scores use different datasets, splits, and metrics; none is an overall System One score. One training seed was used. No separate classifier is used at inference.
4
+
5
+ ## Emotion and typed decisions
6
+
7
+ | Fixed evaluation | Frozen step-420 reference | Aural One preview | Scope |
8
+ |---|---:|---:|---|
9
+ | CREMA-D macro-F1 | 0.361 | **0.391** | 445 actor-held-out development clips |
10
+ | CREMA-D listener-vote cross-entropy ↓ | 1.535 | **1.473** | Same development set |
11
+ | CREMA-D soft-vote Brier ↓ | 0.263 | **0.232** | Same development set |
12
+ | CREMA-D 10-bin soft-vote ECE ↓ | 0.070 | **0.044** | Same development set |
13
+ | SUBESCO macro-F1 | 0.134 | **0.295** | 700 clips from two training-unseen development speakers |
14
+ | SUBESCO listener-vote cross-entropy ↓ | 2.063 | **1.605** | Same development set |
15
+ | SUBESCO soft-vote Brier ↓ | 0.647 | **0.523** | Same development set |
16
+ | SUBESCO 10-bin soft-vote ECE ↓ | 0.064 | **0.035** | Same development set |
17
+ | Typed choice correct | 188/200 | **189/200** | Fixed development examples |
18
+ | Typed No/Null correct | 175/200 | **174/200** | Fixed development examples |
19
+ | Typed ordinal score correct | 189/200 | **191/200** | Fixed development examples |
20
+
21
+ The **SUBESCO reserved test** was opened once after checkpoint selection. On four further unseen speakers, Aural One scored **0.344 macro-F1 over 1,012 unanimously perceived clips** among 1,400 recordings; listener-vote cross-entropy was **1.700** over all 1,400. Per-class F1 on unanimous clips was angry 0.469, disgusted 0.222, fearful 0.047, happy 0.358, neutral 0.581, sad 0.334, surprised 0.394. This is a speaker-held-out acted-emotion result, not a natural-conversation estimate.
22
+
23
+ On 40 same-speaker, same-words development contrasts, the selected model put the target relative emotion margin in the correct direction on **35/40**, versus **23/40** for the frozen reference. Both clips received the correct top seven-way emotion on **2/40** versus **1/40**. This indicates audio sensitivity while leaving room to improve complete class decisions.
24
+
25
+ ## Other audio checks
26
+
27
+ | Check | Frozen reference | Aural One preview | Interpretation |
28
+ |---|---:|---:|---|
29
+ | Previously opened German EmoDB macro-F1 | 0.078 | **0.102** | 195 clips; exploratory transfer, not a blind test |
30
+ | Italian Emozionalmente development macro-F1 | **0.232** | 0.203 | 262 unanimous clips, 69 development actors across all 1,202 clips |
31
+ | Italian listener-vote cross-entropy ↓ | **2.492** | 2.545 | All 1,202 Italian development clips |
32
+ | Original 24-clip sound screen correct | 13/24 | **14/24** | Eight each of baby cry, laugh, cough; source-confounded |
33
+ | Draft ~60-second speech questions correct | — | **14/18** | Six development audiobook clips, three questions each |
34
+ | New natural 45–60-second MMAU-Pro speech questions correct | 16/39 | **18/39** | 39 distinct held-out waveforms; one question each |
35
+
36
+ On the small sound screen, Aural One answered baby cry **6/8**, laughter **8/8**, cough **0/8**. A subsequent **blind human review of a different, source-diverse 24-preview candidate set** confirmed 17 clear labels, found six mixed clips and one different sound. Those human-reviewed previews have **not been scored as model accuracy**. A second 14-clip replacement review is prepared. None of this candidate media is bundled with the release.
37
+
38
+ The 39-clip natural long-speech slice scored **18/39 with the real recording** and **11/39 after audio was swapped**. The two-answer gain over the frozen reference had a paired interval spanning zero; the median maximum choice probability on incorrect real-audio answers was 0.911. The earlier six-clip draft speech slice kept **18/18 top-choice parity** when three question suffixes shared audio processing, though three probability vectors exceeded a predeclared 0.03 parity tolerance (worst 0.0378). These sets establish task-specific behavior, not general 60-second audio understanding.
39
+
40
+ ## Warm serving and cost
41
+
42
+ The fixed request here was a **57.89-second Opus recording (238 KB JSON)** with **three named choice questions**. Measurements used ten warm calls per single-request route after two warmups, and 12 calls per concurrency condition. The fast path processed shared audio once and right-batched question suffixes.
43
+
44
+ | Route | Client or pod-local wall p50 / p95 | Model p50 / p95 | Scope |
45
+ |---|---:|---:|---|
46
+ | RTX PRO 4500, Romania pod-local HTTP | **0.508 / 0.579 s** | 0.242 / 0.290 s | Same pod to authenticated app |
47
+ | Tokyo → Romania direct pod proxy | 2.086 / 2.498 s | ~0.24 / ~0.29 s | Internet client end to end |
48
+ | Tokyo → Japan H100 direct pod proxy | 1.635 / 1.952 s | ~0.21 s median | Different GPU/runtime and route |
49
+
50
+ At the Romania proxy, four concurrent clients achieved **2.029 successful requests/s** over 12 fixed requests with **2.211 s client p95** and 0.115 s queue p95; all top-choice vectors matched. At the Japan H100 proxy, four clients achieved **2.368 successful requests/s**, client p95 **1.888 s**, and queue p95 0.048 s over 12 calls. These short bursts are not sustained-capacity benchmarks.
51
+
52
+ At the observed Romania four-client throughput, allocating the displayed **$0.72/hour GPU price** gives **about $0.000099 per successful request**, or **$0.000061 per 1,000 shared input positions** for this request. This is a GPU-only allocation at sustained throughput; idle, disk, transfer, cold starts and network are excluded. It is not a token-billed price. The Japan H100 allocation was about $0.000409/request at $3.49/hour and the observed 2.368 RPS. The selected Stage-B run reported **11,006 MiB peak allocated GPU memory** during training and evaluation on the 32 GB RTX PRO 4500; minimum inference VRAM was not measured.
53
+
54
+ The Tokyo end-to-end sub-second target is **still a research target**. These HTTP tests used temporary authenticated pods, not a public serverless endpoint. Cold start, sustained load, broader schema choices, and production p95 remain to be measured.
55
+
56
+ The **public reference loader** was separately smoke-tested on a 24 GB RTX PRO 6000 Blackwell partition with PyTorch 2.13.0+cu130. It loaded the pinned release files, verified their hashes, and returned valid distributions for three named questions on both a short speech clip and a synthetic 58-second two-chunk audio input. The three sequential forwards took **5.128 s** on the first short run and **3.669 s** on the later 58-second run after model/base caching; PyTorch reported **9,774 MiB allocated**. These are functional smokes, not an accuracy or warm-service benchmark, and this simple loader does not implement the shared-audio fast path measured above.
57
+
58
+ ## Training follow-ups
59
+
60
+ Two controlled 100-update follow-ups were explored after selecting this release checkpoint. Training all twelve Conformer blocks improved Italian development listener-vote cross-entropy to **2.404** but reached only **0.207** unanimous macro-F1, with sad **0/44** and fearful **1/17** correct; it was not promoted. A separate per-clip top-margin arm reached Italian CE **2.417** and F1 **0.211** but reduced the fixed CREMA-D screen macro-F1 **0.343 → 0.288**; it was rejected. The published weights are the earlier Stage-B checkpoint, not either follow-up.
61
+
62
+ ## Reproducibility notes
63
+
64
+ The base revision, delta hashes, seed, and update count are pinned in [`release.json`](../release.json). The metric scopes above come from the frozen development, reserved-test, long-speech, and warm-serving runs. Audio and row-level private results are not redistributed here because the source datasets have their own terms. The public code exposes the choice-scoring path; the optimized HTTP endpoint used for timing remains experimental and is not represented as a production deployment.
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright [yyyy] [name of copyright owner]
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
NOTICE ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ Aural One E2B Preview
2
+ Copyright 2026 Paras Sharma
3
+
4
+ The Aural One adapter and acoustic delta modify Google Gemma 4 E2B Instruct.
5
+ The base model is not included in this repository and is available from
6
+ https://huggingface.co/google/gemma-4-E2B-it under Apache License 2.0.
7
+
8
+ The release code, adapter, and acoustic delta are distributed under Apache
9
+ License 2.0. See LICENSE. The training and evaluation datasets have separate
10
+ licenses and are not redistributed here.
README.md ADDED
@@ -0,0 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ - bn
5
+ license: apache-2.0
6
+ library_name: transformers
7
+ pipeline_tag: audio-text-to-text
8
+ base_model: google/gemma-4-E2B-it
9
+ tags:
10
+ - audio
11
+ - gemma
12
+ - structured-decisions
13
+ - speech-emotion-recognition
14
+ - research-preview
15
+ ---
16
+
17
+ # Aural One E2B Preview
18
+
19
+ **Native audio in, structured decisions out.** Aural One adapts Gemma 4 E2B to score supplied choices from a recording, a written state, and named questions. One model handles the audio and the decision; its inference path does not require speech-to-text or a separate sound classifier.
20
+
21
+ This **v0.1.0 research preview** publishes a frozen adapter and **54.3 million updated acoustic/projection weights**. The [pinned Gemma 4 E2B base](https://huggingface.co/google/gemma-4-E2B-it) is downloaded separately. The language model was frozen during this final training stage.
22
+
23
+ ## What we measured
24
+
25
+ | Evaluation | Aural One preview |
26
+ |---|---:|
27
+ | English CREMA-D actor-held-out development, macro-F1 | **0.391** on 445 clips |
28
+ | Bangla SUBESCO speaker-held-out development, macro-F1 | **0.295** on 700 clips, up from 0.134 reference |
29
+ | Bangla SUBESCO four-speaker reserved test, macro-F1 | **0.344** on 1,012 consensus clips |
30
+ | Structured choice / No-Null / ordinal development | **189 / 174 / 191** correct out of 200 each |
31
+ | Warm pod-local HTTP, ~58-second Opus with three questions | **0.508 s p50 / 0.579 s p95** |
32
+
33
+ These numbers have different test scopes. The full [evaluation report](EVALUATION.md) covers the frozen reference, listener-vote cross-entropy and calibration, same-words contrasts, sound events, Italian/German transfer, long audio, warm client latency, concurrency, and GPU-only cost. Aural One is an early release with room to improve rare emotions and natural sound transfer. The sub-second number above is measured **inside the warm GPU pod**; Tokyo end-to-end sub-second latency is still a research goal.
34
+
35
+ ## Use the model
36
+
37
+ The public [GitHub repository](https://github.com/Parassharmaa/aural-one) contains the pinned loader and JSON example. Use Python 3.12 with an NVIDIA GPU; install a compatible PyTorch build first, then:
38
+
39
+ ```bash
40
+ pip install git+https://github.com/Parassharmaa/aural-one.git
41
+ ```
42
+
43
+ ```python
44
+ from aural_one import load_aural_one
45
+
46
+ model = load_aural_one("blazeofchi/Aural-One-E2B")
47
+ result = model.score(
48
+ audio="/absolute/path/to/your-audio.wav",
49
+ state={"task": "listen to the voice"},
50
+ questions={
51
+ "emotion": {
52
+ "question": "Which emotion is most evident in the speaker's voice?",
53
+ "options": ["angry", "disgusted", "fearful", "happy", "neutral", "sad", "surprised"],
54
+ },
55
+ "baby_cry": {
56
+ "question": "Is a baby crying audible?",
57
+ "options": ["No", "Yes"],
58
+ },
59
+ },
60
+ )
61
+ print(result)
62
+ ```
63
+
64
+ The preview scores **2–8 options** per named question. A binary or ordinal value is represented by its supplied options. Probabilities are normalized over those options and are **not calibrated confidence estimates**. The simple public loader scores questions separately; the warm HTTP timing above used an experimental shared-audio serving path. The reference loader passed short and synthetic 58-second smoke tests on a **24 GB Blackwell GPU partition**, with 9.8 GiB PyTorch allocation. A lower minimum has not been established.
65
+
66
+ ## Model and data
67
+
68
+ The base revision is `3e22461f65e89153144f8adb70e3b8c2cc9845a7`. The acoustic delta SHA-256 is `ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96` and the adapter SHA-256 is `e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3`. [`release.json`](release.json) pins them for the loader.
69
+
70
+ The selected run used crowd-voted [CREMA-D](https://github.com/CheyneyComputerScience/CREMA-D) English acted speech, listener-voted [SUBESCO](https://zenodo.org/records/4526477) Bangla acted speech, and project typed-decision examples. Training kept the language model and prior adapter frozen while updating the last two audio Conformer blocks and the two audio projections. Details and source terms are in [TRAINING.md](TRAINING.md). No training or evaluation audio is uploaded here.
71
+
72
+ Code and Aural One weight deltas are Apache 2.0. The separately downloaded Gemma 4 E2B base is also Apache 2.0. Aural One is independent of TypeSafe AI and Jev.
TRAINING.md ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Training and data
2
+
3
+ Aural One E2B Preview uses [Gemma 4 E2B Instruct](https://huggingface.co/google/gemma-4-E2B-it) at revision `3e22461f65e89153144f8adb70e3b8c2cc9845a7`. The release is a delta: a frozen prior adapter plus 54,294,272 updated weights in the final two audio Conformer blocks, audio output projection, and audio-to-language projection. The language model weights were frozen in this stage.
4
+
5
+ The selected run used seed `20260928`, **1,000 updates** with eight examples per update and a fixed audio/question schedule. Its 8,000 exposures comprised 3,200 CREMA-D perceived-emotion, 2,400 SUBESCO perceived-emotion, and 2,400 typed-decision examples. Half of the SUBESCO training updates included same-speaker, same-words recordings with different perceived emotion. Supervision combined listener-vote cross-entropy with a crossed audio/option margin (weight 0.25, margin 0.5). The run used BF16 forward weights, FP32 master AdamW, Conformer learning rate 2e-6, projection learning rate 5e-6, weight decay 0.01, and gradient norm clipping at 1.0.
6
+
7
+ The training data sources were:
8
+
9
+ | Source | Role | Source terms |
10
+ |---|---|---|
11
+ | [CREMA-D](https://github.com/CheyneyComputerScience/CREMA-D) | Crowd-voted acted English emotion audio | Dataset: ODbL 1.0; individual contents: DBCL 1.0 |
12
+ | [SUBESCO](https://zenodo.org/records/4526477) | Listener-voted acted Bangla emotion audio | CC BY 4.0 |
13
+ | Project typed spoken-question/intent examples | Structured-choice retention | Project research examples; no raw examples are included in this release |
14
+
15
+ The audio files and speaker-level manifests are **not included** in either public repository. Evaluation datasets have their own terms. For source-disjoint speaker protocols, sample counts, and external checks, see [EVALUATION.md](EVALUATION.md).
16
+
17
+ The published delta is SHA-256 `ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96` for the acoustic weights and `e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3` for the adapter. `release.json` pins the base, file hashes, update count, and seed. The base weights are downloaded directly from Google's repository.
acoustic/acoustic_weights.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96
3
+ size 108594768
adapter/adapter.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3
3
+ size 5997752
configs/stage_b_paired.yaml ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ model:
2
+ repo: google/gemma-4-E2B-it
3
+ revision: 3e22461f65e89153144f8adb70e3b8c2cc9845a7
4
+ answer_labels: [A, B, C, D, E, F, G, H]
5
+ inference_dtype: bfloat16
6
+ train:
7
+ seed: 20260928
8
+ lora_rank: 16
9
+ lora_alpha: 32
10
+ audio_lora_rank: 8
11
+ audio_lora_alpha: 16
12
+ audio_lora_last_layers: 4
13
+ inference_dtype: bfloat16
14
+ gradient_checkpointing: true
15
+ full_audio_last_layers: 2
16
+ block_lr: 0.000002
17
+ projection_lr: 0.000005
18
+ weight_decay: 0.01
19
+ grad_clip_norm: 1.0
20
+ microbatches_per_update: 8
21
+ max_updates: 1000
22
+ warmup_updates: 20
23
+ pair_weight: 0.25
24
+ pair_margin: 0.5
25
+ pair_manifest_sha256: 9e76d3fb02b41e99131343bcea3efc9e2d71aa96876a47f9aa06fcbcbd8a4836
26
+ screen_eval_steps: [0, 100, 250, 500, 750, 1000]
27
+ full_eval_steps: [0, 1000]
28
+ checkpoint_steps: [50, 100, 250, 500, 750, 1000]
29
+ schedule_sha256: cc8c4460fa41c78a641871cc9444a7ce7b70db63313a2dc7cbb7535e5c7eed73
30
+ sound_screen_floor_drop: 1
31
+ typed_screen_floor_drop_per_type: 1
32
+ typed_full_floor_drop_per_type: 4
33
+ crema_macro_f1_floor_drop: 0.02
34
+ subesco_consensus_macro_f1_min_gain: 0.05
release.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "aural-one-e2b-preview-v1",
3
+ "release": "v0.1.0-preview",
4
+ "base_model": "google/gemma-4-E2B-it",
5
+ "base_revision": "3e22461f65e89153144f8adb70e3b8c2cc9845a7",
6
+ "base_weight_sha256": "2db5482b20d746879bb3ef79b5203e9075a2e2b98f54ec7c2f281c1477ddc550",
7
+ "config_sha256": "447358c5fc576d1414c4c52f9d6bb6aab7dc61a70379155bfa5a235d2cba8a88",
8
+ "adapter_sha256": "e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3",
9
+ "acoustic_sha256": "ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96",
10
+ "acoustic_parameter_count": 54294272,
11
+ "training_updates": 1000,
12
+ "seed": 20260928
13
+ }