Audio-Text-to-Text
Transformers
English
Bengali
audio
gemma
structured-decisions
speech-emotion-recognition
research-preview
Instructions to use blazeofchi/Aural-One-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blazeofchi/Aural-One-E2B with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("blazeofchi/Aural-One-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Embed Aural One introduction in model card
Browse files
README.md
CHANGED
|
@@ -19,6 +19,8 @@ tags:
|
|
| 19 |
|
| 20 |
# Aural One E2B Preview
|
| 21 |
|
|
|
|
|
|
|
| 22 |
**Native audio in, structured decisions out.** Aural One adapts Gemma 4 E2B to score supplied choices from a recording, a written state, and named questions. One model handles the audio and the decision; its inference path does not require speech-to-text or a separate sound classifier.
|
| 23 |
|
| 24 |
This **v0.1.0 research preview** publishes a frozen adapter and **54.3 million updated acoustic/projection weights**. The [pinned Gemma 4 E2B base](https://huggingface.co/google/gemma-4-E2B-it) is downloaded separately. The language model was frozen during this final training stage.
|
|
@@ -81,3 +83,5 @@ Use this preview for research and prototyping of audio-grounded, state-condition
|
|
| 81 |
Code and Aural One weight deltas are Apache 2.0. The separately downloaded Gemma 4 E2B base is also Apache 2.0.
|
| 82 |
|
| 83 |
The [public release checklist](https://github.com/Parassharmaa/aural-one/blob/main/docs/RELEASE_CHECKLIST.md) records the final model-card, code, chart, hash, and GPU smoke checks.
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
# Aural One E2B Preview
|
| 21 |
|
| 22 |
+
<video src="https://huggingface.co/blazeofchi/Aural-One-E2B/resolve/main/assets/demo/aural-one-intro.mp4" poster="https://huggingface.co/blazeofchi/Aural-One-E2B/resolve/main/assets/demo/aural-one-intro-poster.png" controls width="100%"></video>
|
| 23 |
+
|
| 24 |
**Native audio in, structured decisions out.** Aural One adapts Gemma 4 E2B to score supplied choices from a recording, a written state, and named questions. One model handles the audio and the decision; its inference path does not require speech-to-text or a separate sound classifier.
|
| 25 |
|
| 26 |
This **v0.1.0 research preview** publishes a frozen adapter and **54.3 million updated acoustic/projection weights**. The [pinned Gemma 4 E2B base](https://huggingface.co/google/gemma-4-E2B-it) is downloaded separately. The language model was frozen during this final training stage.
|
|
|
|
| 83 |
Code and Aural One weight deltas are Apache 2.0. The separately downloaded Gemma 4 E2B base is also Apache 2.0.
|
| 84 |
|
| 85 |
The [public release checklist](https://github.com/Parassharmaa/aural-one/blob/main/docs/RELEASE_CHECKLIST.md) records the final model-card, code, chart, hash, and GPU smoke checks.
|
| 86 |
+
|
| 87 |
+
Audio in the introduction: [CC0 recording credits](assets/demo/ATTRIBUTION.md).
|