--- language: - en - bn license: apache-2.0 library_name: transformers pipeline_tag: audio-text-to-text base_model: google/gemma-4-E2B-it datasets: - yijingwu/HeySQuAD_human - PolyAI/minds14 tags: - audio - gemma - structured-decisions - speech-emotion-recognition - research-preview --- # Aural One E2B Preview **Native audio in, structured decisions out.** Aural One adapts Gemma 4 E2B to score supplied choices from a recording, a written state, and named questions. One model handles the audio and the decision; its inference path does not require speech-to-text or a separate sound classifier. This **v0.1.0 research preview** publishes a frozen adapter and **54.3 million updated acoustic/projection weights**. The [pinned Gemma 4 E2B base](https://huggingface.co/google/gemma-4-E2B-it) is downloaded separately. The language model was frozen during this final training stage. ## What we measured | Evaluation | Aural One preview | |---|---:| | English CREMA-D actor-held-out development, macro-F1 | **0.391** on 445 clips | | Bangla SUBESCO speaker-held-out development, macro-F1 | **0.295** on 700 clips, up from 0.134 reference | | Bangla SUBESCO four-speaker reserved test, macro-F1 | **0.344** on 1,012 consensus clips | | Structured choice / No-Null / ordinal development | **189 / 174 / 191** correct out of 200 each | | Warm pod-local HTTP, ~58-second Opus with three questions | **0.508 s p50 / 0.579 s p95** | These numbers have different test scopes. The full [evaluation report](EVALUATION.md) covers the frozen reference, listener-vote cross-entropy and calibration, same-words contrasts, sound events, Italian/German transfer, long audio, warm client latency, concurrency, and GPU-only cost. [Release validation](VALIDATION.md) records the public-Hub load test. Aural One is an early release with room to improve rare emotions and natural sound transfer. The sub-second number above is measured **inside the warm GPU pod**; Tokyo end-to-end sub-second latency is still a research goal. ![Aural One audio evaluation chart: speech emotion macro-F1 by dataset and separate audio-question correct counts](assets/audio-evaluation.png) ## Use the model The public [GitHub repository](https://github.com/Parassharmaa/aural-one) contains the pinned loader, JSON example, [full acoustic fine-tuning code and data schema](https://github.com/Parassharmaa/aural-one/blob/main/docs/FINE_TUNING.md), and evaluation entry point. Use Python 3.12 with an NVIDIA GPU; install a compatible PyTorch build first, then: ```bash pip install git+https://github.com/Parassharmaa/aural-one.git ``` ```python from aural_one import load_aural_one model = load_aural_one("blazeofchi/Aural-One-E2B") result = model.score( audio="/absolute/path/to/your-audio.wav", state={"task": "listen to the voice"}, questions={ "emotion": { "question": "Which emotion is most evident in the speaker's voice?", "options": ["angry", "disgusted", "fearful", "happy", "neutral", "sad", "surprised"], }, "baby_cry": { "question": "Is a baby crying audible?", "options": ["No", "Yes"], }, }, ) print(result) ``` The preview scores **2–8 options** per named question. A binary or ordinal value is represented by its supplied options. Probabilities are normalized over those options and are **not calibrated confidence estimates**. The simple public loader scores questions separately; the warm HTTP timing above used an experimental shared-audio serving path. The reference loader passed short and synthetic 58-second smoke tests on a **24 GB Blackwell GPU partition**, with 9.8 GiB PyTorch allocation. A lower minimum has not been established. ## Model and data The base revision is `3e22461f65e89153144f8adb70e3b8c2cc9845a7`. The acoustic delta SHA-256 is `ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96` and the adapter SHA-256 is `e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3`. [`release.json`](release.json) pins them for the loader. The selected run used crowd-voted [CREMA-D](https://github.com/CheyneyComputerScience/CREMA-D) English acted speech, listener-voted [SUBESCO](https://zenodo.org/records/4526477) Bangla acted speech, [HeySQuAD Human](https://huggingface.co/datasets/yijingwu/HeySQuAD_human), [MInDS-14](https://huggingface.co/datasets/PolyAI/minds14), and original synthetic speech for typed decisions. Training kept the language model and prior adapter frozen while updating the last two audio Conformer blocks and the two audio projections. Details and source terms are in [TRAINING.md](TRAINING.md). No training or evaluation audio is uploaded here. ## Intended use Use this preview for research and prototyping of audio-grounded, state-conditioned choices. Acted-speech scores may not transfer uniformly to spontaneous speech, accents, or recording conditions. Avoid using its emotion judgments as high-stakes assessments of people. Code and Aural One weight deltas are Apache 2.0. The separately downloaded Gemma 4 E2B base is also Apache 2.0. The [public release checklist](https://github.com/Parassharmaa/aural-one/blob/main/docs/RELEASE_CHECKLIST.md) records the final model-card, code, chart, hash, and GPU smoke checks. Audio in the introduction: [CC0 recording credits](assets/demo/ATTRIBUTION.md).