Audio-Text-to-Text
Transformers
English
Bengali
audio
gemma
structured-decisions
speech-emotion-recognition
research-preview
Instructions to use blazeofchi/Aural-One-E2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use blazeofchi/Aural-One-E2B with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("blazeofchi/Aural-One-E2B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add public release validation and final model card
Browse files- README.md +1 -1
- VALIDATION.md +10 -0
README.md
CHANGED
|
@@ -30,7 +30,7 @@ This **v0.1.0 research preview** publishes a frozen adapter and **54.3 million u
|
|
| 30 |
| Structured choice / No-Null / ordinal development | **189 / 174 / 191** correct out of 200 each |
|
| 31 |
| Warm pod-local HTTP, ~58-second Opus with three questions | **0.508 s p50 / 0.579 s p95** |
|
| 32 |
|
| 33 |
-
These numbers have different test scopes. The full [evaluation report](EVALUATION.md) covers the frozen reference, listener-vote cross-entropy and calibration, same-words contrasts, sound events, Italian/German transfer, long audio, warm client latency, concurrency, and GPU-only cost. Aural One is an early release with room to improve rare emotions and natural sound transfer. The sub-second number above is measured **inside the warm GPU pod**; Tokyo end-to-end sub-second latency is still a research goal.
|
| 34 |
|
| 35 |
## Use the model
|
| 36 |
|
|
|
|
| 30 |
| Structured choice / No-Null / ordinal development | **189 / 174 / 191** correct out of 200 each |
|
| 31 |
| Warm pod-local HTTP, ~58-second Opus with three questions | **0.508 s p50 / 0.579 s p95** |
|
| 32 |
|
| 33 |
+
These numbers have different test scopes. The full [evaluation report](EVALUATION.md) covers the frozen reference, listener-vote cross-entropy and calibration, same-words contrasts, sound events, Italian/German transfer, long audio, warm client latency, concurrency, and GPU-only cost. [Release validation](VALIDATION.md) records the public-Hub load test. Aural One is an early release with room to improve rare emotions and natural sound transfer. The sub-second number above is measured **inside the warm GPU pod**; Tokyo end-to-end sub-second latency is still a research goal.
|
| 34 |
|
| 35 |
## Use the model
|
| 36 |
|
VALIDATION.md
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# v0.1.0 preview validation
|
| 2 |
+
|
| 3 |
+
The public release files were checked before publication and again through an anonymous download from the public [Hugging Face model repo](https://huggingface.co/blazeofchi/Aural-One-E2B).
|
| 4 |
+
|
| 5 |
+
- The published acoustic file has SHA-256 `ef80763236b2467a886d52fba51769de4dcfbdce909dd320803b6d2d2d41db96`, **47 tensors**, and **54,294,272 parameters**. The adapter SHA-256 is `e2b53154b40cd67faf3c9a57226f09c187b060569a7894ee2b3630a4e88937c3`. The base is pinned to `google/gemma-4-E2B-it@3e22461f65e89153144f8adb70e3b8c2cc9845a7` and checked by the loader.
|
| 6 |
+
- The release loader ran on an **NVIDIA RTX PRO 6000 Blackwell Server Edition MIG 1g.24gb** with **PyTorch 2.13.0+cu130**, Transformers 5.17.0, and PEFT 0.21.0. It returned valid choice distributions for three named questions on a short CREMA-D speech sample and a synthetic 58-second two-chunk input. No sample audio is bundled.
|
| 7 |
+
- The public-Hub load gave **exactly the same three answer distributions** as the same release files loaded from a local staging directory. PyTorch reported **9,773.9 MiB allocated**. The simple sequential loader took 1.778 seconds to score three short-input questions on the public-Hub repeat; the synthetic 58-second local-file smoke took 3.669 seconds for three questions. These are functional checks under different cache conditions, not warm-service latency measurements.
|
| 8 |
+
- The GPU pod was deleted after testing. The optimized shared-audio HTTP timings in [EVALUATION.md](EVALUATION.md) come from a separate staged serving path and should not be attributed to this reference loader.
|
| 9 |
+
|
| 10 |
+
The GitHub repository contains only code, docs, examples, configuration, and hashes. Hugging Face contains the adapter and acoustic delta plus the same documentation. Neither contains training/evaluation audio, the base model weights, optimizer state, credentials, or private row-level predictions.
|