Text Classification
PEFT
Safetensors
Transformers
English
lora
audio-question-answering
correctness-assessment
orca
Instructions to use BUT-FIT/orca-llama-3.2-3b-it-multinomial with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use BUT-FIT/orca-llama-3.2-3b-it-multinomial with PEFT:
Task type is invalid.
- Transformers
How to use BUT-FIT/orca-llama-3.2-3b-it-multinomial with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="BUT-FIT/orca-llama-3.2-3b-it-multinomial")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("BUT-FIT/orca-llama-3.2-3b-it-multinomial", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Trim model card, point to GitHub for usage
Browse files
README.md
CHANGED
|
@@ -17,14 +17,10 @@ pipeline_tag: text-classification
|
|
| 17 |
|
| 18 |
# ORCA β Llama-3.2-3B-Instruct (Multinomial, seed 99)
|
| 19 |
|
| 20 |
-
ORCA (**O**pen-ended **R**esponse **C**orrectness **A**ssessment)
|
| 21 |
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
**Paper:** [ORCA: Open-ended Response Correctness Assessment for Audio Question Answering](https://arxiv.org/abs/2512.09066)
|
| 25 |
-
Accepted to *Transactions of the Association for Computational Linguistics (TACL), 2026*.
|
| 26 |
-
|
| 27 |
-
**Code:** [github.com/BUT-FIT/orca](https://github.com/BUT-FIT/orca)
|
| 28 |
**Training data:** [BUT-FIT/orca-audio-qa-annotations](https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations)
|
| 29 |
|
| 30 |
## Model details
|
|
@@ -33,84 +29,20 @@ Accepted to *Transactions of the Association for Computational Linguistics (TACL
|
|
| 33 |
|---|---|
|
| 34 |
| Base model | `meta-llama/Llama-3.2-3B-Instruct` |
|
| 35 |
| LoRA rank / alpha | 128 / 128 |
|
| 36 |
-
| LoRA target modules | q, k, v, o, gate, up, down projections |
|
| 37 |
| Loss function | Multinomial log-likelihood (5-class Likert) |
|
| 38 |
| Training seed | 99 |
|
| 39 |
-
| Training curriculum | Stage 1 (synthetic) β Stage 2 (LLM-judge) β Stage 3 (human
|
| 40 |
| Precision | bfloat16 |
|
| 41 |
|
| 42 |
-
##
|
| 43 |
-
|
| 44 |
-
```
|
| 45 |
-
model/
|
| 46 |
-
βββ config.yaml # ORCA config (score_type: multinomial)
|
| 47 |
-
βββ model_minus_lm.pt # Linear scoring head weights
|
| 48 |
-
βββ lm/
|
| 49 |
-
βββ adapter_config.json # LoRA adapter config
|
| 50 |
-
βββ adapter_model.safetensors
|
| 51 |
-
tokenizer/ # Saved tokenizer (same as base model)
|
| 52 |
-
```
|
| 53 |
-
|
| 54 |
-
## Usage
|
| 55 |
-
|
| 56 |
-
### Install
|
| 57 |
|
| 58 |
```bash
|
| 59 |
-
pip install git+https://github.com/
|
| 60 |
-
# or, after cloning:
|
| 61 |
-
pip install -e .
|
| 62 |
-
```
|
| 63 |
-
|
| 64 |
-
### Download this model
|
| 65 |
-
|
| 66 |
-
```bash
|
| 67 |
-
pip install huggingface_hub
|
| 68 |
hf download BUT-FIT/orca-llama-3.2-3b-it-multinomial --local-dir orca-llama-3b
|
|
|
|
| 69 |
```
|
| 70 |
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
Input is a JSONL file, one item per line:
|
| 74 |
-
|
| 75 |
-
```json
|
| 76 |
-
{"id": "example_1", "question": "How many instruments are played in this recording?", "reference": "Three instruments: piano, violin, and cello.", "candidate": "I hear a piano and violin being played.", "rationale": "The candidate identifies two of the three instruments but misses the cello.", "ratings": []}
|
| 77 |
-
```
|
| 78 |
-
|
| 79 |
-
Fields:
|
| 80 |
-
- `question`: the audio QA question
|
| 81 |
-
- `reference`: the ground-truth reference answer
|
| 82 |
-
- `candidate`: the model response to evaluate
|
| 83 |
-
- `rationale`: an LLM-generated explanation of why the candidate is correct or incorrect
|
| 84 |
-
- `ratings`: list of human ratings (integers 1β5); leave as `[]` for unlabeled data
|
| 85 |
-
|
| 86 |
-
### Run inference
|
| 87 |
-
|
| 88 |
-
```bash
|
| 89 |
-
orca-infer \
|
| 90 |
-
--model_path orca-llama-3b/model \
|
| 91 |
-
--data_jsonl your_data.jsonl \
|
| 92 |
-
--output_dir results/
|
| 93 |
-
```
|
| 94 |
-
|
| 95 |
-
The output `results/your_data/final_result.jsonl` contains one row per input with added fields:
|
| 96 |
-
- `rating_orca`: correctness score in [0, 1]
|
| 97 |
-
- `variance_orca`: uncertainty estimate
|
| 98 |
-
- `params`: raw distribution parameters
|
| 99 |
-
|
| 100 |
-
If your JSONL contains `ratings`, evaluation metrics (Spearman, Kendall, MAE) are printed and saved.
|
| 101 |
-
|
| 102 |
-
### Download evaluation splits (seed 99)
|
| 103 |
-
|
| 104 |
-
```bash
|
| 105 |
-
hf download BUT-FIT/orca-audio-qa-annotations \
|
| 106 |
-
--type dataset --local-dir ./data \
|
| 107 |
-
--include "stage3_human/seed_99/*"
|
| 108 |
-
|
| 109 |
-
orca-infer \
|
| 110 |
-
--model_path orca-llama-3b/model \
|
| 111 |
-
--data_jsonl data/stage3_human/seed_99/test.jsonl \
|
| 112 |
-
--output_dir results/
|
| 113 |
-
```
|
| 114 |
|
| 115 |
## Citation
|
| 116 |
|
|
@@ -129,4 +61,4 @@ orca-infer \
|
|
| 129 |
|
| 130 |
## License
|
| 131 |
|
| 132 |
-
MIT License. See the [repository LICENSE](https://github.com/
|
|
|
|
| 17 |
|
| 18 |
# ORCA β Llama-3.2-3B-Instruct (Multinomial, seed 99)
|
| 19 |
|
| 20 |
+
ORCA (**O**pen-ended **R**esponse **C**orrectness **A**ssessment) scores the correctness of open-ended audio QA responses. Given a question, reference answer, candidate answer, and an LLM-generated rationale, it outputs a correctness score in [0, 1] and an uncertainty estimate.
|
| 21 |
|
| 22 |
+
**Paper:** [ORCA: Open-ended Response Correctness Assessment for Audio Question Answering](https://arxiv.org/abs/2512.09066) β accepted to *TACL 2026*
|
| 23 |
+
**Code & usage:** [github.com/BUTSpeechFIT/ORCA](https://github.com/BUTSpeechFIT/ORCA)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
**Training data:** [BUT-FIT/orca-audio-qa-annotations](https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations)
|
| 25 |
|
| 26 |
## Model details
|
|
|
|
| 29 |
|---|---|
|
| 30 |
| Base model | `meta-llama/Llama-3.2-3B-Instruct` |
|
| 31 |
| LoRA rank / alpha | 128 / 128 |
|
|
|
|
| 32 |
| Loss function | Multinomial log-likelihood (5-class Likert) |
|
| 33 |
| Training seed | 99 |
|
| 34 |
+
| Training curriculum | Stage 1 (synthetic) β Stage 2 (LLM-judge) β Stage 3 (human) |
|
| 35 |
| Precision | bfloat16 |
|
| 36 |
|
| 37 |
+
## Quick start
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
```bash
|
| 40 |
+
pip install git+https://github.com/BUTSpeechFIT/ORCA.git
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
hf download BUT-FIT/orca-llama-3.2-3b-it-multinomial --local-dir orca-llama-3b
|
| 42 |
+
orca-infer --model_path orca-llama-3b/model --data_jsonl your_data.jsonl --output_dir results/
|
| 43 |
```
|
| 44 |
|
| 45 |
+
See the [repository](https://github.com/BUTSpeechFIT/ORCA) for full usage, evaluation scripts, and the `download_and_infer.py` convenience script.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
## Citation
|
| 48 |
|
|
|
|
| 61 |
|
| 62 |
## License
|
| 63 |
|
| 64 |
+
MIT License. See the [repository LICENSE](https://github.com/BUTSpeechFIT/ORCA/blob/main/LICENSE) for details.
|