---
language:
- zh
- en
license: apache-2.0
pipeline_tag: translation
base_model:
- netease-youdao/Confucius4-T3PO
tags:
- gguf
- llama.cpp
- simultaneous-translation
- streaming-translation
- zh-en
- en-zh
---
Confucius4-T3PO: simulTaneous Translation via pareTo Policy Optimization
[](./README-zh.md) [](https://huggingface.co/netease-youdao/Confucius4-T3PO) [](https://github.com/netease-youdao/Confucius4-T3PO)
# Confucius4-T3PO-GGUF
GGUF conversions of [`netease-youdao/Confucius4-T3PO`](https://huggingface.co/netease-youdao/Confucius4-T3PO),
a Chinese–English bidirectional streaming simultaneous translation model.
Refer to the [original model card](https://huggingface.co/netease-youdao/Confucius4-T3PO)
for the streaming protocol, the latency operating points, and evaluation results.
## Files
| File | Output type | Size |
| --- | :---: | ---: |
| `Confucius4-T3PO-F16.gguf` | `F16` | 29.5 GB |
| `Confucius4-T3PO-Q6_K.gguf` | `Q6_K` | 12.1 GB |
| `Confucius4-T3PO-Q5_K_M.gguf` | `Q5_K_M` | 10.5 GB |
The low-bit variants are quantized from the F16 GGUF. Start with `Q6_K` for a
close match to F16 quality at well under half the size; `Q5_K_M` trades a little
more quality for the smallest footprint. No GGUF splitting has been applied, so
the files run as-is. `SHA256SUMS` and `CONVERSION_INFO.md` record the checksums
and the exact conversion commands.
## Use with llama.cpp
Compile and install [llama.cpp](https://github.com/ggml-org/llama.cpp) first.
Single-shot generation:
```bash
llama-cli -m Confucius4-T3PO-Q6_K.gguf -no-cnv --temp 0 -n 128 -p "..."
```
OpenAI-compatible server:
```bash
llama-server -m Confucius4-T3PO-Q6_K.gguf --host 127.0.0.1 --port 8010
```
```bash
curl http://127.0.0.1:8010/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"temperature": 0,
"max_tokens": 128,
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": ""}
]
}'
```
The `user` message must follow the streaming protocol from the original model
card: the task prompt followed by the `` and
`` blocks. An empty response means `WAIT`; a non-empty one is the
next translation segment.