---
language:
- zh
- en
license: apache-2.0
pipeline_tag: text-generation
base_model:
- netease-youdao/Confucius4-T3PO
tags:
- gguf
- llama.cpp
- simultaneous-translation
- streaming-translation
- zh-en
- en-zh
---
Confucius4-T3PO:基于帕累托前沿策略优化的流式翻译
[](./README.md) [](https://huggingface.co/netease-youdao/Confucius4-T3PO) [](https://github.com/netease-youdao/Confucius4-T3PO)
# Confucius4-T3PO-GGUF
[`netease-youdao/Confucius4-T3PO`](https://huggingface.co/netease-youdao/Confucius4-T3PO)
的 GGUF 转换版本,该模型是一个中英双向流式同传翻译模型。流式协议、延迟档位与评测结果请参见[原始模型卡](https://huggingface.co/netease-youdao/Confucius4-T3PO)。
## 文件
| 文件 | 精度 | 大小 |
| --- | :---: | ---: |
| `Confucius4-T3PO-F16.gguf` | `F16` | 29.5 GB |
| `Confucius4-T3PO-Q6_K.gguf` | `Q6_K` | 12.1 GB |
| `Confucius4-T3PO-Q5_K_M.gguf` | `Q5_K_M` | 10.5 GB |
低比特版本由 F16 GGUF 量化得到。`Q6_K` 以不到一半的权重接近 F16 的质量,推荐优先尝试;
`Q5_K_M` 体积最小,质量略有损失。模型未做 GGUF 分片,下载后可直接运行。`SHA256SUMS`
与 `CONVERSION_INFO.md` 记录了校验和与完整的转换命令。
## 使用 llama.cpp
请先编译安装 [llama.cpp](https://github.com/ggml-org/llama.cpp)。
单次生成:
```bash
llama-cli -m Confucius4-T3PO-Q6_K.gguf -no-cnv --temp 0 -n 128 -p "..."
```
启动 OpenAI 兼容服务:
```bash
llama-server -m Confucius4-T3PO-Q6_K.gguf --host 127.0.0.1 --port 8010
```
```bash
curl http://127.0.0.1:8010/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"temperature": 0,
"max_tokens": 128,
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "<任务 prompt 与两个协议区块>"}
]
}'
```
`user` 消息需遵循原始模型卡中的流式协议:任务 prompt 之后依次是
`` 与 `` 两个区块。空响应表示 `WAIT`,非空响应即为
下一段增量译文。