--- license: other base_model: ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16 library_name: gguf pipeline_tag: text-generation model-index: - name: MiniCPM5-1B Agentic Tooluse Nemotron DPO GGUF results: - task: type: text-generation name: Tool calling dataset: name: External ToolACE-derived first-call evaluation of source model type: Team-ACE/ToolACE metrics: - type: parseable_rate value: 0.9933333333333333 name: Source model parseable tool call rate - type: expected_name_rate value: 0.9266666666666666 name: Source model expected tool name rate tags: - gguf - llama.cpp - llama-cpp - minicpm - minicpm5 - minicpm5-1b - tool-calling - function-calling - function-call - tool-use - agentic - xml-tool-calling - q4_k_m - q8_0 - f16 - quantized - vllm - unsloth - openbmb - nemotron - dpo --- # MiniCPM5-1B Agentic Tooluse Nemotron DPO GGUF ### 25 Likes and I'll finetune this further (Work In Progress) Current GGUF exports of the July 2026 Nemotron SFT+DPO repaired [`MiniCPM5-1B-Agentic-Tooluse-Merged-FP16`](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16). These files were generated from the latest evaluated merged checkpoint. The previous warning that this repository had not been regenerated is no longer applicable. ## Model Family | Use case | Repository | |---|---| | Current F16, Q8_0, and Q4_K_M GGUF files | This repository | | Standalone merged Hugging Face model | [`MiniCPM5-1B-Agentic-Tooluse-Merged-FP16`](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16) | | PEFT/LoRA adapter | [`MiniCPM5-1B-Agentic-Tooluse-QLoRA-v2`](https://huggingface.co/ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-QLoRA-v2) | ## Current Files | File | Size | Recommended use | |---|---:|---| | `MiniCPM5-1B-Agentic-Tooluse-Nemotron-DPO.F16.gguf` | 2.17 GB | Highest-fidelity GGUF and reference conversion | | `MiniCPM5-1B-Agentic-Tooluse-Nemotron-DPO.Q8_0.gguf` | 1.15 GB | High fidelity with lower memory use | | `MiniCPM5-1B-Agentic-Tooluse-Nemotron-DPO.Q4_K_M.gguf` | 688 MB | Recommended local size/speed balance | | `CONVERSION_METADATA.json` | Small | Exact source and converter revisions | The pre-Nemotron exports are retained under [`legacy/`](./tree/main/legacy) for reproducibility. Use the root-level `Nemotron-DPO` files for the current model. ## Provenance - Base model: `openbmb/MiniCPM5-1B` - Current adapter: `ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-QLoRA-v2` - Conversion source: `ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16` - Source revision: `67e1cad9b15ffc21d0877bf952d9ece2eeda5d35` - llama.cpp revision: `683f0c72e5b3c07fab90bfd9ec2ce8661d624228` - Quantizations: F16, Q8_0, Q4_K_M The export pipeline checked each output's GGUF magic, format version, size, and successful completion by `llama-quantize`. Full task evaluation was performed on the source merged model, not independently on every quantization. ## Tool-Call Format The model emits XML-style calls: ```xml value ``` A deployment runtime should: 1. Provide the available tool definitions in the training-compatible prompt format. 2. Decode deterministically for tool selection. 3. Stop after the first complete ``. 4. Validate the function name and arguments against the supplied schema. 5. Execute the tool outside the model and provide the result in a new turn. ## llama.cpp Use a recent llama.cpp build: ```bash llama-cli \ -m MiniCPM5-1B-Agentic-Tooluse-Nemotron-DPO.Q4_K_M.gguf \ -p 'Fix the calculator bug. {"name":"run_tests","description":"Run the test suite."} ' \ -n 96 \ --temp 0 ``` The exact chat template and EOS handling matter. If a frontend supplies its own template, verify that it matches the source model's `chat_template.jinja`. ## vLLM As of July 2026, vLLM documents GGUF support as highly experimental and under-optimized. GGUF support also requires the out-of-tree plugin: ```bash uv pip install vllm vllm-gguf-plugin ``` Serve Q4_K_M directly from the Hub: ```bash vllm serve \ ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-GGUF:Q4_K_M \ --tokenizer ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16 \ --hf-config-path ewinregirgojr/MiniCPM5-1B-Agentic-Tooluse-Merged-FP16 ``` The separate tokenizer/config arguments follow vLLM's recommendation to avoid unstable GGUF tokenizer conversion. The merged repository is used because it contains the matching tokenizer, configuration, and chat template. For stable high-throughput vLLM deployment, prefer the merged safetensors repository when memory permits. ## Source-Model Evaluation External evaluation used 300 examples derived from the untouched [`Team-ACE/ToolACE`](https://huggingface.co/datasets/Team-ACE/ToolACE) dataset. The same cases and deterministic first-call evaluator were used for the base and fine-tuned source model. This is not an official ToolACE leaderboard submission. | Metric | Base MiniCPM5-1B | Nemotron repaired source | Delta | |---|---:|---:|---:| | Parseable tool call | 0.0133 | 0.9933 | +0.9800 | | Valid available-tool name | 0.0133 | 0.9700 | +0.9567 | | Expected tool name | 0.0133 | 0.9267 | +0.9133 | | Exact arguments | 0.1500 | 0.6533 | +0.5033 | | Argument-key overlap | 0.0033 | 0.7517 | +0.7484 | | No schema copying | 1.0000 | 1.0000 | +0.0000 | | No repetition | 0.9967 | 1.0000 | +0.0033 | | Clean stop | 0.0000 | 0.1500 | +0.1500 | These scores establish the behavior of the merged source checkpoint. Quantization can alter individual outputs, so evaluate the exact GGUF and runtime configuration used in production. ## Source-Model Improvements and GGUF Scope These GGUF files were converted from the repaired merged checkpoint, whose reported results improved substantially over base MiniCPM5-1B: - Parseable calls: **1.33% -> 99.33%** - Valid available-tool names: **1.33% -> 97.00%** - Expected-tool selection: **1.33% -> 92.67%** - Exact arguments: **15.00% -> 65.33%** - Argument-key overlap: **0.33% -> 75.17%** - No repetition: **99.67% -> 100.00%** - Natural clean termination: **0.00% -> 15.00%** These are source FP16 metrics. The remaining gap to 100% is residual error after improvement, not evidence that the fine-tune made the base model worse. ## Deployment Notes - Q4_K_M and Q8_0 have not yet received an independent 300-case ToolACE run, so their exact scores should not be assumed identical to FP16. - Schema validation and permission checks are universal tool-runtime safety requirements, not model-specific regressions. - Valid-name and exact-argument behavior improved strongly in the source model; rare errors can still occur on unseen tools. - MiniCPM5's parser-based serving path extracts the first completed XML call instead of relying solely on natural EOS. - vLLM GGUF support is experimental; merged safetensors remains the more reliable vLLM format. - llama.cpp behavior depends on build version and the frontend's chat-template handling. ## Version History The current `Nemotron-DPO` files are the July 2026 repaired exports. Legacy pre-Nemotron files are stored under `legacy/` for reproducibility, and earlier repository states remain available from Hugging Face commit history. ## MiniCPM5 Tool-Call Runtime Contract MiniCPM5-1B has native XML-style tool calling. OpenBMB recommends SGLang for tool calling with its built-in `minicpm5` parser: ```bash python -m sglang.launch_server \ --model-path MODEL_PATH \ --tool-call-parser minicpm5 ``` The parser converts a completed `...` block into an OpenAI-compatible `tool_calls` response. A production agent should treat the first completed function block as the action boundary, validate it, execute the tool externally, and send the tool result in a new turn. The reported `stopped_cleanly_rate=0.15` is a strict **natural-termination** metric: only 15% of evaluated generations emitted a completed call and then naturally terminated at that boundary without runtime intervention. It does **not** mean that only 15% produced usable tool calls. On the same 300 cases, the repaired source model achieved: - `parseable_rate=0.9933` - `valid_name_rate=0.9700` - `expected_name_rate=0.9267` - `no_repetition_rate=1.0000` Therefore, natural EOS-at-boundary remains weak, while first-call extraction and selection are much stronger. Use the official parser where supported and configure bounded generation or parser-side stopping after the first complete ``.