granite-4.0-h-tiny-APEX-GGUF / TOOL_CALLING.md
Myric's picture
Upload TOOL_CALLING.md with huggingface_hub
5400ba3 verified
|
Raw
History Blame Contribute Delete
3.07 kB

Tool calling with these APEX quants (reproducible)

Captured against the handroll quant. Turn 2 depends on turn 1's result (a chained call, not a one-off), and a distractor tool (search_web) is offered but correctly never called.

Server

llama-server --jinja --chat-template-file ibm-granite-granite-4.0.jinja --ctx-size 8192 -fa on --temp 0.6 --top-p 0.9 --repeat-penalty 1.05 --model <this-quant>.gguf --host 127.0.0.1 --port 8092

System prompt

You are a helpful assistant with access to tools. When a tool is needed, call it; when it returns a result, answer the user directly using that result.

Tools offered

[
  {
    "type": "function",
    "function": {
      "name": "get_user_location",
      "description": "Get the user's current city.",
      "parameters": {
        "type": "object",
        "properties": {}
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a location.",
      "parameters": {
        "type": "object",
        "properties": {
          "location": {
            "type": "string",
            "description": "City name"
          },
          "unit": {
            "type": "string",
            "enum": [
              "celsius",
              "fahrenheit"
            ]
          }
        },
        "required": [
          "location"
        ]
      }
    }
  },
  {
    "type": "function",
    "function": {
      "name": "search_web",
      "description": "Search the web for a query.",
      "parameters": {
        "type": "object",
        "properties": {
          "query": {
            "type": "string"
          }
        },
        "required": [
          "query"
        ]
      }
    }
  }
]

Conversation

User: What is the weather at my current location? Use celsius.

Turn 1: LLM sent properly formatted JSON to the test server

raw response JSON
{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "type": "function",
      "function": {
        "name": "get_user_location",
        "arguments": "{}"
      },
      "id": "w2dNbZAEilYuMc5S0bET0L4ftPBdf8zf"
    }
  ]
}

Tool result (get_user_location): {"city": "Tokyo"}

Turn 2: LLM sent properly formatted JSON to the test server

raw response JSON
{
  "role": "assistant",
  "content": "",
  "tool_calls": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "arguments": "{\"location\": \"Tokyo\", \"unit\": \"celsius\"}"
      },
      "id": "RZ72zgaiFvGbejS8etV8sqjJSfpZQjDz"
    }
  ]
}

Tool result (get_weather): {"temp_c": 18, "conditions": "cloudy"}

Turn 3: LLM correctly gave a plain-language final answer, no further tool call

raw response JSON
{
  "role": "assistant",
  "content": "The current weather in Tokyo is **18\u00b0C** with cloudy conditions."
}