sabeshbesh commited on
Commit
1c87874
·
verified ·
1 Parent(s): 68d3c18

Add LFM2.5-230M tool-caller as Core AI ANE static .aimodel (int4 palettized, 512 ctx, coreai-torch 0.4.1)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/main.mlirb filter=lfs diff=lfs merge=lfs -text
LICENSE ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ LFM Open License v1.0
2
+
3
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
4
+
5
+ 1. Definitions.
6
+
7
+ "License" shall mean the terms and conditions for use, reproduction, and distribution as defined by this document.
8
+
9
+ "Licensor" shall mean Liquid AI, Inc.
10
+
11
+ "Legal Entity" shall mean the union of the acting entity and all other entities that control, are controlled by, or are under common control with that entity. For the purposes of this definition, "control" means (i) the power, direct or indirect, to cause the direction or management of such entity, whether by contract or otherwise, or (ii) ownership of fifty percent (50%) or more of the outstanding shares, or (iii) beneficial ownership of such entity.
12
+
13
+ "You" (or "Your") shall mean an individual or Legal Entity exercising permissions granted by this License.
14
+
15
+ "Source" form shall mean the preferred form for making modifications, including but not limited to software source code, documentation source, and configuration files.
16
+
17
+ "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types.
18
+
19
+ "Work" shall mean the work of authorship, whether in Source or Object form, made available under the License, as indicated by a copyright notice that is included in or attached to the work.
20
+
21
+ "Derivative Works" shall mean any work, whether in Source or Object form, that is based on (or derived from) the Work and for which the editorial revisions, annotations, elaborations, or other modifications represent, as a whole, an original work of authorship. For the purposes of this License, Derivative Works shall not include works that remain separable from, or merely link (or bind by name) to the interfaces of, the Work and Derivative Works thereof.
22
+
23
+ "Contribution" shall mean any work of authorship, including the original version of the Work and any modifications or additions to that Work or Derivative Works thereof, that is intentionally submitted to Licensor for inclusion in the Work by the copyright owner or by an individual or Legal Entity authorized to submit on behalf of the copyright owner. For the purposes of this definition, "submitted" means any form of electronic, verbal, or written communication sent to the Licensor or its representatives, including but not limited to communication on electronic mailing lists, source code control systems, and issue tracking systems that are managed by, or on behalf of, the Licensor for the purpose of discussing and improving the Work, but excluding communication that is conspicuously marked or otherwise designated in writing by the copyright owner as "Not a Contribution."
24
+
25
+ "Contributor" shall mean Licensor and any individual or Legal Entity on behalf of whom a Contribution has been received by Licensor and subsequently incorporated within the Work.
26
+
27
+ "Commercial Use" shall mean any use of the Work for direct or indirect commercial advantage or monetary compensation.
28
+
29
+ "Qualified Non-Profit Organization" shall mean a Legal Entity that is organized and operated exclusively for religious, charitable, scientific, testing for public safety, literary, or educational purposes, and which is exempt from federal income tax under Section 501(c)(3) of the United States Internal Revenue Code of 1986, as amended, or any equivalent non-profit or charitable organization in a foreign jurisdiction.
30
+
31
+ "Non-Commercial or Research Purposes" shall mean purposes that do not involve any use of the Work or a Derivative Work for Commercial Use.
32
+
33
+ "Threshold" shall mean annual revenue of 10 million United States dollars ($10,000,000) or more.
34
+
35
+ 2. Grant of Copyright License. Subject to the terms and conditions of this License, including the Commercial Use limitation set forth in Section 5, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable copyright license to reproduce, prepare Derivative Works of, publicly display, publicly perform, sublicense, and distribute the Work and such Derivative Works in Source or Object form.
36
+
37
+ 3. Grant of Patent License. Subject to the terms and conditions of this License, including the Commercial Use limitation set forth in Section 5, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable (except as stated in this section) patent license to make, have made, use, offer to sell, sell, import, and otherwise transfer the Work, where such license applies only to those patent claims licensable by such Contributor that are necessarily infringed by their Contribution(s) alone or by combination of their Contribution(s) with the Work to which such Contribution(s) was submitted. If You institute patent litigation against any entity (including a cross-claim or counterclaim in a lawsuit) alleging that the Work or a Contribution incorporated within the Work constitutes direct or contributory patent infringement, then any patent licenses granted to You under this License for that Work shall terminate as of the date such litigation is filed.
38
+
39
+ 4. Redistribution. You may reproduce and distribute copies of the Work or Derivative Works thereof in any medium, with or without modifications, and in Source or Object form, provided that You meet the following conditions:
40
+
41
+ (a) You must give any other recipients of the Work or Derivative Works a copy of this License; and
42
+
43
+ (b) You must cause any modified files to carry prominent notices stating that You changed the files; and
44
+
45
+ (c) You must retain, in the Source form of any Derivative Works that You distribute, all copyright, patent, trademark, and attribution notices from the Source form of the Work, excluding those notices that do not pertain to any part of the Derivative Works; and
46
+
47
+ (d) If the Work includes a "NOTICE" text file as part of its distribution, then any Derivative Works that You distribute must include a readable copy of the attribution notices contained within such NOTICE file, excluding those notices that do not pertain to any part of the Derivative Works, in at least one of the following places: within a NOTICE text file distributed as part of the Derivative Works; within the Source form or documentation, if provided along with the Derivative Works; or, within a display generated by the Derivative Works, if and wherever such third-party notices normally appear. The contents of the NOTICE file are for informational purposes only and do not modify the License. You may add Your own attribution notices within Derivative Works that You distribute, alongside or as an addendum to the NOTICE text from the Work, provided that such additional attribution notices cannot be construed as modifying the License.
48
+
49
+ You may add Your own copyright statement to Your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of Your modifications, or for any such Derivative Works as a whole, provided Your use, reproduction, and distribution of the Work otherwise complies with the conditions stated in this License.
50
+
51
+ 5. Commercial Use Limitation.
52
+
53
+ (a) The rights granted under this License for Commercial Use are conditioned upon You or Your Legal Entity not exceeding the Threshold.
54
+
55
+ (b) Any Commercial Use of the Work or a Derivative Work by a Legal Entity that exceeds the Threshold is not licensed under this Agreement.
56
+
57
+ (c) The Threshold shall not apply to a Qualified Non-Profit Organization's use of the Work or a Derivative Work for Non-Commercial or Research Purposes.
58
+
59
+ 6. Submission of Contributions. Unless You explicitly state otherwise, any Contribution intentionally submitted for inclusion in the Work by You to the Licensor shall be under the terms and conditions of this License, without any additional terms or conditions. Notwithstanding the above, nothing herein shall supersede or modify the terms of any separate license agreement you may have executed with Licensor regarding such Contributions.
60
+
61
+ 7. Trademarks. This License does not grant permission to use the trade names, trademarks, service marks, or product names of the Licensor, except for the reasonable and customary use in describing the origin of the Work and reproducing the content of the NOTICE file.
62
+
63
+ 8. Disclaimer of Warranty. Unless required by applicable law or agreed to in writing, Licensor provides the Work (and each Contributor provides its Contributions) on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied, including, without limitation, any warranties or conditions of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A PARTICULAR PURPOSE. You are solely responsible for determining the appropriateness of using or redistributing the Work and assume any risks associated with Your exercise of permissions under this License.
64
+
65
+ 9. Limitation of Liability. In no event and under no legal theory, whether in tort (including negligence), contract, or otherwise, unless required by applicable law (such as deliberate and grossly negligent acts) or agreed to in writing, shall any Contributor be liable to You for damages, including any direct, indirect, special, incidental, or consequential damages of any character arising as a result of this License or out of the use or inability to use the Work (including but not limited to damages for loss of goodwill, work stoppage, computer failure or malfunction, or any and all other commercial damages or losses), even if such Contributor has been advised of the possibility of such damages.
66
+
67
+ 10. Accepting Warranty or Additional Liability. While redistributing the Work or Derivative Works thereof, You may choose to offer, and charge a fee for, acceptance of support, warranty, indemnity, or other liability obligations and/or rights consistent with this License. However, in accepting such obligations, You may act only on Your own behalf and on Your sole responsibility, not on behalf of any other Contributor, and only if You agree to indemnify, defend, and hold each Contributor harmless for any liability incurred by, or claims asserted against, such Contributor by reason of your accepting any such warranty or additional liability.
68
+
69
+ 11. Termination. This License will terminate automatically and immediately if You fail to comply with any of its terms and conditions. Upon termination, You must cease all use of the Work and any Derivative Works and delete all copies in Your possession.
70
+
71
+ END OF TERMS AND CONDITIONS
README.md ADDED
@@ -0,0 +1,180 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: lfm1.0
4
+ license_link: LICENSE
5
+ base_model: LiquidAI/LFM2.5-230M
6
+ tags:
7
+ - coreai
8
+ - aimodel
9
+ - apple-silicon
10
+ - on-device
11
+ - neural-engine
12
+ - lfm2
13
+ - hybrid
14
+ - tool-calling
15
+ - function-calling
16
+ pipeline_tag: text-generation
17
+ ---
18
+
19
+ # Uncle Rudy — LFM2.5-230M tool-caller — Apple Core AI (`.aimodel`)
20
+
21
+ A **LoRA fine-tune of [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M)** that turns
22
+ a single spoken/typed utterance into a **JSON tool call** for a to-do app — converted to Apple's
23
+ **Core AI** `.aimodel` format and running on the **Neural Engine** (iOS 27 / macOS 27).
24
+
25
+ > [!IMPORTANT]
26
+ > This is a **static-shape / Neural Engine** bundle (4 entrypoints: `load_embeddings`,
27
+ > `gather_embeddings`, `extend_*`, `prompt_opt_*`) → Apple's `EngineFactory` selects the
28
+ > **`static-shape`** engine. It is **not** a `gpu-pipelined` decode bundle, so it will **not**
29
+ > load on the `coreai-pipelined` GPU path that most published Core AI chat bundles use.
30
+ > It also needs a small runtime patch — see [Runtime requirements](#runtime-requirements).
31
+
32
+ ## What it does
33
+
34
+ Give it one utterance; it emits one tool call.
35
+
36
+ ```
37
+ "Remind me to buy milk" → {"name":"create_todo","arguments":{"title":"Buy milk","due":null}}
38
+ "Remind me to call the dentist tomorrow" → {"name":"create_todo","arguments":{"title":"Call the dentist","due":"tomorrow"}}
39
+ "Delete the dentist task" → {"name":"delete_todo","arguments":{"target":"Call the dentist"}}
40
+ ```
41
+
42
+ Tool surface: `create_todo` (`title`, `due`) · `set_status` (`target`, `status`) ·
43
+ `update_todo` (`target`, `title`, `due`) · `delete_todo` (`target`).
44
+
45
+ Apply the bundle's own chat template and feed the raw utterance as the user turn — the fine-tune
46
+ emits the assistant tool call directly. **No system prompt / tool-schema preamble is needed**
47
+ (training used `mask_prompt`, so loss was computed only on the tool-call tokens). Decode **greedily**.
48
+
49
+ ## Bundle
50
+
51
+ ```
52
+ ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/
53
+ ├── uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/ # main.mlirb · main.hash · metadata.json
54
+ ├── tokenizer/ # tokenizer.json · tokenizer_config.json
55
+ │ # chat_template.jinja · generation_config.json
56
+ └── metadata.json # LanguageBundle manifest
57
+ ```
58
+
59
+ | | |
60
+ |---|---|
61
+ | Base | LiquidAI/LFM2.5-230M (14 layers = 8 short-conv + 6 GQA attention, hidden 1024, 16/8 heads, vocab 65536, tied embeddings) |
62
+ | Fine-tune | LoRA r=16, scale 16, 400 iters, `mask_prompt: true` (loss on tool call only), fused |
63
+ | Compression | `4bit_weight_palettized_group32` (embedding table int8, per Core AI's iOS path) |
64
+ | Size | **144 MB** `.aimodel` (~148 MB with tokenizer) |
65
+ | Max context | **512** (prompt + generation) — deliberately small; see [Context length](#context-length) |
66
+ | Engine | `static-shape` (Neural Engine) |
67
+
68
+ ## Runtime requirements
69
+
70
+ 1. **iOS 27 / macOS 27** — Core AI ships with the OS. **Device-only on iOS**: `CoreAI.framework` is
71
+ in the iPhoneOS SDK but **not** the iOS Simulator SDK, so this cannot run in the Simulator.
72
+ 2. **Toolchain ≥ beta 3 era.** This bundle was exported with **`coreai-torch 0.4.1` / `coreai-core 1.0.0b2`**
73
+ so it carries the *versioned-IR* location format the beta-3 on-device compiler requires. (Bundles
74
+ exported with the June-era `coreai-torch 0.4.0` / `coreai-core 1.0.0b1` fail on beta 3 with
75
+ `expected AICode versioned location … Failed to convert to versioned IR … cannot unwrap empty odiec_module_t`.)
76
+ 3. **A conv-cache extra-state patch on the static-shape engine.** LFM2 is a conv+attention hybrid: its
77
+ short-conv layers carry a rolling **conv cache** *in addition to* the KV cache, so the model exports
78
+ **three** states (`key_cache`, `value_cache`, `conv_cache`). Apple's `StaticShapeEngine` hardcodes two.
79
+ The engine must allocate the extra state, **zero-fill it on reset** (unlike KV it is read directly, not
80
+ mask-gated), and bind it by name each step.
81
+
82
+ > Note: the conv state cannot be partially prefix-rewound (it only holds the last `L-1` columns), so a
83
+ > conv model should reprocess each request from scratch rather than reuse a partial KV prefix.
84
+
85
+ ## Use it
86
+
87
+ The bundle is a standard `LanguageBundle` (`.aimodel` + `tokenizer/` + `metadata.json`) — point the
88
+ runtime at the **bundle directory**:
89
+
90
+ ```swift
91
+ import CoreAILanguageModels
92
+
93
+ let bundle = try LanguageBundle(at: bundleDir) // dir containing the .aimodel
94
+ let engine = try await CoreAIRunner(from: bundle).makeInferenceEngine() // → static-shape / ANE
95
+ let tokenizer = try await bundle.loadTokenizer()
96
+
97
+ let generator = try await TextGeneratorBuilder()
98
+ .withInferenceEngine(engine)
99
+ .withTokenizer(tokenizer)
100
+ .withSampling(configuration: .greedy)
101
+ .build()
102
+
103
+ // `.prompt` applies the bundle's chat template; the fine-tune emits the tool call.
104
+ let json = try await generator.generate(input: .prompt("Remind me to buy milk"), maxTokens: 60)
105
+ ```
106
+
107
+ Catalog entry, for apps that pull Core AI bundles from HF by tree path:
108
+
109
+ ```swift
110
+ ModelSpec(
111
+ bundleName: "uncle_rudy_lfm2_230m_toolcaller_int4pal",
112
+ hfRemotePath: "ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal",
113
+ repoURL: "https://huggingface.co/sabeshbesh/uncle-rudy-lfm2-230m-CoreAI",
114
+ label: "Uncle Rudy 230M",
115
+ approxSizeGB: 0.15,
116
+ warmupToken: 1,
117
+ maxContext: 512)
118
+ ```
119
+
120
+ ⚠️ A catalog/`ModelSpec` alone is not sufficient: an app wired to the **`coreai-pipelined`** engine
121
+ must route this bundle to the **static-shape** engine and carry the conv-cache patch above.
122
+
123
+ ## First load is slow — by design
124
+
125
+ The **first** load on a given device compiles the model's Neural Engine graphs **on-device**, inside
126
+ Apple's `AIModel(contentsOf:options:)`. This cannot be shipped precompiled — the compiled program is
127
+ specific to that device + OS. The OS caches the result, so every later load is near-instant.
128
+
129
+ | | |
130
+ |---|---|
131
+ | Cold (first) load, Apple Silicon ANE | **~54 s** |
132
+ | Warm load (cached) | **~0.02–0.1 s** |
133
+
134
+ ### Context length
135
+
136
+ Exported at **512** context on purpose. The iOS static export fans out one specialized ANE graph per
137
+ `(context_bucket × query_length)` per function — at 2048 that is **24** graphs and a **~151 s** cold
138
+ compile; at 512 it is **12** graphs and **~54 s**. A tool-caller only ever sees a short utterance plus
139
+ a ≤60-token call, so 512 is ample. Re-export at a larger `--max-context-length` if you need more, and
140
+ pay the longer one-time compile.
141
+
142
+ ## Measured
143
+
144
+ Greedy, `llm-runner`, **macOS / Apple Silicon Neural Engine** (this bundle):
145
+
146
+ | | |
147
+ |---|---|
148
+ | Prefill | ~80–330 tok/s |
149
+ | Decode | ~73–83 tok/s |
150
+
151
+ > iPhone on-device throughput is **not yet published** — these are Mac-ANE numbers for the same bundle.
152
+ > Treat them as indicative, not as iPhone figures.
153
+
154
+ **Correctness.** The re-authored BC1S / Neural-Engine model was gated against the fp32 Hugging Face
155
+ reference: **100% next-token top-1 match (12/12 positions)**, logits PSNR ~52 dB, on a truncated model
156
+ covering both layer types (conv + attention). The lower-than-macOS PSNR (~70 dB on the GPU/dynamic
157
+ path) is expected: the iOS path int8-quantizes the embedding table and computes attention per-head,
158
+ which reassociates fp16 arithmetic — every argmax still matches.
159
+
160
+ ## Conversion notes
161
+
162
+ Converted from the fused Hugging Face checkpoint via a custom Core AI recipe on top of
163
+ [apple/coreai-models](https://github.com/apple/coreai-models). Two things were needed beyond the
164
+ stock pipeline:
165
+
166
+ - **MLX → PyTorch conv-weight transpose.** `mlx_lm fuse` writes depthwise Conv1d weights in MLX axis
167
+ order `(out, kernel, in)` = `[1024, 3, 1]`; PyTorch's `modeling_lfm2` wants `(out, in, kernel)` =
168
+ `[1024, 1, 3]`. Only the 8 conv layers are affected (Linear/embedding/norm layouts are identical),
169
+ and the fix is an axis swap — verified bit-exact against the base weights.
170
+ - **Re-authoring for the Neural Engine.** BC1S `(B, C, 1, S)` layout, projections as 1×1 `Conv2d`,
171
+ per-head attention (no fused SDPA on ANE), transposed causal mask using `-40000` rather than `-inf`,
172
+ and the short conv as a depthwise `Conv2d(D, D, (1, L), groups=D)` over the sequence axis, with the
173
+ conv cache threaded as a third functional state.
174
+
175
+ ## License
176
+
177
+ Weights derive from [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) and are
178
+ redistributed under the **LFM Open License v1.0** ([LICENSE](LICENSE)) — Apache-style grants, but
179
+ **commercial use is licensed only for entities under US$10M annual revenue** (qualified non-profits
180
+ exempt for non-commercial/research use). Review the LICENSE before any commercial deployment.
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/metadata.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata_version": "0.2",
3
+ "kind": "llm",
4
+ "name": "uncle_rudy_lfm2_230m_toolcaller_int4pal",
5
+ "assets": {
6
+ "main": "uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel"
7
+ },
8
+ "language": {
9
+ "tokenizer": "LiquidAI/LFM2.5-230M",
10
+ "vocab_size": 65536,
11
+ "max_context_length": 512,
12
+ "embedded_tokenizer": true,
13
+ "function_map": {
14
+ "main": [
15
+ "main"
16
+ ]
17
+ }
18
+ },
19
+ "source": {
20
+ "model_definition": "torch",
21
+ "hf_model_id": "LiquidAI/LFM2.5-230M"
22
+ },
23
+ "compression": "4bit_weight_palettized_group32",
24
+ "compilation": {
25
+ "date": "2026-07-13T19:16:31.555725+05:30",
26
+ "targets": []
27
+ }
28
+ }
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/tokenizer/chat_template.jinja ADDED
@@ -0,0 +1,115 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {{- bos_token -}}
2
+ {%- set preserve_thinking = preserve_thinking | default(false) -%}
3
+
4
+ {%- macro format_arg_value(arg_value) -%}
5
+ {%- if arg_value is string -%}
6
+ {{- "'" + arg_value + "'" -}}
7
+ {%- elif arg_value is mapping -%}
8
+ {{- arg_value | tojson -}}
9
+ {%- else -%}
10
+ {{- arg_value | string -}}
11
+ {%- endif -%}
12
+ {%- endmacro -%}
13
+
14
+ {%- macro parse_content(content) -%}
15
+ {%- if content is string -%}
16
+ {{- content -}}
17
+ {%- else -%}
18
+ {%- set _ns = namespace(result="") -%}
19
+ {%- for item in content -%}
20
+ {%- if item["type"] == "image" -%}
21
+ {%- set _ns.result = _ns.result + "<image>" -%}
22
+ {%- elif item["type"] == "text" -%}
23
+ {%- set _ns.result = _ns.result + item["text"] -%}
24
+ {%- else -%}
25
+ {%- set _ns.result = _ns.result + item | tojson -%}
26
+ {%- endif -%}
27
+ {%- endfor -%}
28
+ {{- _ns.result -}}
29
+ {%- endif -%}
30
+ {%- endmacro -%}
31
+
32
+ {%- macro render_tool_calls(tool_calls) -%}
33
+ {%- set tool_calls_ns = namespace(tool_calls=[]) -%}
34
+ {%- for tool_call in tool_calls -%}
35
+ {%- set func_name = tool_call["function"]["name"] -%}
36
+ {%- set func_args = tool_call["function"]["arguments"] -%}
37
+ {%- set args_ns = namespace(arg_strings=[]) -%}
38
+ {%- for arg_name, arg_value in func_args.items() -%}
39
+ {%- set args_ns.arg_strings = args_ns.arg_strings + [arg_name + "=" + format_arg_value(arg_value)] -%}
40
+ {%- endfor -%}
41
+ {%- set tool_calls_ns.tool_calls = tool_calls_ns.tool_calls + [func_name + "(" + (args_ns.arg_strings | join(", ")) + ")"] -%}
42
+ {%- endfor -%}
43
+ {{- "<|tool_call_start|>[" + (tool_calls_ns.tool_calls | join(", ")) + "]<|tool_call_end|>" -}}
44
+ {%- endmacro -%}
45
+
46
+ {%- set ns = namespace(system_prompt="", last_user_index=-1) -%}
47
+ {%- if messages[0]["role"] == "system" -%}
48
+ {%- if messages[0].get("content") -%}
49
+ {%- set ns.system_prompt = parse_content(messages[0]["content"]) -%}
50
+ {%- endif -%}
51
+ {%- set messages = messages[1:] -%}
52
+ {%- endif -%}
53
+ {%- if tools -%}
54
+ {%- set ns.system_prompt = ns.system_prompt + ("\n" if ns.system_prompt else "") + "List of tools: [" -%}
55
+ {%- for tool in tools -%}
56
+ {%- if tool is not string -%}
57
+ {%- set tool = tool | tojson -%}
58
+ {%- endif -%}
59
+ {%- set ns.system_prompt = ns.system_prompt + tool -%}
60
+ {%- if not loop.last -%}
61
+ {%- set ns.system_prompt = ns.system_prompt + ", " -%}
62
+ {%- endif -%}
63
+ {%- endfor -%}
64
+ {%- set ns.system_prompt = ns.system_prompt + "]" -%}
65
+ {%- endif -%}
66
+ {%- if ns.system_prompt -%}
67
+ {{- "<|im_start|>system\n" + ns.system_prompt + "<|im_end|>\n" -}}
68
+ {%- endif -%}
69
+ {%- for message in messages -%}
70
+ {%- if message["role"] == "user" -%}
71
+ {%- set ns.last_user_index = loop.index0 -%}
72
+ {%- endif -%}
73
+ {%- endfor -%}
74
+ {%- for message in messages -%}
75
+ {{- "<|im_start|>" + message.role + "\n" -}}
76
+ {%- if message.role == "assistant" -%}
77
+ {%- generation -%}
78
+ {%- if message.thinking is defined and (preserve_thinking or loop.index0 > ns.last_user_index) -%}
79
+ {{- "<think>" + message.thinking + "</think>" -}}
80
+ {%- endif -%}
81
+ {%- set _cfm_tag = "CONTINUE_FINAL_MESSAGE_TAG " -%}
82
+ {%- set _has_cfm = false -%}
83
+ {%- if message.content is defined -%}
84
+ {%- set content = parse_content(message.content) -%}
85
+ {%- if not (preserve_thinking or loop.index0 > ns.last_user_index) -%}
86
+ {%- if "</think>" in content -%}
87
+ {%- set content = content.split("</think>")[-1] | trim -%}
88
+ {%- endif -%}
89
+ {%- endif -%}
90
+ {%- if message.tool_calls is defined and content.endswith(_cfm_tag) -%}
91
+ {%- set _has_cfm = true -%}
92
+ {%- set _trunc_len = (content | length) - (_cfm_tag | length) -%}
93
+ {{- content[:_trunc_len] -}}
94
+ {%- else -%}
95
+ {{- content -}}
96
+ {%- endif -%}
97
+ {%- endif -%}
98
+ {%- if message.tool_calls is defined -%}
99
+ {{- render_tool_calls(message.tool_calls) -}}
100
+ {%- endif -%}
101
+ {%- if _has_cfm -%}
102
+ {{- _cfm_tag -}}
103
+ {%- endif -%}
104
+ {{- "<|im_end|>\n" -}}
105
+ {%- endgeneration -%}
106
+ {%- else %}
107
+ {%- if message.get("content") -%}
108
+ {{- parse_content(message["content"]) -}}
109
+ {%- endif -%}
110
+ {{- "<|im_end|>\n" -}}
111
+ {%- endif %}
112
+ {%- endfor -%}
113
+ {%- if add_generation_prompt -%}
114
+ {{- "<|im_start|>assistant\n" -}}
115
+ {%- endif -%}
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/tokenizer/generation_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 1,
4
+ "eos_token_id": 7,
5
+ "output_attentions": false,
6
+ "output_hidden_states": false,
7
+ "pad_token_id": 0,
8
+ "do_sample": true,
9
+ "temperature": 0.1,
10
+ "top_k": 50,
11
+ "repetition_penalty": 1.05,
12
+ "transformers_version": "5.2.0",
13
+ "use_cache": true
14
+ }
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/tokenizer/tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/tokenizer/tokenizer_config.json ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<|startoftext|>",
4
+ "clean_up_tokenization_spaces": false,
5
+ "eos_token": "<|im_end|>",
6
+ "is_local": true,
7
+ "legacy": false,
8
+ "model_input_names": [
9
+ "input_ids",
10
+ "attention_mask"
11
+ ],
12
+ "model_max_length": 1000000000000000019884624838656,
13
+ "pad_token": "<|pad|>",
14
+ "sp_model_kwargs": {},
15
+ "spaces_between_special_tokens": false,
16
+ "tokenizer_class": "TokenizersBackend",
17
+ "use_default_system_prompt": false,
18
+ "use_fast": true
19
+ }
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/main.hash ADDED
@@ -0,0 +1 @@
 
 
1
+ �#�*<'������ ����]�&Xr�1��3#
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/main.mlirb ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:912308a32a3c1727f4ababf8ff830bbcb8ab915da50e265872f18c31fbb13323
3
+ size 150825615
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/metadata.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "producer" : "coreai-core 1.0.0b2",
3
+ "assetVersion" : "2.0",
4
+ "creationDate" : "20260713T134631Z"
5
+ }
config.json ADDED
@@ -0,0 +1,7 @@
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "coreai-aimodel",
3
+ "format": "aimodel",
4
+ "framework": "Apple Core AI (iOS 27 / macOS 27)",
5
+ "engine": "static-shape (Neural Engine)",
6
+ "note": "Fine-tuned LFM2.5-230M tool-caller converted to an Apple Core AI .aimodel. This is a STATIC-SHAPE / Neural Engine bundle (4 entrypoints), not a gpu-pipelined decode bundle. See README.md for the bundle layout, the required runtime patch, and run instructions."
7
+ }