Add LFM2.5-230M tool-caller as Core AI ANE static .aimodel (int4 palettized, 512 ctx, coreai-torch 0.4.1)
1c87874 verified | license: other | |
| license_name: lfm1.0 | |
| license_link: LICENSE | |
| base_model: LiquidAI/LFM2.5-230M | |
| tags: | |
| - coreai | |
| - aimodel | |
| - apple-silicon | |
| - on-device | |
| - neural-engine | |
| - lfm2 | |
| - hybrid | |
| - tool-calling | |
| - function-calling | |
| pipeline_tag: text-generation | |
| # Uncle Rudy — LFM2.5-230M tool-caller — Apple Core AI (`.aimodel`) | |
| A **LoRA fine-tune of [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M)** that turns | |
| a single spoken/typed utterance into a **JSON tool call** for a to-do app — converted to Apple's | |
| **Core AI** `.aimodel` format and running on the **Neural Engine** (iOS 27 / macOS 27). | |
| > [!IMPORTANT] | |
| > This is a **static-shape / Neural Engine** bundle (4 entrypoints: `load_embeddings`, | |
| > `gather_embeddings`, `extend_*`, `prompt_opt_*`) → Apple's `EngineFactory` selects the | |
| > **`static-shape`** engine. It is **not** a `gpu-pipelined` decode bundle, so it will **not** | |
| > load on the `coreai-pipelined` GPU path that most published Core AI chat bundles use. | |
| > It also needs a small runtime patch — see [Runtime requirements](#runtime-requirements). | |
| ## What it does | |
| Give it one utterance; it emits one tool call. | |
| ``` | |
| "Remind me to buy milk" → {"name":"create_todo","arguments":{"title":"Buy milk","due":null}} | |
| "Remind me to call the dentist tomorrow" → {"name":"create_todo","arguments":{"title":"Call the dentist","due":"tomorrow"}} | |
| "Delete the dentist task" → {"name":"delete_todo","arguments":{"target":"Call the dentist"}} | |
| ``` | |
| Tool surface: `create_todo` (`title`, `due`) · `set_status` (`target`, `status`) · | |
| `update_todo` (`target`, `title`, `due`) · `delete_todo` (`target`). | |
| Apply the bundle's own chat template and feed the raw utterance as the user turn — the fine-tune | |
| emits the assistant tool call directly. **No system prompt / tool-schema preamble is needed** | |
| (training used `mask_prompt`, so loss was computed only on the tool-call tokens). Decode **greedily**. | |
| ## Bundle | |
| ``` | |
| ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/ | |
| ├── uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/ # main.mlirb · main.hash · metadata.json | |
| ├── tokenizer/ # tokenizer.json · tokenizer_config.json | |
| │ # chat_template.jinja · generation_config.json | |
| └── metadata.json # LanguageBundle manifest | |
| ``` | |
| | | | | |
| |---|---| | |
| | Base | LiquidAI/LFM2.5-230M (14 layers = 8 short-conv + 6 GQA attention, hidden 1024, 16/8 heads, vocab 65536, tied embeddings) | | |
| | Fine-tune | LoRA r=16, scale 16, 400 iters, `mask_prompt: true` (loss on tool call only), fused | | |
| | Compression | `4bit_weight_palettized_group32` (embedding table int8, per Core AI's iOS path) | | |
| | Size | **144 MB** `.aimodel` (~148 MB with tokenizer) | | |
| | Max context | **512** (prompt + generation) — deliberately small; see [Context length](#context-length) | | |
| | Engine | `static-shape` (Neural Engine) | | |
| ## Runtime requirements | |
| 1. **iOS 27 / macOS 27** — Core AI ships with the OS. **Device-only on iOS**: `CoreAI.framework` is | |
| in the iPhoneOS SDK but **not** the iOS Simulator SDK, so this cannot run in the Simulator. | |
| 2. **Toolchain ≥ beta 3 era.** This bundle was exported with **`coreai-torch 0.4.1` / `coreai-core 1.0.0b2`** | |
| so it carries the *versioned-IR* location format the beta-3 on-device compiler requires. (Bundles | |
| exported with the June-era `coreai-torch 0.4.0` / `coreai-core 1.0.0b1` fail on beta 3 with | |
| `expected AICode versioned location … Failed to convert to versioned IR … cannot unwrap empty odiec_module_t`.) | |
| 3. **A conv-cache extra-state patch on the static-shape engine.** LFM2 is a conv+attention hybrid: its | |
| short-conv layers carry a rolling **conv cache** *in addition to* the KV cache, so the model exports | |
| **three** states (`key_cache`, `value_cache`, `conv_cache`). Apple's `StaticShapeEngine` hardcodes two. | |
| The engine must allocate the extra state, **zero-fill it on reset** (unlike KV it is read directly, not | |
| mask-gated), and bind it by name each step. | |
| > Note: the conv state cannot be partially prefix-rewound (it only holds the last `L-1` columns), so a | |
| > conv model should reprocess each request from scratch rather than reuse a partial KV prefix. | |
| ## Use it | |
| The bundle is a standard `LanguageBundle` (`.aimodel` + `tokenizer/` + `metadata.json`) — point the | |
| runtime at the **bundle directory**: | |
| ```swift | |
| import CoreAILanguageModels | |
| let bundle = try LanguageBundle(at: bundleDir) // dir containing the .aimodel | |
| let engine = try await CoreAIRunner(from: bundle).makeInferenceEngine() // → static-shape / ANE | |
| let tokenizer = try await bundle.loadTokenizer() | |
| let generator = try await TextGeneratorBuilder() | |
| .withInferenceEngine(engine) | |
| .withTokenizer(tokenizer) | |
| .withSampling(configuration: .greedy) | |
| .build() | |
| // `.prompt` applies the bundle's chat template; the fine-tune emits the tool call. | |
| let json = try await generator.generate(input: .prompt("Remind me to buy milk"), maxTokens: 60) | |
| ``` | |
| Catalog entry, for apps that pull Core AI bundles from HF by tree path: | |
| ```swift | |
| ModelSpec( | |
| bundleName: "uncle_rudy_lfm2_230m_toolcaller_int4pal", | |
| hfRemotePath: "ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal", | |
| repoURL: "https://huggingface.co/sabeshbesh/uncle-rudy-lfm2-230m-CoreAI", | |
| label: "Uncle Rudy 230M", | |
| approxSizeGB: 0.15, | |
| warmupToken: 1, | |
| maxContext: 512) | |
| ``` | |
| ⚠️ A catalog/`ModelSpec` alone is not sufficient: an app wired to the **`coreai-pipelined`** engine | |
| must route this bundle to the **static-shape** engine and carry the conv-cache patch above. | |
| ## First load is slow — by design | |
| The **first** load on a given device compiles the model's Neural Engine graphs **on-device**, inside | |
| Apple's `AIModel(contentsOf:options:)`. This cannot be shipped precompiled — the compiled program is | |
| specific to that device + OS. The OS caches the result, so every later load is near-instant. | |
| | | | | |
| |---|---| | |
| | Cold (first) load, Apple Silicon ANE | **~54 s** | | |
| | Warm load (cached) | **~0.02–0.1 s** | | |
| ### Context length | |
| Exported at **512** context on purpose. The iOS static export fans out one specialized ANE graph per | |
| `(context_bucket × query_length)` per function — at 2048 that is **24** graphs and a **~151 s** cold | |
| compile; at 512 it is **12** graphs and **~54 s**. A tool-caller only ever sees a short utterance plus | |
| a ≤60-token call, so 512 is ample. Re-export at a larger `--max-context-length` if you need more, and | |
| pay the longer one-time compile. | |
| ## Measured | |
| Greedy, `llm-runner`, **macOS / Apple Silicon Neural Engine** (this bundle): | |
| | | | | |
| |---|---| | |
| | Prefill | ~80–330 tok/s | | |
| | Decode | ~73–83 tok/s | | |
| > iPhone on-device throughput is **not yet published** — these are Mac-ANE numbers for the same bundle. | |
| > Treat them as indicative, not as iPhone figures. | |
| **Correctness.** The re-authored BC1S / Neural-Engine model was gated against the fp32 Hugging Face | |
| reference: **100% next-token top-1 match (12/12 positions)**, logits PSNR ~52 dB, on a truncated model | |
| covering both layer types (conv + attention). The lower-than-macOS PSNR (~70 dB on the GPU/dynamic | |
| path) is expected: the iOS path int8-quantizes the embedding table and computes attention per-head, | |
| which reassociates fp16 arithmetic — every argmax still matches. | |
| ## Conversion notes | |
| Converted from the fused Hugging Face checkpoint via a custom Core AI recipe on top of | |
| [apple/coreai-models](https://github.com/apple/coreai-models). Two things were needed beyond the | |
| stock pipeline: | |
| - **MLX → PyTorch conv-weight transpose.** `mlx_lm fuse` writes depthwise Conv1d weights in MLX axis | |
| order `(out, kernel, in)` = `[1024, 3, 1]`; PyTorch's `modeling_lfm2` wants `(out, in, kernel)` = | |
| `[1024, 1, 3]`. Only the 8 conv layers are affected (Linear/embedding/norm layouts are identical), | |
| and the fix is an axis swap — verified bit-exact against the base weights. | |
| - **Re-authoring for the Neural Engine.** BC1S `(B, C, 1, S)` layout, projections as 1×1 `Conv2d`, | |
| per-head attention (no fused SDPA on ANE), transposed causal mask using `-40000` rather than `-inf`, | |
| and the short conv as a depthwise `Conv2d(D, D, (1, L), groups=D)` over the sequence axis, with the | |
| conv cache threaded as a third functional state. | |
| ## License | |
| Weights derive from [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) and are | |
| redistributed under the **LFM Open License v1.0** ([LICENSE](LICENSE)) — Apache-style grants, but | |
| **commercial use is licensed only for entities under US$10M annual revenue** (qualified non-profits | |
| exempt for non-commercial/research use). Review the LICENSE before any commercial deployment. | |