sabeshbesh's picture
Add LFM2.5-230M tool-caller as Core AI ANE static .aimodel (int4 palettized, 512 ctx, coreai-torch 0.4.1)
1c87874 verified
|
Raw
History Blame Contribute Delete
8.8 kB
---
license: other
license_name: lfm1.0
license_link: LICENSE
base_model: LiquidAI/LFM2.5-230M
tags:
- coreai
- aimodel
- apple-silicon
- on-device
- neural-engine
- lfm2
- hybrid
- tool-calling
- function-calling
pipeline_tag: text-generation
---
# Uncle Rudy — LFM2.5-230M tool-caller — Apple Core AI (`.aimodel`)
A **LoRA fine-tune of [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M)** that turns
a single spoken/typed utterance into a **JSON tool call** for a to-do app — converted to Apple's
**Core AI** `.aimodel` format and running on the **Neural Engine** (iOS 27 / macOS 27).
> [!IMPORTANT]
> This is a **static-shape / Neural Engine** bundle (4 entrypoints: `load_embeddings`,
> `gather_embeddings`, `extend_*`, `prompt_opt_*`) → Apple's `EngineFactory` selects the
> **`static-shape`** engine. It is **not** a `gpu-pipelined` decode bundle, so it will **not**
> load on the `coreai-pipelined` GPU path that most published Core AI chat bundles use.
> It also needs a small runtime patch — see [Runtime requirements](#runtime-requirements).
## What it does
Give it one utterance; it emits one tool call.
```
"Remind me to buy milk" → {"name":"create_todo","arguments":{"title":"Buy milk","due":null}}
"Remind me to call the dentist tomorrow" → {"name":"create_todo","arguments":{"title":"Call the dentist","due":"tomorrow"}}
"Delete the dentist task" → {"name":"delete_todo","arguments":{"target":"Call the dentist"}}
```
Tool surface: `create_todo` (`title`, `due`) · `set_status` (`target`, `status`) ·
`update_todo` (`target`, `title`, `due`) · `delete_todo` (`target`).
Apply the bundle's own chat template and feed the raw utterance as the user turn — the fine-tune
emits the assistant tool call directly. **No system prompt / tool-schema preamble is needed**
(training used `mask_prompt`, so loss was computed only on the tool-call tokens). Decode **greedily**.
## Bundle
```
ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal/
├── uncle_rudy_lfm2_230m_toolcaller_int4pal.aimodel/ # main.mlirb · main.hash · metadata.json
├── tokenizer/ # tokenizer.json · tokenizer_config.json
│ # chat_template.jinja · generation_config.json
└── metadata.json # LanguageBundle manifest
```
| | |
|---|---|
| Base | LiquidAI/LFM2.5-230M (14 layers = 8 short-conv + 6 GQA attention, hidden 1024, 16/8 heads, vocab 65536, tied embeddings) |
| Fine-tune | LoRA r=16, scale 16, 400 iters, `mask_prompt: true` (loss on tool call only), fused |
| Compression | `4bit_weight_palettized_group32` (embedding table int8, per Core AI's iOS path) |
| Size | **144 MB** `.aimodel` (~148 MB with tokenizer) |
| Max context | **512** (prompt + generation) — deliberately small; see [Context length](#context-length) |
| Engine | `static-shape` (Neural Engine) |
## Runtime requirements
1. **iOS 27 / macOS 27** — Core AI ships with the OS. **Device-only on iOS**: `CoreAI.framework` is
in the iPhoneOS SDK but **not** the iOS Simulator SDK, so this cannot run in the Simulator.
2. **Toolchain ≥ beta 3 era.** This bundle was exported with **`coreai-torch 0.4.1` / `coreai-core 1.0.0b2`**
so it carries the *versioned-IR* location format the beta-3 on-device compiler requires. (Bundles
exported with the June-era `coreai-torch 0.4.0` / `coreai-core 1.0.0b1` fail on beta 3 with
`expected AICode versioned location … Failed to convert to versioned IR … cannot unwrap empty odiec_module_t`.)
3. **A conv-cache extra-state patch on the static-shape engine.** LFM2 is a conv+attention hybrid: its
short-conv layers carry a rolling **conv cache** *in addition to* the KV cache, so the model exports
**three** states (`key_cache`, `value_cache`, `conv_cache`). Apple's `StaticShapeEngine` hardcodes two.
The engine must allocate the extra state, **zero-fill it on reset** (unlike KV it is read directly, not
mask-gated), and bind it by name each step.
> Note: the conv state cannot be partially prefix-rewound (it only holds the last `L-1` columns), so a
> conv model should reprocess each request from scratch rather than reuse a partial KV prefix.
## Use it
The bundle is a standard `LanguageBundle` (`.aimodel` + `tokenizer/` + `metadata.json`) — point the
runtime at the **bundle directory**:
```swift
import CoreAILanguageModels
let bundle = try LanguageBundle(at: bundleDir) // dir containing the .aimodel
let engine = try await CoreAIRunner(from: bundle).makeInferenceEngine() // → static-shape / ANE
let tokenizer = try await bundle.loadTokenizer()
let generator = try await TextGeneratorBuilder()
.withInferenceEngine(engine)
.withTokenizer(tokenizer)
.withSampling(configuration: .greedy)
.build()
// `.prompt` applies the bundle's chat template; the fine-tune emits the tool call.
let json = try await generator.generate(input: .prompt("Remind me to buy milk"), maxTokens: 60)
```
Catalog entry, for apps that pull Core AI bundles from HF by tree path:
```swift
ModelSpec(
bundleName: "uncle_rudy_lfm2_230m_toolcaller_int4pal",
hfRemotePath: "ane-static/uncle_rudy_lfm2_230m_toolcaller_int4pal",
repoURL: "https://huggingface.co/sabeshbesh/uncle-rudy-lfm2-230m-CoreAI",
label: "Uncle Rudy 230M",
approxSizeGB: 0.15,
warmupToken: 1,
maxContext: 512)
```
⚠️ A catalog/`ModelSpec` alone is not sufficient: an app wired to the **`coreai-pipelined`** engine
must route this bundle to the **static-shape** engine and carry the conv-cache patch above.
## First load is slow — by design
The **first** load on a given device compiles the model's Neural Engine graphs **on-device**, inside
Apple's `AIModel(contentsOf:options:)`. This cannot be shipped precompiled — the compiled program is
specific to that device + OS. The OS caches the result, so every later load is near-instant.
| | |
|---|---|
| Cold (first) load, Apple Silicon ANE | **~54 s** |
| Warm load (cached) | **~0.02–0.1 s** |
### Context length
Exported at **512** context on purpose. The iOS static export fans out one specialized ANE graph per
`(context_bucket × query_length)` per function — at 2048 that is **24** graphs and a **~151 s** cold
compile; at 512 it is **12** graphs and **~54 s**. A tool-caller only ever sees a short utterance plus
a ≤60-token call, so 512 is ample. Re-export at a larger `--max-context-length` if you need more, and
pay the longer one-time compile.
## Measured
Greedy, `llm-runner`, **macOS / Apple Silicon Neural Engine** (this bundle):
| | |
|---|---|
| Prefill | ~80–330 tok/s |
| Decode | ~73–83 tok/s |
> iPhone on-device throughput is **not yet published** — these are Mac-ANE numbers for the same bundle.
> Treat them as indicative, not as iPhone figures.
**Correctness.** The re-authored BC1S / Neural-Engine model was gated against the fp32 Hugging Face
reference: **100% next-token top-1 match (12/12 positions)**, logits PSNR ~52 dB, on a truncated model
covering both layer types (conv + attention). The lower-than-macOS PSNR (~70 dB on the GPU/dynamic
path) is expected: the iOS path int8-quantizes the embedding table and computes attention per-head,
which reassociates fp16 arithmetic — every argmax still matches.
## Conversion notes
Converted from the fused Hugging Face checkpoint via a custom Core AI recipe on top of
[apple/coreai-models](https://github.com/apple/coreai-models). Two things were needed beyond the
stock pipeline:
- **MLX → PyTorch conv-weight transpose.** `mlx_lm fuse` writes depthwise Conv1d weights in MLX axis
order `(out, kernel, in)` = `[1024, 3, 1]`; PyTorch's `modeling_lfm2` wants `(out, in, kernel)` =
`[1024, 1, 3]`. Only the 8 conv layers are affected (Linear/embedding/norm layouts are identical),
and the fix is an axis swap — verified bit-exact against the base weights.
- **Re-authoring for the Neural Engine.** BC1S `(B, C, 1, S)` layout, projections as 1×1 `Conv2d`,
per-head attention (no fused SDPA on ANE), transposed causal mask using `-40000` rather than `-inf`,
and the short conv as a depthwise `Conv2d(D, D, (1, L), groups=D)` over the sequence axis, with the
conv cache threaded as a third functional state.
## License
Weights derive from [LiquidAI/LFM2.5-230M](https://huggingface.co/LiquidAI/LFM2.5-230M) and are
redistributed under the **LFM Open License v1.0** ([LICENSE](LICENSE)) — Apache-style grants, but
**commercial use is licensed only for entities under US$10M annual revenue** (qualified non-profits
exempt for non-commercial/research use). Review the LICENSE before any commercial deployment.