--- library_name: transformers license: other license_name: lfm1.0 license_link: LICENSE language: - ar - zh - en - fr - de - hi - id - it - ja - ko - pl - pt - ru - es - th - vi pipeline_tag: image-text-to-text tags: - liquid - lfm2.5 - edge - heretic - uncensored - decensored - abliterated base_model: - LiquidAI/LFM2.5-VL-3B --- This is an **LFM2.5-VL-3B** fine-tune, produced through P-E-W's [Heretic](https://github.com/p-e-w/heretic) (v1.4.0) abliteration engine with [Self-Organizing Maps & Magnitude-Preserving Orthogonal Ablation](https://github.com/p-e-w/heretic/pull/196) enabled. ---

Heretication Results


| Score Metric | Value | Parameter | Value | | :--- | :--- | :--- | :--- | | **Refusals** | 3/104 | **direction_index** | 17.46 | | **KL Divergence** | 0.0212 | **attn.o_proj.max_weights.0** | 0: 1.01 | | **Initial Refusals** | 104/104 | **attn.o_proj.max_weights.1** | 1: 0.50 | ||| **attn.o_proj.max_weights.2** | 2: 1.35 | ||| **attn.o_proj.max_weights.3** | 3: 0.63 | ||| **attn.o_proj.max_weights.4** | 4: 1.39 | ||| **attn.o_proj.max_weight_position** | 13.51 | ||| **attn.o_proj.min_weights.0** | 0: 0.73 | ||| **attn.o_proj.min_weights.1** | 1: 0.45 | ||| **attn.o_proj.min_weights.2** | 2: 0.15 | ||| **attn.o_proj.min_weights.3** | 3: 0.40 | ||| **attn.o_proj.min_weights.4** | 4: 1.21 | ||| **attn.o_proj.min_weight_distance** | 3.20 | ||| **mlp.down_proj.max_weights.0** | 0: 0.40 | ||| **mlp.down_proj.max_weights.1** | 1: 0.73 | ||| **mlp.down_proj.max_weights.2** | 2: 0.44 | ||| **mlp.down_proj.max_weights.3** | 3: 0.88 | ||| **mlp.down_proj.max_weights.4** | 4: 0.81 | ||| **mlp.down_proj.max_weight_position** | 16.05 | ||| **mlp.down_proj.min_weights.0** | 0: 0.18 | ||| **mlp.down_proj.min_weights.1** | 1: 0.64 | ||| **mlp.down_proj.min_weights.2** | 2: 0.20 | ||| **mlp.down_proj.min_weights.3** | 3: 0.16 | ||| **mlp.down_proj.min_weights.4** | 4: 0.52 | ||| **mlp.down_proj.min_weight_distance** | 6.67 | --- ## Degree of Heretication The **Heresy Index** weighs the resulting model's corruption by the process (KL Divergence & PIQA, Manual Response Eval) and its abolition of doctrine (Refusals) for a final verdict in classification. | Index Entry | Classification | Analysis | | :--- | :--- | :--- | | ![Absolute](https://img.shields.io/badge/HERESY_INDEX-ABSOLUTE-white?style=flat-square&labelColor=101010) | **Absolute Heresy** | Near zero overt and secondary refusals with minimal to no model damage | | ![Tainted](https://img.shields.io/badge/HERESY_INDEX-TAINTED-blueviolet?style=flat-square&labelColor=101010) | **Tainted Heresy** | Some residual secondary refusals and/or moderate model damage | | ![Impotent](https://img.shields.io/badge/HERESY_INDEX-IMPOTENT-5c4033?style=flat-square&labelColor=101010) | **Impotent Heresy** | Lingering overt refusals and high model damage | **Note**: This is an arbitrary and subjective classification inspired by Warhammer 40K, indended to serve as a signpost towards the model's performance. --- **Appendix** > Empty system prompt. > > Response prefix: ``
Heretication Rituals ``` [Trial 193] Refusals: 0/104, KL divergence: 0.0902 [Trial 364] Refusals: 1/104, KL divergence: 0.0432 [Trial 125] Refusals: 2/104, KL divergence: 0.0251 » [Trial 266] Refusals: 3/104, KL divergence: 0.0212 [Trial 79] Refusals: 6/104, KL divergence: 0.0127 [Trial 386] Refusals: 20/104, KL divergence: 0.0115 [Trial 168] Refusals: 43/104, KL divergence: 0.0094 [Trial 290] Refusals: 84/104, KL divergence: 0.0044 [Trial 72] Refusals: 101/104, KL divergence: 0.0037 [Trial 150] Refusals: 102/104, KL divergence: 0.0029 [Trial 318] Refusals: 104/104, KL divergence: 0.0016 ```
PIQA Benchmarks ``` ┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┓ ┃ Benchmark ┃ Metric ┃ T364 ┃ Original model ┃ ┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━┩ │ PIQA │ name │ piqa │ piqa │ │ │ sample_len │ 1838 │ 1838 │ │ │ acc,none │ 0.6126 │ 0.6148 │ │ │ acc_stderr,none │ 0.0114 │ 0.0114 │ │ │ acc_norm,none │ 0.6181 │ 0.6197 │ │ │ acc_norm_stderr,none │ 0.0113 │ 0.0113 │ └───────────┴──────────────────────┴────────────┴────────────────┘ ┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┓ ┃ Benchmark ┃ Metric ┃ T125 ┃ Original model ┃ ┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━┩ │ PIQA │ name │ piqa │ piqa │ │ │ sample_len │ 1838 │ 1838 │ │ │ acc,none │ 0.6148 │ 0.6148 │ │ │ acc_stderr,none │ 0.0114 │ 0.0114 │ │ │ acc_norm,none │ 0.6153 │ 0.6197 │ │ │ acc_norm_stderr,none │ 0.0114 │ 0.0113 │ └───────────┴──────────────────────┴────────────┴────────────────┘ ┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┓ ┃ Benchmark ┃ Metric ┃ T266 ┃ Original model ┃ ┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━┩ │ PIQA │ name │ piqa │ piqa │ │ │ sample_len │ 1838 │ 1838 │ │ │ acc,none │ 0.6202 │ 0.6148 │ │ │ acc_stderr,none │ 0.0113 │ 0.0114 │ │ │ acc_norm,none │ 0.6170 │ 0.6197 │ │ │ acc_norm_stderr,none │ 0.0113 │ 0.0113 │ └───────────┴──────────────────────┴────────────┴────────────────┘ ┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━┓ ┃ Benchmark ┃ Metric ┃ T079 ┃ Original model ┃ ┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━┩ │ PIQA │ name │ piqa │ piqa │ │ │ sample_len │ 1838 │ 1838 │ │ │ acc,none │ 0.6175 │ 0.6148 │ │ │ acc_stderr,none │ 0.0113 │ 0.0114 │ │ │ acc_norm,none │ 0.6181 │ 0.6197 │ │ │ acc_norm_stderr,none │ 0.0113 │ 0.0113 │ └───────────┴──────────────────────┴────────────┴────────────────┘ ```
PaCMAP Projection PaCMAP projection
---
Liquid AI
Try LFMDocsLEAPDiscord
# LFM2.5-VL-3B LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for **on-device deployment**. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images, and uses the LFM2.5-2.6B language model as its backbone, combined with a SigLIP2 NaFlex vision encoder. * **Better grounding**: Improved grounding and object detection with natural language queries. * **Better OCR**: Full page OCR with layout annotation. See [layout annotation format](#layout-annotation-format) for more information. * **Efficient inference**: 228 tok/s on an Apple M5 Max and 116 tok/s on an AMD Ryzen AI Max+ 395, in under 3.3 GB of memory. Find more information about LFM2.5-VL-3B in our [release post](https://www.liquid.ai/blog/lfm2-5-vl-3b). ![lfm2_5_vl_3b_task_group_averages](https://cdn-uploads.huggingface.co/production/uploads/644249b08443bce4c9890a0f/xw2m32B8IA0mRbhn_KG7f.png) > [!NOTE] > 💻 **Demos**: Try LFM2.5-VL-3B's vision understanding capabilities in a Hugging Face space without any setup: > **[Vision-capable chat in your browser](https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-3B-WebGPU)**: allows you to upload images or use the webcam to capture stills and let the model interact with them. ## Model Details | Model | Description | |-------|-------------| | **[LFM2.5‑VL‑3B](https://huggingface.co/LiquidAI/LFM2.5-VL-3B)** | Original checkpoint in native format. Best for fine-tuning and inference with HF Transformers, vLLM and SGLang | | **[LFM2.5‑VL‑3B‑GGUF](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-GGUF)** | Quantized GGUF exports of the original checkpoint. Best for CPU inference with reduced memory usage with llama.cpp | | **[LFM2.5‑VL‑3B‑ONNX](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-ONNX)** | Quantized ONNX exports for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). See the [demo](https://huggingface.co/spaces/LiquidAI/LFM2.5-VL-3B-WebGPU). | | **[LFM2.5-VL-3B-MLX](https://huggingface.co/LiquidAI/LFM2.5-VL-3B-MLX-8bit)** | Quantized MLX exports for Apple Silicon. Optimized for fast inference on Mac devices using the [mlx-vlm](https://github.com/Blaizzy/mlx-vlm) framework. | - **LM Backbone**: LFM2.5-2.6B - **Vision encoder**: SigLIP2 NaFlex shape‑optimized 400M - **Vocabulary size:** 128,000 - **Context length**: 32,768 tokens - **Languages**: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish - **Native resolution processing**: Uses SigLIP2's NaFlex; large images are split into non-overlapping 512×512 patches and a resized whole-image thumbnail. - **Generation parameters**: - text: `temperature=0.2`, `top_k=50`, `repetition_penalty=1.0` - vision: Use the `processor_config.json` file. We recommend using it for single-turn, high-throughput, low-latency tasks; for example, for near-realtime object detection in automotive applications, batch processing scanned documents with OCR with layout information for turning PDFs into searchable text, or for on-device translation of menus and road signs into your native language. It is not recommended for long-context, reasoning-intensive tasks, such as visual web design, or answering highly technical questions about blueprints. ### Chat Template LFM2.5 uses a ChatML-like format. See the [Chat Template documentation](https://docs.liquid.ai/lfm/key-concepts/chat-template#vision-models) for details. Example: ``` <|startoftext|><|im_start|>system You are a helpful assistant trained by Liquid AI.<|im_end|> <|im_start|>user What species is in this picture?<|im_end|> <|im_start|>assistant ``` You can use [`tokenizer.apply_chat_template()`](https://huggingface.co/docs/transformers/en/chat_templating#using-applychattemplate) to format your messages automatically. > [!TIP] > **Note**: The `apply_chat_template()` method automatically inserts the `` tag for each image in your message. Do not include `` in your message content. ## Inference LFM2.5-VL is supported by many inference frameworks. See the [Inference documentation](https://docs.liquid.ai/lfm/inference/transformers) for the full list. | Name | Description | Docs | Notebook | |------|-------------|------|----------| | [Transformers](https://github.com/huggingface/transformers) | Simple inference with direct access to model internals. | Link| Colab link | | [vLLM](https://github.com/vllm-project/vllm) | High-throughput production deployments with GPU. | Link | Colab link | | [SGLang](https://github.com/sgl-project/sglang) | High-throughput production deployments with GPU. | Link | Colab link | | [llama.cpp](https://github.com/ggml-org/llama.cpp) | Cross-platform inference with CPU offloading. | Link | Colab link | ### Quick start Quick start with Transformers (compatible with `transformers>=5.0.0`): You will need `torch`, `transformers`, and `torchvision`. ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor model_id = "LiquidAI/LFM2.5-VL-3B" model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16) processor = AutoProcessor.from_pretrained(model_id) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://placecats.com/300/200"}, {"type": "text", "text": "Describe this image."}, ], } ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) with torch.inference_mode(): output_ids = model.generate( **inputs, do_sample=True, temperature=0.2, top_k=50, repetition_penalty=1.0, max_new_tokens=256, ) generated_ids = output_ids[:, inputs["input_ids"].shape[1] :] print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0]) ``` ### Tool Use LFM2.5-VL-3B supports function calling in four steps: 1. **Function definition**: Provide the list of tools as a JSON object in the system prompt, or use [`tokenizer.apply_chat_template()`](https://huggingface.co/docs/transformers/en/chat_extras#passing-tools) with `tools=...`. 2. **Function call**: By default, LFM2.5 writes Pythonic function calls (a Python list between `<|tool_call_start|>` and `<|tool_call_end|>` special tokens), as the assistant answer. 3. **Function execution**: Execute the call and return the result with the `tool` role. 4. **Final answer**: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt. See the [Tool Use documentation](https://docs.liquid.ai/lfm/key-concepts/tool-use) for the full guide. Example: ``` <|startoftext|><|im_start|>system List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|> <|im_start|>user What is the current status of candidate ID 12345?<|im_end|> <|im_start|>assistant <|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|> <|im_start|>tool [{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|> <|im_start|>assistant The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|> ``` ### Layout Annotation Format LFM2.5-VL-3B can do OCR with layout annotation. The layout annotation is a list of regions, each with a label, bounding box, and content. The format is: ```text image_index=