--- language: - en license: apache-2.0 library_name: gguf tags: - gguf - qwen2 - instruct - vision - multimodal - affectively - edgework - aether - distributed-inference - edge-deployment base_model: Qwen/Qwen2.5-VL-3B-Instruct base_model_relation: quantized pipeline_tag: image-text-to-text --- # Qwen2.5 VL 3b Instruct (GGUF, Q4_K_M) > **Production-ready** GGUF quantization of [Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) for distributed visual understanding and reasoning — powered by the [Aether](https://github.com/forkjoin-ai/aether) edge inference runtime on [Edgework.ai](https://edgework.ai). ## Model Details | Property | Value | |----------|-------| | Base model | [Qwen/Qwen2.5-VL-3B-Instruct](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) | | Parameters | 3B | | Architecture | Qwen2 | | Quantization | Q4_K_M | | Format | GGUF | | Size | ~1.8 GB | | License | apache-2.0 | ## Usage ### With llama.cpp ```bash ./llama-cli -m Qwen2.5-VL-3B-Instruct-Q4_K_M.gguf -p "Your prompt here" -n 256 ``` ### With Aether (Distributed Inference) This model is deployed across the [Aether](https://github.com/forkjoin-ai/aether) distributed inference network. Weights are layer-sharded and distributed across multiple edge nodes for parallel inference. ## Also available: `.knot` (sovereign format) This repo ships [`qwen2.5-vl-3b.knot`](./qwen2.5-vl-3b.knot) — the model weights in the **KNOT** container that the [Aether](https://github.com/forkjoin-ai/aether) distributed-inference runtime loads natively (the GGUF, when present, sits right beside it). A KNOT is a single self-describing file with a JSON table-of-contents, so **any single tensor is one HTTP `Range` request** — ideal for streaming weights to edge nodes. | | GGUF | KNOT | |---|---|---| | Container | format-specific header | single file, JSON table-of-contents | | Per-tensor fetch | whole-file oriented | one tensor = one Range request | | Ecosystem | broad (llama.cpp, …) | Aether / Gnosis runtime | ```bash huggingface-cli download forkjoin-ai/qwen2.5-vl-3b-instruct-gguf qwen2.5-vl-3b.knot --local-dir ./knots ``` Full format spec: [KNOT_FORMAT.md](https://github.com/forkjoin-ai/aether/blob/main/docs/model-cards/KNOT_FORMAT.md). Inspect the header with `bun run open-source/bitwise/scripts/dump-knot.ts qwen2.5-vl-3b.knot`. ## Deployment Architecture This model runs on the **Aether distributed inference runtime** — a custom engine that shards model layers across multiple nodes for parallel execution: 1. **Coordinator** receives requests and manages token generation 2. **Layer nodes** each hold a subset of model layers (2 nodes for this model) 3. **Hidden states flow** between nodes via gRPC 4. **Zero cold start** via warm pool scheduling Deployed via [Edgework.ai](https://edgework.ai) — bringing fast, cheap, and private inference as close to the user as possible. ## About Published by [AFFECTIVELY](https://huggingface.co/forkjoin-ai) · Managed by [@buley](https://huggingface.co/buley) We quantize and publish **production-ready models** for distributed edge inference via the [Aether](https://github.com/forkjoin-ai/aether) runtime. Every release is tested for correctness and stability before publication. - [All models](https://huggingface.co/forkjoin-ai) · [GitHub](https://github.com/forkjoin-ai) · [Edgework.ai](https://edgework.ai)