Instructions to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Use Docker
docker model run hf.co/a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with Ollama:
ollama run hf.co/a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with Docker Model Runner:
docker model run hf.co/a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
- Lemonade
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Commit ·
b6d4978
0
Parent(s):
Squashed: Release 2026-08-15:2
Browse files- .gitattributes +3 -0
- .gitignore +3 -0
- BF16/Qwen3.8-2.4T-A95B-MTP-ONLY-BF16-00001-of-00002.gguf +3 -0
- BF16/Qwen3.8-2.4T-A95B-MTP-ONLY-BF16-00002-of-00002.gguf +3 -0
- LICENSE +16 -0
- Qwen3.8-2.4T-A95B-MTP-ONLY-Q4_K_M.gguf +3 -0
- Qwen3.8-2.4T-A95B-MTP-ONLY-Q5_K_M.gguf +3 -0
- Qwen3.8-2.4T-A95B-MTP-ONLY-Q6_K.gguf +3 -0
- Qwen3.8-2.4T-A95B-MTP-ONLY-Q8_0.gguf +3 -0
- README.md +166 -0
.gitattributes
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
*.gguf filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.md text -whitespace
|
| 3 |
+
/LICENSE text eol=lf -whitespace
|
.gitignore
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
.*
|
| 2 |
+
!.git?*
|
| 3 |
+
*~
|
BF16/Qwen3.8-2.4T-A95B-MTP-ONLY-BF16-00001-of-00002.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:37660e1d16a3192d750952f74fb98885feed045c4748ca2d9e4ea18cb3cbe343
|
| 3 |
+
size 38439156192
|
BF16/Qwen3.8-2.4T-A95B-MTP-ONLY-BF16-00002-of-00002.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:afeb3a9cc6891d78fdd16152c67d9e5f458d3e6e9f317aac1df92a266b004351
|
| 3 |
+
size 22473313600
|
LICENSE
ADDED
|
@@ -0,0 +1,16 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Qwen3.8-Max License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 Qwen
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy of this software, including the model weights, parameters, configuration files, inference code and associated documentation files (collectively, the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, sell, deploy, host, fine-tune, and create derivative works from (collectively, "Use" or "Using") copies of the Software; and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
|
| 6 |
+
|
| 7 |
+
1. The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. If the Software (or any derivative works thereof) is Used for any of the licensee's commercial products or services that have more than 100,000,000 monthly active users or US$ 20,000,000 (or equivalent in other currencies) monthly revenue, respective model name must be prominently displayed on the user interface of such product or service; and,
|
| 8 |
+
|
| 9 |
+
2. If the licensee or any of its affiliates conducts a Model as a Service or AI Work Assistant business, and the aggregate revenue of the licensee and its affiliates exceeds US$50,000,000 (or the equivalent amount in any other currencies) during any consecutive twelve (12) months, the licensee shall obtain a separate license from Qwen before Using the Software or its derivative works for any commercial purpose. The foregoing requirement shall not apply to the licensee's internal Use of the Software, provided that such Use does not make the Software, its outputs, or its underlying model capabilities available to any third party.
|
| 10 |
+
|
| 11 |
+
"Model as a Service" means giving a third party access to language model inference or fine-tuning (e.g., via API or a hosted endpoint) in a manner that allows such third parties to exercise meaningful control over the inputs, parameters, or training data. This does not include the mere relaying of requests to models hosted by other third parties.
|
| 12 |
+
“AI Work Assistant” means an independent AI-powered product primarily designed for AI-assisted coding or office productivity (e.g., Qoder and QwenWork). It does not include: (a) a single-purpose AI tool (such as an AI translation tool); (b) an AI assistant primarily designed for a domain other than coding or office productivity (such as Taobao AI Shopping Assistant or AMap AI Chat); or (c) an AI assistant that is a feature of a product whose primary purpose is not AI-assisted coding or office productivity.
|
| 13 |
+
|
| 14 |
+
THE SOFTWARE AND ANY OUTPUT AND RESULTS THEREFROM ARE PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL QWEN, ITS AFFILIATES OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. THE USE OF THE SOFTWARE MUST COMPLY WITH APPLICABLE LAWS AND REGULATIONS, AND MUST NOT INFRINGE THE INTELLECTUAL PROPERTY RIGHTS OF ANY THIRD PARTY.
|
| 15 |
+
|
| 16 |
+
For any questions regarding this license, please contact model-business@notice.qwencloud.com.
|
Qwen3.8-2.4T-A95B-MTP-ONLY-Q4_K_M.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:13ab30233ec48a61ea18cc809742f67c42db65c816be934e84b2a008b70a86ce
|
| 3 |
+
size 19897255520
|
Qwen3.8-2.4T-A95B-MTP-ONLY-Q5_K_M.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:01550a1504d7edf7d3686a5f852114c211075834a83c2cd895ec55d64717a30b
|
| 3 |
+
size 22371370592
|
Qwen3.8-2.4T-A95B-MTP-ONLY-Q6_K.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0a7cd74c135c238025390c335a15e5b19534d5198cffbe3b81a68fbe4213f28b
|
| 3 |
+
size 25000117856
|
Qwen3.8-2.4T-A95B-MTP-ONLY-Q8_0.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4d6cf03a5e28d46c867a03d5df5c9ffacc4a861d459a15465bfedeaca2a759b8
|
| 3 |
+
size 32372852320
|
README.md
ADDED
|
@@ -0,0 +1,166 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: gguf
|
| 3 |
+
license: other
|
| 4 |
+
license_name: qwen3.8-max
|
| 5 |
+
license_link: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/207bd685a7e3696cfaff12ded7c6a7ea0f88c996/LICENSE
|
| 6 |
+
base_model:
|
| 7 |
+
- Qwen/Qwen3.8-2.4T-A95B
|
| 8 |
+
base_model_relation: quantized
|
| 9 |
+
tags:
|
| 10 |
+
- qwen3.5
|
| 11 |
+
- qwen3.8
|
| 12 |
+
- speculative-decoding
|
| 13 |
+
- mtp
|
| 14 |
+
- mtp only
|
| 15 |
+
- gguf
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
MTP-only GGUF subset of Qwen3.8 models
|
| 19 |
+
=======================================
|
| 20 |
+
|
| 21 |
+
This is a supplement for Qwen3.8-2.4T-A95B-based models (including fine-tunes and
|
| 22 |
+
abliterated models) **without MTP tensors**.
|
| 23 |
+
|
| 24 |
+
This repository contains an **MTP-only subset** of [Qwen/Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)
|
| 25 |
+
which provides the draft model for speculative decoding in the GGUF format.
|
| 26 |
+
|
| 27 |
+
It accelerates token generation using speculative decoding with the draft model
|
| 28 |
+
from the original Qwen model. In most cases, this is sufficient to accelerate
|
| 29 |
+
Qwen-based derivative models even if this draft model is not trained from them.
|
| 30 |
+
|
| 31 |
+
Note that however, the performance metrics heavily depend on the derivative
|
| 32 |
+
model you use, your machine and your MTP settings.
|
| 33 |
+
|
| 34 |
+
**Benchmark it before blindly trusting it.**
|
| 35 |
+
|
| 36 |
+
|
| 37 |
+
Using this Model
|
| 38 |
+
-----------------
|
| 39 |
+
|
| 40 |
+
It can be used in two ways:
|
| 41 |
+
|
| 42 |
+
1. As a separate draft model file (Method 1)
|
| 43 |
+
2. As a donor for grafting the draft model into a Qwen-based model
|
| 44 |
+
(Method 2; Recommended)
|
| 45 |
+
|
| 46 |
+
### Method 1: Separate Draft Model File
|
| 47 |
+
|
| 48 |
+
It is easy to begin with but memory-inefficient as it does not share
|
| 49 |
+
some tensors with the original model.
|
| 50 |
+
|
| 51 |
+
If you find the draft model can accelerate a Qwen-based model you use, grafting
|
| 52 |
+
the draft model (Method 2) is recommended (note: switching to Method 2 may
|
| 53 |
+
slightly change the acceptance rate).
|
| 54 |
+
|
| 55 |
+
If you use `llama-server`, you may configure like this:
|
| 56 |
+
|
| 57 |
+
```
|
| 58 |
+
llama-server \
|
| 59 |
+
--model Qwen3.6-27B-finetune-Q4_K_M.gguf \
|
| 60 |
+
--model-draft Qwen3.6-27B-MTP-ONLY-Q4_K_M.gguf \
|
| 61 |
+
... \
|
| 62 |
+
--spec-type draft-mtp
|
| 63 |
+
--spec-draft-n-max 4
|
| 64 |
+
```
|
| 65 |
+
|
| 66 |
+
`--model` specifies the original Qwen-based model and
|
| 67 |
+
*new* `--model-draft` specifies a file from this repository.
|
| 68 |
+
|
| 69 |
+
You also need `--spec-type draft-mtp` to enable the draft model.
|
| 70 |
+
|
| 71 |
+
Once the draft model is enabled, you may configure the rest of MTP options
|
| 72 |
+
as you like (in this example, custom `--spec-draft-n-max` is specified).
|
| 73 |
+
|
| 74 |
+
### Method 2: Grafting the Draft Model (Using this as a Donor)
|
| 75 |
+
|
| 76 |
+
This is recommended.
|
| 77 |
+
|
| 78 |
+
First, download [`convert.py`](https://gist.github.com/buzz/1c439684d5e3f36492ae9f64ef7e3f67/7e1d6f929e141f4977f89f99e6a7df39f41eeffa)
|
| 79 |
+
written by [@buzz](https://github.com/buzz) to transplant MTP tensors.
|
| 80 |
+
You may also need to install some dependencies required by this script.
|
| 81 |
+
|
| 82 |
+
Then, you can run this script like:
|
| 83 |
+
|
| 84 |
+
```sh
|
| 85 |
+
# ./convert.py INPUT MTP OUTPUT
|
| 86 |
+
python3 convert.py \
|
| 87 |
+
Qwen3.6-27B-finetune-Q4_K_M.gguf \
|
| 88 |
+
Qwen3.6-27B-MTP-ONLY-Q4_K_M.gguf \
|
| 89 |
+
Qwen3.6-27B-finetune-Q4_K_M+MTP.gguf
|
| 90 |
+
```
|
| 91 |
+
|
| 92 |
+
The second argument of `convert.py` is a GGUF file (donor) from this repository.
|
| 93 |
+
|
| 94 |
+
Once the grafted GGUF file is created, you may use this like a regular
|
| 95 |
+
Qwen model with embedded draft model.
|
| 96 |
+
|
| 97 |
+
This is an example for `llama-server` users.
|
| 98 |
+
|
| 99 |
+
```
|
| 100 |
+
llama-server \
|
| 101 |
+
--model Qwen3.6-27B-finetune-Q4_K_M+MTP.gguf \
|
| 102 |
+
... \
|
| 103 |
+
--spec-type draft-mtp
|
| 104 |
+
--spec-draft-n-max 4
|
| 105 |
+
```
|
| 106 |
+
|
| 107 |
+
The output of `convert.py` must be specified as the model file name
|
| 108 |
+
(**DO NOT** use `--model-draft` in this case).
|
| 109 |
+
|
| 110 |
+
`--spec-type draft-mtp` enables the draft model transplanted into the main one.
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
Additional Quantization (`Q4_K_M`, `Q5_K_M`, `Q6_K` and `Q8_0`)
|
| 114 |
+
----------------------------------------------------------------
|
| 115 |
+
|
| 116 |
+
Quantized GGUF files are provided so that deploying the draft model easier.
|
| 117 |
+
|
| 118 |
+
*It is not required to match the quantization level.*
|
| 119 |
+
For instance, you may pair `Q6_K`-quantized draft model with
|
| 120 |
+
the `Q4_K_S`-quantized main model.
|
| 121 |
+
|
| 122 |
+
|
| 123 |
+
Conversion Process
|
| 124 |
+
-------------------
|
| 125 |
+
|
| 126 |
+
* Tools: [llama.cpp](https://github.com/ggml-org/llama.cpp) ([b10380](https://github.com/ggml-org/llama.cpp/releases/tag/b10380))
|
| 127 |
+
* With: Patched `conversion/base.py`
|
| 128 |
+
The `if` block right after `# verify tensor name presence and identify potentially missing files` is commented out.
|
| 129 |
+
|
| 130 |
+
This modification is performed because the author of this repository downloaded
|
| 131 |
+
only a subset of the full Qwen model while the original `convert_hf_to_gguf.py`
|
| 132 |
+
expects the full model.
|
| 133 |
+
|
| 134 |
+
The `--mtp` option of `convert_hf_to_gguf.py` is the crucial part of this
|
| 135 |
+
conversion process because this option does exactly what the author expects:
|
| 136 |
+
create an MTP-only GGUF subset.
|
| 137 |
+
|
| 138 |
+
For additional quantization, the `llama-quantize` tool (llama.cpp) is used as-is.
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
License and Copyright
|
| 142 |
+
----------------------
|
| 143 |
+
|
| 144 |
+
For all GGUF files under this repository,
|
| 145 |
+
[the license terms of the original Qwen model (Qwen3.8-Max License)](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/207bd685a7e3696cfaff12ded7c6a7ea0f88c996/LICENSE)
|
| 146 |
+
applies (as the author of this repository did not perform any changes
|
| 147 |
+
significant enough for own copyright):
|
| 148 |
+
|
| 149 |
+
> Copyright (c) 2026 Qwen
|
| 150 |
+
|
| 151 |
+
This README file is licensed under the terms of [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/).
|
| 152 |
+
|
| 153 |
+
> Copyright 2026 a4lg.
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
Links: Sources
|
| 157 |
+
---------------
|
| 158 |
+
|
| 159 |
+
* The Original Model: [Qwen/Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/tree/207bd685a7e3696cfaff12ded7c6a7ea0f88c996)
|
| 160 |
+
|
| 161 |
+
|
| 162 |
+
Links: All MTP Subsets: Qwen3.8 (incl. non-open-source one)
|
| 163 |
+
------------------------------------------------------------
|
| 164 |
+
|
| 165 |
+
* [Qwen/Qwen3.8-27B](https://huggingface.co/a4lg/Qwen3.8-27B-MTP-ONLY-GGUF)
|
| 166 |
+
* [Qwen/Qwen3.8-2.4T-A95B](https://huggingface.co/a4lg/Qwen3.8-2.4T-A95B-MTP-ONLY-GGUF)
|