Text Generation
Transformers
Safetensors
English
forgelm
keystack
mla
Mixture of Experts
mrl
quarot
rotorquant
mtp
airllm
weight-transform
training-free
code
conversational
custom_code
Instructions to use GRKTheGreat/ForgeLM-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GRKTheGreat/ForgeLM-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="GRKTheGreat/ForgeLM-v1", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import ForgeLM model = ForgeLM.from_pretrained("GRKTheGreat/ForgeLM-v1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GRKTheGreat/ForgeLM-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GRKTheGreat/ForgeLM-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/GRKTheGreat/ForgeLM-v1
- SGLang
How to use GRKTheGreat/ForgeLM-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GRKTheGreat/ForgeLM-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GRKTheGreat/ForgeLM-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GRKTheGreat/ForgeLM-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use GRKTheGreat/ForgeLM-v1 with Docker Model Runner:
docker model run hf.co/GRKTheGreat/ForgeLM-v1
File size: 10,446 Bytes
72c2207 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 | ---
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
tags:
- keystack
- mla
- moe
- mrl
- quarot
- rotorquant
- mtp
- airllm
- weight-transform
- training-free
- code
language:
- en
pipeline_tag: text-generation
---
# ForgeLM v1
> **A training-free architectural port of Qwen2.5-Coder-1.5B-Instruct with KeyStack transforms.**
ForgeLM v1 is **not a trained model**. It is a **weight-transformed port** of [Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct), created by applying a series of closed-form mathematical transforms (called "Keys") that restructure the architecture without losing the original model's knowledge.
---
## ⚠️ Disclaimers
### Port Notice
**This model is a port of Qwen2.5-Coder-1.5B-Instruct, not an independently trained model.** All of the model's knowledge, capabilities, and limitations come from the original Qwen model. The KeyStack transforms are mathematically lossless or near-lossless at initialization — they restructure the architecture but do not add new knowledge. The original Qwen2.5-Coder-1.5B-Instruct model is licensed under Apache 2.0 by Alibaba/Qwen Team.
### Vibe-Coded Notice
This model and its entire codebase were **vibe-coded via [Devin](https://devin.ai) Desktop** — an AI coding agent by Cognition. No human wrote the transform code, inference engine, or model architecture. The project was directed through natural language prompts and the AI agent implemented everything: weight porting, KeyStack transforms, inference engine, fine-tuning scripts, and this documentation.
### System Support Disclaimer
- **Tested on**: Windows 11, NVIDIA RTX 5070 (12GB VRAM), CUDA 13.1, Python 3.11, PyTorch 2.8
- **GPU**: Requires NVIDIA GPU with ≥8GB VRAM for fast path. CPU-only or <8GB VRAM uses AirLLM layer-streaming (slow but works).
- **OS**: Windows 11 tested. Linux/macOS should work but are untested.
- **torch.compile**: Works on GQA attention. MoE + torch.compile hits a Triton bug on Windows (use without compile for MoE).
- **No warranty**: This is a research project. Do not use in production without thorough testing.
---
## Architecture
| Component | Original (Qwen) | ForgeLM v1 | Transform |
|-----------|----------------|------------|-----------|
| **Attention** | GQA (12 Q heads, 2 KV heads) | MLA (d_c=512) | SVD-based GQA→MLA |
| **FFN** | Dense SwiGLU (8960 intermediate) | MoE (4 routed + 1 shared, d_ff=1792) | Weight splitting |
| **Norm** | RMSNorm | RMSNorm (unchanged) | Direct copy |
| **Embedding** | 151936 vocab, 1536 dim | Same (unchanged) | Direct copy |
| **LM Head** | Tied with embedding | Same (unchanged) | Direct copy |
| **RoPE** | θ=1,000,000 | Same (unchanged) | Direct copy |
### KeyStack Transforms Applied
All transforms are **lossless at initialization** (cosine similarity ≥ 0.9998 vs original Qwen):
| Key | Type | Effect | Cosine Sim |
|-----|------|--------|------------|
| **MLA** | FULL | GQA→MLA via SVD, 4x KV cache compression | 1.0000 (100% energy) |
| **MoE** | FULL | Dense FFN→5 experts, 60% active FLOPs | 1.0000 (exact split) |
| **MRL** | FULL | Matryoshka dimension reordering | 1.0000 (permutation) |
| **QuaRot** | FULL | Hadamard rotation on V/O | 1.0000 (rotation) |
| **ValueResidual** | FULL | V0 stored + 28 gates (init=0) | 1.0000 (no-op at init) |
| **RotorQuant** | FULL | Givens rotation matrices stored | 1.0000 (no-op at init) |
| **MTP** | FULL | 4 prediction heads from LM head | 1.0000 (identity init) |
| **AirLLM** | TRIVIAL | Streamable flag for low-VRAM | N/A (runtime) |
### What Was NOT Applied (and why)
| Technique | Reason |
|-----------|--------|
| **GQA→MQA** | Catastrophic degradation without 5% pretraining compute |
| **Wanda pruning** | Needs calibration data post-MRL |
| **SliceGPT** | Needs calibration data |
| **BitNet** | Needs native ternary training |
| **ShortGPT** | Available as key, not applied to v1 |
---
## Performance
### Quality
- **Cosine similarity vs Qwen2.5-Coder-1.5B**: 0.9998–1.0000 across all layers
- **Generation**: Identical output to original Qwen on all test prompts
- **No quality loss** from any KeyStack transform
### Speed (NVIDIA RTX 5070, 12GB VRAM)
| Configuration | tok/s | KV Compression | Notes |
|---------------|-------|----------------|-------|
| Standard | 21 | 1x | Baseline |
| Hadamard INT4 KV | 21 | 4x | Lossless |
| Streaming KV | 12 | ∞ (sinks+window) | Infinite context |
| RotorQuant KV | 13 | 3.88x | 0.94% error |
| MTP self-spec | 11 | 1x | Correct output, needs batch verify |
### Memory
- **Checkpoint**: 3.4 GB (bf16)
- **GPTQ INT4**: 2.3 GB (63% of original, cos=0.99)
- **VRAM (inference)**: ~6 GB (model + KV cache)
- **Min VRAM**: 4 GB (AirLLM streaming, very slow)
---
## KeyStack System
The KeyStack is a system of **Keys** — closed-form mathematical transforms that convert weights between architectures without training. Each Key implements:
- `forward(data) → weights` — replace training with instant transform
- `reverse(weights) → data` — extract what was learned
- `classify() → KeyClass` — FULL (reversible+composable), PARTIAL (lossy), TRIVIAL (no weights)
### Key Classification
| Class | Criteria | Examples |
|-------|----------|----------|
| **FULL** | Reversible + data→weight + composable | MLA, MoE, MRL, QuaRot, ValueResidual, RotorQuant, MTP |
| **BI** | Both directions, round-trip identity | Embedding, RMSNorm, LM Head, RoPE, CausalMask |
| **PARTIAL** | One direction only (lossy) | GPTQ, ShortGPT, Wanda, SliceGPT |
| **TRIVIAL** | No weights (runtime/formula) | AirLLM, QK-Norm, LogitCap, StreamingLLM |
---
## Runtime Features
ForgeLM v1 includes a unified inference engine (`ForgeEngine`) with pluggable strategies:
### KV Cache Strategies
- `standard` — Basic tensor cache
- `streaming` — StreamingLLM (attention sinks + sliding window, infinite context)
- `snapkv` — Observation-window eviction (8x compression)
- `hadamard_int4` — Block-diagonal Hadamard + INT4 (4x compression, lossless)
- `rotorquant` — Givens rotation + Lloyd-Max quantization (3.88x compression)
- `compressed` — H2O heavy-hitter eviction + KV quantization
- `paged` — vLLM-style paged memory
### Decoding Strategies
- `standard` — Autoregressive token-by-token
- `mtp_selfspec` — MTP self-speculative decoding (draft + verify)
- `speculative` — Draft model speculative decoding
- `medusa` — Medusa multi-head decoding
- `dspark` — DSpark semi-autoregressive decoding
### Acceleration
- `airllm_streaming` — Smart layer-streaming (only when VRAM insufficient)
- `cuda_graph` — CUDA graph capture for decode step
- `torch.compile` — Inductor compilation (1.3-2x on GQA)
- `prefix_cache` — KV cache reuse for repeated prompt prefixes
---
## Usage
### Quick Start
```python
from research.inference.forge_engine import ForgeEngine
# Load ForgeLM v1
engine = ForgeEngine.from_checkpoint(
checkpoint="forgelm_v1.safetensors",
config_name="forgelm_v1",
device="cuda",
)
# Activate with runtime strategies
engine.activate(
kv_cache="hadamard_int4", # 4x KV compression, lossless
decoding="standard",
use_prefix_cache=True,
)
# Generate
output = engine.generate("def fibonacci(n):", max_new_tokens=100)
print(output)
```
### With Streaming KV (infinite context)
```python
engine.activate(
kv_cache="streaming", # 4 sinks + 512 window
decoding="standard",
)
output = engine.generate(long_prompt, max_new_tokens=1000)
```
### With MTP Self-Speculative Decoding
```python
from research.mtp import MTPHead
from safetensors.torch import load_file
# Load fine-tuned MTP head
mtp_state = load_file("forgelm_v1_mtp.safetensors")
mtp_head = MTPHead(d_model=1536, vocab_size=151936, n_predict=4).cuda()
mtp_head.load_state_dict(mtp_state)
engine.model.mtp_head = mtp_head
engine.activate(
kv_cache="standard",
decoding="mtp_selfspec", # Speculative decoding
)
```
### AirLLM Streaming (low VRAM)
```python
engine = ForgeEngine.from_checkpoint(
checkpoint="forgelm_v1.safetensors",
config_name="forgelm_v1",
device="cuda",
)
engine.activate(
kv_cache="standard",
decoding="standard",
acceleration="airllm_streaming", # Auto-detects: streams only if VRAM < model size
)
```
---
## Files
| File | Description |
|------|-------------|
| `forgelm_v1.safetensors` | Model weights (MLA + MoE, 3.4 GB) |
| `config.json` | HF-compatible model config |
| `tokenizer.json` | Qwen2.5 tokenizer (unchanged) |
| `tokenizer_config.json` | Tokenizer config |
| `vocab.json` | Vocabulary |
| `merges.txt` | BPE merges |
| `README.md` | This file |
---
## Technical Details
### MLA Conversion
GQA K/V projections are factored through a shared low-rank bottleneck via SVD:
```
W_KV = [W_K; W_V] → SVD → W_c (compress) + W_KC, W_VC (decompress)
```
With d_c=512 (rank of GQA KV), 100% energy is retained.
### MoE Conversion
Dense SwiGLU FFN is split into 5 experts (4 routed + 1 shared):
```
Dense: Y = W_down(silu(W_gate(X)) * W_up(X)) [intermediate=8960]
MoE: Y = sum_i gate_i * Expert_i(X) + Shared(X)
Each expert: d_ff = 8960 / 5 = 1792
```
With top-4 routing (all experts active), output is identical to dense. Top-2 routing (60% FLOPs) requires router fine-tuning.
### MTP Fine-tune
MTP heads' shared trunk was fine-tuned for 200 steps (122 seconds):
- **Trainable params**: 2.36M (trunk only, heads frozen from LM head)
- **Loss**: 10.6 → 2.0 (81% reduction)
- **Result**: Correct speculative decoding output
---
## Citation
If you use ForgeLM v1, please cite both the original Qwen model and this work:
```bibtex
@misc{forgelm_v1,
title={ForgeLM v1: Training-Free Architectural Port of Qwen2.5-Coder via KeyStack Transforms},
author={ForgeAI Project},
year={2026},
note={Vibe-coded via Devin Desktop (Cognition)}
}
@misc{qwen25_coder,
title={Qwen2.5-Coder Technical Report},
author={Qwen Team},
year={2024},
url={https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct}
}
```
---
## License
Apache 2.0 (inherited from Qwen2.5-Coder-1.5B-Instruct)
---
## Acknowledgments
- **Qwen Team** (Alibaba) for the original Qwen2.5-Coder-1.5B-Instruct model
- **Devin** (Cognition) for the AI coding agent that implemented everything
- **KeyStack research** — the concept of closed-form weight transforms replacing training
|