Qwen3-1.7B-Distilled-30B-A3B GGUF

GGUF quantizations of reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B — a 1.7B-parameter Qwen3 causal language model distilled from Qwen3-30B-A3B-Instruct using discrepancy-informed proof-weighted knowledge distillation on 6,122 STEM chain-of-thought samples.

Quantized by tinyopsec.


Model Details

Attribute Value
Architecture Qwen3ForCausalLM
Parameters ~2.03B
Base model Qwen/Qwen3-1.7B
Teacher model Qwen/Qwen3-30B-A3B-Instruct-2507
Training precision BF16
Context length 1024 tokens (training)
License Apache 2.0
Developer Convergent Intelligence LLC: Research Division

About the Original Model

Qwen3-1.7B-Distilled-30B-A3B is trained with Discrepancy-Informed Knowledge Distillation (DISC v3) — a method that goes beyond standard KL divergence by treating per-token teacher-student divergence as a structured signal:

  • Discrepancy-Weighted KD — identifies reasoning pivot tokens via local KL jump detection and amplifies their distillation weight
  • DG-Limit Smoothing — stabilizes high-entropy student tokens with neighborhood averaging before KD is applied
  • Gap Energy Regularization — monitors and penalizes structural drift at reasoning transitions independent of mean token loss
  • Proof-Weighted Cross-Entropy — applies 2.5× → 1.5× decaying weight on derivation spans (Proof: to Final Answer:)

The model is particularly suited for STEM reasoning, mathematical derivations, and proof-style explanation.


Available Quantizations

File Bits Size (approx.) Use Case
model_f16.gguf 16 ~3.8 GB Reference, maximum quality
model_q8_0.gguf 8 ~2.0 GB High quality, fast on capable hardware
model_q6_k.gguf 6 ~1.6 GB Near-lossless, good balance
model_q5_k_m.gguf 5 ~1.4 GB Recommended for quality-focused use
model_q5_k_s.gguf 5 ~1.3 GB Slightly smaller Q5 variant
model_q4_k_m.gguf 4 ~1.2 GB Best balance of size and quality
model_q4_k_s.gguf 4 ~1.1 GB Smaller Q4 variant
model_q3_k_l.gguf 3 ~1.0 GB Low memory, acceptable quality
model_q3_k_m.gguf 3 ~0.9 GB Compact, moderate quality loss
model_q3_k_s.gguf 3 ~0.8 GB Minimum size Q3
model_q2_k.gguf 2 ~0.7 GB Extreme compression, lowest quality

VRAM Requirements

Quantization VRAM (approx.)
F16 ~4.0 GB
Q8_0 ~2.2 GB
Q6_K ~1.8 GB
Q5_K_M ~1.6 GB
Q4_K_M ~1.4 GB
Q3_K_M ~1.1 GB
Q2_K ~0.9 GB

Usage

llama.cpp

./llama-cli -m model_q4_k_m.gguf \
  -p "Solve the following problem carefully and show a rigorous derivation.\n\nProblem:\nFind the eigenvalues of [[2,1],[1,2]].\n\nProof:\n" \
  -n 512 --temp 0.0

llama-cpp-python

from llama_cpp import Llama

llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=2048)

output = llm(
    "Solve the following problem carefully and show a rigorous derivation.\n\nProblem:\nProve that sqrt(2) is irrational.\n\nProof:\n",
    max_tokens=512,
    temperature=0.0,
)
print(output["choices"][0]["text"])

LM Studio

  1. Download any .gguf file from this repository
  2. Open LM Studio → Load Model → select the file
  3. Use the prompt format below in the chat or playground

Ollama

ollama run hf.co/tinyopsec/Qwen3-1.7B-Distilled-30B-A3B-GGUF

Prompt Format

For best results, use the training format:

Solve the following problem carefully and show a rigorous derivation.

Problem:
{your problem here}

Proof:

For general instruction use, standard chat template also works:

<|im_start|>user
{your question}<|im_end|>
<|im_start|>assistant

Intended Uses

  • Mathematical derivations and proof-style explanation
  • STEM problem solving (physics, engineering, linear algebra, differential equations)
  • Educational tutoring and worked solutions
  • Lightweight reasoning on edge or CPU-only hardware
  • Generator component in verifier-generator or RAG reasoning pipelines

Limitations

  • Not a formal proof verifier or symbolic algebra engine
  • Can produce fluent but incorrect derivations
  • Training context is 1024 tokens — very long derivations may degrade
  • Domain coverage is uneven: physics and linear algebra are well-represented; molecular biology and physiology are sparse

Original Model

reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B

Full methodology: Structure Over Scale (DOI: 10.57967/hf/8165)

Downloads last month
1,338
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tinyopsec/Qwen3-1.7B-Distilled-30B-A3B-GGUF

Finetuned
Qwen/Qwen3-1.7B
Quantized
(1)
this model