Qwen3.5-Legal (Q5_K_M GGUF)

Fine-tuned Qwen3.5-35B-A3B for legal document anonymization (LDA), legal drafting, and agentic reasoning tasks.

Model Details

Property Value
Base Model Qwen3.5-35B-A3B (MoE, 256 experts, 3B active)
Fine-Tuning 4-bit LoRA via Unsloth
Training Data 117 examples (59 LDA + 23 drafting + 26 reasoning + 9 instruction)
Training Hardware RunPod RTX PRO 6000 (96GB VRAM)
Training Steps 45 (3 epochs), final loss 1.34
Quantization Q5_K_M (~24GB)
Format GGUF (for Ollama, llama.cpp, etc.)

Benchmark Results

Tested on Mac Mini M4 Pro (32GB RAM). Perfect 30/30 across all tests.

Model LDA Simple LDA Complex Resolution Memo Planning Tool Use Total t/s
Base Ollama 5 1* 5 5 5 5 26/30 17.0
Fine-tuned Q4 5 1* 5 5 5 5 26/30 17.4
Fine-tuned Q5 5 5 5 5 5 5 30/30 15.8
Base MLX 5 1* 5 5 5 5 26/30 48.0

* = Hit 4000 token limit during thinking. Fine-tuned Q5 completed within budget.

Usage with Ollama

# Download the GGUF file, then create a Modelfile:
cat > Modelfile << 'EOF'
FROM ./qwen3.5-legal-q5_k_m.gguf
TEMPLATE {{ .Prompt }}
RENDERER qwen3-vl-thinking
PARSER qwen3-vl-thinking
PARAMETER temperature 1
PARAMETER top_k 20
PARAMETER top_p 0.95
PARAMETER num_ctx 8192
EOF

# Register with Ollama
ollama create qwen3.5-legal -f Modelfile

# Test
ollama run qwen3.5-legal "Anonymize: John Smith, CEO of Acme Corp, signed a contract on January 1, 2025."

CRITICAL: The RENDERER qwen3-vl-thinking and PARSER qwen3-vl-thinking directives are required. Without them, Ollama mixes thinking tokens into the content, producing garbled output.

Hardware Requirements

  • Inference: Apple Silicon Mac with 32GB+ RAM (26GB model + OS headroom)
  • Speed: ~15.8 tokens/sec on Mac Mini M4 Pro

What This Model Does Better

The fine-tuning made the model more efficient at legal document anonymization:

  • Less thinking overhead (fewer tokens wasted on chain-of-thought)
  • More direct, structured output
  • Handles complex multi-entity documents (4+ persons, multiple orgs, addresses, emails) that base models run out of token budget on

Use Case: Rule 1.6-Compliant AI Workflow

This model powers the VibeCodingLegalTools workflow — an open-source pipeline that lets lawyers use consumer AI apps (Claude, ChatGPT) without violating their duty of confidentiality.

The model runs locally and anonymizes documents before they reach any cloud AI service. The cloud AI only sees {COMPANY_1} and {PERSON_1}, never real client data.

License

Apache 2.0 (same as base model)

Downloads last month
19
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Reytian/qwen3.5-legal-q5_k_m-gguf

Quantized
(278)
this model

Evaluation results