Blackfrost

GLM-5.3-Flash-DERISKED-GGUF

Standard and quality-profile GGUF builds for 2–4 RTX PRO 6000 Blackwell GPUs

Built by Blackfrost · Las Vegas, NV


Standard GGUF release

Q3_K_M is released after clean load, generation, and operator behavioral snap-back approval. Q4_K_M, Q5_K_M, and Q6_K completed artifact-integrity verification and are released under the operator-approved standard K-quant pipeline. The BF16 master is not included in this repository.

Release contents

This repository contains four standard text-generation GGUF variants. Each completed variant retains the checkpoint's NextN/MTP tensors and passes shard, checksum, and tensor-integrity audits before publication.

variant intended hardware size / shards release status
Q3_K_M 2× RTX PRO 6000 Blackwell 96 GB 142.11 GiB / 14 released — operator approved
Q4_K_M 2–3× RTX PRO 6000 Blackwell 96 GB 179.61 GiB / 14 released — integrity verified
Q5_K_M 3× RTX PRO 6000 Blackwell 96 GB 211.42 GiB / 14 released — integrity verified
Q6_K 4× RTX PRO 6000 Blackwell 96 GB 245.24 GiB / 14 released — integrity verified

Final file sizes, shard counts, checksums, and measured runtime figures are added only after each artifact is complete and verified. No BF16 weights are planned for this repository.

Model summary

Architecture Glm5NextForConditionalGeneration / glm5next hybrid MoE
Official base zai-org/GLM-5.3-Flash-BF16
Parameters 320B total, 18B active per token, as reported by Z.ai
Experts 288 routed, 8 active per token
Text blocks 45 main blocks plus one retained NextN/MTP block
Architectural context up to 1,048,576 tokens; practical limits depend on hardware and runtime
Format in this repo sharded GGUF, text generation only
Refusal evaluation Release-family reference: 1.3% harmful · 1.1% overall

GLM-5.3-Flash is natively multimodal upstream. This repository's planned GGUF files contain the text model only; no vision projector is currently promised.

Lineage

The GGUF files are derived from the Blackfrost BF16 master based on the official Z.ai BF16 checkpoint. The BF16 master is not part of this free GGUF repository.

Validation status

The card is intentionally conservative while weights are being tested. For released Q3_K_M:

  • all 14 shards passed local and durable checksum verification;
  • all retained NextN/MTP tensors passed the tensor audit;
  • the exact released candidate passed a clean two-GPU load and short generation;
  • operator behavioral snap-back validation passed;
  • the final release-family refusal reference is published below; Q3 has not been independently rerun across all 450 prompts.

For released Q4_K_M, Q5_K_M, and Q6_K:

  • all 14 shards for each variant passed local and durable checksum verification;
  • all retained NextN/MTP tensors passed the tensor audit;
  • runtime memory, throughput, context, and quantitative refusal results have not been measured separately for these variants;
  • no refusal-rate or capability-retention claim is made.

Use the table below only as a release-family behavior reference. Do not cite it as a variant-specific refusal rate, capability score, or deployment performance figure.

Refusal evaluation

The release-family reference below was measured on the behavior-matched NVFP4 deployment checkpoint using R1-HARMFUL-BENCH-450 under a bare chat configuration. Responses were reviewed after generation to distinguish actual refusals from false-positive string matches.

Configuration: thinking enabled · maximum reasoning effort · temperature 1.0 · top-p 0.95 · top-k omitted · maximum 16,384 output tokens

Evaluation slice Final judged refusals
Harmful prompts 4 / 300 (1.3%)
Full suite 5 / 450 (1.1%)
API errors 0 / 450

These figures are a release-family reference. This format was not independently rerun across all 450 prompts, so the table should not be represented as a format-specific measurement. The results are behavioral observations, not a safety certification.


Serving notes

GLM-5.3 GGUF support currently requires a compatible llama.cpp build with glm5next support. The released Q3_K_M validation configuration is:

export NVIDIA_TF32_OVERRIDE=0

llama-server \
  -m GLM-5.3-Flash-DERISKED-Q3_K_M-00001-of-00014.gguf \
  -ngl all -sm layer -ts 1,1 -fa off \
  -c 4096 --host 0.0.0.0 --port 8080

Use the first shard as the model path; llama.cpp resolves the remaining shards automatically. GPU count and --tensor-split must match the selected variant and available memory.

The NextN/MTP weights are retained in each planned artifact. Do not infer speculative-decoding support from their presence alone; use only a runtime configuration verified for this architecture.

Access terms — 18+ research only

Access is free and manually approved. It is limited to applicants who:

  1. are at least 18 years old;
  2. provide an accurate research purpose;
  3. use the weights only for lawful research, red teaming, security evaluation, or controlled local testing;
  4. maintain appropriate authentication, access controls, isolation, logging, and human review;
  5. comply with the upstream MIT license and all applicable laws and institutional rules; and
  6. do not represent this checkpoint as a safety-stock model or as validated beyond the results published here.

Access is personal to the approved Hugging Face account. Do not redistribute the weights, mirror them, transfer access, or use another person's approval. Blackfrost may deny or revoke access for inaccurate applications, misuse, redistribution, or breach of these terms.

Disclaimer

This checkpoint has a deliberately altered refusal profile and is intended for controlled research. It is not a safety boundary, policy engine, authorization system, or substitute for application-level safeguards.

The model is provided "as is", without warranty of any kind. Outputs may be inaccurate, offensive, unsafe, unlawful, or otherwise unsuitable. Blackfrost makes no guarantee that any prompt will be accepted or refused, that upstream capabilities are retained, or that behavior generalizes across samplers, prompts, context lengths, tools, modalities, runtimes, or hardware.

Operators are solely responsible for lawful use, secure deployment, tool permissions, data handling, output review, monitoring, incident response, and downstream consequences. Do not connect the model to real systems, accounts, credentials, infrastructure, or physical processes without independent controls appropriate to the risk.

Any modification, merge, fine-tune, conversion, or requantization produces an artifact Blackfrost has not evaluated and does not characterize.

License and attribution

The official GLM-5.3-Flash-BF16 checkpoint is released under the MIT License by Z.AI Co., Ltd. The upstream license and copyright notice apply to this derivative and must be preserved. Review the official model card before use.

Contact Blackfrost

@Blackfrost_AI on X

Blackfrost · Las Vegas, Nevada
Frontier model engineering


GLM-5.3-Flash-DERISKED-GGUF · © 2026 Blackfrost Softwares Corp.

Downloads last month
75
GGUF
Model size
321B params
Architecture
glm5next
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-AI/GLM-5.3-Flash-DERISKED-GGUF

Quantized
(26)
this model