How to use from
Lemonade
Pull the model
# Download Lemonade from https://lemonade-server.ai/
lemonade pull Baekpica/K2-Horizon-375B-A23B-GGUF:BF16
Run and chat with the model
lemonade run user.K2-Horizon-375B-A23B-GGUF-BF16
List all available models
lemonade list
Quick Links

K2-Horizon-375B-A23B GGUF intermediates

Public, reproducible intermediate GGUF artifacts converted directly from IFM/K2-Horizon-375B-A23B at revision d33e3ae45281865ebf9f044b12d3635b1d1e17fe.

This repository contains both:

  • BF16: the lossless GGUF conversion used as the sole quantization source.
  • Q8_0: the high-precision intermediate generated directly from that BF16 GGUF.

The memory-targeted mixed quant is published separately at Baekpica/K2-Horizon-375B-A23B-Mixed-Quant-GGUF.

Provenance

  • Source parameters: original BF16 checkpoint; the official FP8 checkpoint was deliberately not used, avoiding quantization-on-quantization.
  • Source parameter count: 379,167,159,168.
  • Converter: IFM's K2-Horizon llama.cpp branch, commit 35999d101cf2233fc54f09c3c8d599da7303ce02.
  • BF16 split target: 30G per shard.
  • Q8 policy: all 2-D weights Q8_0; router weights/biases, normalization weights, and other 1-D control tensors remain F32.

Machine-readable structural audits are stored in validation/BF16.audit.json and validation/Q8_0.audit.json. Content hashes are in the corresponding *-SHA256SUMS files; both 30-shard sets were checked against the Hub's LFS object IDs after every upload closed.

Status

  • BF16: complete — 30/30 shards, 758,484,189,216 bytes (706.393448 GiB), 842 tensors (BF16=603, F32=239).
  • BF16 structural audit: passed.
  • BF16 remote inventory: passed — 30/30 objects and aggregate byte count match.
  • Q8_0: complete — 30/30 shards, 403,079,839,776 bytes (375.397354 GiB), with tensor payload 403,068,341,760 bytes (375.386646 GiB).
  • Q8_0 structural audit: passed — 842 tensors (Q8_0=603, F32=239).
  • Q8_0 remote verification: passed — all 30 Hub LFS SHA-256 values and the aggregate byte count match Q8_0-SHA256SUMS.

Each artifact is usable from its 00001-of-00030 file when all 30 numbered siblings are present in the same directory.

Downloads last month
862
GGUF
Model size
379B params
Architecture
k2-horizon
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Baekpica/K2-Horizon-375B-A23B-GGUF

Quantized
(2)
this model