How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
# Run inference directly in the terminal:
llama cli -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
# Run inference directly in the terminal:
llama cli -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
# Run inference directly in the terminal:
./llama-cli -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Yingyaeliae/grok-oss-Apollyon-24B-heretic:
Use Docker
docker model run hf.co/Yingyaeliae/grok-oss-Apollyon-24B-heretic:
Quick Links

Grok-OSS Apollyon 24B - GGUF Quantizations

Educational Purpose Notice

This model and its respective quantizations are provided strictly for educational, research, and technical evaluation purposes.

Critical Instructions for Users

  1. Upstream Terms: Users are explicitly required to thoroughly read, understand, and consider the original model's licensing instructions, safety guidelines, and terms of use provided by the creator at c4tdr0ut/grok-oss-Apollyon-24B.
  2. User Responsibility & Liability: By downloading, hosting, or interacting with this model, the user assumes full and sole responsibility for any and all content generated. The quantizer accepts no liability for misuse, harmful outputs, or secondary deployments of this artifact.

Available Quantizations

Optimized for local CPU/GPU inference via llama.cpp. Use Q6_K_M or Q5_K_M for the best performance-to-size balance.

Downloads last month
867
Safetensors
Model size
24B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Yingyaeliae/grok-oss-Apollyon-24B-heretic