Instructions to use RthItalia/Rth-lm-25b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use RthItalia/Rth-lm-25b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: llama cli -hf RthItalia/Rth-lm-25b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: llama cli -hf RthItalia/Rth-lm-25b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: ./llama-cli -hf RthItalia/Rth-lm-25b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf RthItalia/Rth-lm-25b # Run inference directly in the terminal: ./build/bin/llama-cli -hf RthItalia/Rth-lm-25b
Use Docker
docker model run hf.co/RthItalia/Rth-lm-25b
- LM Studio
- Jan
- Ollama
How to use RthItalia/Rth-lm-25b with Ollama:
ollama run hf.co/RthItalia/Rth-lm-25b
- Unsloth Desktop
- Docker Model Runner
How to use RthItalia/Rth-lm-25b with Docker Model Runner:
docker model run hf.co/RthItalia/Rth-lm-25b
- Lemonade
How to use RthItalia/Rth-lm-25b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull RthItalia/Rth-lm-25b
Run and chat with the model
lemonade run user.Rth-lm-25b-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
ollama patches?
Hi,
your ollama changes are somewhat obscure. Could you please simply push your version to github?
Hi Matthias, Thanks for the flag! You were right—the custom kernels were missing.
I've just pushed the full C++ implementation of the Fractal TCN operators (
rth_tcn_ops.cpp
) and the GGUF converter to our GitHub. You can find the patches and build instructions here: https://github.com/rthgit/ZetaGrid
Since this is a novel architecture (pure gated convolutions, no attention), your feedback would be incredibly valuable! If you manage to build it or have any thoughts on the kernels, please let me know here or open an issue on GitHub. (And if you like the project, a Star on the repo is always appreciated! ⭐)
Cheers!
Thanks. Will look into it soon(ish).