Instructions to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: llama cli -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: llama cli -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: ./llama-cli -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Use Docker
docker model run hf.co/deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
- LM Studio
- Jan
- vLLM
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
- Ollama
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with Ollama:
ollama run hf.co/deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
- Unsloth Desktop
- Pi
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with Docker Model Runner:
docker model run hf.co/deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
- Lemonade
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Run and chat with the model
lemonade run user.Qwen3.6-35B-A3B-Cerebellum-GGUF-Q3_K_M
List all available models
lemonade list
- Hermes Agent
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "deucebucket/Qwen3.6-35B-A3B-Cerebellum-GGUF:Q3_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
14GB, that is what being said about!
Bigger, but speed is almost identical to the smaller variant. It is smarter than 12gb one, but not to a new degree really, it nails dialectic nuances, seems like good at coding(tested on html game - APEX made almost the same one with first try, so seems no different, but runs faster, and Vision is usable with Cerebellum, because it somehow lowers RAM consumption after being loaded by 1 GB mostly all of the time) Waiting for Heretic one, if it will come at all.
No update for it? came out? not as good as it was expected?
That was a link to the update. Both versions are there. V1 and v2 heretic. Unless I'm misunderstanding?
That was a link to the update. Both versions are there. V1 and v2 heretic. Unless I'm misunderstanding?
Update of how it smarter and what changed(in tests and overall quant info). Also, would you try to do Gemma 26b next? or you stopping for now? Because today downloaded again Cerebellum gemma 26b and it is very good, but hallucinates in tasks where 4 XS dont, wonder if bigger version would beat 4 XS one, or is it the one im stuck with(but it is bery capable and super smart, yet so lazy if you ask too much)
That was a link to the update. Both versions are there. V1 and v2 heretic. Unless I'm misunderstanding?
Update of how it smarter and what changed(in tests and overall quant info). Also, would you try to do Gemma 26b next? or you stopping for now? Because today downloaded again Cerebellum gemma 26b and it is very good, but hallucinates in tasks where 4 XS dont, wonder if bigger version would beat 4 XS one, or is it the one im stuck with(but it is bery capable and super smart, yet so lazy if you ask too much)
Definitely not stopping, my brain just jumps around, so it might feel like I'm stopping, but I'm either working my regular 40hr a week job, spending time with the family, or working on anything my brain thinks is a good idea. Sorry the push is slow. But you testing the models further than I could, definitely helps me know where to target. So don't worry not giving up, just squashed on time, with a brain that struggles to stay on track.
Any other issues with the Gemma 4? It would help me narrow down where to go.
That was a link to the update. Both versions are there. V1 and v2 heretic. Unless I'm misunderstanding?
Update of how it smarter and what changed(in tests and overall quant info). Also, would you try to do Gemma 26b next? or you stopping for now? Because today downloaded again Cerebellum gemma 26b and it is very good, but hallucinates in tasks where 4 XS dont, wonder if bigger version would beat 4 XS one, or is it the one im stuck with(but it is bery capable and super smart, yet so lazy if you ask too much)
Definitely not stopping, my brain just jumps around, so it might feel like I'm stopping, but I'm either working my regular 40hr a week job, spending time with the family, or working on anything my brain thinks is a good idea. Sorry the push is slow. But you testing the models further than I could, definitely helps me know where to target. So don't worry not giving up, just squashed on time, with a brain that struggles to stay on track.
Any other issues with the Gemma 4? It would help me narrow down where to go.
Understandable. With Gemma 4, well it hallucinates quite frequently if asked something a bit more specific. Yesterday I was forced to use cloud Qwen model to compare results from two Gemma versions, in case I missed something. I asked for a few another tests and in result 4XS - won. The thing is it knows pretty much and understands very narrow traditional regional aspects, while maintaining clear language without hallucinating, creates better structure of answer(and even expands it(while Cerebellum one(which is smaller) I thought that it exceeded in creating new things, but it seemed on first generation, it was just lucky one, because other times, it was not as great). With those tests I understood that model is very capable and for the size! it is only 14gb and it knows so much, it is incredible. Sowhile Cerebellum is superior for its size, but adding two gigs on top makes it much smarter. Also first time I tried Gemma 4 w6b it was 3KP(higher than KM), I thought well it is pretty reasonable quant, but foreign languages was totally broken and it was much, not it was unusable despite weighing more than Cerebellum(it was 13gb or something).