Instructions to use ramgpt/MN-Aura-12B-v1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ramgpt/MN-Aura-12B-v1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ramgpt/MN-Aura-12B-v1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ramgpt/MN-Aura-12B-v1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ramgpt/MN-Aura-12B-v1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
- Ollama
How to use ramgpt/MN-Aura-12B-v1-GGUF with Ollama:
ollama run hf.co/ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use ramgpt/MN-Aura-12B-v1-GGUF with Docker Model Runner:
docker model run hf.co/ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
- Lemonade
How to use ramgpt/MN-Aura-12B-v1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ramgpt/MN-Aura-12B-v1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.MN-Aura-12B-v1-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
MN-Aura-12B-v1 GGUF
A local-friendly GGUF of EldritchLabs/MN-Aura-12B-v1, a 12B Mistral-Nemo creative-writing merge built for fiction, roleplay, scene continuation, dialogue, atmosphere, and long-form storytelling.
This repository currently provides a Q4_K_M quant at about 6.96 GiB, making it practical to run with llama.cpp and other GGUF-compatible frontends without downloading the full BF16 model.
What this model is for
The source model combines several Mistral-Nemo 12B writing and roleplay models using EldritchLabs' AURA merge method. Its model card targets:
- creative and fiction writing
- roleplay and character dialogue
- story and scene continuation
- plot and subplot generation
- romance, science fiction, horror, and other genres
- vivid prose and conversational writing
In my local smoke tests, the Q4 generated complete, coherent prose without looping or broken output. It behaved more like a free-form writer than a precision instruction model: it was happy to elaborate, but it was less reliable when asked to obey exact word counts, exact phrase counts, or other rigid formatting constraints.
So the practical expectation is:
Good fit: open-ended writing, RP, scene continuation, brainstorming, dialogue, descriptive prose.
Less ideal: prompts where exact length, exact wording, or strict structural compliance matters.
Download
| File | Quant | Size |
|---|---|---|
MN-Aura-12B-v1-Q4_K_M.gguf |
Q4_K_M | 6.96 GiB |
Run with llama.cpp
llama-cli \
-m MN-Aura-12B-v1-Q4_K_M.gguf \
-c 8192 \
-ngl 999 \
-cnv
Or with the server:
llama-server \
-m MN-Aura-12B-v1-Q4_K_M.gguf \
-c 8192 \
-ngl 999
Use the source model's chat template carried inside the GGUF.
Prompting tip
This model is better treated as a writer than as a deterministic formatter. Give it a clear scene, characters, tone, point of view, and narrative goal. If you need a hard word limit or exact phrase count, plan to validate or trim the output afterward.
Technical notes
Conversion
- Source: EldritchLabs/MN-Aura-12B-v1
- Architecture: Mistral / Mistral-Nemo family
- Quant: Q4_K_M
- File size: 7,477,326,720 bytes
- SHA256:
81f3fdd28506365f546b7af17e6aa9839535d48e00f4ba66d561a282e7cb3393
This is a text-model GGUF conversion. No MTP/NextN mismatch was detected during conversion.
Runtime validation
Before upload, this exact Q4 file passed a real llama-server /v1/chat/completions smoke gate.
Validated behaviors included:
- successful model loading
- system/user chat-template separation
- exact short-response generation
- normal
finish_reason=stoptermination - basic arithmetic generation
- no pseudo-role continuation
- no raw chat-control-token leakage
Generic runtime smoke: PASS.
Writer constraint smoke
I also ran three deliberately strict creative-writing prompts to test instruction adherence rather than prose quality.
Result: 0/3 passed every exact constraint.
However:
- all 3 generations ended normally
- none hit the token cap
- observed 6-gram repetition rate was 0.000 in all 3 samples
- there was no looping, corruption, or broken chat termination
The misses were mostly exact-constraint failures: exceeding the requested word range, repeating a required phrase too many times, omitting an explicitly requested detail, or slightly violating a requested ending/subtext constraint.
This small smoke test should not be read as a prose-quality benchmark. It is mainly a warning that this model appears more comfortable with free-form creative generation than with rigid instruction compliance.
Source integrity
Source revision used for conversion:
f7e5bc5451d53a6841e0c4259c0b19a4e0a8a455
Source HEAD checked immediately before publish:
f7e5bc5451d53a6841e0c4259c0b19a4e0a8a455
The source weights did not change between conversion and publication.
Credits
All model and merge credit belongs to EldritchLabs and the authors of the models included in the original AURA merge.
This repository only provides the GGUF conversion and local validation notes.
- Downloads last month
- 102
4-bit
Model tree for ramgpt/MN-Aura-12B-v1-GGUF
Base model
EldritchLabs/MN-Aura-12B-v1