MN-Aura-12B-v1 GGUF

A local-friendly GGUF of EldritchLabs/MN-Aura-12B-v1, a 12B Mistral-Nemo creative-writing merge built for fiction, roleplay, scene continuation, dialogue, atmosphere, and long-form storytelling.

This repository currently provides a Q4_K_M quant at about 6.96 GiB, making it practical to run with llama.cpp and other GGUF-compatible frontends without downloading the full BF16 model.

What this model is for

The source model combines several Mistral-Nemo 12B writing and roleplay models using EldritchLabs' AURA merge method. Its model card targets:

  • creative and fiction writing
  • roleplay and character dialogue
  • story and scene continuation
  • plot and subplot generation
  • romance, science fiction, horror, and other genres
  • vivid prose and conversational writing

In my local smoke tests, the Q4 generated complete, coherent prose without looping or broken output. It behaved more like a free-form writer than a precision instruction model: it was happy to elaborate, but it was less reliable when asked to obey exact word counts, exact phrase counts, or other rigid formatting constraints.

So the practical expectation is:

Good fit: open-ended writing, RP, scene continuation, brainstorming, dialogue, descriptive prose.

Less ideal: prompts where exact length, exact wording, or strict structural compliance matters.

Download

File Quant Size
MN-Aura-12B-v1-Q4_K_M.gguf Q4_K_M 6.96 GiB

Run with llama.cpp

llama-cli \
  -m MN-Aura-12B-v1-Q4_K_M.gguf \
  -c 8192 \
  -ngl 999 \
  -cnv

Or with the server:

llama-server \
  -m MN-Aura-12B-v1-Q4_K_M.gguf \
  -c 8192 \
  -ngl 999

Use the source model's chat template carried inside the GGUF.

Prompting tip

This model is better treated as a writer than as a deterministic formatter. Give it a clear scene, characters, tone, point of view, and narrative goal. If you need a hard word limit or exact phrase count, plan to validate or trim the output afterward.


Technical notes

Conversion

  • Source: EldritchLabs/MN-Aura-12B-v1
  • Architecture: Mistral / Mistral-Nemo family
  • Quant: Q4_K_M
  • File size: 7,477,326,720 bytes
  • SHA256: 81f3fdd28506365f546b7af17e6aa9839535d48e00f4ba66d561a282e7cb3393

This is a text-model GGUF conversion. No MTP/NextN mismatch was detected during conversion.

Runtime validation

Before upload, this exact Q4 file passed a real llama-server /v1/chat/completions smoke gate.

Validated behaviors included:

  • successful model loading
  • system/user chat-template separation
  • exact short-response generation
  • normal finish_reason=stop termination
  • basic arithmetic generation
  • no pseudo-role continuation
  • no raw chat-control-token leakage

Generic runtime smoke: PASS.

Writer constraint smoke

I also ran three deliberately strict creative-writing prompts to test instruction adherence rather than prose quality.

Result: 0/3 passed every exact constraint.

However:

  • all 3 generations ended normally
  • none hit the token cap
  • observed 6-gram repetition rate was 0.000 in all 3 samples
  • there was no looping, corruption, or broken chat termination

The misses were mostly exact-constraint failures: exceeding the requested word range, repeating a required phrase too many times, omitting an explicitly requested detail, or slightly violating a requested ending/subtext constraint.

This small smoke test should not be read as a prose-quality benchmark. It is mainly a warning that this model appears more comfortable with free-form creative generation than with rigid instruction compliance.

Source integrity

Source revision used for conversion:

f7e5bc5451d53a6841e0c4259c0b19a4e0a8a455

Source HEAD checked immediately before publish:

f7e5bc5451d53a6841e0c4259c0b19a4e0a8a455

The source weights did not change between conversion and publication.

Credits

All model and merge credit belongs to EldritchLabs and the authors of the models included in the original AURA merge.

This repository only provides the GGUF conversion and local validation notes.

Downloads last month
102
GGUF
Model size
12B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ramgpt/MN-Aura-12B-v1-GGUF

Quantized
(3)
this model