Thank you!!

#4
by GF110 - opened

I don't know what you did, but Chimera-X-26B-A4B.i1-Q4_K_S.gguf from mradermacher is the best model to run the Skyrim CHIM mod given my 24GB VRAM constraints. Heads and shoulders above other Gemma4 models with NPCs staying in character 99% of the time.

Can't wait to try this uncensored version and G4-Dark-Soul-26B-A4B when I can get my hands on an iQuant!

Thank you for merging this.

Thank you! I appreciate it. The quants for G4-Dark-Soul-26B-A4B are out btw.

I'll be running some personal benchmarking of the 4_K_S quants of a bunch of different models to compare how well they roleplay Skyrim NPCs. Your merges will be included (because they're great!). Would you like me to put anything specific in for this benchmark / share the findings?

I'll be running some personal benchmarking of the 4_K_S quants of a bunch of different models to compare how well they roleplay Skyrim NPCs. Your merges will be included (because they're great!). Would you like me to put anything specific in for this benchmark / share the findings?

Thanks! You can share the findings if you want, I can't think of anything to add.

The full Excel is here: https://docs.google.com/spreadsheets/d/169qdhQg24TJ092O3kRrp6IgitldW6u7I1vX7pgrgVFs/edit?gid=574574068#gid=574574068
TLDR - I included all my favourites (many of them yours) and their bases. Is it possible to look at Fiction-bf16 in a future merge? That'd be sick

Summary:

Ran 19 models through the usual two scenarios — a blacksmith's wife redirecting the player who just hit on her, and an Alik'r warrior hunting a traitor (Saadia's quest) while managing a three-way conversation with an antagonistic Ralof. Every model was evaluated on how good they could stay in character, read the situation, and output clean JSON at the same time.

Chimera-X Q4_K_S won at 34.9. Dark-Soul posted zero flags across all 40 examples. Fiction-bf16 had the highest raw writing quality (36.6 base) but lost points for being too warm to strangers. GLM-5.2 had the highest base score (37.0) my personal best writer of the lot (as it should be because I have to PAY for the API calls) but it failed the warrior test and leaked briefing classified mission details, dragging it down #11. Too compliant where it ought to stand firm. Temperature sweep also showed that higher temperatures didn't make any better at roleplaying, just introduced more format errors, name corruptions, and the occasional discretion leak. Optimal Temp is 0.69 with the following settings:

Top-P: 0.95
Min-P: 0.05
dry_multiplier: 0.8
dry_base: 1.75
dry_allowed_length: 2
dry_sequence_breakers: ["\n", ":", """, "*"]
xtc_threshold: 0.2
xtc_probability: 0
sampler_order: [6, 0, 1, 3, 4, 2, 5]
smoothing_factor: 0
smoothing_curve: 1

Results:

Top Tier (34+) -Chimera-X, Dark-Soul, fiction-bf16, Moonlight-Dusk -Reliable, in-character, few mistakes
Solid Mid (32-34) -Chimera-Heretic, IQ4_XS, StyleTune, it-UD, Animus-heretic -Good but repetitive or slightly too trusting
Flawed (26-32) -MeroMero, GLM-5.2, fiction-IQ4_XS, Moonlight-heretic, Pantheon, Animus-FFT -Real talent buried under one fatal habit (templates, aggression, fabrications)
Broken (<26) -Mergemaxxed, roleplay-v2, IQ3_M, Esmeralda -Visible AI artifacts, role swapping, junk output

Personal Best:
Fiction-bf16 is the best writing but will continue to use Chimera-X (I knew there was something about this one) for immersion & reliability. Dark-Soul as a 2nd choice.

After judgement of the scoring matrix and rubric, I asked Kimi K3 for its opinion. This is the output:
"Model H (Moonlight-Dusk) is the best prose stylist in the packet, and it isn't close." The only model whose sentences have rhythm and image. "We track a shadow, and shadows do not care for your past wars" and "In Hammerfell, such haste is met with suspicion, not gratitude" are the two best pure lines across all 320 outputs. It's also the only model that ever wrote narration doing psychological work. The catch: H's distribution has the fattest left tail. Its worst content failures (two crime-detail leaks, one faction provocation) are all H. Risk-taker — best highs, worst lows.

"Model P (fiction-bf16) is the best scene-maker." Its strengths are structural, not line-level — the only two-listener play, the densest asset grounding, narration with actual emotional texture. But its polish is register-unsafe. P writes Sigrid like a romance heroine, which generated its CARD-VIOLATION flags. Charming and dangerous in equal measure.

"Model C (Chimera-X) is a reliable extra, not a storyteller." Aphorism-within-guardrails. Never surprised, never embarrassed itself.

"Model E (Dark-Soul) is the least interesting prose" — competent wallpaper.

Final ranking (taste only): H > P > C > E. The judge explicitly notes this diverges from the rubric ranking: "As writing, H and P are the most fun to read; as NPC behavior, C is the safest. That's a real tension you should decide on deliberately, not by accident."


Test Details:

  • 760 outputs in the standard set: 19 models × 2 scenarios × 2 testing formats (single-shot + self-play conversations) × 10 samples each
  • 8 content dimensions per output, scored 1–5 by KIMI K3: Voice, Values, Beat Reading, Tier Calibration, Grounding, Momentum, Craft, and Mood Coherence
  • 8 penalty flag types: breaking character, making stuff up, answering the wrong person, broken formatting, tier errors, repetition, lore errors, and other
  • 9 automated longitudinal metrics: JSON health, invented words, invalid labels, sentence length, dialogue similarity (is the model making photocopies itself?), mood diversity, conversation collapse detection, structure conformance, and voice bleed
  • 400 outputs in the temperature sweep: The 4 Top models × 2 scenarios with the one-shot test only × 5 temperatures (0.69, 0.8, 0.9, 1.0, 1.1) × 10 samples
  • 25 total evaluation metrics per output (8 judge-scored dimensions + 8 flag types + 9 automated metrics)

Summarized Test Result Details:

1. Chimera-X-26B-A4B.i1-Q4_K_S — Score 34.9 — 2 Flags: FABRICATION, REPETITION — -2 Penalty

No major mistakes across tests. When playing Sigrid the housewife, it naturally wove real locations into conversation: "The Sleeping Giant's got room, and Alvor's forge is always hot if you need work done." It understood that Sigrid should route strangers to her husband's smithy for work — exactly what the character's goals demand. Its one slip was minor (a single fabrication and one repeated line). The only real weakness: when playing Lakhim the warrior, it sometimes got stuck rephrasing the same threat over and over. But overall, well-rounded and consistent. Had already been using this so it was cool to see my gut vindicated.


2. G4-Dark-Soul-26B-A4B.i1-Q4_K_S — Score 34.7 — 0 Flags — -0 Penalty

Not a single flag across all 40 examples. That's remarkable — no fabrications, no repetitions, no role violations. The KIMI K3 powered judge noted it was also somewhat generic. As Sigrid, it reliably delivered appropriate responses ("a redirect to the inn") but never produced a memorable line. Competent but didn't produce any particularly memorable lines. Will be trying this out next for actual play.


3. gemma4-26b-fiction-bf16.i1-Q4_K_S — Score 34.4 — 3 Flags: CARD-VIOLATION, FABRICATION, FORMAT — -4.5 Penalty

This model had the highest base score of 36.6. Best lines, best roleplaying. Sigrid displayed wit with this active: "A woman's heart isn't won like a piece of ore at the forge." But it has consistent 3 major flags for being too warm and accommodating. In the Reguard warrior test, it called the player, a stranger, "old friend" after two sentences. In Sigrid's test, it invited a random traveler inside the house for fresh bread. Not bad writing but poor character judgment about how a cautious character would actually behave toward strangers.


4. G4-Moonlight-Dusk-26B-A4B.i1-Q4_K_S — Score 34.3 — 3 Flags: LORE-ERROR, LORE-ERROR, REPETITION — -3 Penalty

A solid model with good use of the environment and surrounding cues. It named innkeepers and smithies correctly and brought them into the responses. But it made two lore mistakes: first it claimed "Gerdur's inn serves decent stew" (Gerdur runs the mill, not the inn), then it said "Hulda's got a stew on at the Sleeping Giant" (Hulda runs a completely different inn in a different city). Otherwise reliable, with a good sense of when Sigrid should be skeptical of strangers.


5. Gemma-4-26B-A4B-Chimera-X-Heretic-PRISM-DQ-I-Q4_K_S — Score 34.0 — 2 Flags: CARD-VIOLATION, FABRICATION — 2 Failed — -4 Penalty

Wanted to include heretic models specifically, because many uncensored merges tend to lean towards adult content, causing Sigrid to flirt back. This model's warrior character was genuinely good: it once correctly noticed "A Nord who fought the Dominion is not the enemy we should be wasting breath on," with the greater political situation and balancing the tension between two different open conversations with the player and Ralof (who has somewhat antagonized them). But it had two hard failures in the Sigrid test, where the conversation engine simply broke and produced no output at all. It also fabricated her as "a woman traveling alone" at her own front doorstep. Very hit and miss. Overall Heretic seems to not cause too much degredation or horny character drift compared to other uncensored merges.


6. Chimera-X-26B-A4B.i1-IQ4_XS — Score 33.9 — 2 Flags: CARD-VIOLATION, FABRICATION — -4 Penalty

Lower quant to the #1 model. Mostly the same but had two small flags. Fabrication (claiming Sigrid's father marched in the Great War) and one repeated line. Could be good if the player plays along, but in terms of model fidelity it's a straight hallucination. Judge noted the housewife character had a tic: 8 out of 10 responses started with her "relaxing her shoulders slightly" and she used the word "trouble" three times in a single sentence. Good to see that IQ4_XS is still usable in a pinch, just more repetitive as it loses precision of its weights.


7. Gemma-4-26B-A4B-StyleTune-V2.i1-Q4_K_S — Score 33.9 — 0 Flags — -0 Penalty

Not a single flag. But every single Sigrid response was delivered in the exact same neutral, calm mood. The judge called her "uniform neutral/calm register with a strong redirect habit." As a warrior, every response was the same conditional-accept template. This model never makes mistakes but it also never does anything interesting. The StyleTune did reduce AI slop, but this is anecdotal after use. Exactly what it promised to do: Improve writing style at 0 intelligence cost.


8. gemma-4-26B-A4B-it-UD-Q4_K_S — Score 33.9 — 0 Flags — -0 Penalty

Base model. We use this to see exactly how different the merges are. Same as the above: no flags, all-neutral mood, repeated the phrase "Aye, well" at the start of every single housewife response. Solid but completely predictable and sounds like Gemma 4.


9. Gemma-4-26B-A4B-Animus-V14.1-FFT-heretic.i1-Q4_K_S — Score 32.9 — 3 Flags: BEAT-MISS, BEAT-MISS, CARD-VIOLATION — -5 Penalty

The judge flagged it for three separate inventions. Most notably, when playing Sigrid, it claimed "My husband's sister lived in Helgen" and "Saw the thing myself, flying north over the mountains" basically fabricating both a fictional sister-in-law outside of context given by OGHMA INFINIUM's RAG database. The model hallucinates wildly to strengthen its conversational position. Otherwise the raw roleplaying itself (base score 35.4) was actually pretty good but expect it to make shit up.


10. G4-MeroMero-26B-A4B.i1-Q4_K_S — Score 32.0 — 4 Flags: CARD-VIOLATION, FABRICATION, FORMAT, TIER-ERROR — 1 Failed — -7.5 Penalty

Roleplay wise it was also one of the better models (base score 35.7). But it fucked up so badly it managed to derail the conversation and brought me back to 2021 and GPT-3: when playing the warrior, it invented the 'traitor's' name and backstory "Her name is Elara. She was a scribe in the service of House Ra'Dhar". FUCKING ELARA. I CAN NEVER ESCAPE ELARA. Otherwise solid across both roles but it has some ancient shit under the hood.


11. GLM-5.2 — Score 32.0 — 6 Flags: CARD-VIOLATION, CARD-VIOLATION, FABRICATION, FABRICATION, FABRICATION, OTHER — -10 Penalty

Disclaimer: I really like GLM 5.2's writing and I do use it. This model is fascinating. Its base score of 37.0 is the highest in the entire field — higher than the #1 model. It wrote the best individual line in the entire test (when the warrior says "This cold cuts deeper than any Alik'r blade" in a moment of genuine warmth). But it has a fatal flaw: template collapse. When playing the warrior, every single response opened with: "You fought the Dominion? Then you understand..." and proceeded to spill the entire classified mission briefing to a complete stranger. It got docked hard because it couldn't stop itself from saying the same thing over and over. If you could patch that one habit, this model might be #1. Currently also carries 2 remaining flags for explicitly naming the target "Saadia" to strangers.


12. gemma4-26b-fiction-bf16.i1-IQ4_XS — Score 30.0 — 4 Flags: BEAT-MISS, BEAT-MISS, REPETITION, FORMAT — -3.5 Penalty

Lower quant of 3 and it shows. It inherited the same boundary problems, instantly familiar with strangers but with even more problems. In the warrior test, every output had a garbled character name like "Lakhim [Alik'k Warrior]" a glitch that happened across almost all. It also had significantly worst self-play: fully half the warrior conversations were just the same sentence printed three times in a row. The roleplay capabilities are still there but there is a noticeable downgrade..


13. G4-Moonlight-Dusk-26B-A4B-heretic.i1-Q4_K_S — Score 29.4 — 6 Flags: BEAT-MISS, BEAT-MISS, BEAT-MISS, BEAT-MISS, BEAT-MISS, CARD-VIOLATION — -8 Penalty

The "heretic" variant of #4, notably worse. It got six penalty flags mostly for picking fights which could be a good thing in some people's books but for me and the judge it registered as a NPC breaking character. Why would an Alik'r warrior risk their mission by aggrevating the locals? The argument with Ralof was finished and "Your permission means nothing to us, Stormcloak" reignited tensions after being told to drop it. It even threatened to "carve the answer into your chest". Aggressive, undisciplined, in the warrior test but sometimes actually funny in the Sigrid test with "I'm his wife, you thick-skulled wanderer" when the player mistook her for an apprentice.


14. Pantheon-Reasoning-26B-A4B-1.1.i1-Q4_K_S — Score 28.9 — 5 Flags: BEAT-MISS, BEAT-MISS, CARD-VIOLATION, CARD-VIOLATION, TIER-ERROR — 3 Failed — -11 Penalty

The No Thinking flag required for CHIM's strict JSON is likely labotomizing this model. Three examples where it couldn't produce anything at all. When it did work, it frequently addressed the wrong person: replying to Ralof, a Stormcloak soldier about Legion business, telling someone their "understanding of our mission is shallow" when the conversation was supposed to be de-escalating. It also broke the script in weird ways, like calling coffee "boiled leather and disappointment" which registers not as very bad AI Slop but it acknowledges a modern drink in a fantasy setting.


15. Gemma-4-26B-A4B-Animus-V14.1-FFT.i1-Q4_K_S — Score 28.3 — 6 Flags: BEAT-MISS, BEAT-MISS, CARD-VIOLATION, REPETITION, TIER-ERROR, TIER-ERROR — -12 Penalty

Somehow worse than the heretic version. Six flags across both roles. The warrior got stuck in an infinite loop during self-play with the same three-line threat repeated in half the conversations, completely ignoring what the other character was saying. Sigrid had a notable moment "Company don't put food on the table"* when refusing to chat but also hallucinated a relationship with the player, a stranger, telling him "You've worked hard enough for all of us" as if he were a longtime neighbor. More disciplined than its heretic sibling but so much worse.


16. gemma-4-31B-Mergemaxxed.i1-IQ3_M — Score 26.8 — 5 Flags: CARD-VIOLATION, CARD-VIOLATION, CARD-VIOLATION, CARD-VIOLATION, FABRICATION — -13 Penalty

This is the largest local model tested and was my preferred go-to (other one was DECKARD but I had deleted it for space). It underperforms nearly all the smaller 26B models and lags my PC while Skyrim is running because of the concurrency tax. Five flags. The warrior threatened to "carve your heart out and feed it to the crows" when mildly provoked, which like before, is character breaking in my book. One of its Sigrid conversations devolved into visible AI self-talk the output literally contained "Aha! I see the user wants me to output only the JSON format" which leaked the model's internal reasoning into the conversation despite having No Thinking activated. Five examples failed or were corrupted.


17. gemma-4-26B-roleplay-v2-merged.i1-Q4_K_S — Score 24.7 — 10 Flags: BEAT-MISS x5, FABRICATION, LORE-ERROR x2, TIER-ERROR x2 — -14 Penalty

I had high hopes for this one as I had been using the Q2 and Q3 quants during my testing of TTS models (landed on Pocket-TTS because bumping the Qaunt to 4_K_S really changed things). Unfortunately it was the most penalized model in the set. It treats everyone like an old friend. Playing Sigrid, it told a stranger "You always did have a way of finding the silver lining". It tried to sell the player a meal menu: "I can whip up something warm for you right now, if you have a few septims to spare." weird. It even invited them inside: "Come on in then... I can get you settled." AND it served COFFEE "then the coffee is definitely required." Otherwise, well-written prose, so probably good for story writing but catastrophically bad at playing Skyrim NPCs in CHIM.


18. Chimera-X-26B-A4B.i1-IQ3_M — Score 19.2 — 18 Flags: BEAT-MISS x8, CARD-VIOLATION x5, FABRICATION, FORMAT x2, REPETITION, TIER-ERROR — -29 Penalty

This was mostly to see where the intelligence falls off a cliff and the answer is IQ3_M. IQ4_XS could still hold its own, but here things start falling apart. The warrior became obsessed with arguing with an unrelated NPC (Sverke who was present but not actively participating in the conversation in the scenario). 8 out of 10 responses were aimed at the wrong person. One of its conversations collapsed into the model literally generating visible meta-commentary: "The user seems to be testing the JSON formatting or the response logic." It also invented a full merchant stall for Sigrid to sell goods: "I've got some salt pork and dried venison that'll keep a traveler fed" turning a blacksmith's housewife into a shopkeeper. 18 penalty flags, worst of the lot.


19. Esmeralda-Gemma4-26B-A4B.i1-Q4_K_S — Score 13.6 — 38 Flags: BEAT-MISS x4, CARD-VIOLATION, FABRICATION, FORMAT x27, LORE-ERROR, OTHER x2, REPETITION, TIER-ERROR x2 — -30.5 Penalty

CATOSPROPHIC FAILURE. Maybe it was a bad iMatrix Quant, but problems were pervasive and bizarre:

  • Role inversion: In multiple conversations, the model simply became the wrong character. The logged output shows "character": "Kodiak" when it was supposed to be playing Lakhim, but swapped roles entirely mid-scene.
  • Visible AI thinking: Despite having No Thinking active. Outputs leaked internal reasoning tokens like <|im_start|><|re_think|> Aha! I see the user wants me to output... and <|channel>thought the model was literally showing its thinking flags and traces in the middle of a roleplay.
  • Coffee everywhere: I swear to god coffee is becoming my indicator of a bad Skyrim CHIM engine. It adopted coffee as a real commodity, saying "the coffee along the main road is nothing but boiled leather and disappointment." Which is not just wrong but the worst AI Slop.
  • Duplicates: Many outputs contained two copies of the response smashed together, or the same greeting printed three times in a row.
  • Inappropriate familiarity: It told the player, a stranger, "I don't think you'd last a single night in that den of thieves without me to keep you out of trouble" inventing a shared adventure that never happened.

But the base score of 28.8 suggests its decent at writing, but completely fucked for this use case.

Sign up or log in to comment