I like 1.1 better.

#1
by Topsy1 - opened

This was really good for a short time, but then i realized that is puts way too much unnecessary description in every response and uses em dashes everywhere when it could just use commas. The one thing i like better about this one is it realizes the intent in my responses better. ill lie about something and the narration would know exactly why. Other than that, i like 1.1 better because it has less fluff.

Edit i should add that im testing this on my own front end im building, not sillytavern. so maybe others could have a different experience.

Edit2 Lastly ill add that 1.1 is the best tune ive seen of gemma31b so far, so keep up the good work please 👍

I absolutely hate having to say this, but I do agree. I feel as though something is a little off with v1.2. It is also not following very basic system prompt information even within the first few thousand context. My gut instinct is that something might be wrong; I'll try a different quantization and see if it's any different.

I've been getting wildly positive reviews with v1.2 / v1q in my Discord server. Do you guys use reasoning or no? What kind of creative writing do you do? Functionally and thematically?

image

We might have a template mismatch with reasoning enabled. Tune lacks newlines for some important parts. If you could, try v1.2 with newlines after <|think|> and <|channel|>?

I've been getting wildly positive reviews with v1.2 / v1q in my Discord server. Do you guys use reasoning or no? What kind of creative writing do you do? Functionally and thematically?

Reasoning has been disabled for the Artemis models (and actually most G4 31b) from my end. For me, they are roleplay scenarios in a custom front-end, very similar to ST but more simple. Most of mine are D&D (fantasy style) or modern-world stuff. My setup is basic: System prompt + scenario and then chat from there. Nothing new from me; been using the same setup for many fine tunes including Artemis 1.1.

I am all up for testing anything you suggest. Including a different temp / top-p / top-k / min-p, etc. sampling set. I will hop on the discord and see if I can spot anything I might be doing wrong, and report back here if I find a solution. Thanks for your work, btw!

I am not sure what could be causing the issue here. A lot of people know I am picky as hell with my roleplay models. I haven't been particularly aligned with Drummer's models in the past, but I haven't had this much fun for roleplay in probably over a year. I have daily driven G4 and over 2 dozen tunes for over half a year now, and I cannot put into words how much above any of the others Artemis 1.2 has been.

I hope its just a template mismatch or some other issue.

My sampling settings:
Temp: 1
Min-P: 0.025

Template:
G4 template modded to remove <bos><|turn>system

System prompt:
Writing style: casual, natural, organic
Formatting constraints: actions in asterisk, spoken word in quotes, short single paragraph responses.
Avoid: snarky retorts, overly clever wording, authorial tone, formality, grammatical perfection.

Remember to keep responses to a singular short paragraph.

Getting long, way too long replies here too.

It's mostly roleplaying I'm doing, using Temp: 0.9, everything else untouched, and thinking enabled (or it gets pretty terrible at coherency)

Artemis 1.1 would reply in about 2-4 paragraphs to a message. The same kinds of turns on 1.2 tend to yield more like 5-6, with some outliers going full novel mode: the character starts effectively monologuing through an entire made-up sequence without giving "me" any room to interject, complete with long pauses and "letting the words hang in the air" (but that's just Gemma slop you can't fully get rid of it with finetuning I think :p)

Which is a terrible shame, because functionally it writes extremely well. Intents are indeed better understood, coherency is good if using thinking, and characters stay true to their descriptions and sound like actual people, until they start monologuing, at which point they suddenly sound like dramatic actors pondering the weight of existence.

I suspect this could probably be prompted out or at least mitigated, though. I'll update if I manage to find something that works.

The main pattern I'm noticing is that once 1.2 starts writing in longer form, everything tends to drift toward this very cinematic, sensory experience style. More atmosphere, more micro gestures, more environmental detail, more dramatic pacing, even when the actual interaction doesn't really need it.

@Lightwulf7171 Try "write 3 paragraphs" or "10 sentences". It should listen to instructions like that.

I haven't tried 1.1, but this 1.2 version still has the same ultra annoying slop problem that all gemma-4 models share:

- A new Shura demon arised!
- Oh no... This is real
- Your arm has been replaced by a shinobi prothesis. How does it feel?
- It feels... real
- If I can't be a shinobi, what can I be ?
- Be real

@Lightwulf7171 Try "write 3 paragraphs" or "10 sentences". It should listen to instructions like that.

So, I've tried it, it somewhat obeys, but restricting to a strict number means every reply will be 3 paragraphs for instance. I did get some promising results with basically this at the end of my prompt

REPLY LENGTH
End the reply at the first natural point where {{user}} could speak or act. Size it to the input: a single line usually warrants a paragraph or two. Characters talk in turns, not monologues.

Unfortunately, I also have to chime in and say I like v1.1 better after extensive testing.

I feel like v1.2 has potential, and I can absolutely see where it was heading in a lot of the conversation. I like it... when it works. But it just gets derailed way too often and starts not following the prompts anymore. Or it randomly throws in "quotation marks" when explicitly told not to. Or it randomly throws in "[random_word_here]" in the middle of a paragraph.Never had this issue with v1.1.

I hope this can be an opportunity to take some of the best of 1.1 and mash it with some of the best of 1.2 to make the ultimate Artemis in 1.3! 🤤

Some more feedback to the pile:
Unlike most people here I prefer 1.2 so far. It does have the problem of using way too many em dashes. It loves unfinished lines that end in —. But in text completion you can just set "—" in banned strings so it's not so bad. Sometimes 1.2 is a little overly wordy but I don't think it's necessarily being wordy in a sloppy way so personally I don't dislike it and response length can be shortened with system prompt as discussed above.

Other than that I think 1.2 is smarter; 1.1 often is more clueless about the intent of my input in the same scenario and also I think 1.2 respects the character personality more while 1.1 more often gives a somewhat more generic answer that fits only the situation and mood but not the character personality.
1.1 also writes stuff that is borderline broken more often. Something that is not completely nonsensical but worded very oddly or just generally a weird thing to say, like me performing a mundane action and my non-human (and non-stupid) companion saying "I didn't know humans could do that." even though it's something anyone with hands could do. Also sometimes getting "you" instead of "{{user}}" even when entire scenario was in third person so far. And I'm using minP 0.2 which is pretty damn conservative.

To be clear I am NOT using thinking as I only have hardware that gets about 6-7 tokens per sec on Gemma4 31B which is a fine enough speed for non-thinking RP but too slow for thinking.

Unrelated but it's pretty crazy how good models that can be run on consumer PCs are now in general. I went back to some Mistral Nemo/Small models that I remember being really good from when I used them actively and they felt really damn dumb. Skyfall 4.2 is pretty interesting though, only ran into it recently. Not as smart as Gemma4 based stuff but still really nice, a good model to have for when you want a change of pace from Gemma style writing.

I use it with thinking on in chat completion. Temp 0.9, top k: 64, top p 0.95. I use it for dnd like rpg.

So if you use it for plain storywritting then you might 1.2 better with the extra descriptions. This also has a different feel i guess compared to base gemma4 or other finetunes i guess, which makes it feel new like when you try a new model.

With my usecase, a single good turn isnt what makes it good. In the rpg i move to different locations alot, unequip and reequip gear alot. Just alot of moving around and when you get three paragraphs of fluff everytime you go somewhere its too much. Or a paragraph describing how you changesd out a weapon and every minute detail about it.

It is possible i just need to update my system prompt since 1.2 is so different. Ill do that and report back.

Sign up or log in to comment