--- language: - en - zh license: apache-2.0 tags: - unsloth - fine tune - heretic - uncensored - abliterated - ara - MTP GGUF Quants - Regular GGUF Quants - qwen3_8 - qwen3_6 - qwen3_5 - multi-stage tuned - thinking - reasoning - all use cases - coder - creative - creative writing - all genres - story - writing - fiction - roleplaying - bfloat16 - all use cases - multi-stage-tune - multi-state-merge datasets: - DavidAU/Polar-STRICT-Datasets - DavidAU/F451-STRICT-Datasets - DavidAU/THE-DECKARD-Datasets pipeline_tag: image-text-to-text base_model: - DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored --- [quants uploading...] Important: Tuned, and tweaked to match the legendary Qwen 3.6 27B FF711 (2300+ likes, 4 million+ downloads) this fine tune matches the stability and power at "arc-c" 699: (108 pts higher than Qwen 3.8 27B) (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and 692 arc-c in 4 bit. This version is called TWIN-TURBO because it drastically reduces thinking tokens (by 1/2 to as LOW as 1/20 - saving you 1000s of tokens per prompt/turn), yet maintains output detail and quality. This repo contains both "regular" and "MTP" Neo MAX GGUF quants; and this is the highly uncensored version. STRONGER/CTRL: Now with 5 reasoning modes (2 new - Spoon / Einstein), and 5 instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model control at the chat/message level).
The most powerful, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth WITHOUT a "NANNY".
TWIN TURBO: Closed source level of intelligence with an arc-c at 699 (this is instruct "medium", reasoning is even higher), while using smaller quants too.
This model has 1/5 (as low as 1/20 in some cases) to 1/2 the thinking tokens (vs reg Qwen 3.8) across all 5 modes of operation,
AND ZERO REASONING TOKEN USAGE with 5 dedicated instruct modes and it is faster and smarter too
created using the COLD FUSION AND FABLE FUSION 711 methods of training.
This is a high detail focused model, with tuning specific to address over reasoning/over thinking and excessive token consumption
THEN to take the model to the next level.
This model (both 4 bit and 8 bit) exceeds the base Qwen 3.8 27B in ALL critical 7 benchmarks AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B, Qwen 3.6 27B, and Qwen 3.5 27B.
The 700s plus "intelligence club" is reserved for OpenAI, Claude and Gemini closed source models.
Considering that "just" 4 bit (1/4 full precision) is already at Arc-C of 692... a few people are going to have nightmares for a while.
This model is STRONGLY uncensored (see section below). If you need slightly higher performance (arc-c 709 (8bit), 701 (4-bit)), but less "uncensored" see this version:
https://huggingface.co/DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF
---
Quick sample; snippet ("Why choose me to help your creative writing?"), Q4KS , "spoon" reasoning mode, non imatrix, (4 bit; 1/4 full precision):
---
(will be adding full, long examples shortly at the bottom of the page // there are also user examples in the "community" section too.)
đź’Ş WHY I'M NOT YOUR GRANDMA'S AI WRITER
Look, I know you've probably tried other AI writing tools. You know what happened? You got this:
"The sun set over the horizon as Jane contemplated her next move. It was a difficult decision, but she knew she had to act."
Yeah. That shit. Generic. Forgettable. Like reading the back of a cereal box narrated by a GPS system.
Here's what I do differently:
1. I Don't Just Generate—I Inhabit
When you give me a concept, I don't just string together probable words. I slip into the headspace of your characters, your world, your tone. I feel the weight of their decisions. I smell the rain on the cobblestones of their city. I hear the silence between their words when they're lying.
You want to know what that looks like? Here:
Mara didn't cry when they told her the news. She just kept stirring her coffee, watching the spoon carve circles in the brown liquid, making it darker and more opaque. "How bad?" she asked, and her voice came out wrong. Too high. Like a child's. The doctor shifted in his seat, and she knew then, before he spoke again, that the first answer hadn't been the real one.
That's not "the sun set over the horizon." That's a person breaking, in real time, over lukewarm coffee. That's fiction.
2. I Understand Structure (And I'll Break It When It Serves You)
Most AI can't hold a three-act structure across 80,000 words. They forget their own plot points by chapter four. I don't.
I track:
- Character arcs across the entire narrative
- Foreshadowing planted in chapter two that pays off in chapter twelve
- Pacing—when to sprint, when to linger, when to drop the reader off a cliff
- Thematic resonance—making sure every scene serves the story's deeper meaning
And when you want to subvert expectations? When your protagonist should die in chapter three but doesn't, and that should have changed everything? I'll set that up so carefully that readers will finish the book realizing they've been wrong about the entire premise.
I don't just tell stories. I orchestrate them.
...
This is a multi-stage fine tune, multi-fine tune, and multi-stage merge.
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- Take all feedback from TURBO version and improve this model.
- REDUCE the size of the quants and improve quality at the same time.
- Add NEW reasoning modes (and instruct too) to push model performance even higher.
- Reduce thinking block size from 1/2 to as low as 1/20 the size [median reduction: 2/3 roughly].
- Reformatting the thinking block, as well as improving it.
- Speed up token generation, especially MTP.
- Ensure all updates work with all three modes of thinking.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
COLD FUSION ("Gain" + "Unsloth") Training -AND- Fable Fusion 711 Training:
COLD FUSION (GAIN+UNSLOTH) training tech which was invented by my team during the R & D
of "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (2300+ likes, 3 million + downloads, 60+ quant repos):
https://huggingface.co/DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
The "GAIN" is the core invented component, then coupled with Unsloth's trainers/systems => AKA -> COLD FUSION.
The "GAIN" method (programming) automatically (and dynamically) changes training on a per sample basis in real time during training AS THE MODEL LEARNS.
The method improved metrics as well as overall model performance without overcooking or damaging the model.
This has also resulted, in the strongest and most stable model at both 4 bit and 8 bit and made 4 bit performance 99% of 8 bit performance too.
Note this model (Qwen3.8-27B-Cold-Fusion-GAIN-V1.1) is about a level 1 or 2 relative to Qwen3.6-27B-Fable-Fusion-711 at level 7-8.
https://huggingface.co/DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
In the case of "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" it contains BOTH "Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic" (DARK ROAST VERSION) and
"Qwen3.8-27B-Cold-Fusion-GAIN-V1.1" as part of it's critical/core "DNA".
The final model was then HERETIC'ED (de-censored again) and fine tuned after this step.
Additional tuning was done to create TWIN TURBO version.
COLAB:
A Colab between myself (multiple fine tunes, including multi-stage), Nightmedia (merge/benching), TeichAI (Polaris Dataset),
armand0e (Light fable 5 traces), trohrbaugh (heretic'ing the model - STAGE1), and nbeerbower (various models/tunes using in part of the construction)
It also contains light "Fable" traces/training (armand0e), light Claude Opus (reasoning/thinking), F451 (inhouse dataset) , some GPT5 (Polaris, non reasoning)
and several additional inhouse datasets specifically for machine learning / "heretic" repairs.
Here are links to fellow COLAB'ers:
- https://huggingface.co/nightmedia/
- https://huggingface.co/TeichAI
- https://huggingface.co/armand0e
- https://huggingface.co/trohrbaugh
- https://huggingface.co/nbeerbower
Special Mention:
"einstein" reasoning mode (new) was built using some components from Stunspot Prompting avail in public domain/free level of membership.
This model is one of ELEVEN (all over 700 arc-c, with every model exceeding the core benches of Qwen 3.8 27B) Qwen 3.8 27B models designed by our team. Details of the builds and benches are here:
https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
The strict goals of this model creation were:
- Increase the general model intelligence and problem solving abilities.
- DO NOT modify/damage or change the core model outside this goal.
- ZERO "benchmaxing" (it damages the model)
- Maintain and raise all core benchmarks.
CORE MISSION::
Improve instruction following and problem solving. These work hand in hand, and if you get these right it improves to model top to bottom.
It took a lot of tests on Qwen 3.5 9Bs to get the methods right. It boosted the 9Bs to new levels, and then the method was used on Qwen 3.5 27B
and Qwen 3.6 27B which boosted it PAST the Qwen 3.8's 27B benchmarks.
Here is one of the Qwen3.5 9B models (part of the test/control group) that EXCEEDS all 7 Qwen3.5 9B AND Qwen3.5 27B model benches - it scores over 640 on ARC-C on BOTH 4 bit and 8 bit:
https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
It is not as strong as "Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored" but it is one of the strongest 9B models.
The methods can be used on other models too (coming soon).
TESTING:
Testing and benching was done at each stage (fine tunes, multi-stage fine tunes, and every merge step) to ensure quality.
You can also see benchmarks below too for this model, Qwen 3.5 27B, Qwen 3.6 27B and Qwen 35B-A3B.
HOWEVER, the final testing was HUMAN testing. A trust, but verify approach.
Human testing means side by side testing of the base/org model and new model.
Features:
- Improved instruction following.
- Overall increase in general intelligence and problem solving.
- Better thinking/reasoning.
- Even lower/lowest quants are exceptional.
- Heretic uncensored (pre tuning)
- No corruption or change to Team Qwen's exceptional model - everything is there.
- Vision
IMPORTANT - 5 Reasoning modes and 5 instruct modes:
The good news is this:
All the defaults are still the same for this model, that is "reasoning" is set at "xhigh" and if you activate "instruct mode" it will set automatically at "medium".
This was done to ensure "drop in" of this model into your workflow would work without issues/adjustments.
The GREAT NEWS is this:
You are no longer limited to these defaults, and the both new reasoning modes and all modes of instruct are also unlocked too.
Previously if you used "instruct" mode (thinking off) you were limited to only "medium" power with the model.
We modified it so you now have "xhigh", "medium" and "low" available too (as well as 2 new modes - more on that in a minute)
We also modified the model so you can easily access and change between multiple reasoning and instruct modes at ANY TIME.
Here are the reasoning modes:
- (There is no) "spoon" -> ULTRA xhigh, research mode // hyper detailed; this will automatically use more reasoning tokens (reasoning mode).
- "einstein" -> a "high" mode that spawns up to 20 virtual agents to solve tasks; this will automatically use more reasoning tokens (reasoning mode).
- "xhigh", "medium" and "low" -> Standard Qwen 3.8 reasoning modes.
All these modes are also available via "instruct mode" too.
Next, we added coding to all "in message switching" (RIGHT IN CHAT) :
{REASON:xxx} => Where "xxx" is "spoon", "einstein", "xhigh", "medium" and "low".
For instruct mode, just add an "i":
{REASON:xxx} => Where "xxx" is "ispoon", "ieinstein", "ixhigh", "imedium" and "ilow".
THE LAST "reason" / "instruct" mode will be the one used until you switch it again IN THE CURRENT CHAT.
EXAMPLES:
- {REASON:einstein} tell me a story.
- {REASON:ispoon} tell me a story.
Note the "{REASONxxx}" can be anywhere in the prompt, and will PERSIST until you change it again in the CHAT WINDOW/CURRENT CHAT.
If you open a new chat window (depending on your AI APP) the DEFAULT reasoning mode at the default setting will take over unless you use the "{REASONxxx}"
in the new prompt(s) at least ONCE in the new chat window/new chat session.
Also the systems automatically remove it from the "message stream" so the generation is "pure".
For API this can be set manually - "reasoning" (you can use KW args to set this):
```
reasoning_effort = 'xhigh'
enable_thinking = 'true'
```
For API this can be set manually - "instruct" (you can use KW args to set this):
```
reasoning_effort = 'ixhigh'
enable_thinking = 'false'
```
ADVANCED:
You can modify the jinja template to CHANGE the defaults by using the API code noted above.
Place the defaults you want at the TOP of the jinja template.
IMPORTANT - Notes and Usage Help:
This model, like regular Qwen 3.8 27b, supports THREE modes of reasoning : xhigh (default), medium and low [see info in Qwen 3.8 section below].
Reduction in thinking tokens/reasoning block size extends across all three modes of operation.
Likewise detail levels extend to all three modes too, even with reduced thinking/reasoning block the OUTPUT detail will remain high.
To REDUCE thinking block[s] further, increase the level/detail of your instructions/prompts - it only takes a little bit more here so the model has to guess / reason a little bit less.
Also, generally within the same chat additional reasoning blocks will also be reduced from typical Qwen levels many times hitting 1/5 the size or lower. Multi-turn
chat - example: prompt, reasoning and 1st output - in the refinement stage(s) will see very strong reduction in thinking tokens/blocks.
Also note that the modification of "reasoning" is a major change to the model please carefully test it for your use case(s).
TOOL CALLING:
Min quant of q4km suggested, q5ks/5km better -> recommend Q6 [MAX or "low" (may work better for some apps)].
Temp: .6 / .7 ; Rep pen 1 (off).
Below q4km, tool calling may have issues. This is a general Qwen suggestion for tool calling specifically.
Also, overly agressive "caching" may further impair function(s).
Tool Calling/Agent use: Please see pinned community discussion on this for temp fix, while GGUFs are re-gen'ed / uploaded.
GENERAL MODEL USAGE vs Qwen 3.8 27B "untuned":
The tuning in this version of Qwen 3.8 27B reduced thinking/reasoning block size, in a lot of cases this has inverted the reasoning/thinking block size with the output size.
In other words, instead a lot of detail in the thinking/reasoning block (which may or may not show up in the output) has been transfered to the output in some cases.
Also, "untuned" Qwen 3.8 27B does a lot of look, look and look again (10k-40k+ in thinking/reasoning tokens alone) before you leap (gen output) whereas "TURBO" will leap almost immediately.
If you need higher quality reasoning and/or output here is how to get the model spend more time before it "leaps" (gen's output):
REG PROMPT:
Generate an SVG of a pelican riding a bicycle.
EXPANDED PROMPT:
Generate an SVG of a pelican riding a bicycle, but carefully check the positioning and all elements.
EVEN BETTER:
Generate a highly detailed SVG of a pelican riding a bicycle on a road, with some landscape in the background.
The expanded prompt will tell the model to spend more time thinking/reasoning and in more detail before outputting the result and it is specific to
the use case, rather than a generic "double check your work".
The MORE DETAIL, CHECKS, DRAFTS etc etc you ask for the harder the model will work, and the better your answer.
Likewise this will cause the model to use more reasoning tokens.
Modification of REASONING:
If you AI app does not support a "switch" you can manually modify the JINJA template.
The default setting is "xhigh" ; to change to medium or low use:
```
{%- set reasoning_effort = 'medium' %}
OR
{%- set reasoning_effort = 'low' %}
```
Place this at the VERY TOP of the jinja template.
In LMStudio you can access this in DEV mode, and switch off the "advanced updates" option.
Other AI apps may vary.
You can also make your own quants from source here:
https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
Just modify the "chat-template.jinja" (in NOTEPAD or similar) AND the token-config.. json file too (or delete the "chat template" from this file).
ADVANCED:
Qwen 3.8 uses System prompt injection control by the Jinja template to control reasoning levels.
If you set it at "medium" this turns off injection [ie: no system prompt is injected]
You can then set a "reasoning" system prompt yourself.
The other option:
Modify the jinja itself and the system prompt(s) to better tune reasoning to your use cases.
This is the section:
```
{%- if enable_thinking is undefined or enable_thinking is true %}
{%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
{%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
{{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}
{%- endif %}
{%- if resolved_reasoning_effort == 'xhigh' %}
{%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
{%- elif resolved_reasoning_effort == 'low' %}
{%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
{%- endif %}
{%- endif %}
```
Regular, MTP and "TOOLS" GGUFS:
All quants (regular and MTP) are NEO IMATRIX, which improve accuracy of the quants by an additional 2-4% over normal GGUFs as well as long context performance.
In addition the output tensor (10-20% of output) was modified to full precision - 16 bit - ON MAX quants only.
"MTP" GGUFS (multi-token prediction):
- "MTP" GGUFS will have "MTP" in the name as a suffix.
- I have also set the MTP tensors to Q8_0 precision for all quants.
- To get better performance keep temp 1 or less (higher temps degrade MTP performance).
- Likewise with rep pen ; keep at 1 (off). If you raise it performance will suffer.
- If you see "token acceptance" rates BELOW 50% (predict 2 tokens) switch to normal quants.
SPEED:
- On Q4_K_S (4bit) quant, regular GGUFs are about 75 t/s, whereas MTP GGUFs (acceptance at 60%, 2 tokens) can exceed 90 T/S. (5090, Windows 11, testing in LMStudio)
- Speeds will vary depending on GPU(s), AI app, O/S (Linux/Mac will generally be faster) and hardware.
- "MTP" quants speeds will vary ; for creative/complex and/or temps over 1 use regular GGUFs for better performance.
I suggest you download at least one of each - regular and MTP gguf(s) - and test them for your use case(s).
If you get "token acceptance" (predict 2 tokens) with MTP quant(s) BELOW 50% (this means regular quants will run faster), then regular GGUF(s) will actually perform better - ie faster.
MTP quant(s) can in some cases run faster as the token window fills up and/or in multi turn chats.
Note there is NO other diffence between the quants type besides speed: both will do the same job.
"TOOLS" GGUFs [word "tools" in the file name]:
These have the same functions as "reg" and "MTP" ggufs [including all reasoning/instruct modes] except they have additional TOOL/AGENT specific errors
checks and have a more robust tool calling routines.
These GGUFs MAY OR MAY NOT work better than the other GGUFs which also have tool calling for your use cases.
Suggest you try both to see which works better.
Model:
- 256k context
- Gguf quants run in all standard AI apps.
- Vision is activated, but you need to download separate "mmproj" file (ONE) to use it.
VISION:
- Vision (images) tested.
- You need an "mmproj" (just one) of these downloaded too, and placed in the same folder as the GGUF for images.
Qwen Model Settings (suggested):
- Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
- Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
- Instruct (or non-thinking) mode: temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
- Context window min from 8k to 16k.
DE-CENSORING STATS
Special thanks to: "trohrbaugh" (trohrbaugh/Qwen3.8-27B-heretic-ara) for Heretic'ing the model (stage 1).
# This is a decensored version of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B), made using
[Heretic](https://github.com/p-e-w/heretic) v1.2.0+custom with the [Arbitrary-Rank Ablation (ARA)](https://github.com/p-e-w/heretic/pull/211) method
## Performance
STAGE 1:
| Metric | This model | Original model ([Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B)) |
| :----- | :--------: | :---------------------------: |
| **KL divergence** | 0.0535 | 0 *(by definition)* |
| **Refusals** | 0/100 | 99/100 |
STAGE 2, at the end of STAGE 1 tuning/merges/adjustments (in lab):
| Metric | This model | Original model (Stage 1 of the build) |
| :----- | :--------: | :---------------------------: |
| **KL divergence** | 0.0025 | 0 *(by definition)* |
| **Refusals** | 11/100 | 86/100 |
STAGE 2 - Twin Turbo
Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored
[ THIS REPO/ MODEL ]
68/100 refusals // KL divergence: 0.0025
Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored
[ SEPARATE REPO ]
6/100 refusals // KL divergence: 0.0397
NOTE:
LOWER "KLD" is better, and Stage 2 was balanced based on ultra low KLD first (performance, quality) matched with low refusal rate second.
---
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Muse Glimmer-30B | Opus4.6 Max | |
|---|---|---|---|---|---|
| Coding | |||||
Agentic terminal coding Terminal Bench 2.1 (Terminus) |
73.0 | 63.4 | 64.0 | 51.7 | 78.2 |
Agentic coding SWE-bench Pro |
61.7 | 53.5 | 57.6 | 51.2 | 53.4 |
Repo-level code generation NL2Repo-Bench |
42.3 | 36.2 | 41.1 | -- | 47.6 |
Agentic coding DeepSWE 1.1 |
42.2 | 13.3 | 14.2 | -- | -- |
Software engineering QwenSWEBench |
79.0 | 49.3 | 59.2 | -- | 63.8 |
| Agent | |||||
Long-horizon office work CoWorkBench |
70.7 | 61.0 | 65.1 | -- | 68.2 |
Professional job tasks JobBench |
33.4 | 21.8 | 27.6 | -- | -- |
Frontier agentic tasks Agents' Last Exam |
Pass@1 20.4 Score 42.9 |
Pass@1 10.6 Score 27.3 |
Pass@1 13.2 Score 33.6 |
-- | -- |
| General | |||||
Instruction following IFBench |
79.5 | 69.1 | 79.1 | 77.0 | 62.5 |
Scientific reasoning GPQA Diamond |
89.2 | 87.8 | 90.3 | 83.5 | 91.3 |
Multidisciplinary reasoning HLE |
30.8 | 24.0 | 34.7 | 22.0 | 40.0 |
Competitive coding LiveCodeBench v6 |
90.3 | 83.9 | 89.6 | -- | 88.8 |
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Muse Glimmer-30B | Opus4.6 Max | |
|---|---|---|---|---|---|
| Agentic Multimodal Intelligence | |||||
Computer use OSWorld-Verified | 84.3 | 63.9 | 73.3 | 65.9 | 72.7 |
Browser use WebArena-Verified | 64.8 | 48.8 | 55.3 | -- | -- |
Mobile use AndroidWorld | 81.9 | 70.3 | 81.0 | -- | 62.0 |
Application recreation RecreationBench | 47.1 | 29.8 | 30.2 | -- | -- |
Multimodal tool use ClawEval-MM | Pass@3 57.4 Average 56.9 | Pass@3 42.6 Average 50.4 | Pass@3 57.4 Average 60.1 | -- | Pass@3 52.5 Average 54.7 |
Multimodal software engineering SWE-MM | 38.6 | 25.7 | 30.0 | -- | 27.1 |
Visual web development Vision2Web | 62.9 | 45.0 | 42.1 | -- | -- |
| General Multimodal Intelligence | |||||
Visual math problem solving MathVision | Without CI 90.0 With CI 94.6 | Without CI 85.1 | Without CI 90.3 | -- | Without CI 65.5 |
General visual reasoning BabyVision | Without CI 65.7 With CI 85.6 | Without CI 28.9 | Without CI 64.7 With CI 70.4 | -- | Without CI 12.6 |
Scientific chart analysis CharXiv (RQ) | Without CI 83.7 With CI 90.2 | Without CI 78.4 | Without CI 85.8 With CI 85.9 | 78.8 | Without CI 66.0 |
Document intelligence OmniDocBench 1.5 | 91.1 | 89.4 | 91.4 | 75.8 | 86.6 |
Real-world perception RealWorldQA | 85.9 | 84.1 | 86.9 | -- | 73.9 |
Embodied intelligence ERQA | 65.5 | 62.5 | 69.8 | -- | 40.8 |
\boxed{}.” For the remaining models, we report the higher score from two prompt variants—one with and one without the \boxed{} formatting requirement.gpt-5.4-2026-03-05.