--- language: - el - en license: apache-2.0 base_model: Qwen/Qwen3.6-27B pipeline_tag: image-text-to-text library_name: transformers tags: - greek - instruction-tuned - lora - non-thinking - qwen3.6 model-index: - name: Sophea-Titan-1 results: - task: type: multiple-choice name: General Greek benchmarks (9) dataset: type: greek-nlu-suite name: General Greek Benchmark Suite metrics: - type: accuracy name: General Greek benchmarks value: 0.7369 - task: type: multiple-choice name: GreekMMLU dataset: type: greekmmlu name: GreekMMLU metrics: - type: accuracy name: Accuracy value: 0.8543 - task: type: multiple-choice name: English Retention (5 benchmarks) dataset: type: english-nlu-suite name: English Retention Suite metrics: - type: accuracy name: English retention value: 0.8783 - task: type: visual-question-answering name: Vision (MMStar, 1500 items) dataset: type: MMStar name: MMStar metrics: - type: accuracy name: Accuracy value: 0.714 ---
KIEFERSA
Sophea-Titan-1
Greek fine-tuned multimodal Qwen3.6-27B — non-thinking chat
**Sophea-Titan-1** is a 27B Greek fine-tuned multimodal chat model built on **Qwen3.6-27B**, with register control and a fixed assistant identity — while retaining English ability and vision. - **Creator:** Kiefer SA - **Base model:** Qwen3.6-27B (multimodal, 64 layers, hidden 5120) - **Languages:** Greek (primary), English (retained) - **Decoding:** **non-thinking** (`enable_thinking=false`) — see recommended sampling under Usage > **Serve non-thinking.** Set `enable_thinking=false` in the chat template. Over an OpenAI-compatible > vLLM endpoint, pass `extra_body={"chat_template_kwargs": {"enable_thinking": false}}`. With > thinking left on, Greek output quality degrades sharply. ## Intended use - General-purpose Greek conversational assistant / chat, with formal ↔ informal register control - Greek-knowledge QA; English retained as a secondary language Not evaluated for safety-critical or legal/medical decisions. Fine-tuned from Qwen3.6-27B. --- ## Evaluation Non-thinking, greedy (temperature 0). **Sophea-Titan-1** in bold. **Models compared:**
column model what it is identity-tuned?
Base Qwen3.6-27B the base Sophea-Titan-1 is built on no
Sophea-Titan-1 this model Qwen3.6-27B + general-purpose Greek conversational SFT yes
Sophea-K1 KIEFERSA/sophea-k1 KIEFERSA production multimodal Greek model (sophea.ai) yes
Qwen3-30B-A3B Qwen/Qwen3-30B-A3B text-only MoE (30B/3B-active), not fine-tuned no
Qwen3.6-35B-A3BQwen/Qwen3.6-35B-A3Bmultimodal MoE (35B / 3B-active), not fine-tunedno
Krikri-8B ilsp/Llama-Krikri-8B-Instruct Llama-3.1-8B Greek LLM (external baseline) no
`n/a` = not applicable (text-only) · `n.t.` = not identity-tuned. ### Headline scorecard
AxisBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
General Greek benchmarks — macro (9)0.71570.73690.72400.63460.69830.5977
English retention — macro (5)0.86310.87830.87020.76310.83620.7423
Vision — MMStar (1500)0.65330.71400.6620n/a0.590n/a
### At a glance **Capability profile** — General Greek benchmarks, English and vision across open-weight models (Sophea-Titan-1 in bold; frontier MoEs GLM-5.2 / MiniMax-M3 shown for reference, no vision): ![Open-source capability profile — General Greek benchmarks / English / Vision](profile.png) **General Greek benchmarks vs. model size** — Sophea-Titan-1 has the strongest general-Greek score of the ~27B open models, approaching far larger frontier MoEs at a fraction of the size: ![General Greek benchmarks vs. model size — open-weight models](size.png) ## Per-benchmark detail Three benchmark families in one table — **General Greek benchmarks (9)**, **English retention (5)**, and **Vision — MMStar (1500)** — on the same five models. `n/a` = not applicable (text-only, no vision).
BenchmarkBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
General Greek benchmarks (9)
greekmmlu0.8490.8540.8550.7530.8380.675
mmlu_greek0.7980.7970.7820.6280.7700.519
hellaswag0.6110.6800.6220.4100.5740.572
medical_mcqa0.3120.3870.3680.2320.3120.287
winogrande0.5900.6260.5990.5590.5640.609
arc_challenge0.9440.9500.9450.8670.9160.690
arc_easy0.9730.9730.9710.9310.9690.832
belebele0.9410.9500.9380.8870.9260.771
truthfulqa0.4230.4150.4370.4460.4160.425
MACRO (9)0.71570.73690.72400.63460.69830.5977
English retention (5)
arc_challenge0.9720.9790.9770.9320.9590.765
arc_easy0.9900.9920.9940.9810.9900.892
hellaswag0.7640.8050.7800.5000.7340.751
mmlu0.8580.8580.8580.7450.8400.608
winogrande0.7320.7580.7420.6580.6580.695
MACRO (5)0.86310.87830.87020.76310.83620.7423
Vision — MMStar (1500)
coarse perception0.7440.7280.728n/a0.704n/a
fine-grained perception0.6360.6080.628n/a0.588n/a
instance reasoning0.7560.7920.792n/a0.740n/a
logical reasoning0.6680.7640.636n/a0.572n/a
math0.4640.7080.500n/a0.372n/a
science & technology0.6520.6840.688n/a0.564n/a
OVERALL0.6530.7140.662n/a0.590n/a
### greekmmlu — per-subject (31)
subjectBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
Accounting0.8640.8640.8590.7610.8480.663
Agriculture0.8530.8700.8530.7210.8190.707
Art0.7480.7560.7700.6210.7570.625
Biology0.8640.8680.8590.7940.8560.672
Chemistry0.8270.7530.7780.6540.7900.519
Civil Engineering0.7930.8270.8190.6780.7850.637
Clinical Knowledge0.8100.8200.7950.6890.7900.686
Computer Networks & Security0.7300.7620.7140.5400.6670.476
Computer Science0.8830.8800.8660.8270.8690.768
Driving Rules0.8370.8200.8290.7560.8110.631
Economics0.9110.9260.9260.8400.8940.670
Education0.8570.8670.8950.7620.9010.687
Electrical Engineering0.8300.8220.8440.6810.8170.538
General Knowledge0.7690.7920.8230.7010.7810.630
Geography0.9430.9640.9550.8700.9500.872
Government and Politics0.9440.9410.9610.9010.9440.899
Greek History0.8710.8770.8780.7110.8840.816
Greek Literature0.7140.7140.6430.4290.5000.500
Greek Mythology0.8700.8570.8450.7440.8490.697
Greek Traditions0.8800.8910.8800.7550.8670.710
Law0.7180.7030.6970.5800.6770.504
Management0.8240.8410.8420.7190.8140.694
Maritime Safety & Rescue0.7030.6690.6620.6080.6820.507
Mathematics0.8970.9240.9040.8440.8570.516
Medicine0.8930.8950.8860.7570.8820.663
Modern Greek Language0.9160.9180.9280.8360.9110.780
Physics0.8380.8510.8510.7690.8380.684
Prehistory1.0001.0000.9680.9840.9840.905
World History0.9500.9500.9000.8500.9500.800
World Religions0.7680.7550.7870.6520.8000.684
OVERALL0.8490.8540.8550.7530.8380.675
## Frontier & other API models (reference) Evaluated via API (thinking disabled where supported): GPT-5.5, Claude Opus 4.8, Gemini 3.5, GLM-5.2, MiniMax-M3. API-model Greek/English benchmarks use letter-answer scoring (local models use log-likelihood). Kimi-K3 added 2026-07-29: kimi-k3 via Moonshot API (temperature=1 forced, thinking cannot be disabled). Inkling-Small added 2026-07-31: thinkingmachines/Inkling-Small-NVFP4 (276B/12B-active open-weight MoE, Apache-2.0), self-hosted via vLLM; thinking on, temperature=1 (same protocol caveat as noted for reasoning-locked APIs). Qwen3.8-Max added 2026-08-04: qwen3.8-max via Qwen Cloud API (thinking disabled, temperature=0 — matches the card protocol).
Axis Sophea-Titan-1 GPT-5.5 Claude-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
General Greek benchmarks macro (9) 0.737 0.909 0.922 0.872 0.810 0.821 0.916 0.881 0.898
English macro (5) 0.878 0.925 0.940 0.911 0.892 0.890 0.947 0.922 0.935
### General Greek benchmarks (9)
benchmark Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
greekmmlu 0.854 0.905 0.900 0.898 0.805 0.837 0.902 0.872 0.898
mmlu_greek 0.797 0.888 0.884 0.874 0.769 0.773 0.932 0.898 0.862
hellaswag 0.680 0.891 0.909 0.816 0.693 0.741 0.790 0.666 0.794
medical_mcqa 0.387 0.928 0.917 0.912 0.787 0.833 0.944 0.926 0.902
winogrande 0.626 0.781 0.813 0.641 0.705 0.662 0.880 0.850 0.860
arc_challenge 0.950 0.968 0.966 0.952 0.904 0.924 0.972 0.960 0.966
arc_easy 0.973 0.984 0.982 0.973 0.949 0.960 0.984 0.972 0.974
belebele 0.950 0.953 0.947 0.934 0.913 0.911 0.956 0.954 0.942
truthfulqa 0.415 0.885 0.977 0.845 0.766 0.749 0.882 0.832 0.888
MACRO 0.7369 0.9094 0.9216 0.8717 0.8102 0.8212 0.9158 0.8811 0.8984
### English — 5 benchmarks
benchmark Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
arc_challenge 0.979 0.976 0.974 0.958 0.968 0.967 0.978 0.968 0.976
arc_easy 0.992 0.993 0.994 0.979 0.989 0.986 0.990 0.982 0.986
hellaswag 0.805 0.918 0.944 0.913 0.858 0.854 0.876 0.806 0.892
mmlu 0.858 0.908 0.910 0.896 0.845 0.852 0.946 0.934 0.914
winogrande 0.758 0.830 0.877 0.811 0.801 0.793 0.946 0.922 0.906
MACRO 0.8783 0.9249 0.9399 0.9113 0.8922 0.8903 0.9472 0.9224 0.9348
### greekmmlu
model Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
greekmmlu 0.854 0.905 0.900 0.898 0.805 0.837 0.902 0.872 0.898
## Quantized variants & other formats Smaller-footprint and alternate-runtime builds of this model: - **GPU / vLLM** (text-only — vision tower not included): [**Sophea-Titan-1-FP8**](https://huggingface.co/KIEFERSA/Sophea-Titan-1-FP8) — lossless, ≈29 GB · [**Sophea-Titan-1-NVFP4**](https://huggingface.co/KIEFERSA/Sophea-Titan-1-NVFP4) — 4-bit, ≈19 GB - **llama.cpp** — [**Sophea-Titan-1-GGUF**](https://huggingface.co/KIEFERSA/Sophea-Titan-1-GGUF) - **Apple Silicon** — [**Sophea-Titan-1-mlx**](https://huggingface.co/KIEFERSA/Sophea-Titan-1-mlx) The trade-off plot below covers the vLLM builds scored on the Greek suite; GGUF / MLX are format conversions. ![Quantization trade-off — VRAM vs. General Greek benchmarks (Sophea-Titan-1 family)](quant_tradeoff.png)

VRAM vs. Greek benchmark score across the Sophea-Titan-1 quantized family (bf16 → FP8 → NVFP4).

## Usage **Serve with vLLM** (OpenAI-compatible; 27B bf16 ≈ 54 GB — add `--tensor-parallel-size N` for multi-GPU): ```bash vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \ --enable-auto-tool-choice --tool-call-parser qwen3_coder \ --reasoning-parser qwen3 ``` For **text-only** serving (skips the vision tower — much faster), add `--language-model-only`. **Serve with vision enabled** — raise the per-request image budget explicitly: ```bash vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \ --limit-mm-per-prompt '{"image": 4}' \ --enable-auto-tool-choice --tool-call-parser qwen3_coder \ --reasoning-parser qwen3 ``` ```python resp = client.chat.completions.create( model="sophea-titan-1", messages=[{"role": "user", "content": [ {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}}, {"type": "text", "text": "Περίγραψε την εικόνα στα ελληνικά."}, ]}], temperature=0, extra_body={"chat_template_kwargs": {"enable_thinking": False}}, # required: non-thinking ) print(resp.choices[0].message.content) ``` **Recommended sampling** (instruct / non-thinking): `temperature=0.7`, `top_p=0.80`, `top_k=20`, `min_p=0.0`, `presence_penalty=1.5`, `repetition_penalty=1.0`. **Client (OpenAI SDK)** — remember non-thinking: ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") resp = client.chat.completions.create( model="sophea-titan-1", messages=[{"role": "user", "content": "Ποια είναι η πρωτεύουσα της Ελλάδας;"}], temperature=0, extra_body={"chat_template_kwargs": {"enable_thinking": False}}, # required: non-thinking ) print(resp.choices[0].message.content) ``` **Transformers** (multimodal — load with the image-text-to-text head to keep vision): ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor proc = AutoProcessor.from_pretrained("KIEFERSA/Sophea-Titan-1", trust_remote_code=True) model = AutoModelForImageTextToText.from_pretrained( "KIEFERSA/Sophea-Titan-1", torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True) messages = [{"role": "user", "content": [{"type": "text", "text": "Ποια είναι η πρωτεύουσα της Ελλάδας;"}]}] text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False) inputs = proc(text=[text], return_tensors="pt").to(model.device) out = model.generate(**inputs, max_new_tokens=256, do_sample=False) print(proc.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0]) ``` ## Speculative decoding (MTP) This model ships the **multi-token-prediction head** — 15 `mtp.*` tensors (~0.85 GB, bf16) in `model-mtp.safetensors`. This is the single-layer draft stack that `config.json` has always declared through `text_config.mtp_num_hidden_layers: 1`. Earlier revisions did **not** contain it. The LoRA merge loaded the base through `AutoModelForCausalLM`, and `transformers` declares `_keys_to_ignore_on_load_unexpected = [r"^mtp.*"]` for this architecture, so the head was discarded at load and never written back — the config advertised a module the weights did not contain. The head has been restored tensor-for-tensor from the base, leaving every other weight untouched; the addition is purely additive, so existing serve commands keep working unchanged. To pin the previous bytes, use `revision="726f410a802dcf7a720dca898da9c7b91943911e"`. **Enable it with vLLM** (≥ 0.23.0): ```bash vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \ --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":1}' ``` > **Provenance.** These are the **base** Qwen3.6-27B MTP weights. The Greek SFT and DPO LoRA never > targeted `mtp.*`, so no fine-tuned draft head exists. This cannot affect output quality: > speculative decoding verifies every drafted token against the main model, so a stale drafter > changes throughput only, never the output distribution. The `-FP8` and `-NVFP4` variants ship the same head and support the same flag; on those builds the MTP Linears are held in bf16 and listed in `quantization_config.ignore`. ## License Inherits the **Qwen3.6-27B** base-model license. Verify base-model terms before use.