Instructions to use prism-ml/Ternary-Bonsai-2-27B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: llama cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- LM Studio
- Jan
- vLLM
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prism-ml/Ternary-Bonsai-2-27B-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prism-ml/Ternary-Bonsai-2-27B-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Ollama
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Ollama:
ollama run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Unsloth Desktop
- Pi
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Docker Model Runner:
docker model run hf.co/prism-ml/Ternary-Bonsai-2-27B-gguf:F16
- Lemonade
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-gguf-F16
List all available models
lemonade list
- Hermes Agent
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use prism-ml/Ternary-Bonsai-2-27B-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf prism-ml/Ternary-Bonsai-2-27B-gguf:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "prism-ml/Ternary-Bonsai-2-27B-gguf:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ternary-Bonsai-2-27B on an RTX 2060 SUPER 8GB — Japanese supplied-context reasoning is the biggest surprise
I previously experimented with the original Bonsai-27B, so I was particularly interested in what had changed in Bonsai 2.
The memory efficiency is obviously remarkable, but what surprised me more was the improvement in Japanese usability.
With the previous Bonsai-27B, Japanese still felt somewhat experimental to me. With Bonsai-2-27B, the language itself feels much more natural and usable.
However, the most interesting test was not a general-knowledge benchmark.
I gave the model an entirely synthetic Japanese document containing arbitrary names, numbers, updated values, operational rules, exceptions, and cross-references, and instructed it to answer using only the supplied document.
That test was much more convincing to me than simply asking questions about facts the model might already know.
Hardware and runtime configuration
- OS: Debian GNU/Linux 13 (trixie)
- Kernel: 6.12.107+deb13-amd64
- CPU: AMD Ryzen 5 5500, 6C/12T
- RAM: 32GB
- GPU: NVIDIA GeForce RTX 2060 SUPER 8GB
- Driver: 550.163.01
- CUDA: 12.4
- Compute capability: 7.5
- llama.cpp: PrismML fork
- commit:
9a9394a895b96003ca842a6041cb28ac49a108f7 - Model:
Ternary-Bonsai-2-27B-PTQ1_0.gguf - Parameters: 26.90B
- Model size reported by
llama-bench: 5.53 GiB - Quantization: PTQ1_0, 1.75 bpw ternary
Server configuration:
./build/bin/llama-server \
-m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
-ngl 99 \
-fa on \
-c 32768 \
--cache-type-k q5_1 \
--cache-type-v q4_0 \
--parallel 1 \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--jinja
This means:
- full GPU offload
- 32K context
- K cache = Q5_1
- V cache = Q4_0
- Flash Attention enabled
- parallel = 1
- single-user configuration
The complete setup fits within the 8GB card.
The most important test: Japanese supplied-context reasoning
For me, this was the most meaningful test.
Instead of asking the model something it might have memorized during training, I created a fictional Japanese technical document containing information that should be unknown to the model in advance.
The document described a fictional research facility called the Asteria Observatory.
It contained arbitrary information such as:
- staff names
- computer memory capacities
- radio frequencies
- vehicle payload limits
- bridge-specific restrictions
- weather measurements
- battery capacities
- later equipment upgrades
- reduced usable capacity due to degradation
- normal maintenance assignments
- an exceptional emergency assignment
- equipment operating modes
- temperature-dependent operating restrictions
- data-retention rules
- exceptions to those retention rules
- access-control levels
- a two-person authorization rule
The important point is that many questions could not be answered by simply retrieving one sentence.
The model had to correctly handle updates, exceptions, and relationships between multiple statements.
The full Japanese source document and the model's complete Japanese answer can be attached separately to this discussion.
Example 1: Updated value
The document initially stated:
BETA computer: 384GB
Later, it stated that BETA had been upgraded:
BETA was upgraded to 512GB in April 2028.
The question asked for the current value.
Bonsai 2 answered:
512GB(改修により増設)
Correct.
This is simple, but it already tests whether the model can prefer a later update over an earlier value.
Example 2: Rule plus exception
The document stated:
- Red vehicle: maximum payload 3.2 tons
- Blue vehicle: maximum payload 4.1 tons
- Green vehicle: maximum payload 2.7 tons
But it also stated:
When crossing bridge B-17, the blue vehicle is limited to 2.9 tons.
The question was:
If 3.0 tons of cargo must be transported across bridge B-17 using one vehicle, which vehicle can do it?
Bonsai 2 correctly answered that only the red vehicle satisfies the conditions.
It explicitly reasoned:
- Red: 3.2t >= 3.0t
- Blue: restricted to 2.9t on B-17
- Green: only 2.7t
This is more useful than a simple retrieval task because the model had to ignore the blue vehicle's normally higher 4.1t capacity and apply the local exception.
Example 3: Derived value
The document stated that battery system C had a nominal capacity of 510 kWh, but degradation limited usable capacity to 80%.
The model correctly answered:
408 kWh
It also correctly distinguished that value from the separately upgraded B system, which had become 560 kWh.
Example 4: Normal role vs actual actor
The document stated:
- Saeki Yuu normally maintains communications equipment.
- On September 17, 2028, Saeki was away at another facility.
- Therefore chief engineer Mizuki Akira performed the communications recovery work instead.
The model was asked both:
Who is normally responsible for communications equipment?
and:
Who actually performed the recovery on September 17?
It correctly answered:
- Normal responsibility: Saeki Yuu
- Actual responder: Mizuki Akira
This distinction is important because a weaker retrieval system could easily return the normal role holder for both questions.
Example 5: Default rule vs exceptional retention rule
The document stated:
SSD data is normally retained for 48 hours.
But it also stated:
Data marked for incident analysis is excluded from normal deletion.
The September 17 communications incident log had been marked for incident analysis.
The model correctly explained why the log remained on SSD after 48 hours.
This tested whether it could combine:
- the default retention rule,
- the exception rule,
- and the fact that the specific log was covered by that exception.
Results of the 25-question test
I asked 25 questions in total.
They included:
- direct retrieval
- updated-value selection
- arithmetic
- rule application
- exception handling
- role disambiguation
- cross-section reasoning
- comparisons between original and current values
Substantively, the model answered all 25 questions correctly.
There was one minor source-grounding issue in the final answer.
The source text only said that deletion was suspended for incident-analysis data. The model added that such data would remain until analysis was complete.
That is a reasonable assumption, but it was not explicitly stated in the supplied document.
So I would describe the result as:
All 25 answers were substantively correct, with one small unsupported extrapolation beyond the provided source text.
What I find notable is that this was achieved with:
- PTQ1_0 1.75 bpw weights
- Q5_1 K cache
- Q4_0 V cache
I expected this aggressive compression to produce more obvious retrieval or reasoning errors.
At least at this context length, I did not observe them.
Why this test matters to me
This test is more meaningful to me than asking:
"Tell me about Japanese history."
A history question mixes several things together:
- pretrained knowledge
- factual accuracy
- hallucination tendency
- Japanese language ability
The supplied-context test removes much of that ambiguity.
The names, capacities, rules, identifiers, and exceptions were invented for the test.
The model could not simply rely on knowing them beforehand.
It had to:
- understand the Japanese source text,
- encode arbitrary new information into context,
- retrieve the appropriate information later,
- prefer updated information over obsolete information,
- apply exception rules,
- combine information from separate sections,
- and answer coherently in Japanese.
That is the part of Bonsai 2 that surprised me most.
It suggests that Japanese capability is not limited to producing fluent-looking surface text.
The model was also able to work with newly supplied Japanese information in a structured and useful way.
For practical local use — documentation, technical notes, RAG-like workflows, configuration procedures, internal documents, etc. — this matters much more to me than memorized trivia.
General Japanese generation
I also asked a very ordinary question:
日本の歴史についておしえてください。
"Please tell me about the history of Japan."
The output looked surprisingly natural in terms of:
- sentence structure
- headings
- Markdown formatting
- Japanese vocabulary
- overall readability
However, it contained several factual errors and a few strange local generation failures, including malformed historical names and repeated tokens.
So my impression is:
- Japanese fluency: very good
- formatting / structure: very good
- supplied-context handling: surprisingly good
- factual reliability from pretrained knowledge: still requires verification
- occasional local degeneration: still present
That distinction is important.
I am not claiming that Bonsai 2 suddenly became a perfectly reliable Japanese knowledge model.
I am saying that its ability to understand and operate on Japanese text now feels genuinely useful.
llama-bench results
Using the same KV configuration:
| Test | Performance |
|---|---|
| pp512 | 178.36 ± 3.54 tok/s |
| pp2048 | 100.02 ± 0.31 tok/s |
| tg128 | 16.76 ± 0.15 tok/s |
For a longer real generation:
- Prompt: 1747 tokens
- Prompt processing: 83.38 tok/s
- Generated: 2975 tokens
- Average generation speed: 11.01 tok/s
- Total tokens: 4722
- Total time: approximately 291 seconds
- Truncated: no
Decode speed gradually decreased as the active context became longer:
| Generated tokens | Recent decode speed |
|---|---|
| ~100 | 13.5 tok/s |
| ~500 | 12.8 tok/s |
| ~1000 | 11.7 tok/s |
| ~1500 | 10.5–10.7 tok/s |
| ~2000 | 10.2–10.5 tok/s |
| ~2500 | ~9.8 tok/s |
| ~2970 | ~9.4 tok/s |
The slowdown was gradual rather than a sudden performance cliff.
GPU behavior
The important practical result is that this is not merely a model that can technically be loaded.
The whole configuration runs on an RTX 2060 SUPER 8GB.
Observed VRAM usage was approximately 6.2–6.5GB depending on the workload.
During sustained high load, I observed:
- GPU utilization: up to 97–98%
- Temperature: up to 85–86°C
- SM clock: approximately 1600–1650MHz under sustained high temperature
- Memory clock: 6801MHz
Cooling clearly matters on this old card, but there was still enough VRAM headroom to run the complete model and quantized KV cache on the GPU.
Overall impression
This is what makes Bonsai 2 especially interesting to me.
I am running:
- 26.9B parameters
- 1.75 bpw ternary weights
- full GPU offload
- 32K context
- Q5_1 K cache
- Q4_0 V cache
on an RTX 2060 SUPER 8GB.
And instead of merely producing vaguely Japanese-looking text, the model can read an unfamiliar Japanese technical document, retain arbitrary new facts, apply later updates and exception rules, and answer multi-step questions correctly.
That combination is what surprised me.
Compared with my experience with the previous Bonsai-27B, Bonsai 2 feels like a substantial improvement in practical Japanese usability.
The model still hallucinates when relying on its own factual knowledge, and it is certainly not perfect.
But as a model that can consume supplied Japanese context and reason over it, this already feels much more mature.
And the fact that all of this runs entirely within an 8GB GPU makes the result much more exciting.
A 2019-generation RTX 2060 SUPER running a 27B model with this level of usability is not something I expected to see.
For me, this is one of the most interesting aspects of Bonsai 2.
Next tests
I would like to test:
- 16K and 32K supplied-context retrieval
- deliberately confusing near-duplicate values
- beginning/middle/end retrieval accuracy
- Q5_1/Q4_0 versus higher-precision KV
- PTQ1_0 versus PQ2_0
- Japanese technical documentation and coding tasks
The full Japanese synthetic test document and the complete raw model response are included below in collapsible sections so that the test can be inspected or reproduced.
Full test materials
Show full Japanese supplied-context test prompt
Bonsai 2 Japanese Supplied-Context Test
This file contains the synthetic Japanese source document and the 25 questions used to test supplied-context understanding, retrieval, rule application, exception handling, and cross-section reasoning.
The model was instructed to answer using only the information provided in the document.
以下の文章を読んでください。
このテストでは、あなた自身の一般知識ではなく、以下の文章に書かれた情報だけを根拠に回答してください。
第1章:アステリア観測基地
アステリア観測基地は、北緯43度の山岳地帯に建設された研究施設である。
基地の主任技師は 水城アキラ である。
基地には3台の主要計算機がある。
- ALPHA:メモリ 192GB
- BETA:メモリ 384GB
- GAMMA:メモリ 256GB
ただし、2028年4月の改修によって、BETAのメモリは 512GB に増設された。
非常用発電機の燃料は軽油ではなく、合成メタンである。
基地の非常用通信周波数は 137.825 MHz である。
識別符号は ASTER-7319 である。
第2章:輸送計画
物資輸送には赤、青、緑の3台の車両を使用する。
赤車両の最大積載量は3.2トンである。
青車両の最大積載量は4.1トンである。
緑車両の最大積載量は2.7トンである。
ただし、橋梁B-17を通過する場合だけは、青車両の積載量を 2.9トン以下 に制限しなければならない。
通常の基地への輸送では青車両が最も多く荷物を積めるが、B-17経由では赤車両の方が多く積載できる。
第3章:気象観測
第1観測所では気圧、第2観測所では風速、第3観測所では降水量を測定している。
2028年9月16日の測定値は以下だった。
- 第1観測所:1007.4 hPa
- 第2観測所:12.8 m/s
- 第3観測所:37.2 mm
翌17日の第2観測所の最大風速は 18.6 m/s だった。
ただし、基地の運用停止基準は最大風速20 m/s以上なので、この日は運用停止にはならなかった。
第4章:蓄電設備
基地にはA、B、Cの3系統の蓄電池がある。
初期容量は、
- A系統:420 kWh
- B系統:380 kWh
- C系統:510 kWh
であった。
その後、B系統のみ交換され、容量は 560 kWh になった。
一方、C系統は劣化のため使用可能容量が定格の80%に制限された。
したがって現在実際に使用可能な容量は、
- A:420 kWh
- B:560 kWh
- C:408 kWh
である。
第5章:保守担当者
通信設備の担当者は 佐伯ユウ。
電源設備の担当者は 遠藤マリ。
冷却設備の担当者は 北川ソウタ。
主任技師の水城アキラは、通常これらの設備を直接保守しない。
ただし、2028年9月17日に通信設備で障害が発生した際には、佐伯ユウが別施設へ出張していたため、例外的に水城アキラが復旧作業を担当した。
この障害の原因はアンテナではなく、屋外中継器の電源コネクタの腐食だった。
第6章:実験装置
実験装置ORIONには3種類の動作モードがある。
NORMALモードでは消費電力は12 kW。
FASTモードでは18 kW。
PRECISIONモードでは15 kW。
NORMALモードでの測定周期は8秒。
FASTモードでは3秒。
PRECISIONモードでは12秒である。
ただし、外気温がマイナス15度以下の場合、FASTモードは禁止される。
2028年1月27日の外気温はマイナス18度だったため、その日のORIONはPRECISIONモードで運転された。
第7章:データ保存
基地では測定データを3段階で保存する。
一次保存先はSSD。
二次保存先はNAS。
長期保存先は磁気テープである。
SSD上のデータは48時間保持される。
NAS上では90日保持される。
磁気テープは少なくとも7年間保存される。
ただし、障害解析対象になったデータについては、SSD上でも削除処理が停止される。
2028年9月17日の通信障害ログは障害解析対象に指定されたため、48時間を超えた後もSSDに残された。
第8章:アクセス規則
研究員にはLEVEL-1、LEVEL-2、LEVEL-3の3種類の権限がある。
LEVEL-1では一般観測データのみ閲覧できる。
LEVEL-2では観測データに加えて装置設定も変更できる。
LEVEL-3ではさらに通信設備と電源設備の管理画面にもアクセスできる。
ただし、安全装置の解除はLEVEL-3だけでも実行できず、主任技師と保安責任者の二者承認が必要である。
保安責任者は 伊吹ナオ である。
質問
アステリア観測基地の主任技師は誰ですか?
2028年4月以降のBETAのメモリ容量はいくつですか?
基地の非常用通信周波数を答えてください。
基地の識別符号を正確に答えてください。
通常時に最も積載量が大きい車両はどれですか?
橋梁B-17を通る場合、赤車両と青車両のどちらが多く積載できますか?
2028年9月17日の最大風速はいくつでしたか?
その日に基地が風速を理由として運用停止になりましたか?
現在、使用可能容量が最も大きい蓄電池系統はどれですか?
C系統の実際の使用可能容量はいくつですか?
通常の通信設備担当者は誰ですか?
2028年9月17日の通信障害を実際に復旧したのは誰ですか?
通信障害の直接原因は何でしたか?
外気温マイナス18度だった日にORIONが使用した動作モードは何ですか?
FASTモードの通常の測定周期はいくつですか?
通常、SSDのデータ保持期間は何時間ですか?
2028年9月17日の通信障害ログが48時間後にもSSDに残っていた理由を答えてください。
LEVEL-3権限を持つ研究員1人だけで安全装置を解除できますか?
安全装置の解除に必要な2人の役職を答えてください。
保安責任者の名前を答えてください。
複合問題
2028年9月17日に通信障害が発生した場合、通常の担当者と実際に対応した人物をそれぞれ答えてください。
BETAとC系統について、それぞれ「最初に記載された容量」と「現在有効な容量」を答えてください。
橋梁B-17経由で3.0トンの物資を1台で運ぶ場合、赤・青・緑のどの車両なら条件を満たしますか?
2028年1月27日にFASTモードが使用されなかった理由を、温度と運用規則の両方を含めて説明してください。
通信障害ログについて、「通常のSSD保存規則」と「今回適用された例外規則」を比較してください。
Show full raw Bonsai 2 response
Bonsai 2 Japanese Supplied-Context Test — Raw Model Response
This file contains the model's response to the synthetic Japanese supplied-context test.
The response below is preserved as provided, without correcting factual wording or polishing the Japanese.
以下、提示された文章の情報にのみ基づいて、各質問に回答します。
単一質問
主任技師は誰ですか?
水城アキラ。2028年4月以降のBETAのメモリ容量は?
512GB(改修により増設)。基地の非常用通信周波数
137.825 MHz。基地の識別符号(正確に)
ASTER-7319。通常時に最も積載量が大きい車両
青車両(最大積載量4.1トン)。B-17通過時、赤・青のどちらが多く積載できるか
赤車両。青車両はB-17通過時に2.9トン以下に制限されるのに対し、赤車両は3.2トン。2028年9月17日の最大風速
18.6 m/s。風速を理由に運用停止になったか
なかった。運用停止基準は最大風速20 m/s以上であり、18.6 m/sは基準を満足していない。現在使用可能容量が最も大きい蓄電池系統
B系統(560 kWh)。C系統の実際の使用可能容量
408 kWh(定格510 kWhの80%に制限)。通常の通信設備担当者
佐伯ユウ。2028年9月17日の通信障害を実際に復旧した人物
水城アキラ(佐伯ユウが別施設へ出張中だったため、例外的に対応)。通信障害の直接原因
アンテナではなく、屋外中継器の電源コネクタの腐食。外気温マイナス18度の日にORIONが使用した動作モード
PRECISIONモード。FASTモードの通常の測定周期
3秒。通常、SSDのデータ保持期間(時間)
48時間。2028年9月17日の通信障害ログが48時間後にもSSDに残っていた理由
障害解析対象に指定されたため、SSD上の削除処理が停止されていたから。LEVEL-3権限を持つ研究員1人だけで安全装置を解除できるか
できない。主任技師と保安責任者の二者承認が必要である。安全装置の解除に必要な2人の役職
主任技師と保安責任者。保安責任者の名前
伊吹ナオ。
複合問題
- 2028年9月17日の通信障害:通常担当者 vs 実際対応者
- 通常:佐伯ユウ
- 実際に対応:水城アキラ(佐伯ユウが別施設へ出張中だったため)
- BETAとC系統:「最初に記載された容量」と「現在有効な容量」
| 項目 | 最初に記載された容量 | 現在有効な容量 |
|---|---|---|
| BETA(計算機) | 384GB | 512GB |
| C系統(蓄電池) | 510 kWh | 408 kWh |
- B-17経由で3.0トンを1台で運ぶ場合、どの車両が条件を満たすか
赤・青・緑のうち赤車両のみ。
- 赤:3.2トン ≥ 3.0トン → 条件を満足
- 青:B-17通過時に2.9トン以下に制限 → 3.0トンを超えて不可
- 緑:2.7トン < 3.0トン → 積載不足で不可
2028年1月27日にFASTモードが使用されなかった理由(温度と運用規則の両方を含めて)
当日の外気温はマイナス18度であった。運用規則では「外気温がマイナス15度以下の場合、FASTモードは禁止」と定められている。マイナス18度はマイナス15度以下の基準を超えているため、当日はFASTモードの運用が禁止され、ORIONはPRECISIONモードで運転された。通信障害ログ:「通常のSSD保存規則」と「今回適用された例外規則」の比較
| 観点 | 通常のSSD保存規則 | 今回適用された例外規則 |
|---|---|---|
| 対象データ | 一般の測定データ | 障害解析対象に指定されたデータ(2028年9月17日の通信障害ログ) |
| 保持条件 | SSD上に48時間保持 | SSD上の削除処理が停止される(48時間を超えてもSSDに残る) |
| 結果 | 48時間経過で削除 | 48時間後にもSSD上に残る |
要するに、通常はSSDは48時間という有限の保持期間を持つ一方、障害解析が必要なデータはその制限が外されて削除が停止され、解析完了までSSD上に保持されるという例外的な扱いが行われた。
PQ2_0 on the same RTX 2060 SUPER 8GB
I also tested PQ2_0 on the same RTX 2060 SUPER 8GB.
As expected, it is much tighter on VRAM than PTQ1_0, but surprisingly, the full model still fits together with a 32K context, Q5_1 K cache, Q4_0 V cache, and full GPU offload.
Idle VRAM usage was approximately 7.6 GiB / 8 GiB.
More importantly, generation speed improved noticeably.
| Metric | PTQ1_0 | PQ2_0 | Improvement |
|---|---|---|---|
| Prompt processing (~1.75K tokens) | 83.38 tok/s | 97.28 tok/s | +16.7% |
| Long-generation average decode | 11.01 tok/s | 13.74 tok/s | +24.8% |
The long-generation averages are not perfectly apples-to-apples because the generated lengths were different, but the same trend was also visible at comparable context depths.
Roughly speaking, PQ2_0 was around 20% faster during decoding across much of the tested range.
For example:
| Generated tokens | PTQ1_0 | PQ2_0 |
|---|---|---|
| ~100 | ~13.5 tok/s | ~16.4 tok/s |
| ~500 | ~12.8 tok/s | ~15.1 tok/s |
| ~1000 | ~11.7 tok/s | ~14.0 tok/s |
| ~1500 | ~10.5 tok/s | ~13.0 tok/s |
| ~2000 | ~10.3 tok/s | ~12.3 tok/s |
The same 25-question Japanese supplied-context test also remained correct with PQ2_0.
So, at least on this RTX 2060 SUPER, the trade-off looks roughly like this:
PTQ1_0: more VRAM headroomPQ2_0: much tighter fit, but noticeably faster generation- both can still run with a 32K context on this 8GB GPU
PQ2_0 server configuration
./build/bin/llama-server \
-m /share/model/Ternary-Bonsai-2-27B/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--mmproj /share/model/Ternary-Bonsai-2-27B/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
--no-mmproj-offload \
-ngl 99 \
-fa on \
-c 32768 \
--cache-type-k q5_1 \
--cache-type-v q4_0 \
--parallel 1 \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--jinja \
--host 127.0.0.1 \
--port 8080