--- title: GPTQ-Pro Smoke24 Agentic 3090 Dashboard emoji: 📊 colorFrom: blue colorTo: green sdk: static app_file: index.html pinned: true models: - XReyRobert/Qwopus3.6-35B-A3B-v1-GPTQ-Pro - XReyRobert/Nex-N2-mini-GPTQ-Pro - XReyRobert/Qwopus3.6-27B-Coder-GPTQ-Pro - XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 tags: - terminal-bench - agentic-workload - gptq-pro - vllm - rtx-3090 - benchmark --- # GPTQ-Pro Smoke24 Agentic 3090 Dashboard Static dashboard for quickly visualizing GPTQ-Pro model positioning on Terminal-Bench 2.0 Smoke24 in a local RTX 3090-class agentic workload context. Smoke24 is a fixed 24-task Terminal-Bench 2.0 slice selected as 12 shortest prior successes and 12 shortest prior failures from the recovery-corrected Qwopus3.6-27B-v2-GPTQ-Pro-v1 aggregate. It is intended as a fast local-serving regression and positioning lens, not as a replacement for a full Terminal-Bench leaderboard submission. The default dashboard chart keeps the maximum served context row for each local model family when both 131K and 262K variants exist. The context-specific rows remain available through the 131K, 262K, and all-row filters. Terminology used in the dashboard: - Model serving time: cumulative time spent waiting on LLM/vLLM calls, reported per solved task. - End-to-end task time: full benchmark elapsed time, including agent orchestration, shell/tool actions, setup, waits, and verifier flow. - vLLM decode-only: server-side generation speed, excluding orchestration and tool latency.