qwen3-32b-p300x2

Qwen3-32B is a dense instruction and reasoning model for general chat, coding, and tool use, served on a four-chip Tenstorrent Blackhole mesh.

Runs on p300x2 (mesh P300x2) โ€” 131,072-token context, up to 32 concurrent sequences.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  stisiTT/qwen3-32b-p300x2 --with-weights
tt-model serve stisiTT/qwen3-32b-p300x2

pull --with-weights downloads the Docker image and the Qwen/Qwen3-32B weights at 9216db5781bf21249d130ec9da846c4624c16137 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal 7d54a363f0cb5ccb0fc0ff8a63fb35398a4ba0b4 (dirty tree โ€” the image includes uncommitted changes)
vLLM v0.26.0
vllm-tt-plugin a local checkout โ€” commit not published
code/ digest 2c329baf1a4b0bea (sha256, first 16 hex digits)
built 2026-09-10T01:37:02+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support