# Roadmap beyond v1.0.0 v1.0.0 covers MiniMax-M3 block-sparse prefill on a single trn2.3xlarge up to 1M seq_len. Forward-looking directions: ## Beyond 1M seq_len on a single trn2.3xlarge Q-parallel DP=4 reaches 1M on one trn2.3xlarge. For longer sequences: - **trn2.48xlarge** (more logical cores): higher DP degree, or a DP x CP topology. - **Ring attention**: keep each rank's query slice local and rotate KV shards via `nki.collectives.collective_permute` (cf. `nkilib.experimental.attention .ring_attention_fwd`), applying online-softmax state updates as each KV shard arrives. This avoids materializing the FP32 `(o, l, m)` state externally, which is what limits the current CP configurations. Would require adding a `kv_block_indices` argument to the ring wrapper and partitioning it per ring step. ## Kernel performance The kernel is currently vector-engine bound (MFU ~16% of a ~49% ceiling on the reference shape). Remaining headroom: - Precompute `K^T` for all `TOPK` slots once per Q-block instead of transposing each slot inside the loop. - Rotating SBUF buffers to overlap the KV gather, the two matmuls, and the softmax. - A modernized KV gather primitive (`nisa.tensor_copy` with `.indirect()`) where the KV cache fits in SBUF. ## Precision - BF16 KV cache is the default. FP8 KV would roughly double the single-instance seq_len ceiling; not yet implemented. ## References - `nkilib/experimental/attention/ring_attention_fwd.py` - `examples/05_qparallel.py` -- the recommended long-seq path (reaches 1M) - `examples/07_cp_ondevice.py` -- multi-process CP with on-device state gather