# Changelog ## v1.0.0 Initial release. - `block_sparse_gqa_prefill` -- SPMD (LNC=2) block-sparse GQA prefill, normalized output. - `block_sparse_gqa_prefill_state` -- SPMD prefill emitting `(o, l, m)` flash state for context-parallel merge. - `block_sparse_attention_single_q` -- single-Q-block MHA reference kernel. - Host helpers: `build_block_kv_indices`, `build_causal_block_mask`, `dense_gqa_reference`, `cp_merge`. - Q-parallel DP=4 driver reaching 1M seq_len prefill on a single trn2.3xlarge. - Context-parallel state-kernel drivers (CP=2, CP=4) with on-device state gather. Verified on trn2.3xlarge, SDK 2.31 (DLAMI 20260708), PyTorch Native Beta 4 (torch-neuronx 2.11.3.0, neuronx-cc 2.26.6360, NKI 0.5.0). Correctness-verified under NKI 0.6.0 (SDK 2.32 pre-release).