\documentclass[11pt,a4paper]{article} % ---------- arXiv-compatible packages ---------- \usepackage[utf8]{inputenc} \usepackage[T1]{fontenc} \usepackage{hyperref} \usepackage{amsmath,amssymb,amsthm} \usepackage{mathtools} \usepackage{booktabs} \usepackage{graphicx} \usepackage{xcolor} \usepackage{algorithm} \usepackage{algpseudocode} \usepackage{geometry} \usepackage{microtype} % enumitem not available — using basic LaTeX lists instead % caption/subcaption not available — using basic figure handling \geometry{margin=1in} % ---------- metadata for arXiv ---------- \hypersetup{ colorlinks=true, linkcolor=blue, citecolor=blue, urlcolor=blue, pdftitle={UraionSpec: A Faithful Implementation of Confidence-Scheduled Speculative Decoding}, pdfauthor={Uraion Labs}, pdfkeywords={speculative decoding, DSpark, semi-autoregressive, confidence scheduling, LLM inference} } % ---------- convenience macros ---------- \newcommand{\UraionSpec}{\textsc{UraionSpec}} \newcommand{\DSpark}{\textsc{DSpark}} \newcommand{\DFlash}{\textsc{DFlash}} \newcommand{\DeepSpec}{\textsc{DeepSpec}} \newcommand{\Eagle}{\textsc{Eagle}} \newcommand{\MTP}{\textsc{MTP}} \DeclareMathOperator{\softmax}{softmax} \DeclareMathOperator{\sigmoid}{sigmoid} \DeclareMathOperator{\CumProd}{CumProd} \DeclareMathOperator{\SPS}{SPS} \begin{document} % ====================================================================== \title{\textbf{\UraionSpec}: A Faithful Implementation of \\ Confidence-Scheduled Speculative Decoding} \author{% \normalsize Uraion Labs \\[4pt] \texttt{uraionlabs@gmail.com} \\[2pt] \texttt{https://huggingface.co/UraionLabs/UraionSpec} } \date{\normalsize Technical Report -- ICML 2026 Systems Track} \maketitle % ====================================================================== \begin{abstract} Speculative decoding accelerates autoregressive language model inference by generating multiple candidate tokens from a lightweight draft model, then verifying them in parallel with the target model. \DSpark (Cheng et al.,~2026) introduces two key innovations: semi-autoregressive generation via a parallel backbone with a lightweight sequential head, and confidence-scheduled verification that dynamically adjusts verification length based on predicted acceptance probabilities and engine throughput. We present \UraionSpec, a clean, modular, and verified implementation of the \DSpark algorithm. Our implementation faithfully reproduces the core components---the Markov and RNN sequential heads, confidence head, Sequential Temperature Scaling (STS) calibration, hardware-aware prefix scheduler (Algorithm~1), and the three-term training objective---while adding a \DFlash-style parallel backbone with target model KV injection that was missing from the initial release. With 80 passing unit tests, a fully configurable training pipeline, comprehensive evaluation tools, and end-to-end smoke tests, \UraionSpec is designed for reproducible research and practical experimentation. The codebase is publicly available on GitHub and HuggingFace under the MIT license. \end{abstract} % ====================================================================== \section{Introduction} \label{sec:introduction} The rapid growth of large language models (LLMs) has made inference latency a critical bottleneck in production deployments. Autoregressive decoding --- generating one token at a time with sequential dependence on all previous tokens --- is inherently bandwidth-bound on modern hardware because the attention mechanism requires loading the full model parameters from memory for each generated token~\cite{pope2023efficiently, leviathan2023fast}. Speculative decoding~\cite{leviathan2023fast, chen2023accelerating} breaks this bottleneck by introducing a two-phase process: a lightweight draft model proposes multiple candidate tokens in one forward pass, and the target model verifies them in parallel, yielding a theoretical speedup proportional to the acceptance rate. This approach preserves the exact output distribution of the target model, making it a lossless acceleration technique. The draft model design determines the fundamental trade-off: more accurate draft models produce longer accepted prefixes but incur higher drafting cost, while simpler draft models are cheaper but suffer from low acceptance rates. \DSpark~\cite{cheng2026dspark} addresses this trade-off through two key contributions: \begin{enumerate} \item \textbf{Semi-autoregressive generation}: A parallel backbone (the \DFlash architecture) processes all proposal positions in a single forward pass, while a lightweight sequential head (Markov or RNN) injects inter-token dependency. This combines the speed of parallel drafters~\cite{stern2018blockwise, xia2022medusa} with the quality of autoregressive ones~\cite{miao2025eagle3}. \item \textbf{Confidence-scheduled verification}: A trained confidence head predicts per-position acceptance probabilities, and a hardware-aware scheduler dynamically selects verification lengths to maximize expected throughput. This prevents wasting target-model compute on low-confidence suffix tokens under heavy load. \end{enumerate} \paragraph{Our contribution.} We present \UraionSpec, a faithful, modular, and verified implementation of \DSpark. Our implementation includes all core components listed in the original paper, plus a \DFlash-style parallel backbone with target model KV injection that was identified as a critical gap in initial implementations. The codebase is designed for clarity and reproducibility rather than production scale, making it suitable for academic research, ablation studies, and educational purposes. Key features include: \begin{itemize} \item Complete implementation of all \DSpark components (80 unit tests) \item \DFlash-style backbone with target model KV injection \item Markov, gated Markov, and RNN sequential heads \item Confidence head with analytical acceptance-rate supervision \item Sequential Temperature Scaling (STS) calibration \item Hardware-aware prefix scheduler (Algorithm~1) \item Training pipeline with three-term objective \item Comprehensive evaluation and benchmarking tools \item All code under MIT license on GitHub and HuggingFace \end{itemize} % ====================================================================== \section{Background and Related Work} \label{sec:background} \subsection{Speculative Decoding} Standard speculative decoding~\cite{leviathan2023fast, chen2023accelerating} works as follows: given a target model $p_t$ and a draft model $p_d$, at each decoding step the draft model proposes $\gamma$ candidate tokens $x_1, \ldots, x_\gamma$ left-to-right. The target model then verifies all $\gamma$ tokens in a single forward pass. Each token $x_k$ is accepted with probability \begin{equation} \textstyle P(\text{accept } x_k) = \min\left(1,\; \frac{p_t(x_k \mid x_{ verify -> accept loop training/ dataset.py -- Anchor-block dataset preparation losses.py -- CE + TV + confidence (Eq. 12) train_drafter.py -- Training loop (frozen target) cache_targets.py -- Target logit cache generation calibration/ sts.py -- Sequential Temperature Scaling evaluation/ eval_acceptance.py -- Acceptance rate/length metrics benchmark_latency.py -- Vanilla vs speculative latency \end{verbatim} \end{minipage} \caption{\UraionSpec package structure. Each module is independently testable and documented.} \label{fig:architecture} \end{figure} \subsection{Sequential Heads} \label{subsec:sequential-heads} \UraionSpec implements three sequential head variants, each with different trade-offs between expressiveness and computational cost. \paragraph{Vanilla Markov head.} The simplest sequential head implements a low-rank transition bias: \begin{equation} B(x_{k-1}, x_k) = W_1[x_{k-1}] \cdot W_2, \qquad W_1 \in \mathbb{R}^{V \times r},\; W_2 \in \mathbb{R}^{r \times V}, \end{equation} where $r = 256$ is the rank and $V$ is the vocabulary size. This head is memoryless: the bias depends only on the immediately preceding token, not on full prefix history. \paragraph{Gated Markov head.} The gated variant modulates the Markov bias using the backbone hidden state: \begin{equation} g_k = \sigma\bigl(W_g [h_k ;\; W_1[x_{k-1}]]\bigr), \qquad B_k = W_2\bigl(g_k \odot W_1[x_{k-1}]\bigr). \end{equation} This allows the backbone representation to influence how much Markov bias to apply, providing context-dependent transition modeling. \paragraph{RNN head.} For full prefix history, we implement a GRU-like recurrent head: \begin{align} z_k &= [s_{k-1} ;\; W_1[x_{k-1}] ;\; h_k], \\ s_k &= \sigma(W_g z_k) \odot s_{k-1} + (1 - \sigma(W_g z_k)) \odot \tanh(W_c z_k), \\ B_k &= W_2^{\mathsf{T}} \tanh(W_o z_k). \end{align} The recurrent state $s_k$ accumulates information across the entire prefix $x_{