--- library_name: uraionspec license: mit language: - en pipeline_tag: text-generation tags: - speculative-decoding - dspark - deepseek - llm-inference - model-optimization - transformer - pytorch - efficient-llm - inference-acceleration - draft-model - torch - uraion-labs - uraion - systems-research - icml-2026 - acceptance-scheduling - semi-autoregressive - confidence-prediction - calibration sdk: docker sdk_version: "1.0" ---

Uraion Labs

Uraion Labs
Foundational systems research.

UraionSpec
Faithful DSpark-style Speculative Decoding — modular, runnable, verified.

--- **UraionSpec** is a clean, modular, and runnable implementation of [**DSpark**](https://www.alphaxiv.org/abs/2026.dspark) — Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation — accepted at **ICML 2026**. It faithfully reproduces the core DSpark algorithm: a semi-autoregressive draft model that combines a parallel backbone with a lightweight sequential Markov/RNN head, a confidence head for per-position acceptance prediction, and a hardware-aware prefix scheduler that dynamically tailors verification length based on engine throughput profiles. This is **research infrastructure** — not a model checkpoint. It provides the training, evaluation, calibration, and decoding pipeline so you can train and evaluate draft models for speculative decoding on your own target models and data. **Intelligence is a systems problem.** This codebase is one piece of that system. --- ## What is DSpark? DSpark is a speculative decoding framework from DeepSeek-AI that introduces two key innovations over existing methods: 1. **Semi-Autoregressive Generation** — A parallel backbone handles bulk computation while a lightweight sequential head injects inter-token dependency, combining the speed of parallel drafters with the quality of autoregressive ones. This fixes the "multi-modal collision" problem where parallel drafters produce incoherent combinations like "of problem" instead of "of course". 2. **Confidence-Scheduled Verification** — A confidence head predicts per-position acceptance probabilities, and a hardware-aware scheduler (Algorithm 1) dynamically tailors verification length based on estimated prefix survival probabilities and engine-specific throughput profiles `SPS(B)`. This prevents wasting compute on tokens with high rejection risk under heavy load. ### Architecture ``` UraionSpec/ ├── src/uraionspec/ │ ├── models/ # Draft model architecture │ │ ├── markov_head.py # Low-rank transition bias (r=256) │ │ ├── rnn_head.py # GRU-like recurrent sequential head │ │ ├── confidence_head.py # Per-position acceptance predictor │ │ └── draft_model.py # Combined parallel backbone + heads │ ├── decoding/ # Speculative decoding core │ │ ├── acceptance.py # Lossless rejection sampling (min ratio) │ │ ├── scheduler.py # Algorithm 1: Hardware-aware prefix scheduler │ │ └── speculative.py # Orchestration: draft → verify → accept │ ├── training/ # Training pipeline │ │ ├── dataset.py # Anchor-block dataset preparation │ │ ├── losses.py # CE + TV + Confidence (position-weighted) │ │ ├── train_drafter.py # Training loop (frozen target) │ │ └── cache_targets.py # Target logit cache generation │ ├── calibration/ # Sequential Temperature Scaling │ │ └── sts.py # Left-to-right ECE minimization │ ├── evaluation/ # Evaluation & benchmarking │ └── utils/ # HF helpers, logging, seeding ├── scripts/ # Runnable entry points ├── tests/ # 55 unit & integration tests └── docs/ # Implementation notes, reports ``` --- ## Quick Start ```bash # Install pip install git+https://huggingface.co/UraionLabs/UraionSpec # Run tests pytest tests/ -v # Smoke train a draft model python scripts/smoke_train.py \ --target Qwen/Qwen2.5-0.5B-Instruct \ --samples 32 --steps 5 # Evaluate acceptance python scripts/smoke_eval.py \ --target Qwen/Qwen2.5-0.5B-Instruct \ --gamma 7 --steps 5 # Benchmark vs vanilla decoding python scripts/run_benchmark.py \ --target Qwen/Qwen2.5-0.5B-Instruct \ --prompts examples/prompts.jsonl ``` --- ## Key Design | Component | Paper Reference | Implementation | |---|---|---| | Draft distribution | Eq. 4: `p_k(v) = softmax(U_k(v) + B_k(v))` | `draft_model.py` + `markov_head.py` | | Markov transition bias | Eq. 5: `B = W1[x_{k-1}] W2` | `VanillaMarkov`, `GatedMarkovHead` | | RNN sequential head | Eq. 6: GRU-like state `s_k` | `RNNHead` | | Confidence head | Eq. 7: `c_k = σ(w^T[h_k; W1[x_{k-1}]])` | `ConfidenceHead` | | Acceptance rate target | Eq. 8: `c* = 1 - ½‖p_d - p_t‖₁` | `compute_accept_rate()` | | Prefix scheduler | Algorithm 1: `argmax τ · SPS(B)` | `hardware_aware_prefix_scheduler()` | | STS calibration | §3.2.1: per-position ECE minimization | `STSCalibrator` | | Loss function | Eq. 12: `L = 0.1·L_ce + 0.9·L_tv + 1.0·L_conf` | `compute_dspark_loss()` | --- ## Verification | Check | Status | |---|---| | 55 unit & integration tests | ✅ All passing | | Package imports | ✅ Clean | | Linting (ruff) | ✅ All checks passed | | Smoke training (CPU, Qwen2.5-0.5B) | ✅ 3 steps, all losses decreasing | | Confidence head gradient flow | ✅ Supervised by analytical acceptance rate | | Acceptance rule correctness | ✅ Verified: all-accepted, first-rejected, partial, bonus token | --- ## Relation to DeepSpec UraionSpec is an independent, faithful implementation of the DSpark algorithm described in the [paper](https://www.alphaxiv.org/abs/2026.dspark) and the [DeepSpec](https://github.com/deepseek-ai/DeepSpec) repository (MIT license). While DeepSpec is a production-grade codebase with multi-GPU training, 38 TB target caches, and vLLM integration, UraionSpec focuses on: - **Clarity** — Modular, documented Python with clean separations - **Runability** — Smoke tests that work on a single GPU or CPU - **Completeness** — Every algorithm component from the paper is implemented --- ## License MIT License. Built with reference to [DeepSpec](https://github.com/deepseek-ai/DeepSpec) (MIT) and the DSpark paper. Copyright © 2026 Uraion Labs. --- ## Citation ```bibtex @software{uraionspec2026, author = {Uraion Labs}, title = {UraionSpec: Faithful DSpark-style Speculative Decoding}, year = {2026}, url = {https://huggingface.co/UraionLabs/UraionSpec} } @article{cheng2026dspark, title={DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation}, author={Cheng, Xin and Yu, Xingkai and Shao, Chenze and Li, Jiashi and Xiong, Yunfan and others}, journal={ICML}, year={2026} } ```