|
Download reproduction_report.md from UraionLabs/UraionSpec: direct link, hf CLI and curl.
- Browser
- Download file 5.48 kB
-
https://huggingface.co/UraionLabs/UraionSpec/resolve/main/reproduction_report.md
- Command line
-
hf download hf://UraionLabs/UraionSpec/reproduction_report.md
-
curl -L -o reproduction_report.md https://huggingface.co/UraionLabs/UraionSpec/resolve/main/reproduction_report.md
5.48 kB
UraionSpec β Reproduction Report
Hardware
- Machine: MacBook Pro 12,1 (Intel i5-5257U) running Ubuntu 24 LTS
- Storage: Root partition (47.9G), 7.9G free
- GPU: None (CPU-only testing)
- RAM: 7.8 GB
Models Used
- Target:
Qwen/Qwen3-0.6B(Qwen3-0.6B, ungated, Apache 2.0) - Draft: Custom
DSparkDraftModelwith markov_rank=64, 2 backbone layers - Note: 0.6B target is usable for shape/loss verification but too small for meaningful speculative decoding speedups
Dataset
- Source:
trl-lib/Capybara(subset: 32 samples) - Format: Chat template applied via tokenizer
- Max length: 512 tokens
- Note: Tiny subset for smoke testing only
Commands Run
1. Package Import Test
python3 -c "from uraionspec import ..." # All imports OK
2. Unit Tests (80 tests)
pytest tests/ -v # 80/80 passed in 7.40s
3. Smoke Training (CPU only)
Would run:
python scripts/smoke_train.py --target Qwen/Qwen3-0.6B --samples 32 --steps 5 --device cpu
Blocked: 0.6B model requires ~1.2 GB RAM just to load. With 7.8 GB and CPU-only, this would be extremely slow but should work. The training loop requires target model logits which adds memory pressure.
4. Lint/Type Check
Will run: ruff check . (configured in pyproject.toml)
Results
Test Results (80/80 passed)
| Test Suite | Tests | Status |
|---|---|---|
test_acceptance.py |
8 | β All passed |
test_markov_head.py |
15 | β All passed |
test_backbone.py |
17 | β All passed |
test_sampling.py |
8 | β All passed |
test_scheduler.py |
14 | β All passed |
test_shapes.py |
10 | β All passed |
test_sts.py |
8 | β All passed |
What Was Verified
- β Acceptance rule: all-accepted, first-rejected, partial, batch-independence, bonus token
- β Expected accept length computation
- β Markov head forward/shape/gradient
- β RNN head stateful forward/step
- β Gated Markov head forward
- β Throughput profile lookup/interpolation
- β Static scheduler threshold/length
- β Hardware-aware scheduler: single/multi request, monotonic, zero confidence
- β Draft model forward/sample shapes
- β Loss: CE, TV, confidence gradients
- β Confidence head forward/accept rate computation
- β End-to-end draft β verify β accept β schedule cycle
- β STS: ECE, temperature fitting, calibrator fit/transform
Failures / Blockers
- Full target model training on CPU: Running the 0.6B target model (~600M params) on CPU with 7.8 GB RAM is feasible but slow. The smoke training script was designed for GPU.
- GPU smoke eval: Same limitation β requires GPU for meaningful speculative decoding evaluation.
- Disk space: Only 7.9 GB free β cannot download the 1.3M sample Open-PerfectBlend dataset or store large target caches.
Next Steps to Scale
To Qwen3-1.7B/4B+ (Local or Colab)
- Provision GPU: Use
colab run --gpu A100 --keepfor training - Full training:
colab run -s uraionspec-train --gpu A100 --keep --timeout 28800 \ python scripts/smoke_train.py \ --target Qwen/Qwen3-4B \ --samples 10000 \ --steps 1000 \ --batch-size 8 \ --block-size 7 - Upload to HF:
hf upload UraionLabs/UraionSpec /checkpoints/ - Full evaluation: Run
scripts/run_benchmark.pyon the trained checkpoint
To Match Paper Settings
- Data: Use
mlabonne/open-perfectblend(1.3M samples) - Blocks: Ξ³ = 7, 16 (experiment with both)
- Training: 10 epochs, multi-GPU (8Γ)
- Backbone: 5 layers (paper default),
hidden_sizematching target - Target cache: Generate target logits for all training data (38 TB for full; use
--max_length 2048for smaller) - SPS profiling: Run offline benchmark to get real SPS(B) curve for your hardware
Files Changed
UraionSpec/
βββ pyproject.toml
βββ README.md
βββ src/uraionspec/
β βββ __init__.py
β βββ models/
β β βββ __init__.py
β β βββ markov_head.py
β β βββ rnn_head.py
β β βββ confidence_head.py
β β βββ draft_model.py
β βββ decoding/
β β βββ __init__.py
β β βββ acceptance.py
β β βββ scheduler.py
β β βββ speculative.py
β βββ training/
β β βββ __init__.py
β β βββ dataset.py
β β βββ losses.py
β β βββ train_drafter.py
β β βββ cache_targets.py
β βββ calibration/
β β βββ __init__.py
β β βββ sts.py
β βββ evaluation/
β β βββ __init__.py
β β βββ eval_acceptance.py
β β βββ benchmark_latency.py
β βββ utils/
β βββ __init__.py
β βββ hf.py
β βββ logging.py
β βββ seed.py
βββ scripts/
β βββ smoke_train.py
β βββ smoke_eval.py
β βββ run_benchmark.py
βββ tests/
β βββ test_acceptance.py
β βββ test_markov_head.py
β βββ test_scheduler.py
β βββ test_shapes.py
β βββ test_sts.py
βββ examples/
β βββ prompts.jsonl
βββ docs/
βββ DSpark_implementation_notes.md
βββ reproduction_report.md
βββ model_card_template.md
Total: ~4100 lines of Python, 80 passing tests