license: apache-2.0
tags:
- eagle3
- speculative-decoding
- pearl
- llama-3.1
Pearl EAGLE3.1 R2C best head (2026-06-19)
Selected head: R2C_t7_g115.
Selection rule: best overall 8×H100 serving throughput, ranked by aggregate best-K req/s across per-GPU concurrency C=1/4/16 and prompt-length buckets. K4/K5 vs K3 was diagnostic only; the selected checkpoint is the one with the best overall serving result.
Final rank by aggregate best-K req/s:
- R2C — 11.0288
- D — 11.0108
- I75 — 11.0029
- R4B — 10.9846
- R3B — 10.9305
Full run artifacts, training states, logs, benchmark outputs, and selection manifests were archived to:
s3://pearl-eagle31-artifacts-507286591552-us-east-2/eagle31-final-r2c-20260619T1318Z
The PR branch was also updated with the multi-replica benchmark used for final evaluation:
feat/real-prompt-eagle-mining-bench @ 85861376 feat: add multi-replica EAGLE benchmark
This HF repo contains the selected head files needed for serving/eval (model.safetensors, config.json, config.py, val_metrics.json) plus selection metadata. Optimizer/scheduler training state is in the S3 archive, not this model repo.