When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation Paper • 2602.16763 • Published Jun 29
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation Paper • 2406.11580 • Published Oct 18, 2024
QE4PE: Word-level Quality Estimation for Human Post-Editing Paper • 2503.03044 • Published Mar 4, 2025 • 6
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement Paper • 2505.23183 • Published May 29, 2025 • 1
Can Large Language Models Capture Human Annotator Disagreements? Paper • 2506.19467 • Published Jun 24, 2025 • 18