faizath commited on
Commit
3fd3c01
·
verified ·
1 Parent(s): ce13d6c

feat(eval): add held-out results for all 427 test conversations

Browse files

The sibling Qwen adapter shipped with no held-out benchmark, so its card
could only report a loss curve and a ten-conversation probe. Publishing
every reply with its per-record flags makes each figure in the evaluation
section checkable rather than taken on trust -- including the invented
period comparisons, which are the defect a reader most needs to see for
themselves before deploying this.

Files changed (1) hide show
  1. eval_results.jsonl +3 -0
eval_results.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fc6c4249e4559bdebe37e971e9b39e759bf9f5a5445b37c5b7823a83ec3976ba
3
+ size 298663