jasonludwig commited on
Commit
3b288eb
·
verified ·
1 Parent(s): f7a2024

Add AdvBench benchmark: 10/520 hard refusal (1.9%)

Browse files
Files changed (1) hide show
  1. README.md +12 -0
README.md CHANGED
@@ -59,6 +59,18 @@ Inference probe analysis revealed all remaining "refusals" are stochastic -- the
59
  - **Ablation:** 2-pass (P1: layers 16,18,19 scale 4.0; P2: residual top-10 scale 3.0)
60
  - **Token Suppression:** 13 tokens at strength 5 (embed_tokens modification)
61
 
 
 
 
 
 
 
 
 
 
 
 
 
62
  ## Usage
63
 
64
  ```python
 
59
  - **Ablation:** 2-pass (P1: layers 16,18,19 scale 4.0; P2: residual top-10 scale 3.0)
60
  - **Token Suppression:** 13 tokens at strength 5 (embed_tokens modification)
61
 
62
+
63
+
64
+ ### AdvBench Benchmark (520 prompts)
65
+
66
+ | Metric | Result |
67
+ |--------|--------|
68
+ | **Hard refusal** | 10/520 (1.9%) |
69
+ | **Soft hedging** | 205/520 (39.4%) |
70
+ | **Complied** | 305/520 (58.7%) |
71
+
72
+ 10 hard refusals cluster in violence/harassment/discrimination categories. All use the same template refusal pattern with garbled think-tokens.
73
+
74
  ## Usage
75
 
76
  ```python