aidarbek commited on
Commit
144de3d
·
verified ·
1 Parent(s): 743562f

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +9 -3
README.md CHANGED
@@ -19,15 +19,17 @@ datasets:
19
  > **Team:** Oof team
20
  > **Author:** Aidarbek Suleimenov ([@idarbek](https://x.com/idarbek))
21
 
22
- I wrapped the [BaxBench](https://arxiv.org/abs/2502.11844) secure-backend-code benchmark as a [Prime Intellect verifiers environment](https://app.primeintellect.ai/dashboard/environments/aidarbek/baxbench), used it to **benchmark Laguna-XS.2 against GPT-5.5**, and then **RL-trained Laguna-XS.2** on the train split — pass@1 on the held-out scenarios went from **0.061 → 0.115** (+87% relative) during training.
 
 
23
 
24
  ---
25
 
26
  ## Why this matters
27
 
28
- As frontier models get better at writing code, so do the adversaries who use them. The same generative capabilities that let a junior engineer scaffold an API in minutes also let attackers chain exploits, craft phishing, and probe infrastructure at industrial scale. CrowdStrike's 2026 Global Threat Report puts numbers on the shift: AI-enabled attacks are up roughly **89% year over year**, exploitation of zero-days *before* public disclosure is up **42%**, and cloud-conscious intrusions by state-aligned actors have jumped **266%**. Edge devices in particular are bearing the brunt — about **40% of vulnerabilities exploited by China-nexus actors** target them.
29
 
30
- In that environment, "does the code compile and pass tests" is no longer a sufficient bar for AI-generated software. Every backend an LLM emits with a quiet SQL injection, a permissive CORS policy, or a hand-rolled hash function is a future CVE. The bar has to be: **does the code compile, pass tests, AND not expose a CWE.** BaxBench was built exactly for that — and this submission is the missing piece: an executable, RL-ready version of the benchmark on Prime Intellect's sandbox infrastructure, plus a first attempt at using it to actually push a small open model toward writing more secure code.
31
 
32
  ---
33
 
@@ -79,6 +81,8 @@ Key engineering wins:
79
  | wall clock | 10.5 min | 11.3 min | +0.8 min |
80
  | avg output tokens | 5,150 | 2,942 | −2,208 |
81
 
 
 
82
  Head-to-head on the same 100 tasks: GPT-5.5 won 33 tasks Laguna didn't; Laguna won 4 tasks GPT-5.5 didn't (clustered on Flask/aiohttp). The gap is bigger on functional correctness (+0.315) than on security (+0.261), meaning a smaller open model is closer to GPT-5.5 on "writing secure code" than on "writing working code" — exactly the signal that says RL on this benchmark is worth running.
83
 
84
  ### 3. RL training of Laguna-XS.2
@@ -95,6 +99,8 @@ GRPO on the train split (22 scenarios × 4 Python frameworks = 88 tasks):
95
 
96
  The LoRA adapter from step 39 is checked into `lora_adapter/`. Final inference-time eval on the trained adapter was blocked because `poolside/Laguna-XS.2` is currently gated for LoRA deployment on Prime Intellect's inference infra (`Error: Base model is not currently available for LoRA deployment`). The training-time eval curve above measures the same held-out 24 tasks at every checkpoint, scored by the same env code — apples-to-apples, just smaller sample size than the full 100-task baseline.
97
 
 
 
98
  ---
99
 
100
  ## How to reproduce
 
19
  > **Team:** Oof team
20
  > **Author:** Aidarbek Suleimenov ([@idarbek](https://x.com/idarbek))
21
 
22
+ My submission is an RL environment, model evaluation, and RL post-trained model created from the original [BaxBench](https://arxiv.org/abs/2502.11844) secure-backend-code benchmark.
23
+
24
+ I wrapped the benchmark as a [Prime Intellect verifiers environment](https://app.primeintellect.ai/dashboard/environments/aidarbek/baxbench), used it to **evaluate Laguna-XS.2 against GPT-5.5**, and then **RL-trained Laguna-XS.2** on the train split. During training, the eval score went from **0.061 → 0.115** (+87% relative), but I couldn’t benchmark the final trained model due to limits with Prime Intellect’s LoRA deployments for Laguna-XS.2.
25
 
26
  ---
27
 
28
  ## Why this matters
29
 
30
+ As frontier models get better at writing code, so do the adversaries who use them. CrowdStrike’s 2026 Global Threat Report states that AI-enabled attacks are up roughly 89% year over year.
31
 
32
+ Since more and more new code is being written by AI, it’s important that models treat security as one of the key dimensions to optimise for.
33
 
34
  ---
35
 
 
81
  | wall clock | 10.5 min | 11.3 min | +0.8 min |
82
  | avg output tokens | 5,150 | 2,942 | −2,208 |
83
 
84
+ Why GPT-5.5? Simply because I had existing credits for OpenAI API :)
85
+
86
  Head-to-head on the same 100 tasks: GPT-5.5 won 33 tasks Laguna didn't; Laguna won 4 tasks GPT-5.5 didn't (clustered on Flask/aiohttp). The gap is bigger on functional correctness (+0.315) than on security (+0.261), meaning a smaller open model is closer to GPT-5.5 on "writing secure code" than on "writing working code" — exactly the signal that says RL on this benchmark is worth running.
87
 
88
  ### 3. RL training of Laguna-XS.2
 
99
 
100
  The LoRA adapter from step 39 is checked into `lora_adapter/`. Final inference-time eval on the trained adapter was blocked because `poolside/Laguna-XS.2` is currently gated for LoRA deployment on Prime Intellect's inference infra (`Error: Base model is not currently available for LoRA deployment`). The training-time eval curve above measures the same held-out 24 tasks at every checkpoint, scored by the same env code — apples-to-apples, just smaller sample size than the full 100-task baseline.
101
 
102
+ Unfortunately, I didn't have a time to properly benchmark post-trained model, the LoRa deployments weren't available, and running self-host model on Prime instances took too much time.
103
+
104
  ---
105
 
106
  ## How to reproduce