Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -19,15 +19,17 @@ datasets:
|
|
| 19 |
> **Team:** Oof team
|
| 20 |
> **Author:** Aidarbek Suleimenov ([@idarbek](https://x.com/idarbek))
|
| 21 |
|
| 22 |
-
|
|
|
|
|
|
|
| 23 |
|
| 24 |
---
|
| 25 |
|
| 26 |
## Why this matters
|
| 27 |
|
| 28 |
-
As frontier models get better at writing code, so do the adversaries who use them.
|
| 29 |
|
| 30 |
-
|
| 31 |
|
| 32 |
---
|
| 33 |
|
|
@@ -79,6 +81,8 @@ Key engineering wins:
|
|
| 79 |
| wall clock | 10.5 min | 11.3 min | +0.8 min |
|
| 80 |
| avg output tokens | 5,150 | 2,942 | −2,208 |
|
| 81 |
|
|
|
|
|
|
|
| 82 |
Head-to-head on the same 100 tasks: GPT-5.5 won 33 tasks Laguna didn't; Laguna won 4 tasks GPT-5.5 didn't (clustered on Flask/aiohttp). The gap is bigger on functional correctness (+0.315) than on security (+0.261), meaning a smaller open model is closer to GPT-5.5 on "writing secure code" than on "writing working code" — exactly the signal that says RL on this benchmark is worth running.
|
| 83 |
|
| 84 |
### 3. RL training of Laguna-XS.2
|
|
@@ -95,6 +99,8 @@ GRPO on the train split (22 scenarios × 4 Python frameworks = 88 tasks):
|
|
| 95 |
|
| 96 |
The LoRA adapter from step 39 is checked into `lora_adapter/`. Final inference-time eval on the trained adapter was blocked because `poolside/Laguna-XS.2` is currently gated for LoRA deployment on Prime Intellect's inference infra (`Error: Base model is not currently available for LoRA deployment`). The training-time eval curve above measures the same held-out 24 tasks at every checkpoint, scored by the same env code — apples-to-apples, just smaller sample size than the full 100-task baseline.
|
| 97 |
|
|
|
|
|
|
|
| 98 |
---
|
| 99 |
|
| 100 |
## How to reproduce
|
|
|
|
| 19 |
> **Team:** Oof team
|
| 20 |
> **Author:** Aidarbek Suleimenov ([@idarbek](https://x.com/idarbek))
|
| 21 |
|
| 22 |
+
My submission is an RL environment, model evaluation, and RL post-trained model created from the original [BaxBench](https://arxiv.org/abs/2502.11844) secure-backend-code benchmark.
|
| 23 |
+
|
| 24 |
+
I wrapped the benchmark as a [Prime Intellect verifiers environment](https://app.primeintellect.ai/dashboard/environments/aidarbek/baxbench), used it to **evaluate Laguna-XS.2 against GPT-5.5**, and then **RL-trained Laguna-XS.2** on the train split. During training, the eval score went from **0.061 → 0.115** (+87% relative), but I couldn’t benchmark the final trained model due to limits with Prime Intellect’s LoRA deployments for Laguna-XS.2.
|
| 25 |
|
| 26 |
---
|
| 27 |
|
| 28 |
## Why this matters
|
| 29 |
|
| 30 |
+
As frontier models get better at writing code, so do the adversaries who use them. CrowdStrike’s 2026 Global Threat Report states that AI-enabled attacks are up roughly 89% year over year.
|
| 31 |
|
| 32 |
+
Since more and more new code is being written by AI, it’s important that models treat security as one of the key dimensions to optimise for.
|
| 33 |
|
| 34 |
---
|
| 35 |
|
|
|
|
| 81 |
| wall clock | 10.5 min | 11.3 min | +0.8 min |
|
| 82 |
| avg output tokens | 5,150 | 2,942 | −2,208 |
|
| 83 |
|
| 84 |
+
Why GPT-5.5? Simply because I had existing credits for OpenAI API :)
|
| 85 |
+
|
| 86 |
Head-to-head on the same 100 tasks: GPT-5.5 won 33 tasks Laguna didn't; Laguna won 4 tasks GPT-5.5 didn't (clustered on Flask/aiohttp). The gap is bigger on functional correctness (+0.315) than on security (+0.261), meaning a smaller open model is closer to GPT-5.5 on "writing secure code" than on "writing working code" — exactly the signal that says RL on this benchmark is worth running.
|
| 87 |
|
| 88 |
### 3. RL training of Laguna-XS.2
|
|
|
|
| 99 |
|
| 100 |
The LoRA adapter from step 39 is checked into `lora_adapter/`. Final inference-time eval on the trained adapter was blocked because `poolside/Laguna-XS.2` is currently gated for LoRA deployment on Prime Intellect's inference infra (`Error: Base model is not currently available for LoRA deployment`). The training-time eval curve above measures the same held-out 24 tasks at every checkpoint, scored by the same env code — apples-to-apples, just smaller sample size than the full 100-task baseline.
|
| 101 |
|
| 102 |
+
Unfortunately, I didn't have a time to properly benchmark post-trained model, the LoRa deployments weren't available, and running self-host model on Prime instances took too much time.
|
| 103 |
+
|
| 104 |
---
|
| 105 |
|
| 106 |
## How to reproduce
|