Spaces:
Running
RUNBOOK β End-to-end training
Two paths: AWS EKS (recommended, real cloud cluster) or local kind (free, Docker-based). Both use identical agent code β only the backend swaps.
For detailed env-var wiring see ENV.md.
| Path | Cost | Pros | Cons |
|---|---|---|---|
| AWS EKS | ~$4-5/day while running | real cloud, CloudWatch logs, S3 checkpoints | needs AWS account |
| Local kind | free | offline, fast iterate | no cloud services |
Path A β AWS EKS (recommended)
A.1 Provision
aws configure # one-time
.\scripts\setup_eks.ps1 # ~12 min
.\scripts\load_env.ps1 .env.aws.local
A.2 Install Python deps (isolated venv β recommended)
The global Python install on most machines is polluted with transformers/torch from other projects and will cause version conflicts. Use the isolated venv:
.\scripts\setup_venv.ps1 -Train # creates .venv and installs server + train deps
.\.venv\Scripts\Activate.ps1
Or install globally (only if you know your environment is clean):
pip install -r rl-agent/requirements.txt
pip install -r rl-agent/requirements-train.txt
A.3 Start the env server
python -m uvicorn server:app --host 127.0.0.1 --port 7860 --app-dir rl-agent
Verify: curl.exe http://localhost:7860/k8s/health β "enabled": true.
Open the multi-page dashboard: http://localhost:7860/dashboard.
A.4 Smoke-test real fault injection
curl.exe -X POST http://localhost:7860/k8s/inject `
-H "Content-Type: application/json" `
-d '{\"fault_type\":\"oom_kill\"}'
kubectl -n ic-payments get pods -w
curl.exe -X POST http://localhost:7860/k8s/reset
All 9 fault types: oom_kill, crashloop, image_pull, bad_config,
scale_zero, liveness_probe, resource_quota, secret_mismatch,
wrong_image_tag.
A.5 Training β Actor-Critic PPO
We moved off GRPO (needs a big GPU to batch N generations per prompt) to actor-critic PPO: one rollout per update, a small value head bootstraps V(s). Works on CPU for smoke tests, on 8 GB GPU for real training.
# Pipeline smoke test β scripted policy, no model load, CPU only
python -m training.train_ppo --tiny `
--updates 20 --rollouts-per-update 3 `
--env-url http://localhost:7860 `
--out-dir rl-agent/checkpoints/ppo-tiny
# Real PPO on CPU (slow β good for debugging)
python -m training.train_ppo --cpu --local `
--updates 10 --rollouts-per-update 2 `
--env-url http://localhost:7860 `
--out-dir rl-agent/checkpoints/ppo-cpu
# Real PPO on GPU (8 GB min, Qwen2.5-0.5B + LoRA r=8)
python -m training.train_ppo --local `
--updates 200 --rollouts-per-update 4 `
--env-url http://localhost:7860 `
--out-dir rl-agent/checkpoints/ppo-local
Every run writes metrics.jsonl + summary.json into its --out-dir and
those files drive the Training page of the dashboard
(http://localhost:7860/dashboard/training). Checkpoints are uploaded to
s3://$S3_CHECKPOINT_BUCKET/ppo/<timestamp>/final/ when AWS creds are set.
A.6 Evaluate
cd rl-agent
python -m eval --env-url http://localhost:7860
A.7 Teardown (important β EKS bills hourly)
.\scripts\teardown_eks.ps1
Path B β Local kind (no cloud)
| Tool | Where to get it | Windows quick-install |
|---|---|---|
| Docker Desktop | https://www.docker.com/products/docker-desktop | winget install Docker.DockerDesktop |
| kind | https://kind.sigs.k8s.io | choco install kind (or scoop/winget) |
| kubectl | https://kubernetes.io/docs/tasks/tools/ | choco install kubernetes-cli |
| Python 3.10+ | https://www.python.org | winget install Python.Python.3.11 |
| (optional) NVIDIA GPU w/ CUDA 12 | β | for real GRPO; CPU works for --dry-run |
Verify:
docker info; kind version; kubectl version --client; python --version
1. Spin up the local K8s cluster
From the repo root (E:\meta-rl-hack\incident-commander):
.\scripts\setup_kind.ps1
This creates a cluster named incident-commander with three app namespaces
(ic-payments, ic-frontend, ic-auth) and 5 healthy deployments.
Verify everything is green:
kubectl get pods -A | findstr ic-
Linux/macOS equivalent: bash scripts/setup_kind.sh.
2. Install Python deps
# server + k8s client
pip install -r rl-agent/requirements.txt
# trainer (torch, trl, peft, bitsandbytes). Only needed if you plan to train.
pip install -r rl-agent/requirements-train.txt
3. Start the environment server against the real cluster
$env:REAL_K8S = "true"
cd rl-agent
uvicorn server:app --host 127.0.0.1 --port 7860
Leave this terminal running. Open a second terminal for the next steps.
Sanity-check:
curl.exe http://localhost:7860/k8s/health
# -> {"enabled": true, "health": {"ic-payments": {...}, ...}}
4. Drive the cluster manually (optional smoke test)
Inject a real OOMKill fault:
curl.exe -X POST http://localhost:7860/k8s/inject `
-H "Content-Type: application/json" `
-d '{\"fault_type\":\"oom_kill\"}'
Watch the pod actually crash:
kubectl -n ic-payments get pods -w
# payments-api-xxx 0/1 OOMKilled 0 10s
Have the agent investigate via kubectl:
curl.exe -X POST http://localhost:7860/step `
-H "Content-Type: application/json" `
-d '{\"action_type\":\"exec_kubectl\",\"params\":{\"command\":\"kubectl get pods -n ic-payments\"}}'
Reset everything back to healthy:
curl.exe -X POST http://localhost:7860/k8s/reset
All nine fault types are supported:
oom_kill, crashloop, image_pull, bad_config, scale_zero,
liveness_probe, resource_quota, secret_mismatch, wrong_image_tag.
5. Dry-run training (no GPU, verifies pipeline end-to-end)
cd rl-agent
python -m training.train_grpo --local --dry-run
You should see one rollout with a cumulative reward printed.
6. Real GRPO training (local, ~8GB GPU)
cd rl-agent
python -m training.train_grpo --local --inject-fault oom_kill `
--env-url http://localhost:7860 `
--max-steps 50
What --local does:
- switches to
Qwen/Qwen2.5-0.5B-Instruct - LoRA
r=8, 4-bit quantization (fits in 8GB) num_generations=2-4,max_steps<=60- no Hub push, no vLLM colocate
Checkpoints land in rl-agent/checkpoints/grpo.
Without --local you get the full hackathon config (1.5B, 8 gens, vLLM
colocate β needs >=40GB VRAM).
7. Evaluate
cd rl-agent
python -m eval --env-url http://localhost:7860
8. Teardown
.\scripts\teardown_kind.ps1
Troubleshooting
| Symptom | Fix |
|---|---|
/k8s/health returns {"enabled": false} |
$env:REAL_K8S="true" was not set before uvicorn; restart it. |
Pods stuck ImagePullBackOff on first setup |
Docker Desktop is rate-limited; docker login with any account. |
kind not found after choco install |
Open a new PowerShell so PATH refreshes. |
ModuleNotFoundError: kubernetes |
pip install kubernetes>=29.0.0. |
| bitsandbytes install fails on Windows | pip install bitsandbytes-windows (prebuilt wheel). |
| Training OOMs on GPU | Lower --grad-accum, keep --local, or add --no-hub-push. |
What's real vs mock
| Layer | Mock mode (default / HF Space) | Real mode (REAL_K8S=true) |
|---|---|---|
kubectl get pods |
in-memory dict | live cluster API |
inject_failure |
flips a flag | patches live Deployment |
reset_to_healthy |
restores dict | patch_namespaced_deployment for each tracked deploy |
| Reward signal | heuristic + LLM judge | identical + real pod phase feedback |
The agent code is the same in both modes β the backend swap is transparent.