Training Learnings: What Failed, What Worked, and What We Shipped
Adaptive Clinical Recruitment: What Failed, What Worked, and What We Learned Training a Long-Horizon GRPO Agent
Adaptive Clinical Recruitment is a long-horizon OpenEnv benchmark for data-driven trial-planning decisions. It models the patient funnel over 180 simulated steps, exposes typed observations and actions, and lets agents balance screening, follow-up, site allocation, planning, memory use, and budget pressure from a live environment URL.
This writeup stays on Theme #2 only: what the repo currently supports as a benchmark package, what training evidence is committed today, and what is still missing from the strongest possible submission story.
This post is intentionally conservative. It describes the current repo state after a re-audit, a corrected evaluation pass, and a fresh 5-seed sweep.
TL;DR
- The benchmark has
3public tasks,8implemented action types, and a live judge-facing environment athttps://pratimassaravanan-clinical-recruitment.hf.space. - The final GRPO run that actually learned was HF Job
69ed378ed2c8bd8662bce819: Qwen3-1.7B, NVIDIA L4,30steps,4generations per prompt, about165Ktokens. - The model learned the real funnel order
screen_patient -> recontact -> allocate_to_site, which was the main failure mode in earlier runs. - SFT helped with output format and action diversity, but not with enrollment strategy.
- The biggest unlock was not a larger model. It was the combination of a TRL-compatible tool interface, non-zero reward variance, easy-task curriculum, and HF Jobs memory headroom.
- Trained adapter and completions:
pratimassaravanan/grpo_output - Large traces, adapters, and checkpoints are now stored separately in
pratimassaravanan/clinical-recruitment-artifactsso GitHub stays clean. - After rerunning the corrected sweep,
HCAPOhas the highest mean score at0.2215, but no pairwise comparison reachesp < 0.05.
Recently Updated Files and Artifacts
| File or repo | What changed | Why it matters |
|---|---|---|
README.md |
Submission links, artifact links, and learnings summary refreshed | Gives judges one accurate entry point |
HUGGINGFACE_BLOG.md |
Reframed into a learnings-first article | Explains the real engineering journey, not just the final result |
TRAINING_LEARNINGS.md |
Preserved the full engineering log | Captures failed attempts, fixes, and infra issues |
train_grpo_hfjob.py |
Final HF Jobs training path uses TRL environment_factory plus HTTP env wrapper |
This is the script behind the successful L4 run |
train_grpo_trl.py |
Local TRL reference path kept aligned with the same tool-method pattern | Makes the final training pattern reproducible outside HF Jobs |
notebooks/clinical_recruitment_grpo.ipynb |
Judge-rerunnable Colab notebook refreshed | Gives a lower-friction re-run path on T4 |
scripts/plot_training_curves.py |
Regenerated training evidence plots | Keeps visual evidence tied to current results |
data/training_outputs/grpo_trl_metrics.json |
Real metrics extracted from the completed L4 job | Provides the actual reward/loss/grad trace |
data/training_outputs/sft_grpo_results.json |
Retained the earlier SFT pilot evidence | Shows what improved before GRPO was working |
demo/grpo_loss_curve.png |
Real GRPO loss plot from the successful run | Shows non-zero training signal |
demo/grpo_training_summary.png |
Real GRPO summary plot from the successful run | Shows token count, LR schedule, and completion length |
demo/sft_loss_curve.png |
SFT loss evidence refreshed | Shows that SFT changed format even when it did not solve the task |
demo/llm_before_after.png |
Before/after action behavior figure refreshed | Makes the SFT behavior shift visible |
demo/grpo_reward_design.png |
Reward design visualization refreshed | Explains why the final reward function worked better |
demo/heuristic_comparison.png |
Baseline comparison figure refreshed | Grounds the RL story in a benchmark context |
pratimassaravanan/grpo_output |
Final trained LoRA adapter and 30 completion parquet files | Public artifact for the successful GRPO run |
pratimassaravanan/clinical-recruitment-artifacts |
Large traces, SFT checkpoints, and adapter outputs uploaded | Keeps large training outputs off GitHub while preserving reproducibility |
Why this is an interesting training target
Clinical recruitment is not a one-step prediction problem. It is a workflow-shaped resource-allocation loop:
- which patients to screen first
- when to spend effort on recontact
- which sites deserve scarce capacity
- when to change strategy under budget and dropout pressure
Those decisions play out over weeks or months of simulated time, with delayed effects, constraint pressure, and recovery actions. That makes the environment a better fit for Theme #2 than a short-horizon reward toy.
What the benchmark exposes
At each step, the environment returns a typed Observation with:
- Funnel state and action-specific candidate pools
- Per-site performance metrics
- Milestones and delayed-effects state
- Constraint and uncertainty summaries
- Plan state and indexed-memory summaries
- Token accounting and token-efficiency signals
- Counterfactual hints and simple rollout estimates
The current action interface contains exactly these 8 actions:
screen_patientrecontactallocate_to_siteadjust_strategyplan_next_phasesummarize_and_indexretrieve_relevant_historystop_recruitment
Site negotiation is represented through adjust_strategy values such as negotiate_site_A, not as a separate ninth or tenth action.
What changed during the re-audit
Before regenerating results, we fixed several issues that made the older docs too optimistic.
- The experiment path previously used
available_patientsforrecontactandallocate_to_site, even though those actions should draw from their own candidate pools. - The sweep charts could fail on small seed counts because of malformed error bars.
- The chart refresh path updated
data/sweep_results/but could leavedocs/images/stale. - Several public docs still described a
10-action interface and repeated an outdated significance claim.
The current docs and charts now follow the corrected benchmark path.
Fresh 5-seed sweep
The regenerated report lives in data/sweep_results/benchmark_report.{md,json}.
| Baseline | Mean | Std | 95% CI |
|---|---|---|---|
HCAPO |
0.2215 |
0.0127 |
[0.2100, 0.2303] |
KLong |
0.2152 |
0.0222 |
[0.1977, 0.2286] |
MemexRL |
0.2148 |
0.0270 |
[0.1943, 0.2352] |
MiRA |
0.2094 |
0.0095 |
[0.2023, 0.2165] |
Pairwise tests from the same report show:
HCAPO vs MiRA:p = 0.1823HCAPO vs KLong:p = 0.3849HCAPO vs MemexRL:p = 0.6370- no comparison reaches
p < 0.05
That means the honest headline is not "hierarchical planning wins." The honest headline is:
The repo exposes a real benchmark surface, but the current baseline suite remains too tightly clustered to support a winner claim.
Current training evidence
The repository contains committed training artifacts with rendered plots and a re-runnable training pipeline.
SFT Training (Completed)
data/training_outputs/sft_grpo_results.jsoncaptures a Tesla T4SFT -> GRPOpilot on Qwen3-4B.- SFT loss fell from
0.858to0.745(13.2%reduction) over 9 steps. - The model expanded from
1repeated action to5distinct action types. - JSON parse rate improved from ~0% to ~85%.
- Enrollment after training remained
0โ this is evidence of format learning and action diversification, not full task mastery.
GRPO Training Pipeline (Completed โ Real Run)
A successful 30-step GRPO training run on Qwen3-1.7B (NVIDIA L4 GPU):
- Trained model pushed to Hub:
pratimassaravanan/grpo_output - 30 steps, 4 generations/prompt, ~165K tokens processed, ~27 min on L4
- LoRA adapter (69.8MB) + 30 completion parquets with per-step tool-call traces
- The training loop connected to the live HF Space via HTTP โ TRL executed tool calls against real environment state
- Large supporting artifacts now live separately in
pratimassaravanan/clinical-recruitment-artifacts, includingtraces/sft_traces_5k*.json,lightning_sft_checkpoint/*, andlora_adapter/*
The pipeline uses TRL's environment_factory:
from trl import GRPOTrainer
from tool_env import ClinicalRecruitmentToolEnv, REWARD_FUNCS
trainer = GRPOTrainer(
model="Qwen/Qwen3-0.6B",
reward_funcs=REWARD_FUNCS,
environment_factory=ClinicalRecruitmentToolEnv,
...
)
trainer.train()
train_grpo_hfjob.pyโ GRPO via HF Jobs with HTTP env wrapper (production, completed 30 steps)train_grpo_trl.pyโ Standalone GRPO script with local envnotebooks/clinical_recruitment_grpo.ipynbโ Colab T4 notebook (re-runnable by judges)
GRPO Training Results (HF Jobs, L4 GPU)
GRPO training completed 30 steps on NVIDIA L4 against the live HF Space environment:
| Metric | Early (1-10) | Late (20-30) | Trend |
|---|---|---|---|
| Reward mean | 0.3088 | 0.3153 | Slight upward |
| Loss | 0-0.009 | 0-0.012 | Learning signal present |
| Grad norm | 0-0.29 | 0.20-0.29 | Stable, healthy |
| Tool calls/ep | 6 | 6 | Consistent |
| Failures | 0 | 0 | Clean execution |
What the model learned through GRPO:
- Calls
screen_patientas first action (previously collapsed toadjust_strategywith reward=0) - Follows the correct pipeline: screen โ recontact โ enrollment
- Enrolls ~1 patient per episode on easy_bench
- Generations that move to screening a second patient faster get higher rewards (+0.49 advantage)
What it didn't learn (honest limits):
- Reward variance between 4 generations is only ~0.005, limiting GRPO signal strength
- Reward stays flat at ~0.31 without dramatic upward trend
- With 6 tool calls per episode (1024 token limit), only 1-2 patients processed
Trained adapter: pratimassaravanan/grpo_output (LoRA r=16 on Qwen3-1.7B)
The training journey involved 3 failed attempts before success โ see TRAINING_LEARNINGS.md for the full engineering log including OOM on T4, zero-reward diagnosis, and the progressive reward fix that unblocked learning.
Engineering Learnings That Actually Changed the Result
1. TRL tool discovery is strict
Our earliest GRPO path used a generic step(action) style wrapper. That looked reasonable from an RL perspective, but it was the wrong abstraction for TRL OpenEnv training. The final working path exposed explicit public tool methods like screen_patient, recontact, and allocate_to_site, each with typed arguments and docstrings. That change is the difference between a trainer that can build tool schemas and a trainer that silently learns nothing useful.
2. SFT taught format faster than strategy
The SFT pilot was not wasted effort. It clearly improved loss, JSON formatting, and action diversity. But it did not solve enrollment. The honest learning here is that short SFT runs on strong instruct models mostly teach output shape first. Strategy improvement required an online reward signal from the real environment.
3. Memory headroom mattered more than ambition
The T4 path failed once the full GRPO backward pass hit memory limits. Moving to L4 was not about prestige hardware; it was about getting one stable end-to-end run with enough memory to finish. That run gave us the first real evidence that the model could enroll patients instead of collapsing to adjust_strategy.
4. Reward variance is the real bottleneck for GRPO here
The early zero-reward runs were not just a model problem. They were a signal-design problem. Medium and hard tasks often start with no immediately available patients, so the model defaulted to adjust_strategy, got zero reward, and never separated one generation from another. Switching to easy_bench only and awarding partial credit for screened and consented states created the first usable gradient signal.
5. Some environment bugs were actively suppressing planning and memory
The engineering log uncovered several reward and serving issues: planning actions were over-penalized, consistency penalties were effectively unbounded, current-step events leaked into observations, and session handling in the adapter was too loose. Those were not cosmetic bugs. They changed what behaviors the model could or could not learn.
6. Instruct-model behavior can work against tool calling
Short SFT runs on Qwen3 mostly produced verbose reasoning or copied example IDs. DeepSeek-style reasoning models added <think> blocks that polluted action parsing. The lesson was not that these models are unusable in general. It was that this benchmark strongly rewards clean, low-ceremony tool calls, and the training pipeline has to be designed around that constraint.
7. Storage discipline matters for hackathon delivery
The large trace JSONs and adapter checkpoints were useful artifacts, but they were the wrong thing to keep in GitHub history. Moving them into a dedicated Hugging Face artifacts repo made the project easier to ship and easier to audit. The benchmark code stays lightweight, while the heavy outputs remain public and reproducible.
Heuristic Agent Improvement
The heuristic agent improves +23.9% over random baseline. The optimized agent adds hypothesis-aware and memory strategies for additional gains on hard_bench.
Reward Function Design
Multi-component reward signal with 7 positive components (max +0.90) and a format penalty (-0.25). The TRL reward functions score 5 environment-state dimensions.
Reward Trajectory View
This view is useful because it shows benchmark-side behavior, not just trainer-side loss. The heuristic and optimized agents build cumulative reward faster than the random baseline, which helps explain why imitation-style traces were a sensible starting point even before GRPO was fully working.
Why the live URL matters
The environment is already served at:
https://pratimassaravanan-clinical-recruitment.hf.space
That matters because the benchmark is not only a local code artifact. The repo exposes a judge-facing URL, local FastAPI/OpenEnv serving, and a separate tool_env.py training wrapper for TRL/OpenEnv-style experiments.
What that means
This repo is a benchmark package with real training evidence and a working end-to-end pipeline.
- The environment interface is typed and deterministic.
- The action construction path matches the observation schema.
- The main diagrams and sweep charts are regenerated from code.
- The integration checks pass
30/30acrosseasy_bench,medium_bench, andhard_bench. - SFT training produced measurable improvement (loss, action diversity, JSON format).
- GRPO training produced a model that enrolls patients โ the first successful end-to-end RL result.
- The trained LoRA adapter is published on HF Hub with full training completions.
- Training evidence plots are committed as
.pngfiles with labeled axes.
Trying the benchmark
Local API
uvicorn app:app --host 0.0.0.0 --port 7860
Minimal Python loop
from env import ClinicalRecruitmentEnv
from models import Action
env = ClinicalRecruitmentEnv()
result = env.reset(task="medium_bench")
obs = result.observation
action = Action(
action_type="screen_patient",
patient_id=obs.available_patients[0]["id"],
hypothesis="noise_dominant",
confidence=0.7,
)
result = env.step(action)
print(result.reward)
Reproducing the sweep
python experiments/full_sweep.py --seeds 1 7 21 42 123 --episodes 30 --eval-episodes 5
What this post does not claim
- It does not claim a
10-action interface. - It does not claim that all
50roadmap features are implemented and validated. - It does not claim externally validated reproductions or benchmark-leading status for external named methods.
- It does not claim a statistically significant
HCAPOwin. - It does not claim that GRPO training solved the benchmark โ the model enrolls ~1/80 patients per episode, far from the 80-patient target.
- It does not claim dramatic GRPO learning curves โ reward improved from 0.3088 to 0.3153 (a modest signal).
Key files
README.md: current repo overviewtrain_grpo_hfjob.py: GRPO training script for HF Jobs (production, completed 30 steps)train_grpo_trl.py: GRPO training script (TRLenvironment_factorypattern, local env)notebooks/clinical_recruitment_grpo.ipynb: Colab T4 notebook (re-runnable)TRAINING_LEARNINGS.md: detailed engineering log of failed attempts and fixestool_env.py: TRL-compatible environment wrapper with tool methodsexperiments/grpo_metrics.json: GRPO training metrics (30 steps, reward/loss/grad)data/training_outputs/grpo_trl_metrics.json: extracted metrics from HF Job69ed378ed2c8bd8662bce819docs/theme2_alignment.md: conservative Theme #2 mappingdocs/theme2_completion_checklist.md: reality-based status filedata/sweep_results/benchmark_report.md: fresh benchmark summarydata/training_outputs/sft_grpo_results.json: committed pilot LLM training artifactdemo/sft_loss_curve.png: SFT training loss plotdemo/llm_before_after.png: before/after LLM behavior comparisondemo/grpo_reward_design.png: multi-component reward design visualizationdemo/heuristic_comparison.png: agent performance comparisonpaper/main.pdf: current anonymous paper build
Links
- GitHub repository:
https://github.com/pratimassaravanan/clinical-recruitment - Hugging Face Space:
https://huggingface.co/spaces/pratimassaravanan/clinical-recruitment - Trained model:
https://huggingface.co/pratimassaravanan/grpo_output - Large artifacts repo:
https://huggingface.co/pratimassaravanan/clinical-recruitment-artifacts - Live environment URL:
https://pratimassaravanan-clinical-recruitment.hf.space






