Training Learnings: What Failed, What Worked, and What We Shipped

#1
by pratimassaravanan - opened

Adaptive Clinical Recruitment: What Failed, What Worked, and What We Learned Training a Long-Horizon GRPO Agent

Adaptive Clinical Recruitment is a long-horizon OpenEnv benchmark for data-driven trial-planning decisions. It models the patient funnel over 180 simulated steps, exposes typed observations and actions, and lets agents balance screening, follow-up, site allocation, planning, memory use, and budget pressure from a live environment URL.

This writeup stays on Theme #2 only: what the repo currently supports as a benchmark package, what training evidence is committed today, and what is still missing from the strongest possible submission story.

This post is intentionally conservative. It describes the current repo state after a re-audit, a corrected evaluation pass, and a fresh 5-seed sweep.

TL;DR

  • The benchmark has 3 public tasks, 8 implemented action types, and a live judge-facing environment at https://pratimassaravanan-clinical-recruitment.hf.space.
  • The final GRPO run that actually learned was HF Job 69ed378ed2c8bd8662bce819: Qwen3-1.7B, NVIDIA L4, 30 steps, 4 generations per prompt, about 165K tokens.
  • The model learned the real funnel order screen_patient -> recontact -> allocate_to_site, which was the main failure mode in earlier runs.
  • SFT helped with output format and action diversity, but not with enrollment strategy.
  • The biggest unlock was not a larger model. It was the combination of a TRL-compatible tool interface, non-zero reward variance, easy-task curriculum, and HF Jobs memory headroom.
  • Trained adapter and completions: pratimassaravanan/grpo_output
  • Large traces, adapters, and checkpoints are now stored separately in pratimassaravanan/clinical-recruitment-artifacts so GitHub stays clean.
  • After rerunning the corrected sweep, HCAPO has the highest mean score at 0.2215, but no pairwise comparison reaches p < 0.05.

Recently Updated Files and Artifacts

File or repo What changed Why it matters
README.md Submission links, artifact links, and learnings summary refreshed Gives judges one accurate entry point
HUGGINGFACE_BLOG.md Reframed into a learnings-first article Explains the real engineering journey, not just the final result
TRAINING_LEARNINGS.md Preserved the full engineering log Captures failed attempts, fixes, and infra issues
train_grpo_hfjob.py Final HF Jobs training path uses TRL environment_factory plus HTTP env wrapper This is the script behind the successful L4 run
train_grpo_trl.py Local TRL reference path kept aligned with the same tool-method pattern Makes the final training pattern reproducible outside HF Jobs
notebooks/clinical_recruitment_grpo.ipynb Judge-rerunnable Colab notebook refreshed Gives a lower-friction re-run path on T4
scripts/plot_training_curves.py Regenerated training evidence plots Keeps visual evidence tied to current results
data/training_outputs/grpo_trl_metrics.json Real metrics extracted from the completed L4 job Provides the actual reward/loss/grad trace
data/training_outputs/sft_grpo_results.json Retained the earlier SFT pilot evidence Shows what improved before GRPO was working
demo/grpo_loss_curve.png Real GRPO loss plot from the successful run Shows non-zero training signal
demo/grpo_training_summary.png Real GRPO summary plot from the successful run Shows token count, LR schedule, and completion length
demo/sft_loss_curve.png SFT loss evidence refreshed Shows that SFT changed format even when it did not solve the task
demo/llm_before_after.png Before/after action behavior figure refreshed Makes the SFT behavior shift visible
demo/grpo_reward_design.png Reward design visualization refreshed Explains why the final reward function worked better
demo/heuristic_comparison.png Baseline comparison figure refreshed Grounds the RL story in a benchmark context
pratimassaravanan/grpo_output Final trained LoRA adapter and 30 completion parquet files Public artifact for the successful GRPO run
pratimassaravanan/clinical-recruitment-artifacts Large traces, SFT checkpoints, and adapter outputs uploaded Keeps large training outputs off GitHub while preserving reproducibility

Why this is an interesting training target

Clinical recruitment is not a one-step prediction problem. It is a workflow-shaped resource-allocation loop:

  • which patients to screen first
  • when to spend effort on recontact
  • which sites deserve scarce capacity
  • when to change strategy under budget and dropout pressure

Those decisions play out over weeks or months of simulated time, with delayed effects, constraint pressure, and recovery actions. That makes the environment a better fit for Theme #2 than a short-horizon reward toy.

What the benchmark exposes

At each step, the environment returns a typed Observation with:

  • Funnel state and action-specific candidate pools
  • Per-site performance metrics
  • Milestones and delayed-effects state
  • Constraint and uncertainty summaries
  • Plan state and indexed-memory summaries
  • Token accounting and token-efficiency signals
  • Counterfactual hints and simple rollout estimates

The current action interface contains exactly these 8 actions:

  1. screen_patient
  2. recontact
  3. allocate_to_site
  4. adjust_strategy
  5. plan_next_phase
  6. summarize_and_index
  7. retrieve_relevant_history
  8. stop_recruitment

Site negotiation is represented through adjust_strategy values such as negotiate_site_A, not as a separate ninth or tenth action.

What changed during the re-audit

Before regenerating results, we fixed several issues that made the older docs too optimistic.

  1. The experiment path previously used available_patients for recontact and allocate_to_site, even though those actions should draw from their own candidate pools.
  2. The sweep charts could fail on small seed counts because of malformed error bars.
  3. The chart refresh path updated data/sweep_results/ but could leave docs/images/ stale.
  4. Several public docs still described a 10-action interface and repeated an outdated significance claim.

The current docs and charts now follow the corrected benchmark path.

Fresh 5-seed sweep

The regenerated report lives in data/sweep_results/benchmark_report.{md,json}.

Baseline Mean Std 95% CI
HCAPO 0.2215 0.0127 [0.2100, 0.2303]
KLong 0.2152 0.0222 [0.1977, 0.2286]
MemexRL 0.2148 0.0270 [0.1943, 0.2352]
MiRA 0.2094 0.0095 [0.2023, 0.2165]

Pairwise tests from the same report show:

  • HCAPO vs MiRA: p = 0.1823
  • HCAPO vs KLong: p = 0.3849
  • HCAPO vs MemexRL: p = 0.6370
  • no comparison reaches p < 0.05

That means the honest headline is not "hierarchical planning wins." The honest headline is:

The repo exposes a real benchmark surface, but the current baseline suite remains too tightly clustered to support a winner claim.

Current training evidence

The repository contains committed training artifacts with rendered plots and a re-runnable training pipeline.

SFT Training (Completed)

SFT Training Loss

  • data/training_outputs/sft_grpo_results.json captures a Tesla T4 SFT -> GRPO pilot on Qwen3-4B.
  • SFT loss fell from 0.858 to 0.745 (13.2% reduction) over 9 steps.
  • The model expanded from 1 repeated action to 5 distinct action types.
  • JSON parse rate improved from ~0% to ~85%.
  • Enrollment after training remained 0 โ€” this is evidence of format learning and action diversification, not full task mastery.

Before vs After SFT

GRPO Training Pipeline (Completed โ€” Real Run)

A successful 30-step GRPO training run on Qwen3-1.7B (NVIDIA L4 GPU):

GRPO Loss Curve

GRPO Training Summary

  • Trained model pushed to Hub: pratimassaravanan/grpo_output
  • 30 steps, 4 generations/prompt, ~165K tokens processed, ~27 min on L4
  • LoRA adapter (69.8MB) + 30 completion parquets with per-step tool-call traces
  • The training loop connected to the live HF Space via HTTP โ€” TRL executed tool calls against real environment state
  • Large supporting artifacts now live separately in pratimassaravanan/clinical-recruitment-artifacts, including traces/sft_traces_5k*.json, lightning_sft_checkpoint/*, and lora_adapter/*

The pipeline uses TRL's environment_factory:

from trl import GRPOTrainer
from tool_env import ClinicalRecruitmentToolEnv, REWARD_FUNCS

trainer = GRPOTrainer(
    model="Qwen/Qwen3-0.6B",
    reward_funcs=REWARD_FUNCS,
    environment_factory=ClinicalRecruitmentToolEnv,
    ...
)
trainer.train()
  • train_grpo_hfjob.py โ€” GRPO via HF Jobs with HTTP env wrapper (production, completed 30 steps)
  • train_grpo_trl.py โ€” Standalone GRPO script with local env
  • notebooks/clinical_recruitment_grpo.ipynb โ€” Colab T4 notebook (re-runnable by judges)

GRPO Training Results (HF Jobs, L4 GPU)

GRPO training completed 30 steps on NVIDIA L4 against the live HF Space environment:

Metric Early (1-10) Late (20-30) Trend
Reward mean 0.3088 0.3153 Slight upward
Loss 0-0.009 0-0.012 Learning signal present
Grad norm 0-0.29 0.20-0.29 Stable, healthy
Tool calls/ep 6 6 Consistent
Failures 0 0 Clean execution

What the model learned through GRPO:

  • Calls screen_patient as first action (previously collapsed to adjust_strategy with reward=0)
  • Follows the correct pipeline: screen โ†’ recontact โ†’ enrollment
  • Enrolls ~1 patient per episode on easy_bench
  • Generations that move to screening a second patient faster get higher rewards (+0.49 advantage)

What it didn't learn (honest limits):

  • Reward variance between 4 generations is only ~0.005, limiting GRPO signal strength
  • Reward stays flat at ~0.31 without dramatic upward trend
  • With 6 tool calls per episode (1024 token limit), only 1-2 patients processed

Trained adapter: pratimassaravanan/grpo_output (LoRA r=16 on Qwen3-1.7B)

The training journey involved 3 failed attempts before success โ€” see TRAINING_LEARNINGS.md for the full engineering log including OOM on T4, zero-reward diagnosis, and the progressive reward fix that unblocked learning.

Engineering Learnings That Actually Changed the Result

1. TRL tool discovery is strict

Our earliest GRPO path used a generic step(action) style wrapper. That looked reasonable from an RL perspective, but it was the wrong abstraction for TRL OpenEnv training. The final working path exposed explicit public tool methods like screen_patient, recontact, and allocate_to_site, each with typed arguments and docstrings. That change is the difference between a trainer that can build tool schemas and a trainer that silently learns nothing useful.

2. SFT taught format faster than strategy

The SFT pilot was not wasted effort. It clearly improved loss, JSON formatting, and action diversity. But it did not solve enrollment. The honest learning here is that short SFT runs on strong instruct models mostly teach output shape first. Strategy improvement required an online reward signal from the real environment.

3. Memory headroom mattered more than ambition

The T4 path failed once the full GRPO backward pass hit memory limits. Moving to L4 was not about prestige hardware; it was about getting one stable end-to-end run with enough memory to finish. That run gave us the first real evidence that the model could enroll patients instead of collapsing to adjust_strategy.

4. Reward variance is the real bottleneck for GRPO here

The early zero-reward runs were not just a model problem. They were a signal-design problem. Medium and hard tasks often start with no immediately available patients, so the model defaulted to adjust_strategy, got zero reward, and never separated one generation from another. Switching to easy_bench only and awarding partial credit for screened and consented states created the first usable gradient signal.

5. Some environment bugs were actively suppressing planning and memory

The engineering log uncovered several reward and serving issues: planning actions were over-penalized, consistency penalties were effectively unbounded, current-step events leaked into observations, and session handling in the adapter was too loose. Those were not cosmetic bugs. They changed what behaviors the model could or could not learn.

6. Instruct-model behavior can work against tool calling

Short SFT runs on Qwen3 mostly produced verbose reasoning or copied example IDs. DeepSeek-style reasoning models added <think> blocks that polluted action parsing. The lesson was not that these models are unusable in general. It was that this benchmark strongly rewards clean, low-ceremony tool calls, and the training pipeline has to be designed around that constraint.

7. Storage discipline matters for hackathon delivery

The large trace JSONs and adapter checkpoints were useful artifacts, but they were the wrong thing to keep in GitHub history. Moving them into a dedicated Hugging Face artifacts repo made the project easier to ship and easier to audit. The benchmark code stays lightweight, while the heavy outputs remain public and reproducible.

Heuristic Agent Improvement

Agent Comparison

The heuristic agent improves +23.9% over random baseline. The optimized agent adds hypothesis-aware and memory strategies for additional gains on hard_bench.

Reward Function Design

Reward Components

Multi-component reward signal with 7 positive components (max +0.90) and a format penalty (-0.25). The TRL reward functions score 5 environment-state dimensions.

Reward Trajectory View

Reward Curve

This view is useful because it shows benchmark-side behavior, not just trainer-side loss. The heuristic and optimized agents build cumulative reward faster than the random baseline, which helps explain why imitation-style traces were a sensible starting point even before GRPO was fully working.

Why the live URL matters

The environment is already served at:

  • https://pratimassaravanan-clinical-recruitment.hf.space

That matters because the benchmark is not only a local code artifact. The repo exposes a judge-facing URL, local FastAPI/OpenEnv serving, and a separate tool_env.py training wrapper for TRL/OpenEnv-style experiments.

What that means

This repo is a benchmark package with real training evidence and a working end-to-end pipeline.

  • The environment interface is typed and deterministic.
  • The action construction path matches the observation schema.
  • The main diagrams and sweep charts are regenerated from code.
  • The integration checks pass 30/30 across easy_bench, medium_bench, and hard_bench.
  • SFT training produced measurable improvement (loss, action diversity, JSON format).
  • GRPO training produced a model that enrolls patients โ€” the first successful end-to-end RL result.
  • The trained LoRA adapter is published on HF Hub with full training completions.
  • Training evidence plots are committed as .png files with labeled axes.

Trying the benchmark

Local API

uvicorn app:app --host 0.0.0.0 --port 7860

Minimal Python loop

from env import ClinicalRecruitmentEnv
from models import Action

env = ClinicalRecruitmentEnv()
result = env.reset(task="medium_bench")
obs = result.observation

action = Action(
    action_type="screen_patient",
    patient_id=obs.available_patients[0]["id"],
    hypothesis="noise_dominant",
    confidence=0.7,
)

result = env.step(action)
print(result.reward)

Reproducing the sweep

python experiments/full_sweep.py --seeds 1 7 21 42 123 --episodes 30 --eval-episodes 5

What this post does not claim

  • It does not claim a 10-action interface.
  • It does not claim that all 50 roadmap features are implemented and validated.
  • It does not claim externally validated reproductions or benchmark-leading status for external named methods.
  • It does not claim a statistically significant HCAPO win.
  • It does not claim that GRPO training solved the benchmark โ€” the model enrolls ~1/80 patients per episode, far from the 80-patient target.
  • It does not claim dramatic GRPO learning curves โ€” reward improved from 0.3088 to 0.3153 (a modest signal).

Key files

  • README.md: current repo overview
  • train_grpo_hfjob.py: GRPO training script for HF Jobs (production, completed 30 steps)
  • train_grpo_trl.py: GRPO training script (TRL environment_factory pattern, local env)
  • notebooks/clinical_recruitment_grpo.ipynb: Colab T4 notebook (re-runnable)
  • TRAINING_LEARNINGS.md: detailed engineering log of failed attempts and fixes
  • tool_env.py: TRL-compatible environment wrapper with tool methods
  • experiments/grpo_metrics.json: GRPO training metrics (30 steps, reward/loss/grad)
  • data/training_outputs/grpo_trl_metrics.json: extracted metrics from HF Job 69ed378ed2c8bd8662bce819
  • docs/theme2_alignment.md: conservative Theme #2 mapping
  • docs/theme2_completion_checklist.md: reality-based status file
  • data/sweep_results/benchmark_report.md: fresh benchmark summary
  • data/training_outputs/sft_grpo_results.json: committed pilot LLM training artifact
  • demo/sft_loss_curve.png: SFT training loss plot
  • demo/llm_before_after.png: before/after LLM behavior comparison
  • demo/grpo_reward_design.png: multi-component reward design visualization
  • demo/heuristic_comparison.png: agent performance comparison
  • paper/main.pdf: current anonymous paper build

Links

  • GitHub repository: https://github.com/pratimassaravanan/clinical-recruitment
  • Hugging Face Space: https://huggingface.co/spaces/pratimassaravanan/clinical-recruitment
  • Trained model: https://huggingface.co/pratimassaravanan/grpo_output
  • Large artifacts repo: https://huggingface.co/pratimassaravanan/clinical-recruitment-artifacts
  • Live environment URL: https://pratimassaravanan-clinical-recruitment.hf.space

Sign up or log in to comment