# Project rules — from experience These came out of real mistakes made while iterating on judge/subagent prompts (`agent/skills/pediatry_wiki/scripts/prompts_minimal_judge.py`, `evaluate_judge.py`) and apply project-wide, not just to that work. ## Ask before making a real API call Every subagent draft and every judge call here hits a live model and costs real money and real time (calls in this lab commonly take 5-30s each; a full `evaluate_judge.py` run is hundreds of calls). **Don't run one — including a "quick verification" after a prompt edit — without asking first.** Instead: describe the exact (question, reference, answer) scenarios you want to test and what verdict you expect, and let the user decide whether to run them (themselves via the UI, or by telling you to go ahead). This applies even when the motivation is "just confirming my fix worked." ## Prompt edits cost tokens on every future call, not just this one `MINIMAL_JUDGE_PROMPT` (and the subagent's `GENERATE_ANSWER_PROMPT`) gets resent in full on every LLM call — there's no prompt caching on the current provider setup, so nothing here is free to grow. A rule or example added to fix one disagreement is a permanent tax on every judge call from then on, in both tokens (cost) and latency. - Prefer the smallest addition that fixes the actual disagreement. Don't add a general-purpose rule when a one-line clarification covers the observed case. - When adding a new rule, look for an existing rule/example this makes redundant and remove it — the goal is a prompt that stays roughly constant-size as it's calibrated, not one that only grows. - Before adding a worked example, check whether the existing rule text alone would have been enough — an example is the most expensive way to convey a rule (many tokens for one data point). ## Few-shot examples must not be the literal case under test If a prompt's illustrative example is the exact same scenario as the failing case you're trying to fix, "PASS after the edit" proves nothing — the model can match text spans instead of applying the rule. This actually happened: the first fix to the fever/acetaminophen threshold case used that exact question/reference/answer as the in-prompt example, which would have made re-testing it meaningless. - Illustrate a rule with a **different domain** than any case you'll evaluate against (e.g. use the vaccine-age pattern to teach a rule you're fixing because of a fever/temperature failure). - Prefer synthetic/generic examples (a different illness, or "the vaccine is recommended for X" instead of the real fever wording) over reusing a real gold-set case verbatim as the anchor — save the real case as a held-out regression test instead, so it can actually confirm the rule generalized. ## Keep comments and docstrings short — and self-contained Comments should fit on one line, two at most. Docstrings should ideally not exceed two lines. If an explanation needs more than that, it's a sign the code itself (or the CLAUDE.md rules above) should carry that context instead of a growing comment block. A comment/docstring must also make sense to someone with no memory of the conversation that produced it: name the thing plainly instead of pointing at it. This actually happened — a docstring said "covers rules 4, 6, 21, 29, 30 (without the no-guidance fallback), 33", where those numbers only meant anything inside one chat's numbered list. - Never reference a rule/item by a number or label that only exists in a chat transcript. Say what the rule *is* in a few words instead. - A reader opening the file cold, with no chat history, must be able to understand it.