foam-cfd-unified-14b-private (v3 cycle-2 candidate)

A private v3 candidate of the foam-cfd-unified-14b agent, produced by one cycle of the seven-layer self-evolution pipeline introduced in v3.

This checkpoint = base merged 14B (v2) $+$ one corrective QLoRA epoch (loss $0.35\to0.04$) on a 289-row training mix:

  • 184 rows curated from the rolling capture file (score $\geq 0.65$)
  • 79 rows from the frozen v1 anchor corpus (Layer-5 anchor mix at 30%)
  • 26 rows of fresh active-learning prompts on the weakest solver family from the previous OOD eval (Layer-6 active learning, target family = pimpleFoam)

Why this is private

This is an internal candidate, not the production v2 model. Use the public foam-cfd-unified-14b unless you specifically need to A/B against the v3 cycle-2 changes documented below.

Behaviour vs the v2 baseline

On the held-out 110-prompt OOD evaluation in raw-LLM mode (agent-side keyword guards disabled):

Metric v2 (public) v3 cycle-2 (this)
End-to-end PASS rate $110/110 = 100.0%$ $110/110 = 100.0%$
Solver-pick exact match $106/110 = 96.4%$ $106/110 = 96.4%$
Same 4 borderline mismatches yes yes
Per-prompt FILES_CHANGE n/a (baseline) 4 prompts (3 attributable to active-learning training, 1 noise floor)
Per-prompt regression flags n/a 0

The aggregate eval gate would have promoted both models equally; the per-prompt regression diff (Layer 7) reveals 3 prompts whose case files differ as a result of the active-learning training, with no SUCCESS_FLIP, SOLVER_CHANGE, or SCORE_DROP.

How it was trained

  • Base model: arungovindneelan/foam-cfd-unified-14b (v2 merged, bf16)
  • Method: QLoRA via Unsloth, then bf16 merge for vLLM
  • LoRA configuration: $r = 64$, $\alpha = 128$, dropout $= 0.0$, target $= {q,k,v,o,gate,up,down}_\text{proj}$
  • Schedule: 1 corrective epoch on top of v2, paged AdamW 8-bit, bf16, batch $1 \times$ grad-accum $8$, max sequence length 8192, learning rate $2 \times 10^{-4}$, cosine schedule, 5% warm-up. Reward-weighted loss with weights $w_i = s_i^2 / \overline{s^2}$.
  • Training data: 289-row mix described above
  • Hardware: single H100 80GB
  • Wall time: 5 min training $+$ 1 min adapter merge

Quick start

from vllm import LLM, SamplingParams
llm = LLM(
    model="arungovindneelan/foam-cfd-unified-14b-private",
    dtype="bfloat16",
    max_model_len=8192,
    gpu_memory_utilization=0.55,
)
out = llm.chat(
    [{"role": "system", "content": "You are an expert OpenFOAM CFD engineer..."},
     {"role": "user",   "content": "2D lid-driven cavity Re=1000, 2m square, water"}],
    sampling_params=SamplingParams(temperature=0.0, max_tokens=1024),
)
print(out[0].outputs[0].text)

Related artefacts

Downloads last month
9
Safetensors
Model size
15B params
Tensor type
F32
路
BF16
路
U8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for arungovindneelan/foam-cfd-unified-14b-private