Trinity Nano Base GSM8K AdamW-LoRA profile gate
This repo currently preserves the 8xH200 profile-gate artifacts for the Trinity Nano Base baseline. The full run is intentionally not uploaded yet because the best measured 8-GPU profile reached only about 0.95% estimated MFU at per-device batch size 8, with max memory around 131.8 GiB per H200.
See best_profile.json and the profiles directory for throughput, memory, W&B URLs, and adapter/profile artifacts.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for TokenBender/doramuon-trinity-nano-base-gsm8k-adamw-lora-8xh200
Base model
arcee-ai/Trinity-Nano-Base-Pre-Anneal Finetuned
arcee-ai/Trinity-Nano-Base