Macaron-V1-Tall LoRA specialists β€” BF16 (optional, smaller L2)

Upstream ships L2 (Coding) as FP32, 14.07 GiB, while L0/L1/L3 are already BF16. vLLM loads adapters as bfloat16 regardless, so the FP32 storage buys nothing. This is the set with L2 cast down.

35.16 GiB -> 28.13 GiB on disk. Only L2 changes; L0/L1/L3 are byte-identical to upstream.

⚠️ This saves disk, not VRAM β€” serving footprint is identical either way.

Verified equivalent, not assumed

FP32 vs BF16 L2 through 15 hard problems: 10 algorithmic tasks where the generated code was executed against assertions (LRU cache, edit distance, median of two sorted arrays, hand-rolled regex with . and *, N-queens, trapping rain water, word break, merge-k-sorted, LIS, coin change) plus 5 exact-answer math problems. Greedy decoding, identical prompts and token budgets.

FP32 BF16
Code, executed and passed 8/10 8/10
Math, exact 3/5 3/5
Total 11/15 11/15
Tokens 31,632 31,747

The same problems passed and the same ones failed. Output is not bit-identical β€” greedy decoding diverges on wording β€” but capability is equivalent on an execution-scored test.

Use with kingjones777/Macaron-V1-Tall-NVFP4. Original adapters: mindlab-research/Macaron-V1-Tall.

Other public builds of this model

Compiled from Hugging Face repository metadata β€” file sizes, shipped files, quant variant as named by each repo. No third-party build was run or benchmarked here, so this table makes no speed or quality claim about any of them. It is here so you can see the size and format options at a glance and pick what fits your hardware.

Repository Largest model file Variant Ships Downloads Likes
kingjones777/Macaron-V1-Tall-NVFP4 23.32 GiB NVFP4 safetensors 18 1

Base model: mindlab-research/Macaron-V1-Tall. Generated from Hub metadata; download counts move over time.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for kingjones777/Macaron-V1-Tall-LoRA-BF16

Adapter
(1)
this model