How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity",
		"messages": [
			{
				"role": "user",
				"content": [
					{
						"type": "text",
						"text": "Describe this image in one sentence."
					},
					{
						"type": "image_url",
						"image_url": {
							"url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
						}
					}
				]
			}
		]
	}'
Use Docker
docker model run hf.co/AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity
Quick Links

RL-MPQ High Fidelity — llava-hf/llava-1.5-13b-hf

RL-MPQ fake-quantized VLM evaluation artifacts.

Repo AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity
Scenario High_Fidelity
Avg bits 6.7
Collection RL-MPQ VLM — LLaVA-1.5-13B
Full results dataset

Benchmark table (group comparison)

Model avg_bits group MMMU MMBench ScienceQA Avg ΔMMMU ΔMMBench ΔScienceQA ΔAvg
FP16 (baseline) 16.0 llama13b_aggressive 35.33 63.78 71.24 56.78 0.0 0.0 0.0 0.0
INT4 (bnb NF4) 4.0 llama13b_aggressive 35.56 61.76 71.15 56.16 0.23 -2.02 -0.09 -0.62
RL-MPQ Aggressive 3.75 llama13b_aggressive 34.56 63.0 71.1 56.22 -0.77 -0.78 -0.14 -0.56

Artifacts in this repo

  • artifacts/figures/ — plots for the model group
  • artifacts/benchmark_table.csv — FP16 / INT4 / RL-MPQ accuracies
  • eval_results.json — structured eval metadata (if present)
Downloads last month
21
Safetensors
Model size
13B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity

Finetuned
(5)
this model

Collections including AvoCahDoe/llava-1.5-13b-rlmpq-high-fidelity