How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "oceansweep/mera-mix-4x7B-GGUF"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "oceansweep/mera-mix-4x7B-GGUF",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Use Docker
docker model run hf.co/oceansweep/mera-mix-4x7B-GGUF
Quick Links

New: mera-mix-4x7B GGUF

This is a repo for GGUF quants of mera-mix-4x7B. Currently it holds the FP16 and Q8_0 items only.

Original: Model mera-mix-4x7B

This is a mixture of experts (MoE) model that is half as large (4 experts instead of 8) as the Mixtral-8x7B while been comparable to it across different benchmarks. You can use it as a drop in replacement for your Mixtral-8x7B and get much faster inference.

mera-mix-4x7B achieves 76.37 on the openLLM eval v/s 72.7 by Mixtral-8x7B (as shown here).

You can try the model with the Mera Mixture Chat.

Open LLM Leaderboard Evaluation Results

Detailed results can be found here

Metric Value
Avg. 75.91
AI2 Reasoning Challenge (25-Shot) 72.95
HellaSwag (10-Shot) 89.17
MMLU (5-Shot) 64.44
TruthfulQA (0-shot) 77.17
Winogrande (5-shot) 85.64
GSM8k (5-shot) 66.11
Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results