How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "munzurul/SPEAKLAR-RAG-1.7B" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "munzurul/SPEAKLAR-RAG-1.7B",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "munzurul/SPEAKLAR-RAG-1.7B" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "munzurul/SPEAKLAR-RAG-1.7B",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Quick Links

SPEAKLAR-RAG-1.7B

This is a standalone full checkpoint for retrieval-augmented customer-support responses.

It is intended for retrieval-augmented customer-support responses. Supply concise, relevant retrieved evidence rather than a full policy document. The model was tested with Bengali evidence and produces Bengali answers.

Load

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "munzurul/SPEAKLAR-RAG-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

Important

This model should not be used as the source of truth for changing business facts. Retrieve relevant knowledge first, then use the retrieved evidence as model context. Validate prices, calculations, and policy-critical responses in application code.

Downloads last month
294
Safetensors
Model size
2B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using munzurul/SPEAKLAR-RAG-1.7B 1