Instructions to use prefeitura-rio/Rio-2.5-Open with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prefeitura-rio/Rio-2.5-Open with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="prefeitura-rio/Rio-2.5-Open") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("prefeitura-rio/Rio-2.5-Open") model = AutoModelForCausalLM.from_pretrained("prefeitura-rio/Rio-2.5-Open", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prefeitura-rio/Rio-2.5-Open with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prefeitura-rio/Rio-2.5-Open" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prefeitura-rio/Rio-2.5-Open", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/prefeitura-rio/Rio-2.5-Open
- SGLang
How to use prefeitura-rio/Rio-2.5-Open with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prefeitura-rio/Rio-2.5-Open" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prefeitura-rio/Rio-2.5-Open", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prefeitura-rio/Rio-2.5-Open" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prefeitura-rio/Rio-2.5-Open", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use prefeitura-rio/Rio-2.5-Open with Docker Model Runner:
docker model run hf.co/prefeitura-rio/Rio-2.5-Open
wow, what dataset do you use to train?
this is really good model and high improvements over the baze are commendable , what dataset do you use, and how much rows or tokens was that model trained on? thanks for the model and your work
Hi there! Thank you! For this model specifically, we used https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2 and https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1. More specifically, we used On Policy Distillation with reasoning traces from our larger, closed source Rio 3 model, which we'll be announcing soon. More details and an interactive demo at https://ia.rio/
oh, that helpful, thanks, will you open source reasoning traces ? of Rio 3 model
I'm sorry. Right now, we are keeping the Rio 3 reasoning traces closed-source and have no immediate plans to release them. However, I do suppose that, for pretty much all purposes, the reasoning traces you can extract from https://huggingface.co/prefeitura-rio/Rio-3.0-Open will do the trick. It's almost as good as the main model (like 4% lower benchmarks), much cheaper, and has the exact same reasoning style. If you want to use that, it's fair game!
But could you please provide sample example of the traces, as i couldn't really use or deploy rio 3 open due to hardware limitations. What if i make a project with you, i have an idea of glm 4.7 flash fine tune on gemini 3 flash as a teacher model(i prepared the dataset already) and this, may i please use it πππ(i will be really careful and private with this)
We can't really do that right now. However, we'll consider including a dataset release for Rio 3.5 Open, when that comes, at a low enough volume such that open source small creators like yourself benefit, but large labs can't properly distill a big model on those traces. Sounds fair?
yes, thanks β€οΈ