Instructions to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") model = AutoModelForMultimodalLM.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
- SGLang
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Docker Model Runner:
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
It may generate garbled characters.
I suspect the problem is with the base model.
It could also originate from the Uncensored process.
Not working
Seems to be working in my tests, on a DGX Spark with the aeon-vllm-ultimate image. That said you need to use the engine that works for your system the QuickStart guide is intended specifically for running on a DGX Spark. Running this model on another system would require a different setup and you need to make sure you are using the right tool call and reasoning parsers as well as the right config flags for the model. There are a multitude of variables that could cause it to spit garbled text or work incorrectly. In my tests it's extremely snappy and replies with clean text. Quality of output is dependent an many factors like engine, option flags, and harness configuration if being used.