Instructions to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") model = AutoModelForMultimodalLM.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
- SGLang
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Docker Model Runner:
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
RTX 5090
Hello.
May I ask about quick start to run any uncensored/heretic 35B NVFP4 model with a single 32GB RTX 5090 with any reasonable context like 100k or 200k?
AEON-7 Ornith or some Qwen3.6
Much appreciated 🥺
I generally don't have any problems running 35B .gguf models with llama.cpp, but right now I've spent insane amount of time figuring to run vLLM and without success. But tbf Ornith 35B NVFP4 is suited for DGX Spark 128GB so... Idk
You might be better off running the Qwen3.6-27B-Aeon-Ultimate-Uncensored-NVFP4-MTP-XS I've heard a lot of people run that on 5090 cards with max 256k context window and it runs fast and is great for coding as well as all around function. It will be slightly slower than 35B A3B but it will also be higher quality since it's a dense model and has all weight active vs 35B only having 3B active.
Also you really need to run this in NVFP4 since 5090 supports native NVFP4 tensor decoding you want to make sure cutlass backend is enabled and it should fit, do not try to use the BF16 weights there is no way that will fit on a 5090.
Try using this one and make sure you have cutlass enabled https://huggingface.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 it should fit even if it's a tight fit. Otherwise if you cannot get that to fit with a usable context window try this one https://huggingface.co/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Text-NVFP4-MTP-XS it's more compact overall and will enable max supported 256k context window.