Instructions to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16") model = AutoModelForMultimodalLM.from_pretrained("AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
- SGLang
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16 with Docker Model Runner:
docker model run hf.co/AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
Solid work β testing for mobile deployment
We're always scanning HuggingFace for models that could work on mobile. This one caught our attention.
At Dispatch AI (FZE, Sharjah UAE), we test models on a 40-phone farm (Samsung S20 FE, Snapdragon 865). If this fits the mobile deployment criteria, we'll quantize and benchmark it.
Appreciate the open release. Open weights move the whole field forward.
β Dispatch AI (FZE), Sharjah UAE
so 40 phones working together as one super cluster? i just spun this up on a b200 on runpod... this wont fit on any one phone anyone has...unless you're from the future. Can I borrow your phone to make a call real quick? I'll be quick.
We're always scanning HuggingFace for models that could work on mobile. This one caught our attention.
At Dispatch AI (FZE, Sharjah UAE), we test models on a 40-phone farm (Samsung S20 FE, Snapdragon 865). If this fits the mobile deployment criteria, we'll quantize and benchmark it.
Appreciate the open release. Open weights move the whole field forward.
β Dispatch AI (FZE), Sharjah UAE
This is one of the most unique and unusual things I've heard of to host AI models, 40 phones aggregated together?
If you were trying to quantize and distil this down to 1bit bitmap you might be able to squeeze it on to a very powerful phone. Good luck on the project.
so 40 phones working together as one super cluster? i just spun this up on a b200 on runpod... this wont fit on any one phone anyone has...unless you're from the future. Can I borrow your phone to make a call real quick? I'll be quick.
I imagine they are planning to massively quantize it maybe they have some kind of solution to aggregate the phone compute, can't imagine it's super fast. Never heard of a project like this but there is no way it's fitting on a single phone with the ram capacity those models have even with really deep quantization. Unless they found some quant alien tech.
It seems that the discussion here isn't really related to this model.