Text Generation
GGUF
English
Vietnamese
acne
dermatology
skincare
research
quantization
qwen3.5
imatrix
conversational
Instructions to use Acnoryx/Airy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Acnoryx/Airy with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Acnoryx/Airy:IQ1_M # Run inference directly in the terminal: llama cli -hf Acnoryx/Airy:IQ1_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Acnoryx/Airy:IQ1_M # Run inference directly in the terminal: llama cli -hf Acnoryx/Airy:IQ1_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Acnoryx/Airy:IQ1_M # Run inference directly in the terminal: ./llama-cli -hf Acnoryx/Airy:IQ1_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Acnoryx/Airy:IQ1_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Acnoryx/Airy:IQ1_M
Use Docker
docker model run hf.co/Acnoryx/Airy:IQ1_M
- LM Studio
- Jan
- vLLM
How to use Acnoryx/Airy with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Acnoryx/Airy" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Acnoryx/Airy", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Acnoryx/Airy:IQ1_M
- Ollama
How to use Acnoryx/Airy with Ollama:
ollama run hf.co/Acnoryx/Airy:IQ1_M
- Unsloth Desktop
- Docker Model Runner
How to use Acnoryx/Airy with Docker Model Runner:
docker model run hf.co/Acnoryx/Airy:IQ1_M
- Lemonade
How to use Acnoryx/Airy with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Acnoryx/Airy:IQ1_M
Run and chat with the model
lemonade run user.Airy-IQ1_M
List all available models
lemonade list
- Atomic Chat
File size: 1,740 Bytes
ba4ce94 2009f9e 425b216 2009f9e 425b216 2009f9e 8f415ee 425b216 2009f9e 37bf90d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 | ---
language:
- en
- vi
license: other
library_name: gguf
tags:
- acne
- dermatology
- skincare
- gguf
- research
- quantization
- qwen3.5
pipeline_tag: text-generation
base_model:
- Qwen/Qwen3.5-0.8B
---
# Acnoryx AI Research Bundle
## Overview
- Base model: Qwen/Qwen3.5-0.8B
- Model size: 0.8b
- Research quantizations: Q3_K_M, IQ3_M, Q2_K, IQ2_M, IQ2_XS, IQ2_XXS, IQ1_M, IQ1_S
- Purpose: evaluate quality vs. size trade-offs below the production threshold
## Notes
- IQ1/IQ2 formats require an importance matrix (imatrix).
- These files are more experimental than the release bundle.
- Production-facing use should prefer the release bundle.
- If prompting in Vietnamese, write with full accents for best consistency.
## Evaluation Snapshot
Research GGUFs were continued from the existing results and merged with the latest rerun on the same curated 58-question bilingual benchmark.
| Quant | Think | No-Think | Avg | Status |
|---|---:|---:|---:|---|
| Q3_K_M | 74.1% | 72.4% | 73.2% | Best current research quant |
| IQ3_M | 60.3% | 60.3% | 60.3% | Heavy quality loss |
| IQ2_M | 20.7% | 19.0% | 19.8% | Below usable threshold |
| IQ2_XS | 5.2% | 3.4% | 4.3% | Triggered early-stop for lower bits |
## Research Guidance
- Public research recommendation: **Q3_K_M** only
- **IQ3_M** is still uploadable for experiments, but quality is clearly degraded
- The rerun auto-stopped below **IQ2_XS** because average pass rate fell under 50%, so lower-bit quants should be considered archival artifacts rather than viable deployments
- For any user-facing scenario, prefer the release bundle instead of this research branch
For cross-family ranking and release-vs-research comparison, see `results/COMPARISON.md` in the workspace.
|