Instructions to use qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI
- SGLang
How to use qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI with Docker Model Runner:
docker model run hf.co/qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI
EXAONE-4.0-1.2B-OptAI
EXAONE-4.0 is a large language model series developed by LG AI Research, consisting of a mid-size 32B model optimized for high performance and a small 1.2B model designed for on-device environments.
We have quantized the weights of this model to INT4 (w/o embedding) to optimize it for on-device deployment.
Model Conversion Contributor: Juneyoung Park (OptAI Inc.)
Model Stats:
- Input sequence length for Prompt Processor: 128
- Maximum context length: 4096
- Quantization Type: w4a16 (4-bit weights with 16-bit activations), KV-Cache(INT8)
- Supported languages: English, Korean, ...etc
- TTFT: Time To First Token is the time it takes to generate the first response token. This is expressed as a range because it varies based on the length of the prompt. The lower bound is for a short prompt (up to 128 tokens, i.e., one iteration of the prompt processor) and the upper bound is for a prompt using the full context length (4096 tokens).
- Response Rate: Rate of response generation after the first response token. Measured on a short prompt with a long response; may slow down when using longer context lengths.
Model Details
- Type: Causal Language Models
- Training Stage: Pretraining & Post-training
- Architecture: Transformers with RoPE, QK-Reorder-Norm, and GQA (Grouped Query Attention)
- Number of Parameters: 1.2B
- Number of Parameters (Non-Embedding): 1.07B
- Context Length Support: Up to 4096 tokens (optimized for on-device)
For more details, please refer to the official EXAONE4.0 Blog, GitHub, and Documentation.
Model Performance
| Model | Chipset | Target Runtime | Precision | Primary Compute Unit | Context Length | Response Rate (TPS) | Time to First Token (sec) |
|---|---|---|---|---|---|---|---|
| EXAONE-4.0-1.2B | Snapdragon 8 Elite Mobile | QNN(2.42)-GENIE | W4A16 | NPU | 4096 | 55.6 | 0.04 - 0.9 |
Model Conversion & Inference
If you want faster inference and conversion support for a wider variety of models, feel free to reach out anytime. When running inference with the uploaded model, we recommend using non-reasoning mode. Please refer to the template below.
[|system|]
{SYSTEM_PROMPT}[|endofturn|]
[|user|]
{USER_PROMPT}[|endofturn|]
[|assistant|]
<think>
</think>
Repository Structure
EXAONE-4.0-1.2B-OptAI/
βββ LICENSE
βββ README.md
βββ .gitattributes
βββ EXAONE4.0-1.2B-genie-w4a16-qualcomm_snapdragon_8_elite_OptAI.zip
Internal structure
EXAONE4.0-1.2B-genie-w4a16-qualcomm_snapdragon_8_elite_OptAI/
βββ config.json
βββ exaone4_part_1_of_5.bin
βββ ...
βββ exaone4_part_5_of_5.bin
βββ genie_config.json
βββ tokenizer.json
βββ tokenizer_config.json
βββ ...
References
License
Model tree for qualcomm-ai-hub-community/EXAONE-4.0-1.2B-OptAI
Base model
LGAI-EXAONE/EXAONE-4.0-1.2B