Instructions to use TechnoBaptist/d1-omni-600M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TechnoBaptist/d1-omni-600M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TechnoBaptist/d1-omni-600M-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf TechnoBaptist/d1-omni-600M-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TechnoBaptist/d1-omni-600M-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf TechnoBaptist/d1-omni-600M-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TechnoBaptist/d1-omni-600M-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf TechnoBaptist/d1-omni-600M-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TechnoBaptist/d1-omni-600M-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf TechnoBaptist/d1-omni-600M-GGUF:BF16
Use Docker
docker model run hf.co/TechnoBaptist/d1-omni-600M-GGUF:BF16
- LM Studio
- Jan
- vLLM
How to use TechnoBaptist/d1-omni-600M-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TechnoBaptist/d1-omni-600M-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TechnoBaptist/d1-omni-600M-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/TechnoBaptist/d1-omni-600M-GGUF:BF16
- Ollama
How to use TechnoBaptist/d1-omni-600M-GGUF with Ollama:
ollama run hf.co/TechnoBaptist/d1-omni-600M-GGUF:BF16
- Unsloth Desktop
- Docker Model Runner
How to use TechnoBaptist/d1-omni-600M-GGUF with Docker Model Runner:
docker model run hf.co/TechnoBaptist/d1-omni-600M-GGUF:BF16
- Lemonade
How to use TechnoBaptist/d1-omni-600M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TechnoBaptist/d1-omni-600M-GGUF:BF16
Run and chat with the model
lemonade run user.d1-omni-600M-GGUF-BF16
List all available models
lemonade list
- Atomic Chat
d1-omni-600M-GGUF
d1-omni-600M is a decision model developed by Liquid AI: it answers named, typed questions over a state (text, JSON, images or audio) in one forward pass, with no generated tokens.
Find more details in the original model card: https://huggingface.co/LiquidAI/d1-omni-600m
🏃 How to run d1-omni-600M
Example usage with llama.cpp:
llama-server -hf LiquidAI/d1-omni-600M-GGUF:Q8_0 -b 4096 -ub 4096
A question is read in one batch, so -b and -ub must hold it: 4096 covers any image and states of a few thousand tokens, raise them up to 16384 for longer states.
Then send questions to the /v1/systemone endpoint.
Text
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "I was charged twice this month, please refund one of them.",
"questions": {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]}
}
}'
Image + Text
curl -sL -o cats.jpg http://images.cocodataset.org/val2017/000000039769.jpg # two cats on a sofa
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d @- <<JSON
{
"state": "Photo attached to a pet-sitting request.",
"images": ["data:image/jpeg;base64,$(base64 < cats.jpg | tr -d '\n')"],
"questions": {
"pet": {"type": "choice", "instructions": "Which animals are in the photo?",
"criteria": {"cats": "Cats", "dogs": "Dogs", "birds": "Birds"}},
"sofa": {"type": "noul", "instructions": "Are the animals on a sofa?"}
}
}
JSON
Audio + Text
curl -sL -o jfk.wav https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d @- <<JSON
{
"state": "Recording from a public event.",
"audio": "data:audio/wav;base64,$(base64 < jfk.wav | tr -d '\n')",
"questions": {
"kind": {"type": "choice", "instructions": "What kind of utterance is this?",
"criteria": {"request": "A request to do something",
"question": "A question asking for information",
"speech": "A speech to an audience"}},
"calm": {"type": "score", "instructions": "How calm is the speaker?",
"criteria": ["Agitated", "Neutral", "Calm"]}
}
}
JSON
The state can be null when the images or the audio are the whole state. A request carries images or one audio clip (up to 30 s), not both.
- Downloads last month
- 233
8-bit
16-bit
Model tree for TechnoBaptist/d1-omni-600M-GGUF
Base model
LiquidAI/LFM2.5-350M-Base