Text Generation
Transformers
PyTorch
Catalan
Spanish
English
llama
finetune
chatml
gpt4
catalan
text-generation-inference
Instructions to use xaviviro/FLAMA-0.5-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use xaviviro/FLAMA-0.5-3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="xaviviro/FLAMA-0.5-3B")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("xaviviro/FLAMA-0.5-3B") model = AutoModelForCausalLM.from_pretrained("xaviviro/FLAMA-0.5-3B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use xaviviro/FLAMA-0.5-3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "xaviviro/FLAMA-0.5-3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xaviviro/FLAMA-0.5-3B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/xaviviro/FLAMA-0.5-3B
- SGLang
How to use xaviviro/FLAMA-0.5-3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "xaviviro/FLAMA-0.5-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xaviviro/FLAMA-0.5-3B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "xaviviro/FLAMA-0.5-3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xaviviro/FLAMA-0.5-3B", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use xaviviro/FLAMA-0.5-3B with Docker Model Runner:
docker model run hf.co/xaviviro/FLAMA-0.5-3B
| license: apache-2.0 | |
| base_model: openlm-research/open_llama_3b_v2 | |
| datasets: | |
| - xaviviro/oasst2_ca_gpt | |
| - xaviviro/oasst2_es_gpt | |
| tags: | |
| - finetune | |
| - chatml | |
| - gpt4 | |
| - catalan | |
| model-index: | |
| - name: FLAMA-0.5-3B | |
| results: [] | |
| library_name: transformers | |
| widget: | |
| - text: "<|im_start|>user\n Qui va ser Isaac Newton?<|im_end|>\n<|im_start|>assistant\n" | |
| - text: "<|im_start|>user\n ¿Quién fue Isaac Newton?<|im_end|>\n<|im_start|>assistant\n" | |
| language: | |
| - ca | |
| - es | |
| - en | |
| # FLAMA: Model 3B ChatML en Català i Castellà. Versió 0.5 | |
|  | |
| 👉🏻 [Format GGUF i quantitzat](/xaviviro/FLAMA-0.5-3B-GGUF) | |
| FLAMA és el primer model petit 3B bilingüe en català i castellà. És el resultat de finetunejar el model [open_llama_3b_v2](/openlm-research/open_llama_3b_v2) amb les instruccions d'[OpenAssistant v2](/datasets/OpenAssistant/oasst2) traduïdes automàticament al català i al castellà amb recursos de [Helsinki-NLP](/Helsinki-NLP) i tractades en format ChatML. | |
| ## Novetats de la versió 0.5 | |
| 1. Català millorat | |
| 1. Afegit el Castellà | |
| # Prompt Template | |
| FLAMA usa ChatML com a prompt template: | |
| ``` | |
| <|im_start|>user | |
| Qui va ser Isaac Newton?<|im_end|> | |
| <|im_start|>assistant\n | |
| ``` | |
| ``` | |
| <|im_start|>user | |
| Quien fué Isaac Newton?<|im_end|> | |
| <|im_start|>assistant\n | |
| ``` | |
| [<img src="https://raw.githubusercontent.com/OpenAccess-AI-Collective/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/>](https://github.com/OpenAccess-AI-Collective/axolotl) | |
| ## Referències | |
| ``` | |
| @software{xaviviro2023flama, | |
| author = {xaviviro}, | |
| title = {FLAMA: Model 3B ChatML en Català. Versió 0.5}, | |
| month = January, | |
| year = 2024, | |
| url = {https://huggingface.co/xaviviro/FLAMA-0.5-3B} | |
| } | |
| ``` | |
| ``` | |
| @software{openlm2023openllama, | |
| author = {Geng, Xinyang and Liu, Hao}, | |
| title = {OpenLLaMA: An Open Reproduction of LLaMA}, | |
| month = May, | |
| year = 2023, | |
| url = {https://github.com/openlm-research/open_llama} | |
| } | |
| ``` | |
| ``` | |
| @software{together2023redpajama, | |
| author = {Together Computer}, | |
| title = {RedPajama-Data: An Open Source Recipe to Reproduce LLaMA training dataset}, | |
| month = April, | |
| year = 2023, | |
| url = {https://github.com/togethercomputer/RedPajama-Data} | |
| } | |
| ``` | |
| ``` | |
| @article{touvron2023llama, | |
| title={Llama: Open and efficient foundation language models}, | |
| author={Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth{\'e}e and Rozi{\`e}re, Baptiste and Goyal, Naman and Hambro, Eric and Azhar, Faisal and others}, | |
| journal={arXiv preprint arXiv:2302.13971}, | |
| year={2023} | |
| } | |
| ``` |