--- library_name: transformers datasets: - allenai/llama-3.1-tulu-3-8b-preference-mixture language: - en base_model: - allenai/Llama-3.1-Tulu-3-8B-SFT license: llama3.1 pipeline_tag: text-generation --- # Stackelberg Learning from Human Feedback Preference finetuned model using the Stackelberg Learning from Human Feedback approach for general conversational applications. ## Model Details ### Model Description This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated. - **Developed by:** [Barna Pasztor](https://pasztorb.github.io/), Thomas Kleine Buening, Andreas Krause - **Model type:** A model trained on a mix of publicly available, synthetic and human-created datasets using LLM-as-a-judge (Skywork/Skywork-Critic-Llama-3.1-70B). - **Language(s) (NLP):** Primarily English - **License:** Llama 3.1 Community License Agreement - **Finetuned from model:** meta-llama/Llama-3.1-8B ### Model Sources - **Repository:** [GitHub](https://github.com/lasgroup/stackelberg-learning) - **Paper:** [ArXiv Preprint](https://arxiv.org/abs/2512.16626)