vvijayk commited on
Commit
c3942a8
路
verified 路
1 Parent(s): 6ff2e53

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +13 -11
README.md CHANGED
@@ -3,20 +3,21 @@ language: en
3
  license: apache-2.0
4
  tags:
5
  - text-generation
6
- - mpt
7
  - moe
8
  - question-answering
9
  - natural-questions
10
- base_model: mosaicml/mpt-7b
 
11
  ---
12
 
13
- # MPT-7B-MoE Fine-tuned on Natural Questions
14
 
15
- This model is a fine-tuned version of [mosaicml/mpt-7b](https://huggingface.co/mosaicml/mpt-7b) on the Natural Questions dataset.
16
 
17
  ## Model Description
18
 
19
- - **Base Model**: mosaicml/mpt-7b
20
  - **Architecture**: Mixture-of-Experts (MoE) variant
21
  - **Training Dataset**: Natural Questions (NQ) - annotated for expert routing
22
  - **Task**: Question Answering / Text Generation
@@ -24,7 +25,7 @@ This model is a fine-tuned version of [mosaicml/mpt-7b](https://huggingface.co/m
24
  ## Training Details
25
 
26
  The model was fine-tuned using:
27
- - DeepSpeed with ZeRO optimization
28
  - HuggingFace Accelerate
29
  - Custom MoE expert annotations
30
 
@@ -65,7 +66,7 @@ Training was performed using the configuration in the repository. See `train_mix
65
 
66
  ## Limitations and Biases
67
 
68
- This model inherits limitations and biases from the base MPT-7B model and the Natural Questions dataset.
69
  Users should be aware of potential biases in question-answering outputs.
70
 
71
  ## Citation
@@ -73,9 +74,9 @@ Users should be aware of potential biases in question-answering outputs.
73
  If you use this model, please cite:
74
 
75
  ```bibtex
76
- @misc{mpt-7b-moe-nq,
77
- author = {Your Name},
78
- title = {MPT-7B-MoE Fine-tuned on Natural Questions},
79
  year = {2025},
80
  publisher = {HuggingFace},
81
  howpublished = {\url{https://huggingface.co/vvijayk/mixtral-8x7b-moe-nq-finetuned}}
@@ -84,6 +85,7 @@ If you use this model, please cite:
84
 
85
  ## Acknowledgements
86
 
87
- - Base model by MosaicML
88
  - Natural Questions dataset by Google Research
89
  - Training infrastructure using DeepSpeed and HuggingFace Accelerate
 
 
3
  license: apache-2.0
4
  tags:
5
  - text-generation
6
+ - mistralai
7
  - moe
8
  - question-answering
9
  - natural-questions
10
+ base_model:
11
+ - mistralai/Mixtral-8x7B-v0.1
12
  ---
13
 
14
+ # Mixtral-8x7B MoE Fine-tuned on Natural Questions
15
 
16
+ This model is a fine-tuned version of [MistralAI/Mixtral-8x7B-0.1](https://huggingface.co/mistralai/Mixtral-8x7B-v0.1) on the Natural Questions dataset.
17
 
18
  ## Model Description
19
 
20
+ - **Base Model**: mistralai/Mixtral-8x7B-0.1
21
  - **Architecture**: Mixture-of-Experts (MoE) variant
22
  - **Training Dataset**: Natural Questions (NQ) - annotated for expert routing
23
  - **Task**: Question Answering / Text Generation
 
25
  ## Training Details
26
 
27
  The model was fine-tuned using:
28
+ - DeepSpeed with ZeRO Stage 2 optimization
29
  - HuggingFace Accelerate
30
  - Custom MoE expert annotations
31
 
 
66
 
67
  ## Limitations and Biases
68
 
69
+ This model inherits limitations and biases from the base Mixtral-8x7B model and the Natural Questions dataset.
70
  Users should be aware of potential biases in question-answering outputs.
71
 
72
  ## Citation
 
74
  If you use this model, please cite:
75
 
76
  ```bibtex
77
+ @misc{mixtral-8x7b-moe-nq-finetuned,
78
+ author = {Vijay Venkatraman},
79
+ title = {Mixtral-8x7B-0.1 Fine-tuned on Natural Questions},
80
  year = {2025},
81
  publisher = {HuggingFace},
82
  howpublished = {\url{https://huggingface.co/vvijayk/mixtral-8x7b-moe-nq-finetuned}}
 
85
 
86
  ## Acknowledgements
87
 
88
+ - Base model by Mistral AI
89
  - Natural Questions dataset by Google Research
90
  - Training infrastructure using DeepSpeed and HuggingFace Accelerate
91
+