mesh-ops commited on
Commit
d88810b
·
verified ·
1 Parent(s): ee507bd

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +98 -0
README.md ADDED
@@ -0,0 +1,98 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Trinity Nano Base Pre-Anneal (Q6_K Imatrix)
2
+
3
+ This is a **Q6_K Imatrix** quantization of [arcee-ai/Trinity-Nano-Base-Pre-Anneal](https://huggingface.co/arcee-ai/Trinity-Nano-Base-Pre-Anneal).
4
+
5
+ **Quantization:** Q6_K (6-bit K-Quant)
6
+ **Imatrix:** Optimized using 'ridiculous_tokens' calibration dataset.
7
+
8
+ ---
9
+
10
+ ---
11
+ license: apache-2.0
12
+ language:
13
+ - en
14
+ - es
15
+ - fr
16
+ - de
17
+ - it
18
+ - pt
19
+ - ru
20
+ - ar
21
+ - hi
22
+ - ko
23
+ - zh
24
+ library_name: transformers
25
+ ---
26
+ <div align="center">
27
+ <picture>
28
+ <img
29
+ src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/i-v1KyAMOW_mgVGeic9WJ.png"
30
+ alt="Arcee Trinity Nano"
31
+ style="max-width: 100%; height: auto;"
32
+ >
33
+ </picture>
34
+ </div>
35
+
36
+ # Trinity Nano Base Pre Anneal
37
+
38
+ Trinity-Nano-Base-Pre-Anneal is an Arcee AI 6B MoE model with 1B active parameters. It is the small-sized model in our new Trinity family, a series of open-weight models for enterprise and tinkerers alike.
39
+
40
+ This base model is a pre-anneal checkpoint captured at Adam LR: 0.002, Muon LR: 0.001 before starting learning rate decay on a high-quality data mix.
41
+ While this checkpoint was not exposed to the anneal phase mix containing high proportions of math and code content, it has been trained on significant amounts of such data.
42
+ This checkpoint is not suitable for chatting or general use without further finetuning and should be trained for your specific domain before use.
43
+
44
+ ***
45
+
46
+ Trinity-Nano-Base-Pre-Anneal is trained on 8.8T tokens gathered and curated through a key partnership with [Datology](https://www.datologyai.com/), building upon the excellent dataset we used on [AFM-4.5B](https://huggingface.co/arcee-ai/AFM-4.5B) with additional math and code.
47
+
48
+ Training was performed on a cluster of 512 H200 GPUs powered by [Prime Intellect](https://www.primeintellect.ai/) using HSDP parallelism.
49
+
50
+ More details, including key architecture decisions, can be found on our blog [here](https://www.arcee.ai/blog/the-trinity-manifesto)
51
+
52
+ ***
53
+
54
+ ## Model Details
55
+
56
+ * **Model Architecture:** AfmoeForCausalLM
57
+ * **Parameters:** 6B, 1B active
58
+ * **Experts:** 128 total, 8 active, 1 shared
59
+ * **Context length:** 4K
60
+ * **Learning rate during pretraining**:
61
+ * `adam_lr = 0.0002`
62
+ * `muon_lr = 0.001`
63
+ * **Training Tokens:** 8.8T
64
+ * **License:** [Apache 2.0](https://huggingface.co/arcee-ai/Trinity-Mini#license)
65
+
66
+ ***
67
+
68
+ <div align="center">
69
+ <picture>
70
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/sSVjGNHfrJKmQ6w8I18ek.png" style="background-color:ghostwhite;padding:5px;" width="17%" alt="Powered by Datology">
71
+ </picture>
72
+ </div>
73
+
74
+ ## Try out our reasoning tune of our medium-sized Trinity Mini model
75
+
76
+ Trinity Mini is available today on openrouter:
77
+
78
+ https://openrouter.ai/arcee-ai/trinity-mini
79
+
80
+ ```
81
+ curl -X POST "https://openrouter.ai/v1/chat/completions" \
82
+ -H "Authorization: Bearer $OPENROUTER_API_KEY" \
83
+ -H "Content-Type: application/json" \
84
+ -d '{
85
+ "model": "arcee-ai/trinity-mini",
86
+ "messages": [
87
+ {
88
+ "role": "user",
89
+ "content": "What are some fun things to do in New York?"
90
+ }
91
+ ]
92
+ }'
93
+ ```
94
+
95
+
96
+ ## License
97
+
98
+ Trinity-Nano-Base-Pre-Anneal is released under the Apache-2.0 license.