Hello everyone! Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues. As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. π€£ All dates could be changed at any moment, as we are high school students not full-time ML engineers π . Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos@Timmy6767@MUK-IS-GOAT@smilyai-large-team@Sbui503@Banaxi-Tech@Bc-AI@atom77777@Harley-ml@Datdanboi25@Fishtiks@smartdigitalnetworks@vovaRL@EmetTheGolum@juiceb0xc0de@ProCreations
1. G1 series status. G1 is training nicely, and the loss is dropping nicely. The metrics are publicly available and i made a small space you can use to see the nice graphs: hugging-science/Loss-Plot-G1-Large G1-MINI is a lot slower in converging for reasons unknown yet, but we are investigating it.
2. I have built a small chat app for open SLMs here: ml-intern-explorers/slm-arena Feel free to add your models in a pull request!
Pretrained on 4x more tokens than the previous releases (20b vs 5b). Instruct tuned versions are coming soon. Very interesting models are coming soon too (hint: super long context).
We're excited to release BananaMind 2.1 Pico Preview!
It includes the first preview of our BananaMind 2.1 architecture! This model gets near BananaMind 2 Micro performance at half the parameters and 37.5x less tokens! Thats insane!
The current architectures includes about 500K parameters of the total 1.5M parameters in n-gram embeddings and the layer 2 is run twice.
It also includes XSA and the XSA refresh gate.
We're still going to improve the architecture in the final release.