Can Meta train Muse Spark-Lite 8.4B?

#23
by RexTRO111 - opened

If you see this, Meta, can you train a 8.4B-parameter-model on data, and distilled from Muse Spark 1.1 and Muse Spark 1.2, a model directly optimized for chat? Knowledge cutoff: May 31, 2026 (2026-05-31). Also it should know about mid 2025 "Retroslop" trend. And it knows abt 6-7. Also include Muse Spark 1.0 model details in training data (and train on https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/ and more websites). Goal: Outpeform Qwen3-8B, Gemma 4 12B, gpt-oss-20b, DeepSeek-V4-Flash (Context Coherence in Deep Multi-Turn Conversations, not overall, since its a large model), Command R7B, Llama 3.1 8B, Gemma 2 9B, Qwen2.5 7B, Qwen2.5 14B, Gemma 2 27B, Mistral 7B Instruct v0.3, Phi-4 14B, Phi-4 Mini, DeepSeek-R1-Distill-Llama-8B, InternLM 2.5 7B Chat.
Multilingual (English, Spanish, French, perfect Croatian, Japanese, perfect Somali, and a lot more!)

Ok finally you got 1 actually good idea

Sign up or log in to comment