--- license: apache-2.0 language: - en tags: - conversational - instruction-following - chat - gguf - llama.cpp - ollama - local-llm pipeline_tag: text-generation datasets: - HuggingFaceFW/finepdfs - fka/awesome-chatgpt-prompts - roneneldan/TinyStories - wikimedia/wikipedia - codeparrot/github-code - HuggingFaceFW/finewiki - karpathy/fineweb-edu-100b-shuffle library_name: gguf base_model: - neuralcrew/neutrino ---
# Neutrino-Instruct ![Neutrino](https://ollama.com/assets/fardeen0424/neutrino/742305f0-8c9e-4ae8-acff-a7c2b133d3d8) **A 7B-parameter instruction-tuned language model for conversational AI**

Model card · Quickstart · Architecture · Limitations · License

--- ## Model Card | Field | Value | |---|---| | **Developer** | Fardeen NB | | **Model type** | Autoregressive transformer, decoder-only | | **Parameters** | 7B | | **Base model** | [neuralcrew/neutrino](https://huggingface.co/neuralcrew/neutrino) | | **Fine-tuning** | Instruction tuning for multi-turn dialogue | | **Language(s)** | English | | **Format** | GGUF (quantized for `llama.cpp`, `Ollama`, `llama-cpp-python`) | | **License** | Apache 2.0 | | **Version** | 2.0 | --- ## Overview Neutrino-Instruct is a 7B-parameter language model fine-tuned from the Neutrino base model for conversational and instruction-following tasks. It is distributed in GGUF format for efficient local inference on consumer hardware via `llama.cpp`, `Ollama`, and `llama-cpp-python`. The model is designed to hold coherent, contextual dialogue across multiple turns and to follow natural-language instructions reliably at chat scale, while remaining light enough to run on a single consumer GPU or CPU-only machine. --- ## Intended Use **Primary use cases** - Conversational assistants and chatbots - Instruction-following agents (task completion, Q&A, summarization) - Local/offline research prototypes where data cannot leave the device - Educational and hobbyist LLM experimentation **Out of scope** - Medical, legal, or financial advice, or any use where model error could cause real-world harm - Autonomous decision-making without human review - Generation of content intended to deceive, impersonate, or manipulate - High-stakes classification (e.g., hiring, credit, law enforcement) Neutrino-Instruct is a general-purpose research and hobbyist model. It has **not** been evaluated for production or safety-critical deployment, and Fardeen NB makes no warranty as to its factual accuracy, safety, or fitness for any particular purpose. --- ## Quickstart ### llama.cpp ```bash git clone https://github.com/ggerganov/llama.cpp cd llama.cpp && make # Single prompt ./main -m ./neutrino-instruct.gguf -p "Hello, who are you?" # Interactive chat ./main -m ./neutrino-instruct.gguf -i -p "Let's chat." # Control output length ./main -m ./neutrino-instruct.gguf -n 256 -p "Write a poem about stars." # Adjust temperature ./main -m ./neutrino-instruct.gguf --temp 0.7 -p "Explain quantum computing simply." # GPU offload (if built with CUDA/Metal) ./main -m ./neutrino-instruct.gguf --gpu-layers 50 -p "Summarize this article." ``` ### Ollama ```bash ollama run fardeen0424/neutrino ``` ### Python (llama-cpp-python) ```python from llama_cpp import Llama llm = Llama(model_path="./neutrino-instruct.gguf") response = llm("Who are you?") print(response["choices"][0]["text"]) # Streaming for token in llm("Tell me a story about Neutrino:", stream=True): print(token["choices"][0]["text"], end="", flush=True) ``` --- ## Training Data Neutrino-Instruct's base model was pretrained on a mixture of web, encyclopedic, code, and long-document text, and instruction-tuned on conversational data. Component sources include: | Dataset | Type | Role | |---|---|---| | finepdfs | Long-form documents | Pretraining | | finewiki | Encyclopedic text | Pretraining | | fineweb-edu-100b-shuffle | Educational web text | Pretraining | | wikipedia | Encyclopedic text | Pretraining | | github-code | Source code | Pretraining | | TinyStories | Short narrative text | Fine-tuning / eval | | awesome-chatgpt-prompts | Instruction/persona prompts | Instruction tuning | *Full dataset links are available in the metadata panel on the right side of this page.* ### Training procedure | Field | Value | |---|---| | **Context length** | 32,768 tokens | | **Precision** | bfloat16 | --- ## Architecture Neutrino-Instruct is built on a standard decoder-only Transformer stack. The core computations are as follows. **Scaled dot-product self-attention**, applied per head: $$\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V$$ with multi-head attention combining $h$ heads via a learned output projection: $$\text{MultiHead}(Q,K,V) = \text{Concat}(\text{head}_1, \dots, \text{head}_h)\,W^O, \qquad \text{head}_i = \text{Attention}(QW_i^Q, KW_i^K, VW_i^V)$$ Neutrino-Instruct uses **grouped-query attention (GQA)**: the 32 query heads are partitioned into $g=8$ groups, each group sharing a single key/value projection ($KW^K_g, VW^V_g$) instead of every query head having its own — trading a small amount of expressivity for a much smaller KV cache at inference time. **Position encoding** — rotary position embeddings (RoPE), applied to query/key vectors by rotating consecutive coordinate pairs as a function of token position $m$ and frequency $\theta_i$: $$f(x_m, m) = \big[x_m^{(1)}\cos m\theta_i - x_m^{(2)}\sin m\theta_i,\ \ x_m^{(1)}\sin m\theta_i + x_m^{(2)}\cos m\theta_i\big]$$ **Feed-forward block** (SwiGLU-style gating): $$\text{FFN}(x) = \big(\text{Swish}(xW_1)\otimes xW_3\big)W_2$$ **Pre-normalization** (RMSNorm) around each sub-block: $$\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{n}\sum_{i=1}^n x_i^2 + \epsilon}} \odot g$$ **Training objective** — standard next-token cross-entropy over the vocabulary: $$\mathcal{L} = -\sum_{t=1}^{T} \log P_\theta(x_t \mid x_{