--- license: mit library_name: transformers tags: - text-generation - causal-lm - custom-architecture - core pipeline_tag: text-generation ---

CORe

# CORe Pico V1.5-e CORe Pico V1.5-e is the refined edition of Pico V1.5, a compact 183M-parameter conversational model from CORe Technologies. This revision stays on topic and answers the question you actually asked. Where the original V1.5 could greet "Hi!" with a business email, V1.5-e replies "Hello! How can I help you today?" It is small enough to run on a CPU, carries a working sense of identity, and holds a coherent single-turn conversation. It is not trying to be a giant general assistant; it is a small, fast, self-aware model you can run anywhere. ## What changed from V1.5 - **Stays on topic.** Answers the prompt instead of drifting into unrelated text. - **Clean stopping.** Ends its turn reliably at `<|endoftext|>` instead of running on. - **Same identity, same size.** Still 183M parameters, still knows it is a CORe model. ## What it does well - **Identity questions.** "Who are you", "what model are you", "who made you", "are you ChatGPT" all get correct, consistent answers. - **Short factual answers.** Direct questions get direct replies ("What is the capital of France?" gives "Paris"). - **Brief explanations and chat.** Single-turn requests in plain language. ## What it is not Pico V1.5-e is a 183M model. It will state wrong facts, struggle with arithmetic, and improvise when it does not know something. Treat its answers as a starting point, not ground truth. For anything that matters, verify. ## Quick start This is a custom architecture, so `trust_remote_code=True` is required. Without it `from_pretrained` will fail on the unknown `core` model type. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained( "OpenCOReTechnologies/core-pico-v1-5-e", trust_remote_code=True ) model.eval() tok = AutoTokenizer.from_pretrained("OpenCOReTechnologies/core-pico-v1-5-e") def ask(question, max_new_tokens=200, temperature=0.7): text = f"<|user|>\n{question}\n<|assistant|>\n" enc = tok(text, add_special_tokens=False, return_tensors="pt") out = model.generate(**enc, max_new_tokens=max_new_tokens, temperature=temperature, top_k=40, do_sample=True, pad_token_id=0) return tok.decode(out[0][enc["input_ids"].shape[1]:], skip_special_tokens=True).strip() print(ask("Who are you?")) print(ask("What is the capital of France?")) ``` ## What it says about itself | You ask | It answers | |---|---| | Who are you? | "I'm CORe Flash, a tiny language model developed by CORe Technologies, a US-based company." | | What is the capital of France? | "The capital of France is Paris." | ## Files | File | Size | Use | |---|---|---| | `model.safetensors` | 783 MB | fp32 weights, full precision | | `gguf/CORe-Pico-V1.5-e-f16.gguf` | 368 MB | llama.cpp, full precision | | `gguf/CORe-Pico-V1.5-e-q8_0.gguf` | 197 MB | llama.cpp, 8-bit | | `gguf/CORe-Pico-V1.5-e-q4_k_m.gguf` | 122 MB | llama.cpp, 4-bit, smallest | Run it in llama.cpp, LM Studio, Ollama, or llama-cpp-python: ```bash llama-cli -m CORe-Pico-V1.5-e-q4_k_m.gguf \ -p "<|user|>\nWho are you?\n<|assistant|>\n" -n 60 ``` ## Chat template (important) Pico uses a specific chat format. If your app uses a different template (most default to `Human:`/`AI:` or ChatML), the model will produce rambling nonsense. Always use this exact template: ``` {% for message in messages %}{% if message['role'] == 'user' %}<|user|> {{ message['content'] }} {% elif message['role'] == 'assistant' %}<|assistant|> {{ message['content'] }} <|endoftext|> {% endif %}{% endfor %}{% if add_generation_prompt %}<|assistant|> {% endif %} ``` And set the **stop string** to `<|endoftext|>` so it stops after each answer. ### LM Studio LM Studio does not read the built-in template from the GGUF, so set it manually: 1. Load the model, open the chat settings (the model card or the **Prompt Template** field under "My Models" > model settings). 2. Replace the **Prompt Template** with the Jinja block above. 3. Under **Stop Strings**, add `<|endoftext|>`. 4. Save and start a new chat. If you skip this, LM Studio's default `Human:`/`AI:` template will make Pico output gibberish. That is the template's fault, not the model's. ### Raw prompt (no template engine) If you are feeding a raw string directly: ``` <|user|> Who are you? <|assistant|> ``` Then stop on `<|endoftext|>`. ## Details | | | |---|---| | Architecture | `COReForCausalLM`, custom transformer | | Parameters | 183M | | Layers / heads / width | 24 / 12 / 768 | | Context length | 512 tokens | | Tokenizer | 16,384-token BPE with a chat template (`<\|user\|>`, `<\|assistant\|>`) | | License | MIT | ## Notes - Best on single-turn prompts under a few hundred tokens. - English only. - Identity answers are reliable on common phrasings; very unusual wordings may drift. - Registered as a custom `core` model via `trust_remote_code`, so it loads with plain `transformers` and nothing else. - GGUF uses a pre-existing architecture while we prepare to submit a llama.cpp PR to add our custom architecture to the list.