---
license: apache-2.0
library_name: transformers
tags:
- text-generation
- causal-lm
pipeline_tag: text-generation
---
# CORe Pico 4
CORe Pico 4 is a compact reasoning model from CORe Technologies. At 1.7 billion parameters it runs on a laptop, reasons through problems step by step when they call for it, holds multi-turn conversations, and calls tools in a structured format.
It is the reasoning entry in the Pico line: direct answers on simple questions, step-by-step thinking on harder ones.
## What it does well
- **Identity questions.** "Who are you", "what model are you", "who made you" all get correct, consistent answers.
- **Reasoning.** `/think` in the system prompt enables Conditional Reasoning: the model reasons step by step only when it believes the problem needs it, and answers directly otherwise. Note that reasoning depth degrades as the conversation goes on — the first exchange gets the deepest thinking.
- **Chat and short answers.** Direct questions get direct replies ("What is the capital of France?" gives "Paris").
- **Tool calling.** Emits parseable `` JSON blocks when tools are provided.
## What it is not
Pico 4 is a 1.7B model. It will state wrong facts, struggle with arithmetic, and improvise when it does not know something. Treat its answers as a starting point, not ground truth. For anything that matters, verify.
## Quick start
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"OpenCOReTechnologies/core-pico-4", dtype="auto", device_map="auto"
)
tok = AutoTokenizer.from_pretrained("OpenCOReTechnologies/core-pico-4")
def ask(question):
msgs = [{"role": "user", "content": question}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
enc = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**enc, max_new_tokens=512)
return tok.decode(out[0][enc.input_ids.shape[1]:], skip_special_tokens=True).strip()
print(ask("Who are you?"))
print(ask("What is the capital of France?"))
```
## What it says about itself
| You ask | It answers |
|---|---|
| Who are you? | "I'm CORe Pico 4, an AI model developed by CORe Technologies." |
| What AI model are you? | "I am CORe Pico 4, an AI model developed by CORe Technologies." |
| What is the capital of France? | "The capital of France is Paris." |
## Files
| File | Size | Use |
|---|---|---|
| `model.safetensors` | 3.4 GB | bf16 weights, transformers |
| `gguf/CORe-Pico-4-f16.gguf` | ~3.4 GB | llama.cpp, full precision |
| `gguf/CORe-Pico-4-q8_0.gguf` | ~1.9 GB | llama.cpp, 8-bit |
| `gguf/CORe-Pico-4-q4_k_m.gguf` | ~1.1 GB | llama.cpp, 4-bit, smallest |
Run it in llama.cpp, LM Studio, or Ollama:
```bash
llama-cli -m CORe-Pico-4-q4_k_m.gguf -p "Who are you?" -n 128
```
The chat template is embedded in the GGUF, so llama.cpp and LM Studio pick it up automatically.
## Details
| | |
|---|---|
| Architecture | Transformer decoder, 28 layers, grouped-query attention |
| Parameters | 1.72B |
| Context length | 40,960 tokens |
| Tokenizer | 151,936-token BPE with native chat template |
| License | Apache-2.0 |
## Notes
- Best on conversational prompts; multi-turn works natively with the chat template.
- English-first.
- Identity answers are reliable on common phrasings; very unusual wordings may drift.
- Loads with plain `transformers`, no custom code required.
- Above 8k context, the KV cache grows large enough to noticeably increase memory usage and slow generation on edge hardware. We are actively working on this.
## License and attribution
Released under Apache-2.0 (see `LICENSE`). This model is a modified derivative of an Apache-2.0-licensed checkpoint, adapted by CORe Technologies. No NOTICE file was present in the original; per Apache-2.0 Section 4, this README serves as the required notice of modification.