Berna R7 - Root Model

A 1.06B multilingual language model with a modular cell-based architecture.

This repository will host the Root model of Berna R7, the first large-scale implementation of the DNA-Kernel Plexus architecture from Berna R5.

Status: Training in progress. Weights will be uploaded on or around 5 October 2026.

Zenodo DOI: 10.5281/zenodo.23037065 Mendeley DOI: 10.17632/k24wwx6tp6.1

Model Details

Property Value
Parameters 1,056,265,728 (1.06B)
Architecture Decoder-only Transformer (Berna)
Layers 28
Hidden size 1536
Intermediate size 6144
Attention heads 16 (GQA, 4 KV heads)
Vocab size 32,000 (BPE)
Context length 4,096
RoPE theta 500,000
Normalization RMSNorm
Activation SwiGLU
Precision BF16
Tied embeddings Yes

Training

Property Value
Data 2.66B tokens (English + Math + Code)
Epochs 1
Hardware 1x RTX 5090 (32 GB)
Duration ~6.5 days
Optimizer AdamW8bit
Learning rate 3e-4 (cosine, 500-step warmup)
Batch size 64 x 4096 = 262,144 tokens/step
Gradient checkpointing enabled

Architecture (Berna R5 implementation)

Six-layer structure from R5:

Sensors -> Spinal Cord -> Plexus -> Knowledge Cells -> DNA Kernel -> Registry

  • Knowledge Cells: 112 cells (4 per MLP layer x 28 layers).
  • 6D Knowledge Vector: K = [L, W, H, D, T, E], saturation S = ||K||_omega.
  • DNA Kernel: 100 chromosomes x 4 genes (a, s, p, pi).
  • Plexus: dynamic graph with edges updated from co-activations.
  • Registry: SQL catalog with UUIDs and connection tokens.

Compliance with R5

R5 Component Status
Registry + UUIDs implemented
connection_token implemented
Federation implemented
Knowledge Cells 112 cells
6D K vector K1, K3, K5, K6
DNA Kernel 100x4
S* = 0.85 yes
Splitting (H1) tested
Plexus implemented
Router (Spinal Cord) deferred
Bounded Forgetting (H2) not yet tested

Intended Use

  • Research on continual learning and growing neural networks.
  • Base model for fine-tuning on domain-specific tasks.
  • Reference implementation of the R5 DNA-Kernel Plexus framework.

Out-of-Scope Use

  • Production deployment without evaluation.
  • Safety-critical applications.
  • Uses prohibited by the Berna Research License v1.0.

Limitations

  • Undertrained: 1 epoch over 2.66B tokens (~2.5 tokens/param).
  • Text-only: no vision, no audio.
  • Not instruction-tuned.
  • No benchmarks yet (MMLU, GSM8K, HumanEval pending).
  • English + Math + Code only.

What This Model Does NOT Claim

  • State-of-the-art on any benchmark.
  • Superiority over GPT-4-class models.
  • Scalability beyond 1.5B parameters.
  • Theoretical completeness.

Citation

@software{berna_r7_2026,
  author    = {Muhammed, Mohammed Kamil},
  title     = {Berna R7: A 1.06B Multilingual Language Model},
  year      = {2026},
  publisher = {Berna Labs},
  url       = {https://github.com/Berna-Labs/berna-r7}
}

For the underlying theory:

@software{berna_r5_2026,
  author    = {Muhammed, Mohammed Kamil},
  title     = {DNA-Kernel Plexus},
  year      = {2026},
  doi       = {10.5281/zenodo.23015308}
}

License

Berna Research License v1.0. Research use is free. Commercial use requires a separate license. Contact: info@bernalabs.com

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support