amewebstudio commited on
Commit
47f4ba3
·
verified ·
1 Parent(s): aaf8e8e

SCLM v2 - Stateful Coherent Language Model with EARCP

Browse files
README.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ tags:
7
+ - sclm
8
+ - stateful
9
+ - memory
10
+ - earcp
11
+ - text-generation
12
+ - conversational
13
+ pipeline_tag: text-generation
14
+ base_model: mistralai/Mistral-7B-v0.1
15
+ widget:
16
+ - text: "The wizard Elara lived in Silverwood forest. One day, she discovered"
17
+ example_title: "Fantasy Story"
18
+ - text: "In the year 2050, humanity had finally achieved"
19
+ example_title: "Science Fiction"
20
+ - text: "The detective examined the crime scene carefully. The clues pointed to"
21
+ example_title: "Mystery"
22
+ inference:
23
+ parameters:
24
+ max_new_tokens: 100
25
+ temperature: 0.7
26
+ top_p: 0.9
27
+ repetition_penalty: 1.1
28
+ ---
29
+
30
+ # 🧠 SCLM: Stateful Coherent Language Model
31
+
32
+ **SCLM** adds **persistent latent memory** to transformer language models, enabling better coherence across long conversations and multi-turn generation.
33
+
34
+ ## 🎯 Key Features
35
+
36
+ - **Persistent State**: Memory that evolves across conversation turns
37
+ - **Entity Coherence**: Maintains context about characters, places, and objects
38
+ - **Edit Mode**: Make local changes without affecting global memory
39
+ - **Lightweight**: Only 91.7M additional parameters (2.44% overhead)
40
+
41
+ ## 📊 Architecture: EARCP
42
+
43
+ ```
44
+ EARCP = Encapsulation + Alignment + Revision + Coherence + Propagation
45
+ ```
46
+
47
+ | Component | Function |
48
+ |-----------|----------|
49
+ | **Encapsulation** | GRU-style state update from hidden states |
50
+ | **Alignment** | Cross-attention between state and hidden layers |
51
+ | **Revision** | Drift detection and correction |
52
+ | **Coherence** | Mixture-of-Experts for consistency |
53
+ | **Propagation** | State injection into transformer layers |
54
+
55
+ ## 🔧 Model Details
56
+
57
+ | Parameter | Value |
58
+ |-----------|-------|
59
+ | Base Model | mistralai/Mistral-7B-v0.1 |
60
+ | EARCP Parameters | 91.7M |
61
+ | Latent State Dim | 256 |
62
+ | Injection Layers | [8, 16] |
63
+ | Alpha (injection strength) | 0.02 |
64
+ | Experts | 2 |
65
+
66
+ ## 🚀 Quick Start
67
+
68
+ ```python
69
+ # Note: Full SCLM requires custom loading (see below)
70
+ # The inference widget uses the base model only
71
+
72
+ from transformers import AutoTokenizer
73
+ import torch
74
+
75
+ # Load tokenizer
76
+ tokenizer = AutoTokenizer.from_pretrained("amewebstudio/ananke-sclm")
77
+
78
+ # For full SCLM functionality, load weights separately:
79
+ # 1. Load base Mistral-7B
80
+ # 2. Load EARCP weights from earcp_weights.pt
81
+ # 3. Apply SCLM wrapper
82
+ ```
83
+
84
+ ## 📈 Validation Results
85
+
86
+ | Test | Result |
87
+ |------|--------|
88
+ | Forward Pass | ✅ |
89
+ | State Evolution | ✅ (norm: 0 → 4.6 → 7.5) |
90
+ | Coherent Generation | ✅ |
91
+ | Edit Mode | ✅ |
92
+ | Entity Memory | ✅ (Elara, Nimbus retained) |
93
+
94
+ ## 💡 Use Cases
95
+
96
+ - **Interactive Fiction**: Characters and plot points remain consistent
97
+ - **Long Conversations**: Context persists without growing prompts
98
+ - **Creative Writing**: Maintain story coherence across chapters
99
+ - **Role-Playing**: NPCs remember past interactions
100
+
101
+ ## 📝 Citation
102
+
103
+ ```bibtex
104
+ @article{amega2025sclm,
105
+ title={SCLM: Stateful Coherent Language Models with EARCP Architecture},
106
+ author={Amega, Mike},
107
+ year={2025},
108
+ note={Ame Web Studio}
109
+ }
110
+ ```
111
+
112
+ ## 👤 Author
113
+
114
+ **Mike Amega** - [Ame Web Studio](https://github.com/Volgat)
115
+
116
+ ---
117
+
118
+ *SCLM is an experimental architecture exploring persistent memory in language models.*
earcp_weights.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:76d6e5369d11288e4329e75b59b83ada490fadc367fdf6c9dcb0ecf170e37ec8
3
+ size 366837710
sclm_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 32000,
3
+ "hidden_size": 4096,
4
+ "num_hidden_layers": 32,
5
+ "latent_state_dim": 256,
6
+ "n_experts": 2,
7
+ "n_coherence_heads": 4,
8
+ "expert_intermediate": 1024,
9
+ "state_injection_layers": [
10
+ 8,
11
+ 16
12
+ ],
13
+ "alpha_inject": 0.02
14
+ }
special_tokens_map.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token": {
3
+ "content": "<s>",
4
+ "lstrip": false,
5
+ "normalized": false,
6
+ "rstrip": false,
7
+ "single_word": false
8
+ },
9
+ "eos_token": {
10
+ "content": "</s>",
11
+ "lstrip": false,
12
+ "normalized": false,
13
+ "rstrip": false,
14
+ "single_word": false
15
+ },
16
+ "pad_token": "</s>",
17
+ "unk_token": {
18
+ "content": "<unk>",
19
+ "lstrip": false,
20
+ "normalized": false,
21
+ "rstrip": false,
22
+ "single_word": false
23
+ }
24
+ }
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer.model ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dadfd56d766715c61d2ef780a525ab43b8e6da4de6865bda3d95fdef5e134055
3
+ size 493443
tokenizer_config.json ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_bos_token": true,
3
+ "add_eos_token": false,
4
+ "add_prefix_space": null,
5
+ "added_tokens_decoder": {
6
+ "0": {
7
+ "content": "<unk>",
8
+ "lstrip": false,
9
+ "normalized": false,
10
+ "rstrip": false,
11
+ "single_word": false,
12
+ "special": true
13
+ },
14
+ "1": {
15
+ "content": "<s>",
16
+ "lstrip": false,
17
+ "normalized": false,
18
+ "rstrip": false,
19
+ "single_word": false,
20
+ "special": true
21
+ },
22
+ "2": {
23
+ "content": "</s>",
24
+ "lstrip": false,
25
+ "normalized": false,
26
+ "rstrip": false,
27
+ "single_word": false,
28
+ "special": true
29
+ }
30
+ },
31
+ "additional_special_tokens": [],
32
+ "bos_token": "<s>",
33
+ "clean_up_tokenization_spaces": false,
34
+ "eos_token": "</s>",
35
+ "extra_special_tokens": {},
36
+ "legacy": false,
37
+ "model_max_length": 1000000000000000019884624838656,
38
+ "pad_token": "</s>",
39
+ "sp_model_kwargs": {},
40
+ "spaces_between_special_tokens": false,
41
+ "tokenizer_class": "LlamaTokenizer",
42
+ "unk_token": "<unk>",
43
+ "use_default_system_prompt": false
44
+ }
validation_results.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "forward": true,
3
+ "evolution": true,
4
+ "generation": true,
5
+ "edit_mode": true
6
+ }