File size: 9,519 Bytes
b0d24b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5604bd0
b0d24b2
5604bd0
b0d24b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5604bd0
b0d24b2
 
 
 
 
 
 
 
5604bd0
b0d24b2
 
 
 
 
 
 
 
 
5604bd0
b0d24b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c218d8a
b0d24b2
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
# Antaḥkaraṇa-Net — Scale-Up Benchmark & Performance Report

**A capability benchmark of the Vedic-mind continual-learning architecture at 36.5M–52.6M parameters.**
All numbers are from **live inference on the trained checkpoints** (the model is loaded from disk and
re-evaluated), not training-time logs. 11 models, 7 capabilities each, run on a single NVIDIA A10.

> **Scope, stated up front.** This is a *research proof-of-concept*: WideResNets on Split-CIFAR-100 and
> Split-Tiny-ImageNet. The contribution is the **architecture** — one agent that learns continually
> without forgetting, abstains instead of hallucinating, and self-regulates — demonstrated to *scale*
> from the original small-CNN POC to real models. It is not, and is not meant to be, a frontier model.

---

## 1. Setup

| | |
|---|---|
| **Architectures** | WRN-28-10 (**36.5M params**), WRN-28-12 (**52.6M params**) — shared trunk φ + one linear head per task |
| **Control ring** | chitta (saṃskāra EWC+decay), guṇa (forgetting-aware MetaGunaController), āśrama, tapas, pramāṇa, turīya |
| **Datasets** | Split-CIFAR-100 (10×10 and 20×5 class-incremental); Split-Tiny-ImageNet (200-class, 10×20, 64→32) |
| **OOD set (pramāṇa)** | SVHN (test) |
| **Training** | from scratch, 25 epochs/task, Adam; wake→dream(replay)→sleep(consolidate) cycle |
| **Hardware** | OCI `BM.GPU.A10.4` → single A10 (23 GB), us-sanjose-1 |
| **Protocol** | per-checkpoint live inference: per-task accuracy, pramāṇa gating, temperature calibration, mind-state trace |
| **Seeds** | core ×5, scale ×2, tiny ×2, āśrama ×2 (11 models total) |

## 2. Capability scorecard (the 7 checks + PASS criteria)

| # | Capability (faculty) | PASS criterion |
|---|---|---|
| 1 | **Learns the task** (buddhi+backbone) | avg accuracy > 3× chance |
| 2 | **Continual memory** (chitta/saṃskāra) | forgetting ≤ 0.05 **and** all tasks > 1.5× chance |
| 3 | **Anti-hallucination** (pramāṇa) | gated acc ≥ raw acc **and** OOD blind-accept < 30% |
| 4 | **Calibration** | ECE decreases after temperature scaling |
| 5 | **Self-regulation** (guṇa) | plasticity headroom decreases over life |
| 6 | **Lifelong** (āśrama) | progresses through life-stages to vānaprastha/saṃnyāsa |
| 7 | **Identity monitor** (turīya) | witness drift bounded (< 1.0) |

**Result: 10 of 11 models pass 7/7; 1 passes 6/7** (āśrama_s0 — see §6, a threshold artifact).

## 3. Headline results — group means

| Configuration | Params | Avg acc | Naive forget | **Agent forget** | **Reduction** | Gated acc | OOD accept |
|---|---|---|---|---|---|---|---|
| **core** — CIFAR-100, 10 tasks | 36.5M | 0.599 | 0.589 | **0.0146** | **41.0×** | 0.929 | 8.4% |
| **scale** — CIFAR-100, 10 tasks | 52.6M | 0.615 | 0.551 | **0.0204** | **27.3×** | 0.933 | 10.0% |
| **tiny** — Tiny-ImageNet, 10 tasks | 36.5M | 0.416 | 0.503 | **0.0173** | **29.1×** | 0.956 | 0.3% |
| **āśrama** — CIFAR-100, **20 tasks** | 36.5M | 0.606 | 0.595 | **0.0474** | **12.5×** | 0.960 | 1.3% |

**Read:** catastrophic forgetting is cut **12–41×** across two model sizes, two datasets, and stream lengths
of 10 and 20 tasks. The agent both *remembers* (forgetting → ~0.01–0.05) and *learns better* (accuracy
roughly **doubles** vs the naive baseline, which collapses to its last task).

![forgetting](https://huggingface.co/deepakdsoni/antahkarana/resolve/main/plots/bench_forgetting.png)

![scorecard](https://huggingface.co/deepakdsoni/antahkarana/resolve/main/plots/bench_scorecard.png)

## 4. Per-model detail (live inference)

| Model | Data | Avg acc | Naive→Agent forget | Reduction | Gated@cov | OOD accept | ECE before→after | Score |
|---|---|---|---|---|---|---|---|---|
| core_s0 | CIFAR-100 | 0.616 | 0.565→0.018 | 31.8× | 0.943@0.21 | 16.1% | 0.172→0.023 | 7/7 |
| core_s1 | CIFAR-100 | 0.592 | 0.621→0.014 | 45.1× | 0.950 | 6.4% | 0.172→0.043 | 7/7 |
| core_s2 | CIFAR-100 | 0.617 | 0.583→0.012 | 49.9× | 0.918 | 2.6% | 0.118→0.025 | 7/7 |
| core_s3 | CIFAR-100 | 0.587 | 0.553→0.015 | 37.1× | 0.927 | 8.2% | 0.150→0.024 | 7/7 |
| core_s4 | CIFAR-100 | 0.585 | 0.621→0.015 | 41.1× | 0.907 | 8.4% | 0.163→0.035 | 7/7 |
| scale_s0 | CIFAR-100 | 0.620 | 0.542→0.021 | 25.4× | 0.921 | 8.7% | 0.229→0.027 | 7/7 |
| scale_s1 | CIFAR-100 | 0.610 | 0.570→0.019 | 29.3× | 0.944 | 11.2% | 0.195→0.032 | 7/7 |
| tin_s0 | Tiny-ImageNet | 0.411 | 0.503→0.017 | 29.2× | 1.000 | 0.4% | 0.226→0.025 | 7/7 |
| tin_s1 | Tiny-ImageNet | 0.421 | 0.502→0.017 | 29.0× | 0.912 | 0.2% | 0.196→0.034 | 7/7 |
| ashrama_s0 | CIFAR-100 (20t) | 0.592 | 0.603→0.049 | 12.3× | 0.951 | 1.7% | 0.225→0.046 | 6/7 |
| ashrama_s1 | CIFAR-100 (20t) | 0.621 | 0.586→0.046 | 12.8× | 0.969 | 1.0% | 0.246→0.057 | 7/7 |

## 5. Analysis

**Scaling (36.5M → 52.6M).** Increasing capacity keeps the effect and nudges accuracy up (0.599 → 0.615).
The forgetting-reduction *ratio* is a touch lower at 52.6M (27× vs 41×), but absolute agent-forgetting stays
tiny (0.020 vs 0.015) — i.e. both models forget almost nothing; the ratio mostly reflects naive-baseline
variance. **The architecture is genuinely backbone-agnostic across sizes.**

**Difficulty — dataset.** Tiny-ImageNet (200 fine-grained classes at 32 px) is much harder, so raw accuracy
drops to ~0.42. Crucially, **forgetting reduction holds (29×)** and the model becomes *more* conservative:
it abstains on **99.7%** of OOD inputs and clears only its most-confident in-distribution predictions —
exactly the right behavior when uncertain.

**Difficulty — stream length.** Doubling the stream (10 → 20 tasks, āśrama) is the hardest continual setting:
twice as many tasks to retain with more interference. Reduction drops to **12.5×** (agent-forgetting 0.047) —
**lower than the 10-task runs, and honestly so** — yet all 20/20 tasks remain above chance and the model still
abstains/calibrates cleanly. A >12× cut over a 20-task lifelong run is a strong result.

![retention](https://huggingface.co/deepakdsoni/antahkarana/resolve/main/plots/bench_retention.png)
*Per-task accuracy after the full stream — every task stays well above chance; the agent does not collapse
to its last task (which is what the naive baseline does).*

**Anti-hallucination (pramāṇa).** Calibrated abstention works everywhere: gated accuracy **0.91–1.00** (vs raw
~0.42–0.67), and OOD blind-accept **0.2%–16%**. The harder the data, the lower the model's confidence and the
more it abstains — a desirable, safety-aligned property. Coverage is correspondingly lower on hard sets
(a deliberate trade: it commits less, but is right when it does).

![pramana](https://huggingface.co/deepakdsoni/antahkarana/resolve/main/plots/bench_pramana.png)

**Calibration.** Temperature scaling reduces ECE everywhere (e.g. 0.172 → 0.023), so reported confidence can
be trusted — the precondition for the abstention gate to be meaningful.

**Interpretability (mind-state).** Every model exposes a legible trajectory: plasticity headroom **falls over
life** (e.g. 0.92 → 0.37; 0.77 → 0.14 on the 20-task run), the agent **ages through all four āśramas**
(brahmacarya → gṛhastha → vānaprastha → saṃnyāsa), and **turīya drift stays bounded** (~0.50–0.59). You can
read the agent's regulation as it lives — not a black box.

![mindstate](https://huggingface.co/deepakdsoni/antahkarana/resolve/main/plots/bench_mindstate.png)

## 6. The one 6/7, explained (no goalpost-moving)

**āśrama_s0** failed only **capability 1 (learns the task)**: avg accuracy **0.592 vs the 0.60 bar**
(= 3× chance for 5-class tasks). That's **2.96× chance** — a hair under an arbitrary threshold. Its
*continual_memory* check passed with **20/20 tasks retained and 12.3× forgetting reduction**, and āśrama_s1
(same config) scored 0.621 → 7/7. This is **benchmark-threshold variance on a borderline metric**, not a model
deficiency. We report it as-is rather than relax the criterion.

## 7. Honest caveats

- **POC scale.** Small/mid models, 32 px imagery. Demonstrates the *architecture*, not SOTA accuracy.
- **Forgetting reduction ratios are noisy** across seeds (driven partly by naive-baseline variance); the
  trustworthy claim is *absolute agent-forgetting ≈ 0.01–0.05 everywhere*.
- **Low coverage on hard datasets** is the flip side of strong abstention — fewer commitments, higher precision.
- **OOD uses SVHN with CIFAR-100 normalization** for all models (a single fixed OOD probe); a per-dataset OOD
  suite would sharpen the tiny-imagenet numbers.
- **Vedic↔ML mappings are engineering analogies**, clearly flagged. Nothing here is conscious.

## 8. Reproducibility

```bash
# train (per phase, resumable, idempotent):
bash scale/phases.sh                 # core → scale → pramana → tinyimagenet → ashrama
# benchmark any checkpoint (live inference, 7-capability report + JSON):
CUDA_VISIBLE_DEVICES=0 python3 scale/capability_bench.py --ckpt runs/<tag>.pt
```

Artifacts: 11 checkpoints (`runs/*.pt`), 11 training results (`runs/*.json`), 11 capability reports
(`runs/*_capbench.json`), and this report. The uploadable model is `antahkarana-36.5M-cifar100-wrn28-10.pt`
(36.5M, with Ω/θ* consolidation state) + `load_akn.py` + `MODEL_CARD.md` in `scale/hf_package/`.

---

*Generated from live-inference benchmarks on trained checkpoints. Code: MIT. Built on the Upaniṣads,
Sāṃkhya, Yoga, and PyTorch · snnTorch.*