Instructions to use ryoji-info/Gemma-4-12B-PsiLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ryoji-info/Gemma-4-12B-PsiLM with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Gemma-4-12B-PsiLM ryoji-info/Gemma-4-12B-PsiLM
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Card: final numbers before publishing
Browse files
README.md
CHANGED
|
@@ -164,6 +164,24 @@ The multi-mode bridges are released: they reach 100% in-distribution (n=48, MAE
|
|
| 164 |
|
| 165 |
The prompt runs through Gemma's first 20 layers. The **forward bridge** reads the queried position *x₀* by pooling the hidden states over its tokens (a deterministic span pointer computed by the QA builder, plus a 100-bin classifier) and the initial-condition parameters with a learned pool, after the calibrated per-dimension standardization) and emits (*a*, sin *φ*, cos *φ*) and *x₀*; from these it builds the initial condition on a 128-point grid. The frozen **FNO** evolves it to t = 0.5. A learned periodic lookup kernel reads the field at *x₀*, and the **value-token channel** turns that single number into eight soft tokens through Fourier features. At layer 30 a **gated cross-attention** injects them into the residual stream, capped at 20% of the stream's RMS; the gate is a small MLP on the residual stream, trained to open on physics prompts and close elsewhere. Layers 30–48 and the answer are Gemma's own. Details, ablations and the failure analysis that produced this design are in the paper (`paper/psilm.pdf` in the repository, Section 9 for scaling, the guard-rail and Gemma).
|
| 166 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 167 |
## Limitations
|
| 168 |
|
| 169 |
- **One task family.** The bridges read exactly the three quantities of the trained question and the FNO solves exactly one equation family; a different PDE, boundary condition, viscosity or final time is out of scope, and the gate closing on non-physics text does not mean it can recognize *other* physics. Free-text initial conditions ("a Gaussian bump near the left edge") are not supported.
|
|
|
|
| 164 |
|
| 165 |
The prompt runs through Gemma's first 20 layers. The **forward bridge** reads the queried position *x₀* by pooling the hidden states over its tokens (a deterministic span pointer computed by the QA builder, plus a 100-bin classifier) and the initial-condition parameters with a learned pool, after the calibrated per-dimension standardization) and emits (*a*, sin *φ*, cos *φ*) and *x₀*; from these it builds the initial condition on a 128-point grid. The frozen **FNO** evolves it to t = 0.5. A learned periodic lookup kernel reads the field at *x₀*, and the **value-token channel** turns that single number into eight soft tokens through Fourier features. At layer 30 a **gated cross-attention** injects them into the residual stream, capped at 20% of the stream's RMS; the gate is a small MLP on the residual stream, trained to open on physics prompts and close elsewhere. Layers 30–48 and the answer are Gemma's own. Details, ablations and the failure analysis that produced this design are in the paper (`paper/psilm.pdf` in the repository, Section 9 for scaling, the guard-rail and Gemma).
|
| 166 |
|
| 167 |
+
## Is the answer really coming through the channel?
|
| 168 |
+
|
| 169 |
+
Two controls, and the second is decisive. **Zeroing the injection** while running
|
| 170 |
+
everything else — readout, FNO, value tokens, gate — removes the physics result
|
| 171 |
+
(0% for Qwen3-8B, 10% for Gemma, which is what the reply template alone
|
| 172 |
+
recovers). **Corrupting only the number** — feeding the value encoder another
|
| 173 |
+
question's answer at matched magnitude, with prompt, readout, gate, reply length
|
| 174 |
+
and parsing untouched — makes the frozen model report the corruption: the spoken
|
| 175 |
+
answer lands within ±0.05 of the *injected* value on **99 of 100** held-out
|
| 176 |
+
questions and within ±0.05 of the truth on 9. Accuracy falls 98% → 9% while the
|
| 177 |
+
KL to the base model is unchanged (0.222 either way): the output distribution
|
| 178 |
+
travels just as far, to a different number.
|
| 179 |
+
|
| 180 |
+
Run on non-physics prompts the same swap changes nothing (GSM8K 0.88 both ways,
|
| 181 |
+
MMLU 0.66 both ways, p = 1.00), which separates what the channel does by its
|
| 182 |
+
**presence** from what it does by its **content**. Full sweep and records:
|
| 183 |
+
[`results/bench/leaky_8b_shuf_guardrail_summary.json`](https://github.com/ryoji-info/PsiLM/blob/main/results/bench) and §9.7 of the [paper](https://github.com/ryoji-info/PsiLM/blob/main/paper/psilm.pdf).
|
| 184 |
+
|
| 185 |
## Limitations
|
| 186 |
|
| 187 |
- **One task family.** The bridges read exactly the three quantities of the trained question and the FNO solves exactly one equation family; a different PDE, boundary condition, viscosity or final time is out of scope, and the gate closing on non-physics text does not mean it can recognize *other* physics. Free-text initial conditions ("a Gaussian bump near the left edge") are not supported.
|