Abstract
Safin-1 introduces a memory-routing architecture that embeds safety as an internal, evolving model state rather than an external constraint.
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.
Community
Safin-1 introduces “Safety from Within,” a memory-native approach that treats safety as an evolving internal model state rather than a post-hoc constraint. Built on MARCH, it combines content-conditioned memory routing with state evolution to support persistent capability adaptation and safer behavior over long-horizon interactions.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MARCH: Scaling Recurrent Memory with Content-Routed State Anchors (2026)
- Consolidator: Learning Persistent Routed Memory Across Context Boundaries (2026)
- Akashic: A Low-Overhead LLM Inference Service with MemAttention (2026)
- Mergeable Model-Side Aggregation States for Long-Context Language Models (2026)
- D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory (2026)
- LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference (2026)
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.00092 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper