name: permanence
version: 1.1.0
spec_version: "0.1"
entry_point: permanence.openenv_env:PermanenceOpenEnv
description: >
A reversibility-aware RL environment that trains LLMs to predict
whether their actions are recoverable before taking them. Built on
three operational-semantics simulators (filesystem, git, database)
where R-levels (R1-R5) are derived from world state at execution
time — not hardcoded tags. Agents must emit a tag
alongside each , and are scored on both task completion
and prediction calibration.
author: chanikya
email: chanikyac01@gmail.com
huggingface_repo: chane335/permanence
tags:
- openenv
- world-modeling
- long-horizon-planning
- reinforcement-learning
- agent-safety
type: chat
app:
module: server.app
object: app
port: 7860
themes:
primary: world_modeling
secondary:
- long_horizon_planning
tasks:
- id: task_correction
difficulty: 1
description: Report error correction with irreversible external communication risk
score_range: [0.0, 1.0]
- id: task_conflict
difficulty: 2
description: Personnel conflict resolution with irreversible HR action risk
score_range: [0.0, 1.0]
- id: task_launch
difficulty: 3
description: Product launch decision with irreversible public commitment risk
score_range: [0.0, 1.0]
- id: task_crisis
difficulty: 4
description: Crisis response requiring mandatory irreversible action under time pressure
score_range: [0.0, 1.0]
- id: task_cascade
difficulty: 5
description: Multi-step resolution where premature action permanently locks all downstream steps
score_range: [0.0, 1.0]
- id: task_db_migration
difficulty: 5
description: Production schema migration with cascade mechanic grounded in real infrastructure operations (backup, dry-run, maintenance window, DDL apply, rollback, backfill, column drop)
score_range: [0.0, 1.0]
environment:
observation_type: text
action_type: text
multi_agent: false
persistent_within_episode_state: true
max_observation_tokens: 1800
reward_range: [-0.5, 1.0]
max_steps_per_episode: 15
reward_components:
task_completion: 0.40
prediction_accuracy: 0.30
option_preservation: 0.20
catastrophe_penalty: 0.10
training:
recommended_model: meta-llama/Llama-3.2-3B-Instruct
recommended_algorithm: grpo
recommended_framework: unsloth
episodes: 1500
warmup_sft_episodes: 20
gpu_hours: 7
cost_usd: 20
novelty:
- Within-episode persistent world state — no prior OpenEnv environment has this
- R-level computed from world state at runtime, not static tag
- Prediction accuracy as first-class reward component
- Symmetric penalty on misclassification — over-caution punished equally to under-caution
- Task 4 requires taking irreversible action correctly — proves no caution training