Spaces:
Sleeping
Sleeping
File size: 27,359 Bytes
8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 8ea9f7d 817fe64 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 | ---
title: STEM BIO-AI
emoji: "🧬"
colorFrom: blue
colorTo: green
sdk: gradio
sdk_version: "5.29.0"
python_version: "3.11"
app_file: app.py
pinned: false
---
# STEM BIO-AI
<p align="center">
<img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/logo.png" alt="STEM BIO-AI logo" width="390">
</p>
<p align="center">
<b>Deterministic evidence-surface scanner for bio/medical AI repositories.</b><br>
No LLM. No API key. No model runtime. No secrets sent anywhere.
</p>
<p align="center">
<a href="https://github.com/flamehaven01/STEM-BIO-AI/actions/workflows/python-package.yml"><img src="https://github.com/flamehaven01/STEM-BIO-AI/actions/workflows/python-package.yml/badge.svg" alt="CI"></a>
<a href="CHANGELOG.md"><img src="https://img.shields.io/badge/stable-v1.8.4-informational.svg" alt="v1.8.4"></a>
<a href="pyproject.toml"><img src="https://img.shields.io/badge/python-3.9%2B-blue.svg" alt="Python 3.9+"></a>
<a href="https://pypi.org/project/stem-ai/"><img src="https://img.shields.io/pypi/v/stem-ai.svg" alt="PyPI"></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-blue.svg" alt="Apache 2.0"></a>
<a href="https://huggingface.co/spaces/Flamehaven/stem-bio-ai"><img src="https://img.shields.io/badge/demo-Hugging%20Face%20Space-yellow.svg" alt="HF Space"></a>
<a href="https://doi.org/10.5281/zenodo.20154479"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.20154479.svg" alt="DOI"></a>
</p>
---
**Navigation:**
[Why](#why-stem-bio-ai) •
[Quick Start](#quick-start) •
[Verification](#verification-path) •
[Architecture](docs/ARCHITECTURE.md) •
[Trust Boundary](#runtime--security--compliance-boundary) •
[CLI Reference](docs/CLI_REFERENCE.md) •
[Scoring Rationale](docs/SCORING_RATIONALE.md)
---
## Why STEM BIO-AI
Bio and medical AI repositories vary enormously in evidence quality — from rigorous academic tools to marketing-grade demos that carry clinical language with no data provenance, no reproducibility path, and no clinical-use disclaimer. Manual review is slow and inconsistent.
STEM BIO-AI scans the **observable repository surface** — README, docs, code structure, CI configuration, dependency manifests, changelogs — and maps detected signals to a structured evidence tier (T0–T4). The scan runs in seconds on a local clone, produces machine-readable JSON and PDF reports, and makes every scoring decision traceable to a specific file, line, and pattern.
> A T4 score means strong observable evidence signals. It does not mean the repository is safe for clinical deployment — that requires independent expert validation.
---
## Quick Start
```bash
git clone https://github.com/flamehaven01/STEM-BIO-AI.git
cd STEM-BIO-AI
pip install stem-ai
```
```bash
# editable local install with PDF output support
pip install -e .[pdf]
# fastest path: scan a local repository
stem /path/to/bio-ai-repo
# 8-page full evidence packet with proof trace
stem scan /path/to/bio-ai-repo --level 3 --format all --explain
```
```bash
# workflow-oriented CLI
stem scan /path/to/bio-ai-repo --level 2
stem scan /path/to/bio-ai-repo --policy strict_clinical_adjacency
stem gate /path/to/bio-ai-repo --min-tier T2
stem policy list
stem policy explain strict_clinical_adjacency
stem policy derive --clinical-strictness 4 --code-integrity-priority 3 --reproducibility-priority 2 --structured-limitations-requirement 3
stem policy simulate /path/to/bio-ai-repo --clinical-strictness 4 --code-integrity-priority 3 --reproducibility-priority 2 --structured-limitations-requirement 3
stem policy simulate /path/to/bio-ai-repo --profile-file policy/drafts/scoring_profile.reproducibility_first.v1.json
stem advisory validate /path/to/bio-ai-repo
stem advisory packet /path/to/bio-ai-repo --output advisory_out
stem advisory check-response /path/to/bio-ai-repo --response provider_advisory.json
```
```bash
# backward-compatible shortcuts still work
stem /path/to/bio-ai-repo --level 3 --format all --explain
stem audit /path/to/bio-ai-repo --tier-gate T3 --quiet
```
Clone the target repository first; the CLI operates on local paths only.
Calibration profiles are implemented in `mirror_only` mode in `1.8.4`. `--policy` changes what profile is surfaced in artifacts, while `policy derive` and `policy simulate` provide governed preview lanes without mutating the authoritative deterministic score path. `policy simulate --profile-file <path>` allows local schema-valid profile experiments without registering a new named policy. In the current rule scope, `strict_clinical_adjacency` is the only release-grade named recommendation; stronger reproducibility postures still fall back to `preview_only` simulation deltas rather than a named profile.
Researchers and domain specialists are expected to influence calibration through `derive`, `simulate`, and documented preview/profile proposals. The intent interview uses a governed `1–5` posture scale, while official score-affecting policy changes still require profile promotion rather than direct ad hoc tuning.
Full CLI reference: [`docs/CLI_REFERENCE.md`](docs/CLI_REFERENCE.md)
## Verification Path
Use the same verification surface exposed in CI and package smoke tests:
```bash
pip install -e ".[pdf]"
python -m py_compile stem_ai/cli.py stem_ai/scanner.py stem_ai/render.py stem_ai/app.py
stem --help
python -m stem_ai --help
python -m pytest -q
python -m build
```
Primary references:
- [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md)
- [`docs/API_CONTRACT.md`](docs/API_CONTRACT.md)
- [`docs/SCORING_RATIONALE.md`](docs/SCORING_RATIONALE.md)
- [`docs/ADVISORY_RUNTIME.md`](docs/ADVISORY_RUNTIME.md)
- [`SECURITY.md`](SECURITY.md)
## Document Map
Use these docs by review purpose:
**Core operation**
- [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md)
- [`docs/CLI_REFERENCE.md`](docs/CLI_REFERENCE.md)
- [`docs/DETERMINISTIC_DIAGNOSTICS.md`](docs/DETERMINISTIC_DIAGNOSTICS.md)
- [`docs/UI_HTML_REPORT.md`](docs/UI_HTML_REPORT.md)
**Scoring and evidence**
- [`docs/SCORING_RATIONALE.md`](docs/SCORING_RATIONALE.md)
- [`docs/EXAMPLE_AUDITS.md`](docs/EXAMPLE_AUDITS.md)
- [`docs/CALIBRATION_PROFILE_DESIGN.md`](docs/CALIBRATION_PROFILE_DESIGN.md)
- [`docs/regulatory_basis_registry.v1.json`](docs/regulatory_basis_registry.v1.json)
**Trust boundary and governance**
- [`SECURITY.md`](SECURITY.md)
- [`docs/API_CONTRACT.md`](docs/API_CONTRACT.md)
- [`docs/ADVISORY_RUNTIME.md`](docs/ADVISORY_RUNTIME.md)
- [`docs/ADVISORY_SECRET_HANDLING.md`](docs/ADVISORY_SECRET_HANDLING.md)
- [`docs/REGULATORY_MAPPING.md`](docs/REGULATORY_MAPPING.md)
- [`docs/AIRI_DATA_GOVERNANCE.md`](docs/AIRI_DATA_GOVERNANCE.md)
- [`docs/THIRD_PARTY_DATA.md`](docs/THIRD_PARTY_DATA.md)
**Public proof surfaces**
- Demo: [Hugging Face Space](https://huggingface.co/spaces/Flamehaven/stem-bio-ai)
- Example audits: [`docs/EXAMPLE_AUDITS.md`](docs/EXAMPLE_AUDITS.md)
- Scoring rationale: [`docs/SCORING_RATIONALE.md`](docs/SCORING_RATIONALE.md)
---
## Triage Tiers
- **T0 Rejected (0–39):** insufficient evidence — do not rely on without independent expert validation
- **T1 Quarantine (40–54):** exploratory review only — expert validation required before any use
- **T2 Caution (55–69):** research reference and supervised non-clinical technical review only
- **T3 Supervised (70–84):** supervised institutional review candidate
- **T4 Candidate (85–100):** strong evidence posture — clinical deployment still requires independent validation
Clinical-adjacent repositories without an explicit disclaimer are **hard-capped at T2** (score ≤ 69).
Repositories with unbounded CA-DIRECT claims are **hard-capped at T0** (score ≤ 39).
Tier boundary derivation and calibration gap disclosures: [`docs/SCORING_RATIONALE.md`](docs/SCORING_RATIONALE.md).
---
## Scoring Model
```
Final = (Stage 1 × 0.40) + (Stage 2R × 0.20) + (Stage 3 × 0.40) − C1 Penalty
```
| Stage | Weight | What Is Measured |
|-------|-------:|-----------------|
| **Stage 1** README Evidence | 40% | Bio-domain vocabulary; H1–H6 hype-claim penalties; R1–R5 responsibility signals (limitations, regulatory framing, clinical disclaimer, demographic-bias, reproducibility) |
| **Stage 2R** Repo-Local Consistency | 20% | Vocabulary overlap across README, docs, package metadata, CI, and tests; limitation repetition; contradiction, staleness, and unsupported-workflow deductions |
| **Stage 3** Code/Bio Responsibility | 40% | CI presence; domain test coverage; changelog hygiene (T3); data provenance and IRB/dataset citation (B1); bias/limitation measurement evidence (B2); conflict-of-interest disclosure (B3) |
| **Stage 4** Replication Evidence | Separate lane | Containers; reproducibility targets; dependency locks/pins; dataset and model artifact references; seed, CLI, and citation signals; license/use-scope restrictions |
| **C1–C6** Code Integrity | Penalty / advisory | Hardcoded credentials (C1, −10 pts); dependency pinning and external-service fragility (C2); deprecated patient-adjacent paths (C3); fail-open exception handlers (C4); compliance and clinical-boundary integrity (C5); mock-auth or no-auth local/self-host boundary warnings (C6) |
Stage 4 is reported as `replication_score` / `replication_tier` and does **not** affect `score.final_score`. Full scoring rationale and calibration gap disclosures are in [`docs/SCORING_RATIONALE.md`](docs/SCORING_RATIONALE.md).
---
## Architecture
```mermaid
flowchart LR
A[Target repository] --> B[LOCAL_ANALYSIS scanner]
B --> C[Stage 1\nREADME evidence]
B --> D[Stage 2R\nRepo-local consistency]
B --> E[Stage 3\nCode/bio responsibility]
B --> F[Stage 4\nReplication lane]
B --> K[C1–C6\nCode integrity]
B --> CC[CC1–CC3\nAST contract detectors]
C --> G[Weighted evidence score]
D --> G
E --> G
K --> G
CC --> R[code_contract + AIRI coverage]
F --> H[replication_score / tier]
G --> I[Canonical JSON result]
H --> I
R --> I
I --> L[Evidence ledger]
I --> M[Explain trace]
I --> N[Markdown report]
I --> O[PDF packets 1p / 5p / 8p]
I --> P[Interactive HTML dashboard]
```
Core modules: `stem_ai/scanner.py`, `stem_ai/render.py`, `stem_ai/cli.py`, `stem_ai/detectors.py`, `stem_ai/detector_surface.py`, `stem_ai/detector_ast.py`, `stem_ai/detector_bio.py`, `stem_ai/detector_contract.py`, `stem_ai/detector_stage4.py`, `stem_ai/evidence.py`, `stem_ai/airi_risk_mapping.py`, `stem_ai/app.py`
---
## Output Artifacts
Each run writes to `--out DIR` (default: `stem_output/`).
The plain `stem <repo>` and `stem scan <repo>` path now defaults to `--level 3`, which emits the full 8-page evidence packet unless you select a lower level explicitly.
`audits/` is retained only for historical benchmark and reference artifacts; routine CLI output should land in `stem_output/<repo_slug>/`.
| Level | Pages | Audience | Artifacts |
|-------|------:|---------|-----------|
| `--level 1` | 1 | Executive / triage (legacy) | Score, tier, stage cards, code integrity summary |
| `--level 2` | 5 | Standard audit review | Level 1 + Stage 1/2R/3/4 breakdown, AIRI summary, closeout page |
| `--level 3` | 8 | Full evidence packet | Level 2 + dedicated regulatory traceability page, code integrity deep dive, remediation/AIRI/method page, metadata page |
```
<repo>_experiment_results.json # machine-readable score + full evidence object
<repo>_report.html # interactive 7-section HTML dashboard (v1.7.0+)
<repo>_report.md # human-readable audit report
<repo>_brief_1p.pdf # Level 1 executive dashboard
<repo>_detailed_5p.pdf # Level 2 standard review packet
<repo>_detailed_8p.pdf # Level 3 full review packet
<repo>_explain.txt # --explain: file/line/snippet proof trace
```
---
## HTML Report Dashboard
`--format html` generates a self-contained interactive dashboard (v1.7.0+). Single `.html` file — no network, no external dependencies.
<p align="center">
<img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/html_report_preview.png" alt="STEM BIO-AI interactive HTML dashboard" width="760">
</p>
**Example interactive HTML audit**
- Open in browser: <https://htmlpreview.github.io/?https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/yorkeccak_bio_report.html>
- Raw HTML artifact: [`docs/assets/report-preview/yorkeccak_bio_report.html`](docs/assets/report-preview/yorkeccak_bio_report.html)
**7 sections:** Executive Summary · Decision Path · Code Integrity · Regulatory Traceability · AIRI Risk Triggers · Evidence Detail · Developer Follow-up
Interactive features: sticky scroll-spy nav · repo hyperlink in the hero header · `?` tooltip icons on every metric · click-to-expand integrity cards · covered/gaps + domain filtering for AIRI risks · FAIL/WARN/PASS/INFO filter on the evidence ledger.
Current `1.8.4` HTML semantics:
- `Decision Path` explains score construction and policy posture with `Configured, Not Rewritten`
- `Code Integrity` surfaces the split between `C4` fail-open exceptions, `C5` compliance/boundary integrity, and `C6` mock-auth/no-auth trust boundaries
- `AIRI Risk Triggers` distinguishes the **full local AIRI registry**, the **curated runtime bundle**, and the **detector mapping registry**
- covered AIRI rows carry bounded `why mapped` reasoning derived from detector-trigger evidence plus the local detector-mapping registry
This is a review aid, not a claim that AIRI independently verified the repository.
---
## Report Preview
<p align="center">
<img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-1.png" alt="STEM BIO-AI full 8-page packet — page 1" width="760">
</p>
**Sample PDF:** [Download the 8-page full packet preview](https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/yorkeccak_bio_detailed_8p.pdf)
<details>
<summary>View all 8 full-packet preview pages</summary>
| Page 1 | Page 2 |
|--------|--------|
| <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-1.png" alt="Page 1"> | <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-2.png" alt="Page 2"> |
| Page 3 | Page 4 |
|--------|--------|
| <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-3.png" alt="Page 3"> | <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-4.png" alt="Page 4"> |
| Page 5 | Page 6 |
|--------|--------|
| <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-5.png" alt="Page 5"> | <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-6.png" alt="Page 6"> |
| Page 7 | Page 8 |
|--------|--------|
| <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-7.png" alt="Page 7"> | <img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/report-preview/8p-8.png" alt="Page 8"> |
</details>
---
## Detection Methods
Every scored item maps to a concrete, inspectable detection method. No inference, no LLM judgment.
<details>
<summary>Full detection table</summary>
| Component | Detection Method |
|-----------|-----------------|
| Stage 1 baseline | Non-zero README present (+60 base) |
| Stage 1 domain signal | Bio-domain keyword regex in README and package metadata |
| Stage 1 hype penalties (H1–H6) | Regex: clinical certainty, regulatory approval, autonomous replacement, breakthrough marketing, universal generalization, perfect accuracy claims |
| Stage 1 responsibility signals (R1–R5) | Regex: limitations section, regulatory framework, clinical disclaimer (CA-severity-weighted), demographic-bias disclosure, reproducibility provisions |
| Stage 2R consistency | Vocabulary set intersection across README/docs/package/tests; limitation repetition; clinical-boundary contradiction, version-staleness, and workflow-support deductions |
| Stage 3 T1 CI | `.github/workflows/` contains at least one file |
| Stage 3 T2 domain tests | `tests/` directory text contains bio-domain vocabulary (regex) |
| Stage 3 T3 changelog | CHANGELOG file presence + bug-fix/patch/security entry detection (3-tier: 0/+5/+15) |
| Stage 3 B1 data provenance | Dependency manifest presence + IRB/dataset-citation language detection (3-tier: 0/+10/+15) |
| Stage 3 B2 bias measurement | Bias/limitations vocabulary + quantitative measurement evidence (subgroup analysis, AUROC, demographic parity) (3-tier: 0/+8/+15) |
| Stage 3 B3 COI/funding | Funding, grant, sponsor, conflict-of-interest language in README/docs/FUNDING.md |
| Stage 4 containers | Dockerfile or compose file present |
| Stage 4 reproducibility target | Makefile with reproduce/eval/benchmark/test targets |
| Stage 4 dependency lock | Environment/lock/requirements file; exact pins or hash evidence |
| Stage 4 artifact references | Dataset/model/checkpoint URLs or checksum files |
| Stage 4 citation/interface | CITATION.cff; argparse CLI entry points (AST) |
| Stage 4 license restriction | Non-commercial, research-only, academic-only, no-clinical-use restrictions in LICENSE/README |
| CA severity | Clinical/diagnostic phrase regex in README, docs, and package metadata |
| C1 credentials | AWS `AKIA*`, OpenAI `sk-*`, GitHub `ghp_*`, `api_key=...` patterns; obvious placeholders excluded from penalty |
| C2 dependency pinning | `==` or hash pin vs. loose `>=`, `~=`, `<`, `>` ranges |
| C3 deprecated paths | Patient-metadata patterns in `deprecated/`, `legacy/`, `archive/` directories |
| C4 fail-open | `except Exception: pass` or `except: pass` in Python source (AST) |
| C5 compliance boundary integrity | Unsupported legal/compliance claims or missing clinical-boundary integrity in reviewed sources |
| **CC1** clinical zero default | AST scan of function defaults: keyword-only and positional params named `confidence_threshold`, `score_threshold`, `min_confidence`, etc. defaulted to `0.0` |
| **CC2** API contract | README-declared names cross-checked against `__all__` exports; phantom APIs flagged |
| **CC3** shallow validator | `validate_*` / `check_*` functions using only `len()` (no regex structure check) flagged as insufficient for clinical/PII validation |
Stage 2R and Stage 3 rubric artifacts now surface additive `detector_id` and `decision_basis` fields so reviewers can see which bounded detector or contradiction rule produced a deduction or credit.
</details>
---
## AI Advisory Contract
The advisory system exports a sanitized, provider-neutral handoff packet and validates provider responses — without making any provider API call.
```bash
stem advisory validate /path/to/repo # offline contract check
stem advisory packet /path/to/repo # export sanitized input packet
stem advisory check-response /path/to/repo --response FILE
```
**Non-negotiable rules (enforced by the validator):**
- Provider output cannot override `score.final_score` or `score.formal_tier`
- Every advisory item must cite exact `finding_id` strings from `allowed_finding_ids`
- Raw repository source text is not included in provider packets
- Responses containing clinical safety, efficacy, regulatory, or medical-advice claims are rejected
- `allowed_finding_ids` is capped at 40 entries per packet
**Packet hardening added in v1.5.7:**
- `provider_request` now carries a secret-free request schema plus deterministic argument-validation status
- `contract_schemas` exports the advisory input/output contract shapes for downstream validators
- `packet_contract` confirms allowlist parity, snippet omission, and non-negative omission counts before handoff
**Secret boundary hardening added in v1.5.9:**
- provider-specific environment variables are recognized before the generic advisory key fallback
- provider handoff metadata exports endpoint-policy validation and the expected env-var name, never the key value
- embedded-credential URLs are rejected; cloud providers require `https`; plain `http` is limited to localhost
- `.env` files are ignored by default; `.env.example` documents supported variable names only
- `--advisory call` is now the explicit provider-call boundary, with centralized redaction, logging-policy export, child-env allowlist reporting, and artifact pre-write sanitization
Full contract: [`docs/API_CONTRACT.md`](docs/API_CONTRACT.md)
Secret policy: [`docs/ADVISORY_SECRET_HANDLING.md`](docs/ADVISORY_SECRET_HANDLING.md)
Runtime boundary: [`docs/ADVISORY_RUNTIME.md`](docs/ADVISORY_RUNTIME.md)
---
## The AI Risk Repository (AIRI)
STEM BIO-AI uses local derived data from the MIT **AI Risk Repository (AIRI)** as a broader risk-vocabulary layer around deterministic repository findings.
Upstream references:
- MIT AI Risk Repository: <https://airisk.mit.edu/>
- AI Incident Tracker: <https://airisk.mit.edu/ai-incident-tracker>
How AIRI is used here:
- AIRI does **not** replace the local scoring and audit system
- AIRI does **not** prove harm, causality, clinical safety, or regulatory status
- AIRI helps place local findings into a wider risk vocabulary for review
In the current `1.8.4` line, AIRI is used through three local governed layers:
1. full normalized local registry
2. curated runtime bundle used by deterministic scans
3. detector-to-risk mapping registry plus known-gap tracking
This allows STEM BIO-AI to keep scan behavior local and deterministic while still surfacing broader AI risk language, provenance, and bundle-scope boundaries in runtime artifacts.
License / provenance note:
- Upstream AIRI source license: `MIT`
- Local attribution and usage details: [`docs/AIRI_DATA_GOVERNANCE.md`](docs/AIRI_DATA_GOVERNANCE.md), [`docs/THIRD_PARTY_DATA.md`](docs/THIRD_PARTY_DATA.md)
---
## Runtime / Security / Compliance Boundary
STEM BIO-AI can help teams become more **audit-ready**, but it does not by itself create certification, attestation, or legal compliance.
What can be prepared internally:
- runtime and security evidence review
- control-matrix and evidence-room preparation
- validation-package assembly for electronic records / signature workflows
- gap assessment for logging, access control, change control, retention, and traceability
- independent third-party audit readiness and penetration-test readiness
What still requires external review or attestation:
- SOC 2 report issuance
- ISO 13485 certification
- strong `21 CFR Part 11 compliant` claims
- `independent audit passed` claims
In other words: internal teams can do substantial readiness work, but external claims still require external auditors, certification bodies, or independent assessors.
Related boundary guidance: [`docs/REGULATORY_MAPPING.md`](docs/REGULATORY_MAPPING.md)
---
## MICA Memory Layer
The repository keeps a versioned MICA memory layer under `memory/` for agent-session initialization,
drift control, and release provenance. Historical state is preserved through Git-tagged release history;
the active layer is selected by `memory/mica.yaml`.
The working tree intentionally keeps only the current active MICA trio:
- `memory/stem-ai.mica.v1.8.4.json`
- `memory/stem-ai-playbook.v1.8.4.md`
- `memory/stem-ai-lessons.v1.8.4.md`
Older release-memory snapshots are preserved in Git-tagged history rather than as parallel live files in the visible repo surface.
The active package now follows the non-breaking `MICA v0.2.4` runtime contract:
- `memory/mica.yaml` is the composition contract
- `python tools/mica_pct.py .` validates package integrity
- `python tools/mica_runtime.py . --format text` emits a portable session summary
- `python tools/mica_runtime.py . --format session-report` emits an opening-state gate packet
- `python tools/mica_invoke.py . --mode guided --format json` compiles a host-consumable activation packet
- `mica_invoke.bat . --mode forced` is the Windows forced-preflight entry point
- DI binding remains progressive rather than speculative
critical invariants are not mass-rewritten just to satisfy schema formality
Operational reference: [`docs/MICA_MEMORY.md`](docs/MICA_MEMORY.md)
---
## Web Demo
Live demo: [huggingface.co/spaces/Flamehaven/stem-bio-ai](https://huggingface.co/spaces/Flamehaven/stem-bio-ai)
<p align="center">
<img src="https://raw.githubusercontent.com/flamehaven01/STEM-BIO-AI/main/docs/assets/HF-STEM-BIO_AI.png" alt="STEM BIO-AI Hugging Face Space" width="760">
</p>
The Space runs the same deterministic local scanner on public GitHub repositories. No provider API call is made.
Run locally:
```bash
pip install -e .[demo]
python app.py
```
---
## Repository Structure
```
STEM-BIO-AI/
stem_ai/ # Core Python package
docs/ # API contract, advisory runtime/secret policy, scoring rationale, MICA policy, report previews
memory/ # Versioned MICA archive/playbook/lessons; active layer selected by mica.yaml
audits/ # Historical benchmark/reference artifacts only
stem_output/ # Default live CLI output root (generated, ignored)
scripts/ # Benchmark and validation scripts
tests/ # Regression test suite
app.py # HuggingFace Spaces / Gradio entry point
pyproject.toml # Package metadata and extras
SKILL.md # Universal agent skill definition
CHANGELOG.md # Version history
```
---
## Agent Skill Install
```bash
# Claude Code
git clone --depth 1 https://github.com/flamehaven01/STEM-BIO-AI.git ~/.claude/skills/stem-bio-ai
# Generic agent frameworks
git clone --depth 1 https://github.com/flamehaven01/STEM-BIO-AI.git ~/.agents/skills/stem-bio-ai
```
---
## Contributing
See [CONTRIBUTING.md](CONTRIBUTING.md). High-value areas: rubric discrimination examples, clinical-adjacency trigger refinements, additional bio-domain benchmark repositories, report rendering improvements.
---
## Citation
Preferred citation metadata lives in [`CITATION.cff`](CITATION.cff).
Current concept DOI-backed archive for the `1.8.4` line:
- <https://doi.org/10.5281/zenodo.20154479>
```bibtex
@software{stem-bio-ai,
author = {Yun, Kwansub},
title = {STEM BIO-AI: Deterministic Evidence-Surface Scanner for Bio/Medical AI Repositories},
version = {1.8.4},
year = {2026},
doi = {10.5281/zenodo.20154479},
url = {https://doi.org/10.5281/zenodo.20154479}
}
```
---
## License
Apache 2.0. See [LICENSE](LICENSE).
Maintained by [flamehaven01](https://github.com/flamehaven01)
|