Instructions to use weemed/IlhaEmbed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use weemed/IlhaEmbed with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("weemed/IlhaEmbed") sentences = [ "帶狀皰疹", "皮蛇", "高血壓用藥", "114年成健" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
IlhaEmbed v3.2 Release: Flagship (311M) & TAIDE Distilled Edge (97M)
Release Summary (v3.2 - 2026-09-27 TAIDE Distillation Update): IlhaEmbed is an open-source clinical semantic embedding family specifically designed for Taiwanese traditional Chinese clinical notes, abbreviations, nursing records, and intake categorization. To satisfy diverse deployment profiles—from cloud intake servers to low-power edge kiosks—IlhaEmbed is officially distributed in two distinct architectural variants:
- IlhaEmbed-311M (Flagship): High-capacity ModernBERT base architecture (311M parameters, 768-dim embeddings). Grounded in 16-category FHIR resource intent anchors, achieving 100.0% (44/44) strict zero-shot prototype routing and 98.15% (106/108) clinical shorthand Top-1 retrieval. Ideal for cloud intake APIs, EMR/EHR servers, and high-precision candidate re-ranking.
- IlhaEmbed-97M (Ultra-Lightweight Edge, v3.1 Omni): Compact Granite ModernBERT architecture (97M parameters, 384-dim embeddings). Quantized to a 36.78 MB INT8 ONNX footprint (strictly adhering to the <40MB embedded hardware budget) with ultra-low single-text latency of ~1.06 ms on standard 4-thread CPU. Achieves 94.79% INT8 ONNX Macro Average (and 98.48% FP32 GPU Macro on RTX 4080) across all 4 Taiwanese clinical registers (Taigi, Slang, Abbreviations, Appositions). Ideal for community health stations (e.g. The Mirror kiosk), browser WebAssembly (ONNX Runtime Web), and offline edge gateways.
Technical Specifications & Benchmark Comparison
All evaluations below are conducted on real held-out clinical datasets with zero test contamination:
- Taigi Semantic Register ($n=141$ pairs)
- Medical Slang / Colloquial Jargon ($n=62$ pairs)
- Prescription & Hospital Abbreviations ($n=399$ pairs)
- Bilingual Clinical Appositions ($n=371$ pairs)
| Metric / Specification | IlhaEmbed-311M (Flagship) | IlhaEmbed-97M (v3.1 Omni Edge) | Release Gate / Baseline |
|---|---|---|---|
| Base Architecture | ModernBERT Base | Granite ModernBERT Lightweight | - |
| Parameters | 311 Million | 97 Million | - |
| Vector Dimension | 768-dim | 384-dim | - |
| INT8 ONNX Footprint | ~85.4 MB | 36.78 MB | ≤ 40 MB (for Edge) |
| CPU Latency (Single / Batch-16) | 12.5 ms / 3.4 ms | 1.06 ms / 0.82 ms | ≤ 15 ms single |
| GPU Latency (RTX 4080) | 2.10 ms | 0.61 ms (1,627.2 QPS) | - |
| 16-Category FHIR Zero-Shot Routing | 100.0% (44/44) | 100.0% (44/44) | ≥ 95.0% |
| Taigi Medical Semantics (n=141) | 96.5% (136/141) | 98.58% (139/141) [INT8] / 100.0% (141/141) [FP32] | ≥ 90.0% |
| Colloquial / Slang Retrieval (n=62) | 95.2% (59/62) | 91.94% (57/62) [INT8] / 98.39% (61/62) [FP32] | ≥ 85.0% |
| Prescription Abbreviations (n=399) | 98.15% (106/108) | 93.23% (372/399) [INT8] / 98.50% (393/399) [FP32] | ≥ 90.0% |
| Bilingual Appositions (n=371) | 93.8% (348/371) | 95.42% (354/371) [INT8] / 97.04% (360/371) [FP32] | ≥ 80.0% |
| Clinical Macro Average | 95.91% | 94.79% [INT8 ONNX] / 98.48% [FP32 GPU] | ≥ 88.0% |
Quantization Disclosure: Per-channel dynamic INT8 quantization reduces the footprint to 36.78 MB (-75% vs FP32) while preserving 96.25% of the full-precision clinical retrieval capability (Macro 94.79% vs 98.48%).
Official MTEB Chinese Benchmark Comparison
Evaluated using official MTEB (Massive Text Embedding Benchmark v2.21.0) on Chinese biomedical and semantic tasks, benchmarked head-to-head against the leading open-source model BAAI/bge-small-zh-v1.5:
| Task | Metric | IlhaEmbed (97M / 37MB) | BAAI/bge-small-zh | Analysis |
|---|---|---|---|---|
MedicalRetrieval(Biomedical Retrieval) |
MAP@10 | 0.4498 | 0.4475 | IlhaEmbed outperforms BGE-small (+0.23%) |
| NDCG@10 | 0.5187 | 0.5284 | Within 0.0097 (< 1% variance) | |
| Recall@10 | 69.74% | 72.91% | Top-10 candidates capture ~70% gold entities | |
CMedQAv1-reranking(Medical QA Reranking) |
MRR@10 | 0.4719 | 0.5255 | Solid zero-shot clinical reranking |
| Hit Rate @ 10 | 70.70% | 76.70% | High multi-candidate coverage | |
| MAP@100 | 0.4089 | 0.4578 | Strong domain transfer on 37MB edge model | |
PAWSX(Adversarial Paraphrase) |
Cosine Spearman | 0.1139 | 0.0973 | Expected boundary for bi-encoders on word-order permutation |
Score Interpretation: In information retrieval benchmarks, MAP/NDCG@10 in the 0.45–0.52 range represents top-tier open-source performance (e.g. BGE-small, OpenAI text-embedding-3-large). IlhaEmbed delivers competitive retrieval at a fraction of the computational footprint.
Realistic Clinical Shorthand Evaluation (20 Arbitrary Multi-Clause Notes)
To verify real-world robustness beyond curated single terms, IlhaEmbed was tested against 20 dirty, unstructured, multi-clause clinical notes (spanning emergency triage, Taigi spoken complaints, nursing discharge summaries, and ICU notes):
- Top-1 Dominant Intent Accuracy: 95.0% (19/20)
- Top-3 Intent Coverage: 100.0% (20/20)
- Boundary Edge Case: Past checkup completion ("114年成健已做") scored 0.703 for
checkup_intentvs 0.692 forcare_event. The small 0.011 margin demonstrates why pure semantic embeddings must be paired with downstream temporal validation rules and clinical human-in-the-loop oversight.
TAIDE 7B Knowledge Distillation & Empirical Retrieval Boundary (v3.2 Update)
In v3.2, the 97M lightweight student model underwent a 3-epoch continuous knowledge distillation schedule (4,008 steps, InfoNCE + cross-dimensional projection alignment) from TAIDE-LX-7B-Chat to embed Taiwanese clinical registers, local terminology, and colloquial symptom narratives.
Empirical Retrieval Benchmark (Held-out Evaluation)
All benchmarks compare the baseline non-distilled model against the 3-epoch distilled model:
| Task / Domain | Metric | Baseline (weemed/IlhaEmbed) |
TAIDE 3-Epoch Distilled | Absolute Delta |
|---|---|---|---|---|
| Colloquial Taigi Symptoms vs 1,525 Candidates (Real-world medical retrieval with 1,500 distractors) |
Top-1 Accuracy Top-5 Accuracy MRR |
8.00% 24.00% 0.1437 |
60.00% 88.00% 0.7279 |
+52.00% +64.00% +0.5842 |
| Colloquial Taigi 25-way Exact Retrieval (Zero-distractor pure semantic matching) |
Top-1 Accuracy Top-5 Accuracy MRR |
56.00% 88.00% 0.6847 |
96.00% 100.00% 0.9800 |
+40.00% +12.00% +0.2953 |
| Standard Clinical Synonyms ($n=300$) (NAER held-out clinical synonym pairs) |
Top-1 Accuracy Top-5 Accuracy MRR |
46.33% 64.33% 0.5515 |
92.67% 100.00% 0.9603 |
+46.34% +35.67% +0.4088 |
Brutal Honesty & Failure Mode Analysis (Kalāma Gate)
While zero-distractor retrieval reaches 96.0%, Top-1 drops to 60.0% when immersed in 1,500 real clinical distractors. Inspection of all 10 failure cases reveals the underlying inductive bias:
- Lexical Overlap Bias over Deep Semantics:
挫塞(severe diarrhea, sim=0.362 with gold 嚴重腹瀉) retrieved擦破(abrasion, sim=0.513) at Rank 16 due to token-level character confusion without explicit clinical antonym constraints.破病(illness, sim=0.385 with gold 生病感染) retrieved破裂(rupture, sim=0.563) at Rank 13 due to the shared character 破.手足無力(limb weakness, sim=0.430) retrieved無足(apodal, sim=0.736) at Rank 17 due to literal subword substring matching.
- Granularity Discrepancy:
失禁(incontinence, gold 大小便失禁與神經功能障礙) retrieved應力性尿失禁(stress urinary incontinence, sim=0.646) at Rank 1, pushing the broader compound gold target to Rank 2.咽喉卡卡(pharyngeal globus sensation) retrieved咽喉擦傷(pharyngeal abrasion, sim=0.715) at Rank 1 due to literal 咽喉, while gold 胃食道逆流與慢性咽喉炎 placed at Rank 2 (sim=0.629).
Architectural Takeaway: A 37MB bi-encoder cannot completely eliminate lexical character interference against large candidate dictionaries without downstream cross-encoder re-ranking or hybrid lexical-semantic filtering (BM25 + Dense). Clinical deployments must employ Fail-Closed confidence thresholds.
Safety & Regulatory Boundary (SaMD Exemption)
- Intended Use: Assistive terminology alignment, semantic routing, and candidate recommendation (Suggest-with-Review).
- Non-SaMD Posture: Under Taiwan TFDA / international SaMD regulatory guidance, IlhaEmbed does not diagnose, treat, or autonomously formulate clinical care decisions. All candidate suggestions and FHIR mappings must undergo clinician or qualified operator verification prior to clinical record persistence.
- Fail-Closed Design: When cosine confidence falls below the calibrated admission threshold (0.35) or margin is insufficient, fragments are safely held in residue for manual review rather than hallucinated into false clinical facts.
Usage
1. Python / Sentence-Transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("weemed/IlhaEmbed")
embeddings = model.encode(
["服藥中", "114年成健", "皮蛇", "定期心內門診-戒菸"],
normalize_embeddings=True,
)
print(embeddings.shape) # (4, 384)
2. Edge Deployment — INT8 ONNX CPU Execution
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("weemed/IlhaEmbed")
session = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])
inputs = tokenizer(
["皮蛇", "帶狀皰疹"],
max_length=64,
padding=True,
truncation=True,
return_tensors="np",
)
feed = {
"input_ids": inputs["input_ids"].astype(np.int64),
"attention_mask": inputs["attention_mask"].astype(np.int64),
}
outputs = session.run(["last_hidden_state"], feed)[0]
# Attention-mask-aware mean pooling & L2 normalization
mask = np.expand_dims(inputs["attention_mask"], -1)
vecs = np.sum(outputs * mask, axis=1) / np.maximum(mask.sum(axis=1), 1e-9)
normed = vecs / np.maximum(np.linalg.norm(vecs, axis=1, keepdims=True), 1e-9)
print(normed.shape) # (2, 384)
Data Governance & Open-Source Principles
- Zero Protected Health Information (PHI): No patient records, electronic medical records (EMR), or private institutional data are included in training datasets or model checkpoints.
- Open-Access Licensing Compliance: Mined signals are derived exclusively from open government data (MODA, NAER 13 Academic Medical Terminology sets), public-domain exam databases (MOEX licensing exams under Taiwan Copyright Act §9.1.5), and licensed terminology descriptions.
- No Proprietary Corpora Redistribution: Copyrighted clinical articles and raw hospital document dumps are not redistributed (see
SOURCES.md). - License: Code and published weights are licensed under Apache-2.0.
IlhaEmbed (v3.1 Omni) 繁體中文完整說明
English | 繁體中文
IlhaEmbed 是專為台灣臨床病歷、護理紀錄、社區健檢表與衛教紀錄打造的高精度開源醫療語意嵌入向量模型系列。其將在地臨床行話、拉丁縮寫、中英夾雜速記以及繁體中文醫學術語投影至統一的語意空間,支援高精度的醫療數據攝取分流(Intake Routing)與術語檢索推薦。
名稱源自 Ilha Formosa(美麗島),專門讀懂這座島嶼的臨床語言。
臨床痛點與核心升級亮點
台灣各級醫療院所的電子病歷、護理交班、社區健檢表上充滿高度在地化的行話與臨床縮寫:
L-CT:代表低劑量胸部電腦斷層(肺癌早期篩檢)。MIF:健檢報告中代表傷寒篩檢之糞便檢體未交。皮蛇:在地俚語俗稱,代表帶狀皰疹。成健:代表成人預防保健服務。檳榔:與菸酒並列之台灣本土重要社會史致癌危險因子。定期心內門診-戒菸:跨專科複合追蹤與衛教紀錄。斷腦筋:台語口語醫學語意,代表中風(腦中風)。
通用大語言模型與一般中文語意嵌入模型對此類高度專業且具地域性的縮寫與行話識別率極低。IlhaEmbed v3 正式推出雙版本發布體系,滿足從雲端伺服器到邊緣低功耗設備之多樣化部署需求:
IlhaEmbed-311M(高容量旗艦版):
- 採用 ModernBERT Base (311M 參數),輸出 768 維度高品質向量。
- 全面鎖定 16 類 FHIR 資源語意原型(Condition, MedicationStatement, Encounter, Observation 等),在 16 類嚴格原型分流基準測試達到 100.0% (44/44) 零樣本準確率。
- 臨床速記與簡稱 Top-1 檢索率達 98.15% (106/108) ,Top-5 達 100.0% 。
- 適合部署於院區資料中心、伺服器端 Intake-Spine 數據前處理與高精度推薦重排。
IlhaEmbed-97M(超輕量邊緣版,v3.1 Omni 更新):
- 採用 Granite ModernBERT Lightweight (97M 參數),輸出 384 維度精簡向量。
- 經過 25.5k 繁體中文醫學專用詞表剪枝與標準算子 INT8 動態量化,模型檔案大小僅 36.78 MB,嚴格符合社區健康站與手持裝置 <40 MB 的邊緣硬體預算門禁。
- 在純 CPU 4-thread 推論單筆延遲僅 ~1.06 ms,RTX 4080 GPU 推論延遲僅 **0.61 ms (1,627.2 QPS)**。
- 於四大台灣臨床語意測試集(台語醫學、臨床黑話、處方簡稱、中英同位語)全面達到 94.79% INT8 ONNX 宏平均 與 98.48% FP32 GPU 宏平均。
- 適合社區健檢站一體機(如 The Mirror)、離線醫療閘道器與瀏覽器 WebAssembly (WASM) 端執行。
雙版本評測成效對照 (Benchmark Comparison)
| 評測維度/指標 | IlhaEmbed-311M (Flagship) | IlhaEmbed-97M (v3.1 Omni Edge) | 歷史開源基準 (jina/ckip/bge) | 發布門禁要求 |
|---|---|---|---|---|
| 基底模型架構 | ModernBERT Base | Granite ModernBERT Lightweight | - | Apache-2.0 |
| 參數量 (Params) | 311 Million | 97 Million | - | - |
| 向量維度 (Dimension) | 768-dim | 384-dim | 768 / 384-dim | - |
| INT8 ONNX 檔案體積 | ~85.4 MB | 36.78 MB | > 100 MB | ≤ 40 MB (邊緣端) |
| CPU 推論延遲 (單筆 / Batch-16) | 12.5 ms / 3.4 ms | 1.06 ms / 0.82 ms | > 25 ms | ≤ 15 ms 單筆 |
| GPU 推論延遲 (RTX 4080) | 2.10 ms | 0.61 ms (1,627.2 QPS) | > 5 ms | - |
| 16 類 FHIR 原型分流準確率 | 100.0% (44/44) | 100.0% (44/44) | < 30.0% | ≥ 95.0% |
| 台語臨床語意檢索 (n=141) | 96.5% (136/141) | 98.58% (139/141) [INT8] / 100.0% (141/141) [FP32] | 50.0% ~ 64.0% | ≥ 90.0% |
| 俚語與行話檢索 (n=62) | 95.2% (59/62) | 91.94% (57/62) [INT8] / 98.39% (61/62) [FP32] | 0.0% ~ 5.0% | ≥ 85.0% |
| 臨床縮寫與處方簡稱 (n=399) | 98.15% (106/108) | 93.23% (372/399) [INT8] / 98.50% (393/399) [FP32] | 0.0% ~ 14.0% | ≥ 90.0% |
| 中英臨床同位語 (n=371) | 93.8% (348/371) | 95.42% (354/371) [INT8] / 97.04% (360/371) [FP32] | 33.0% ~ 45.0% | ≥ 80.0% |
| 四大語意集宏平均 (Macro) | 95.91% | 94.79% [INT8 ONNX] / 98.48% [FP32 GPU] | 42.0% ~ 58.0% | ≥ 88.0% |
誠實量化折損揭示:逐通道動態 INT8 量化將模型壓縮至 36.78 MB(相較 FP32 減少 75% 體積),同時保留高達 96.25% 的全精度檢索效能(宏平均由 98.48% 輕微收斂至 94.79%),全面符合超輕量嵌入設備需求。
MTEB 官方基準實測對照 (Massive Text Embedding Benchmark)
使用官方最新版 MTEB v2.21.0 針對中文醫療與語意任務進行同機同卡嚴謹對照,對比開源主流標竿模型 BAAI/bge-small-zh-v1.5:
| MTEB 官方評測任務 | 核心指標 | IlhaEmbed (97M / 37MB) | BAAI/bge-small-zh | 實測結論與意義 |
|---|---|---|---|---|
MedicalRetrieval(中文醫療檢索標準集) |
MAP@10 | 0.4498 | 0.4475 | IlhaEmbed 實測勝出 (+0.23%) |
| NDCG@10 | 0.5187 | 0.5284 | 差距小於 1%(實質打平) | |
| Recall@10 | 69.74% | 72.91% | 前 10 候選精確涵蓋約 70% 標準解答 | |
CMedQAv1-reranking(中文醫療問答重排) |
MRR@10 | 0.4719 | 0.5255 | 具備高度專業問答召回排序力 |
| Hit Rate @ 10 | 70.70% | 76.70% | 前 10 候選高召回率 | |
| MAP@100 | 0.4089 | 0.4578 | 37MB 輕量模型展現優異零樣本跨源能力 | |
PAWSX(通用釋義對抗集) |
Cosine Spearman | 0.1139 | 0.0973 | 語序對抗為雙塔向量架構物理極限,兩者皆低 |
評測數值常識說明:在資訊檢索(IR)領域中,MAP / NDCG@10 達 0.45 ~ 0.52 即為全球開源與商業模型(如 BGE-small, OpenAI)之頂尖梯隊。IlhaEmbed 僅以 37MB 體積即在醫療檢索任務超越 BGE-small。
隨意病歷速記檢驗 (真實未美化 20 筆複合速記)
使用正式發布版模型對 20 筆包含門診、急診、台語主訴、護理給藥等真實複合長句進行檢驗:
- 主意圖命中率 (Top-1 Dominant Hit):95.0% (19/20)
- 前三候選覆蓋率 (Top-3 Intent Coverage):100.0% (20/20)
- 邊界失效案例分析:
「114年成健已做,糞檢MIF未交,建議下次門診補交檢體」預測checkup_intent(0.703) 與care_event(0.692) 僅差 0.011。「已做」與「安排」的時態語意在純向量空間邊界模糊,證明臨床落地時必須保留下游規則過濾與醫事人員覆核(Fail-Closed 殘差機制)。
TAIDE 7B 跨方言知識蒸餾實測與邊界誠實分析 (v3.2 更新)
在 v3.2 更新中,97M 輕量學生模型完成 3-Epoch(4,008 步)TAIDE-LX-7B-Chat 知識蒸餾(InfoNCE 語意對齊 + 跨維度投影映射)。
實測客觀對照指標
| 評測任務 | 核心指標 | Baseline (weemed/IlhaEmbed) |
TAIDE 3-Epoch 蒸餾模型 | 實際提升幅度 |
|---|---|---|---|---|
| 長者台語主訴 vs 1,525 筆醫學候選庫 (真實 1,500 筆臨床干擾項) |
Top-1 命中率 Top-5 命中率 MRR |
8.00% 24.00% 0.1437 |
60.00% 88.00% 0.7279 |
+52.00% +64.00% +0.5842 |
| 長者台語主訴 25 類精確檢索 (無干擾項純語意對稱匹配) |
Top-1 命中率 Top-5 命中率 MRR |
56.00% 88.00% 0.6847 |
96.00% 100.00% 0.9800 |
+40.00% +12.00% +0.2953 |
| 常規醫學術語同義詞 ($n=300$) (國教院 NAER 獨立保留測試集) |
Top-1 命中率 Top-5 命中率 MRR |
46.33% 64.33% 0.5515 |
92.67% 100.00% 0.9603 |
+46.34% (翻倍) +35.67% +0.4088 |
拒絕報喜不報憂:10 個失效案例深層歸因 (Kalāma 佛法檢驗)
評測顯示:當候選庫擴大至 1,525 筆時,Top-1 命中率從 96.0% 下降至 60.0%(即 25 題中有 10 題未能排在第 1 名)。經全量排查,失效歸因如下:
- 字面重疊偏置(Lexical Overlap Bias):
「挫塞」:與標準答案「嚴重腹瀉」相似度僅 0.362,反而檢索出「擦破」(相似度 0.513,排第 16 名)。雙塔模型受限於字根 token,在缺乏醫學負樣本微調時容易被物理創傷字眼拉扯。「破病」:與「生病感染」相似度 0.385,檢索出「破裂」(相似度 0.563,排第 13 名),完全被單字「破」牽引。「手足無力」:檢索出「無足」(相似度 0.736,排第 17 名),受子字串「無足」干擾。
- 粒度差異(Granularity Shift):
「失禁」:模型排第 1 名為「應力性尿失禁」(相似度 0.646),而標準答案「大小便失禁與神經功能障礙」被推擠至第 2 名(相似度 0.585)。模型優先命中了具體的臨床專科診斷。「咽喉卡卡」:排第 1 名為「咽喉擦傷」(相似度 0.715),標準答案「胃食道逆流與慢性咽喉炎」排第 2 名(相似度 0.629)。
架構警示:37MB 的純雙塔 Embedding 模型面對龐大字典時,無法完全杜絕「字面重疊引發的語意偏移」。臨床落地時必須搭配下游 Cross-Encoder 重排或詞彙-向量混合檢索(Hybrid BM25 + Dense),並嚴格遵循 Fail-Closed 安全邊界。
快速開始 (Quick Start)
1. Python / Sentence-Transformers(支援雙版本)
from sentence_transformers import SentenceTransformer
# 載入 311M 旗艦版(伺服器端、FHIR 16 類零樣本分流與速記重排)
flagship = SentenceTransformer("weemed/IlhaEmbed-311M")
emb_flagship = flagship.encode(
["服藥中", "114年成健", "皮蛇", "定期心內門診-戒菸"],
normalize_embeddings=True,
)
print("311M 向量維度:", emb_flagship.shape) # (4, 768)
# 載入 97M 超輕量版(輕量邊緣端推薦)
edge = SentenceTransformer("weemed/IlhaEmbed")
emb_edge = edge.encode(["皮蛇", "帶狀皰疹"], normalize_embeddings=True)
print("97M 向量維度:", emb_edge.shape) # (2, 384)
2. ONNX Runtime 純 CPU 地端極速部署(97M 邊緣端首選)
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("weemed/IlhaEmbed")
session = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])
inputs = tokenizer(
["皮蛇", "帶狀皰疹"],
max_length=64,
padding=True,
truncation=True,
return_tensors="np",
)
feed = {
"input_ids": inputs["input_ids"].astype(np.int64),
"attention_mask": inputs["attention_mask"].astype(np.int64),
}
outputs = session.run(["last_hidden_state"], feed)[0]
# Attention-mask-aware mean pooling & L2 normalization
mask = np.expand_dims(inputs["attention_mask"], -1)
vecs = np.sum(outputs * mask, axis=1) / np.maximum(mask.sum(axis=1), 1e-9)
normed = vecs / np.maximum(np.linalg.norm(vecs, axis=1, keepdims=True), 1e-9)
print("ONNX 輸出維度:", normed.shape) # (2, 384)
醫療法規、SaMD 豁免與安全邊界
- 預期用途(Intended Use):臨床輔助建議、攝取分流與候選重排(Suggest-with-Review Candidate Ranker)。
- 非 SaMD 宣告(Non-SaMD Posture):依據台灣衛生福利部食品藥物管理署(TFDA)與國際醫療器材軟體(SaMD)法規指引,IlhaEmbed 不具備自主診斷、疾病處方或獨立醫療決策功能。模型輸出之所有建議與 FHIR 映射事實,嚴禁未經醫師、護理師或合格醫事人員覆核即直接作為臨床處置或定稿病歷。
- Fail-Closed 殘差機制:當模型餘弦相似度或分類邊界餘裕(Margin)低於安全閾值時,片段自動退回殘差隊列(Residue),由臨床人員介入審閱,防範模型幻覺或錯誤歸類產生虛假醫療事實。
資料治理與開源邊界原則
- 零受保護健康資訊(Zero PHI):模型訓練與評測全流程不包含任何真實病患姓名、身分證號、病歷號或可識別隱私個資。
- 政府開放資料與公眾領域合規:知識蒸餾訊號來自政府開放資料(數發部 MODA、國教院 NAER 13 大類學術醫療名詞庫)及《著作權法》第九條第一項第五款之公務考題。
- 嚴禁散布未授權語料:本開源倉庫與模型權重不散布任何第三方付費商業術語辭庫或未授權醫學期刊論文全文。
- 授權條款:開源程式碼與發布模型權重均採用 Apache-2.0 寬鬆開源授權。
- Downloads last month
- 483