AI & ML interests

Building a globally standardized, AI-driven digital health and long-term care platform for aging societies — FHIR-first, with integrated health data and intelligent decision support for personalized preventive medicine. AI accelerates digitization; it isn't the product.

Recent Activity

gloomcheng  updated a model 1 day ago
weemed/IlhaEmbed
gloomcheng  updated a Space 3 days ago
weemed/ilhaembed-demo
gloomcheng  updated a model 8 days ago
weemed/README
View all activity

Organization Card

WeeMed AI

Building open-source, edge-native clinical AI and FHIR infrastructure for aging societies.

Deployed in real clinics, mobile screening stations, and community eldercare sites across Taiwan.

Live Space Demo · IlhaEmbed Model · GitHub


🇹🇼 Taiwan Sovereign Clinical AI Stack

We engineer hyper-compact, edge-native models designed for the realities of frontline healthcare: constrained hardware, strict air-gapped privacy, localized clinical shorthand, and spoken elderly dialects.

🧭 IlhaEmbed: Edge Biomedical Embedding Engine

37.28 MB INT8 ONNX · 384-dim · 3.97 ms on Vanilla CPU (250+ notes/sec)

  • Dual-Faceted Real-World Clinical Adaptation:
    • Spoken Elderly Vernacular (AST / STT): Resolves raw spoken complaints transcribed from older adults in Taiwanese (Taigi) into international clinical concepts (e.g., 「阿嬤講伊心臟跳真緊,腳頭烏痛,全身軟巡巡沒力氣」 → Condition: Generalized Malaise & Palpitations, 「皮蛇」 → Herpes Zoster).
    • Healthcare Staff Shorthand & NHI Codes: Decodes ultra-fast nursing notes and screening acronyms (e.g., 「114年成健已做,糞檢 MIF 未交」 → DiagnosticReport: Fecal Occult Blood / LOINC 14563-1, 「排 L-CT」 → ServiceRequest: Low-Dose Chest CT).
  • 100% Zero-Shot Intent Routing: Flawlessly categorizes text across 44 standard clinical and administrative anchors (HL7 FHIR Condition, DiagnosticReport, Medication, Observation, ServiceRequest, etc.).
  • MTEB Medical Benchmark: Achieves 0.4498 MAP, outperforming models several times its size (including BGE-small) while consuming a fraction of the compute.
  • Fail-Closed Safety Gate: Explicitly rejects administrative noise and non-clinical requests without hallucinating false diagnosis codes.
  • Interactive Console: Experience real-time inference directly in your browser at weemed/ilhaembed-demo.
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("weemed/IlhaEmbed")

# 1. Elderly spoken complaint (AST transcription)
emb1 = model.encode(["阿嬤講伊心臟跳真緊,全身軟巡巡沒力氣"])

# 2. Nursing shorthand & screening notation
emb2 = model.encode(["114年成健已做,糞檢 MIF 未交,排 L-CT"])

🎙️ Breeze-ASR-26-edge & Taiwanese Tailo ASR

Bilingual Taiwanese Hokkien (台語) + Mandarin Speech Recognition Stack

  • Quantized for edge devices in CTranslate2, ONNX, and GGML runtimes.
  • Multi-format phonetic transcription support (Traditional Chinese Hanzi, Taiwanese Romanization / Tâi-lô, and code-switched clinical dialogue).
  • Built upon MediaTek Research's Breeze-ASR-26 & Whisper architectures, fine-tuned on SuíSiann, TAT_MOE, and real-world multi-speaker clinical intake audio.

Why Edge-Native Clinical AI?

Taiwan is aging at one of the fastest rates in the world. The people providing care — community health workers, outreach nurses, and geriatric case managers — interact with seniors who speak Taigi, working in noisy community centers with handheld tablets or legacy PC kiosks.

  1. Patient Privacy & Air-Gap Compliance: Clinical notes and patient complaints cannot be routed through commercial cloud APIs. Everything must run on-premise.
  2. Vanilla Hardware Realities: Ward carts and rural health stations rarely have high-end GPUs. A model that requires 24GB of VRAM cannot help a rural nurse; a 37MB model running at 3.97ms on an older Intel CPU can.
  3. Honest Scientific Boundaries: Vector embeddings excel at semantic concept clustering, but are naturally insensitive to temporal sequencing ("completed" vs. "pending follow-up"). We position our models as calibrated, fail-closed edge sidecars, leaving final temporal logic to deterministic business rules and clinical professionals.

Open Science & Provenance

  • Permissive Open Source: Model weights, inference scripts, and web demonstrations are published under Apache-2.0.
  • Transparent Data Lineage: Grounded in open public health taxonomies (MODA, NAER, LOINC, SNOMED CT, RxNorm) with explicit attribution.
  • Field-Verified Benchmarks: Evaluated on genuine clinical notes, ASR transcripts, and real multi-speaker recordings—not synthetic noise.

Apache-2.0 License · Crafted in Yunlin & Chiayi, Taiwan · WeeMed AI

datasets 0

None public yet