Thumnail. Overall Framework
π‘
Summary
- Project Timeline: KoSAIM 2025 Poster(π Best Poster Award) β AI-BioX 2025 Poster β Critical Care Manuscript (Under review)
| Duration | 2025.05 - 2026.07 |
|---|---|
| Role | Co-First Author |
| Research Area | Clinical AI Β· Large Language Models |
| Keywords | LLM Β· Clinical AI Β· ICU Β· Explainable AI |
| Research Output | 2 Poster accepted / 1 Prize, |
| 1 Journal - Under review |
| KoSAIM | AI-BIoX | Critical Care | |
|---|---|---|---|
| Prompting | Zero-shot | Zero + Few-shot | 4 Prompting Strategies |
| LLM | GPT-5, Gemini | GPT-5, Gemini | GPT-5, Gemini |
| ML Models | LR, RF, XGBoost | LR, RF, XGBoost | ML + DL (6 models) |
| External Validation | MIMIC-IV | MIMIC-IV | MIMIC-IV + K-MIMIC |
| Explainability | SHAP | SHAP + LLM Reasoning | SHAPβLLM Alignment |
| Trustworthiness | - | - | Clinician Evaluation |
| Publication Stage | Poster | Poster | Journal Submission |
Research Evolution
KoSAIM 2025 | First Research Milestone π Best Poster Award
Conference: Korean Society of Artificial Intelligence in Medicine
Poster title: Developing an Early Mortality Prediction Model for ICU Patients Using LLM-based Approaches
Conference Poster πBest Poster Award - LLM/Agent Track
λ³Έ μ°κ΅¬λ μΌλ° λͺ©μ λκ·λͺ¨ μΈμ΄λͺ¨λΈ(LLM)μ΄ κ΅¬μ‘°νλ μ μμ무기λ‘(Structured EHR)λ§μ νμ©νμ¬ μ€νμμ€(ICU) μ‘°κΈ° μ¬λ§μ μμΈ‘ν μ μλμ§λ₯Ό νμνλ κ²μμ μμλμμ΅λλ€. κΈ°μ‘΄μ λ¨Έμ λ¬λ κΈ°λ° μμΈ‘ λͺ¨λΈκ³Ό GPT-5, Gemini-2.5-Proλ₯Ό λΉκ΅νμ¬ Zero-shot ν둬νν λ§μΌλ‘λ κ²½μλ ₯ μλ μμΈ‘ μ±λ₯μ λ³΄μΌ μ μλμ§λ₯Ό κ²μ¦νμμ΅λλ€. λν SHAP κΈ°λ° μ€λͺ κ°λ₯μ± λΆμκ³Ό LLMμ μΆλ‘ κ²°κ³Όλ₯Ό ν¨κ» λΉκ΅νμ¬ μμμ ν΄μ κ°λ₯μ±μ νκ°νμμ΅λλ€.

π Key Contribution
- General-purpose LLM κΈ°λ° ICU μ‘°κΈ° μ¬λ§ μμΈ‘
- Zero-shot Prompting μ±λ₯ νκ°
- SHAP κΈ°λ° μ€λͺ κ°λ₯μ± λΆμ
- GPT-5μ Gemini λΉκ΅ νκ°
- π KoSAIM 2025 Best Poster Award (LLM & Agent Track)
AI-BIoX 2025 | Second Milestone - Research Extension
Conference: AI-BioX ConfEX Grand Summit
KoSAIMμμ μ μν μ°κ΅¬λ₯Ό κΈ°λ°μΌλ‘ μ€νμ νμ₯νμ¬ Few-shot Promptingμ μΆκ°νκ³ , LLMμ΄ μμ±ν μΆλ‘ κ³Όμ μ μμμ νλΉμ±μ λ³΄λ€ μ¬μΈ΅μ μΌλ‘ λΆμνμμ΅λλ€. λ¨μν μμΈ‘ μ±λ₯ λΉκ΅λ₯Ό λμ΄, λͺ¨λΈμ΄ μ΄λ€ κ·Όκ±°λ‘ μμΈ‘μ μννλμ§μ κ·Έ μΆλ‘ μ΄ μμμ μΌλ‘ ν΄μ κ°λ₯νμ§λ₯Ό ν¨κ» νκ°νμ¬ μ°κ΅¬μ μμ±λλ₯Ό λμμ΅λλ€.

π Key Contribution
- Few-shot Prompting μΆκ° νκ°
- GPT-5μ Gemini μΆλ‘ λΉκ΅
- μ λμ μ±λ₯κ³Ό μ μ±μ μΆλ‘ λΆμ
- μμμ ν΄μ κ°λ₯μ± νκ°
Critical Care Journal Submission | Full Research
Title: Explainable Early ICU Mortality Prediction Using General-Purpose Large Language Models and Structured EHR
Status: Under Review
ν¬μ€ν° λ°νμμ μ»μ κ²°κ³Όλ₯Ό κΈ°λ°μΌλ‘ μ°κ΅¬λ₯Ό λν νμ₯νμ¬ Critical Care μ λ ν¬κ³ μ© λ Όλ¬ΈμΌλ‘ λ°μ μμΌ°μ΅λλ€. κΈ°μ‘΄ μ°κ΅¬μ λ€κ΅κ° ICU μ½νΈνΈ(MIMIC-IV, eICU-CRD, K-MIMIC)λ₯Ό μΆκ°νμ¬ μΌλ°ν μ±λ₯μ κ²μ¦νμμΌλ©°, MLΒ·DLΒ·LLMμ ν¬ν¨ν λ€μν λͺ¨λΈμ λΉκ΅νμμ΅λλ€. λν λ€μν Promptingμ λ΅κ³Ό SHAPβLLM μ λ ¬ λΆμ, μλ£μ§ μ λ’°λ νκ°λ₯Ό ν¬ν¨νμ¬ λͺ¨λΈμ μμΈ‘ μ±λ₯λΏ μλλΌ μ€λͺ κ°λ₯μ±κ³Ό μμμ νμ© κ°λ₯μ±κΉμ§ μ’ ν©μ μΌλ‘ νκ°νμμ΅λλ€.
π― What problem did I solve?
1. Background
Early ICU mortality prediction is essential for timely clinical decision-making, but conventional machine learning models require institution-specific development and often provide limited interpretability. Although recent large language models (LLMs) have shown promising reasoning capabilities, their ability to perform reliable clinical prediction from structured electronic health records remains largely unexplored.
β Research Questions
- Can general-purpose LLMs predict ICU mortality without task-specific fine-tuning?
- Do LLMs generalize across different healthcare systems?
- Does LLM-generated reasoning align with clinically meaningful predictors?
- Can clinicians trust LLM-generated explanations?
π― Objective
To evaluate the predictive performance, generalizability, and clinical trustworthiness of general-purpose LLMs for early ICU mortality prediction using structured EHR data across multinational ICU cohorts.
π§ How did I solve it?
2. Dataset
- Training Cohort: eICU-CRD, 124,020 ICU stays
-
Test Cohort: MIMIC-IV(84,926 ICU stays), K-MIMIC(6,673 ICU stays)
Three independent ICU cohorts from the United States and South Korea were used to evaluate cross-system generalization.
3. Methods
Data Processing
- Structured EHR preprocessing
- Missing value imputation
- Feature engineering
- Standardized preprocessing pipeline
Model Comparison
- Machine Learning: Logistic Regression, Random Forest, XGBoost
- Deep Learning: CNN, MLP, Transformer
- Large Language Models: GPT-5, Gemini-2.5-Pro
Prompt Engineering
- Zero-shot
- Chain-of-Thought
- Instruction-based CoT
- Agentic CoT
Explainability
- SHAP analysis
- SHAPβLLM reasoning alignment
- Clinician trustworthiness evaluation
π What were the results?
4. Results
- Finding 1: General-purpose LLMs achieved competitive discrimination without any task-specific fine-tuning, approaching the performance of supervised machine learning models.
- Finding 2: Zero-shot prompting provided stable and competitive performance, while increasingly complex prompting strategies produced only modest improvements.
- Finding 3: GPT-5 reasoning consistently focused on clinically important predictors identified by SHAP, demonstrating strong agreement between model explanations and statistical feature importance.
- Finding 4: Medical experts rated GPT-5 explanations as generally trustworthy, supporting the clinical interpretability of LLM-generated reasoning despite moderate inter-rater agreement.

Table 1. Main performance
Figure 2. Alignment between XGBoost SHAP β GPT-5 Reasoning txt
5. Discussion & Limitation
Discussion
- General-purpose LLMs demonstrated strong zero-shot capability for structured clinical prediction.
- Model reasoning aligned with clinically relevant predictors, suggesting potential as interpretable clinical decision-support tools.
- LLMs may complement rather than replace conventional machine learning models in real-world ICU settings.
Limitations
- Lower precision-recall performance under severe class imbalance.
- Proprietary LLMs limit reproducibility.
- Clinician trustworthiness evaluation involved only two raters.
- Prospective validation is required before clinical deployment.
π©βπ» What was my contribution?
- Research planning and study design
- Structured EHR preprocessing pipeline
- Prompt engineering for multiple LLM strategies
- LLM inference pipeline development
- Model benchmarking across multinational cohorts
- SHAP-based explainability analysis
- Result interpretation and visualization
- Manuscript writing
π‘ What did I learn?
Technical Insight
μ νν λͺ¨λΈμ ꡬμΆνλ κ²μ νλμ λ¨κ³μ λΆκ³Όν¨μ 체κ°ν μ μμμ΅λλ€. μ¬νκ°λ₯ν νκ° νμ΄νλΌμΈμ μ€κ³νκ³ μ€λͺ
κ°λ₯μ±μ ν΅ν΄ λͺ¨λΈμ λμμ μ΄ν΄νλ κ² λν μ λ’°ν μ μλ AI μμ€ν
μ κ°λ°νλ λ° μμ΄μ λ§€μ° μ€μνλ€λ κ²μ 체κ°ν μ μλ νλ‘μ νΈμμ΅λλ€.
Building an accurate model is only the first step. Designing reproducible evaluation pipelines and understanding model behavior through explainability are equally important for developing reliable AI systems.
Research Insight
μ°κ΅¬λ₯Ό λ°λΌλ³΄λ κ΄μ μ βAIκ° μ νν μμΈ‘μ ν μ μλκ°?βμμ βμμμκ° AIκ° μμ±ν κ²°μ μ μ λ’°ν μ μλκ°?βλ‘ μ ννκ² λμμ΅λλ€. μ΄λ€ λλ©μΈμ΄λ , λλ©μΈμ μ€λ¬΄μκ° μ¬μ©νλ AIλ μ±λ₯ μ§ν λΏλ§ μλλΌ ν¬λͺ
μ±, μΌλ°ν κ°λ₯μ±, κ·Έλ¦¬κ³ μ€μ©μ μΈ μ¬μ©μ±μΌλ‘λ νκ°λμ΄μΌ νλ€λ κ²μ κΉ¨λ¬μμ΅λλ€.
This project shifted my research perspective from "Can AI make accurate predictions?" to "Can clinicians trust AI-generated decisions?". I realized that future clinical AI should be evaluated not only by performance metrics but also by transparency, generalizability, and practical usability.