← Back to portfolio
Journal Β· Best Poster Award

Explainable ICU Mortality Prediction Using General-Purpose LLMs

Co-First AuthorLLM Β· Clinical AIπŸ† KoSAIM 2025 Best PosterUnder Review

Thumnail. Overall Framework

Thumnail. Overall Framework

πŸ’‘

Summary

  • Project Timeline: KoSAIM 2025 Poster(πŸ† Best Poster Award) β†’ AI-BioX 2025 Poster β†’ Critical Care Manuscript (Under review)
Duration 2025.05 - 2026.07
Role Co-First Author
Research Area Clinical AI Β· Large Language Models
Keywords LLM Β· Clinical AI Β· ICU Β· Explainable AI
Research Output 2 Poster accepted / 1 Prize,
1 Journal - Under review
KoSAIM AI-BIoX Critical Care
Prompting Zero-shot Zero + Few-shot 4 Prompting Strategies
LLM GPT-5, Gemini GPT-5, Gemini GPT-5, Gemini
ML Models LR, RF, XGBoost LR, RF, XGBoost ML + DL (6 models)
External Validation MIMIC-IV MIMIC-IV MIMIC-IV + K-MIMIC
Explainability SHAP SHAP + LLM Reasoning SHAP–LLM Alignment
Trustworthiness - - Clinician Evaluation
Publication Stage Poster Poster Journal Submission

Research Evolution

KoSAIM 2025 | First Research Milestone πŸ† Best Poster Award

Conference: Korean Society of Artificial Intelligence in Medicine
Poster title: Developing an Early Mortality Prediction Model for ICU Patients Using LLM-based Approaches

Conference Poster πŸ†Best Poster Award - LLM/Agent Track

λ³Έ μ—°κ΅¬λŠ” 일반 λͺ©μ  λŒ€κ·œλͺ¨ μ–Έμ–΄λͺ¨λΈ(LLM)이 κ΅¬μ‘°ν™”λœ μ „μžμ˜λ¬΄κΈ°λ‘(Structured EHR)λ§Œμ„ ν™œμš©ν•˜μ—¬ μ€‘ν™˜μžμ‹€(ICU) μ‘°κΈ° 사망을 μ˜ˆμΈ‘ν•  수 μžˆλŠ”μ§€λ₯Ό νƒμƒ‰ν•˜λŠ” κ²ƒμ—μ„œ μ‹œμž‘λ˜μ—ˆμŠ΅λ‹ˆλ‹€. 기쑴의 λ¨Έμ‹ λŸ¬λ‹ 기반 예츑 λͺ¨λΈκ³Ό GPT-5, Gemini-2.5-Proλ₯Ό λΉ„κ΅ν•˜μ—¬ Zero-shot ν”„λ‘¬ν”„νŒ…λ§ŒμœΌλ‘œλ„ 경쟁λ ₯ μžˆλŠ” 예츑 μ„±λŠ₯을 보일 수 μžˆλŠ”μ§€λ₯Ό κ²€μ¦ν•˜μ˜€μŠ΅λ‹ˆλ‹€. λ˜ν•œ SHAP 기반 μ„€λͺ…κ°€λŠ₯μ„± 뢄석과 LLM의 μΆ”λ‘  κ²°κ³Όλ₯Ό ν•¨κ»˜ λΉ„κ΅ν•˜μ—¬ μž„μƒμ  해석 κ°€λŠ₯성을 ν‰κ°€ν•˜μ˜€μŠ΅λ‹ˆλ‹€.

μŠ¬λΌμ΄λ“œ1.PNG

πŸ“Œ Key Contribution

AI-BIoX 2025 | Second Milestone - Research Extension

Conference: AI-BioX ConfEX Grand Summit

KoSAIMμ—μ„œ μ œμ•ˆν•œ 연ꡬλ₯Ό 기반으둜 μ‹€ν—˜μ„ ν™•μž₯ν•˜μ—¬ Few-shot Prompting을 μΆ”κ°€ν•˜κ³ , LLM이 μƒμ„±ν•œ μΆ”λ‘  κ³Όμ •μ˜ μž„μƒμ  타당성을 보닀 μ‹¬μΈ΅μ μœΌλ‘œ λΆ„μ„ν•˜μ˜€μŠ΅λ‹ˆλ‹€. λ‹¨μˆœν•œ 예츑 μ„±λŠ₯ 비ꡐλ₯Ό λ„˜μ–΄, λͺ¨λΈμ΄ μ–΄λ–€ 근거둜 μ˜ˆμΈ‘μ„ μˆ˜ν–‰ν•˜λŠ”μ§€μ™€ κ·Έ 좔둠이 μž„μƒμ μœΌλ‘œ 해석 κ°€λŠ₯ν•œμ§€λ₯Ό ν•¨κ»˜ ν‰κ°€ν•˜μ—¬ μ—°κ΅¬μ˜ 완성도λ₯Ό λ†’μ˜€μŠ΅λ‹ˆλ‹€.

μŠ¬λΌμ΄λ“œ1.PNG

πŸ“Œ Key Contribution

Critical Care Journal Submission | Full Research

Title: Explainable Early ICU Mortality Prediction Using General-Purpose Large Language Models and Structured EHR
Status: Under Review

ν¬μŠ€ν„° λ°œν‘œμ—μ„œ 얻은 κ²°κ³Όλ₯Ό 기반으둜 연ꡬλ₯Ό λŒ€ν­ ν™•μž₯ν•˜μ—¬ Critical Care 저널 투고용 λ…Όλ¬ΈμœΌλ‘œ λ°œμ „μ‹œμΌ°μŠ΅λ‹ˆλ‹€. κΈ°μ‘΄ 연ꡬ에 λ‹€κ΅­κ°€ ICU μ½”ν˜ΈνŠΈ(MIMIC-IV, eICU-CRD, K-MIMIC)λ₯Ό μΆ”κ°€ν•˜μ—¬ μΌλ°˜ν™” μ„±λŠ₯을 κ²€μ¦ν•˜μ˜€μœΌλ©°, MLΒ·DLΒ·LLM을 ν¬ν•¨ν•œ λ‹€μ–‘ν•œ λͺ¨λΈμ„ λΉ„κ΅ν•˜μ˜€μŠ΅λ‹ˆλ‹€. λ˜ν•œ λ‹€μ–‘ν•œ Promptingμ „λž΅κ³Ό SHAP–LLM μ •λ ¬ 뢄석, μ˜λ£Œμ§„ 신뒰도 평가λ₯Ό ν¬ν•¨ν•˜μ—¬ λͺ¨λΈμ˜ 예츑 μ„±λŠ₯뿐 μ•„λ‹ˆλΌ μ„€λͺ…κ°€λŠ₯μ„±κ³Ό μž„μƒμ  ν™œμš© κ°€λŠ₯μ„±κΉŒμ§€ μ’…ν•©μ μœΌλ‘œ ν‰κ°€ν•˜μ˜€μŠ΅λ‹ˆλ‹€.

🎯 What problem did I solve?

1. Background

Early ICU mortality prediction is essential for timely clinical decision-making, but conventional machine learning models require institution-specific development and often provide limited interpretability. Although recent large language models (LLMs) have shown promising reasoning capabilities, their ability to perform reliable clinical prediction from structured electronic health records remains largely unexplored.

❓ Research Questions

🎯 Objective

To evaluate the predictive performance, generalizability, and clinical trustworthiness of general-purpose LLMs for early ICU mortality prediction using structured EHR data across multinational ICU cohorts.

🧠 How did I solve it?

2. Dataset

3. Methods

260714_figure_1_overall_framework.svg

Data Processing

Model Comparison

Prompt Engineering

Explainability

πŸ“Š What were the results?

4. Results

Table 1. Main performance

Table 1. Main performance

Figure 2. Alignment between XGBoost SHAP ↔ GPT-5 Reasoning txt

Figure 2. Alignment between XGBoost SHAP ↔ GPT-5 Reasoning txt

5. Discussion & Limitation

Discussion

Limitations

πŸ‘©β€πŸ’» What was my contribution?

πŸ’‘ What did I learn?

Technical Insight

μ •ν™•ν•œ λͺ¨λΈμ„ κ΅¬μΆ•ν•˜λŠ” 것은 ν•˜λ‚˜μ˜ 단계에 λΆˆκ³Όν•¨μ„ 체감할 수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€. μž¬ν˜„κ°€λŠ₯ν•œ 평가 νŒŒμ΄ν”„λΌμΈμ„ μ„€κ³„ν•˜κ³  μ„€λͺ…κ°€λŠ₯성을 톡해 λͺ¨λΈμ˜ λ™μž‘μ„ μ΄ν•΄ν•˜λŠ” 것 λ˜ν•œ μ‹ λ’°ν•  수 μžˆλŠ” AI μ‹œμŠ€ν…œμ„ κ°œλ°œν•˜λŠ” 데 μžˆμ–΄μ„œ 맀우 μ€‘μš”ν•˜λ‹€λŠ” 것을 체가할 수 μžˆλŠ” ν”„λ‘œμ νŠΈμ˜€μŠ΅λ‹ˆλ‹€.
Building an accurate model is only the first step. Designing reproducible evaluation pipelines and understanding model behavior through explainability are equally important for developing reliable AI systems.

Research Insight

연ꡬλ₯Ό λ°”λΌλ³΄λŠ” 관점을 β€œAIκ°€ μ •ν™•ν•œ μ˜ˆμΈ‘μ„ ν•  수 μžˆλŠ”κ°€?β€μ—μ„œ β€œμž„μƒμ˜κ°€ AIκ°€ μƒμ„±ν•œ 결정을 μ‹ λ’°ν•  수 μžˆλŠ”κ°€?β€λ‘œ μ „ν™˜ν•˜κ²Œ λ˜μ—ˆμŠ΅λ‹ˆλ‹€. μ–΄λ–€ 도메인이든, λ„λ©”μΈμ˜ μ‹€λ¬΄μžκ°€ μ‚¬μš©ν•˜λŠ” AIλŠ” μ„±λŠ₯ μ§€ν‘œ 뿐만 μ•„λ‹ˆλΌ 투λͺ…μ„±, μΌλ°˜ν™” κ°€λŠ₯μ„±, 그리고 μ‹€μš©μ μΈ μ‚¬μš©μ„±μœΌλ‘œλ„ ν‰κ°€λ˜μ–΄μ•Ό ν•œλ‹€λŠ” 것을 κΉ¨λ‹¬μ•˜μŠ΅λ‹ˆλ‹€.
This project shifted my research perspective from "Can AI make accurate predictions?" to "Can clinicians trust AI-generated decisions?". I realized that future clinical AI should be evaluated not only by performance metrics but also by transparency, generalizability, and practical usability.