
Thumnail. Overall Framework
💡
Summary

| Duration | 2025.10 - 2026.07 |
|---|---|
| Role | First Author |
| Research Area | Computational Pathology |
| Keywords | MIL · SSL · Digital Pathology · Generalization |
| Status | Accepted, M.S. Thesis |
What problem did I solve?
1. Background

Oncotype DX is expensive and not easily accessible, making treatment decisions difficult in many clinical settings.

❓Research Questions
- Which is more important: Feature Extractor or MIL?
- Does SSL consistently improve performance?
- Which architecture generalizes best across institutions?
🎯 Objective
Systematically evaluate SSL and MIL combinations for robust and generalizable WSI-based Oncotype DX prediction.
How did I solve it?
2. Dataset

| Cohort | Purpose |
|---|---|
| Internal | Training/Test |
| External | External Test |
| BCRNet | External Test, Cross-domain |
| TCGA | External Test, Cross-resource |
3. Method
Pipeline: Whole Slide Image → Feature Extraction → MIL → Prediction



-
Experimental Design
markdown 12 Feature Extractors × 14 MIL Models = 168 Experiments -
Evaluation strategy
markdown 3-fold Cross Validation ↓ 3 Random Seeds ↓ 9 Ensemble Models ↓ 4 Independent Test Cohorts
What were the results?
4. Results
- Finding 1
: Feature Extractor selection had a larger impact than MIL architecture under external validation. - Finding 2
: Attention-based MIL improved cross-institution generalization. - Finding 3
: SSL did not consistently outperform supervised models across all cohorts. - Finding 4
: Low-risk class remained difficult in 3-class prediction.





5. Discussion & Limitation


What was my contribution?
✅ Literature Review
✅ Dataset Curation
✅ WSI Preprocessing
✅ SSL Feature Extraction
✅ MIL Benchmark
✅ Experimental Design
✅ Statistical Analysis
✅ Visualization
✅ Thesis Writing
What did I learn?
Technical Insight
- Model architecture alone does not guarantee generalization.
- Representation learning plays a larger role than expected.
Research Insight
- External validation changes the ranking of models.
- Evaluation protocol is as important as model development.
Future work
- Pathology Foundation Models
- Multimodal Learning
- Better Multiclass Learning
- Clinical Integration
Appendix.
-
Attention Visualization

-
Related works


-
Feature Extractor - Supplementary



-
Multiple Instance Learning - Supplementary

-
Preprocessing: Patch Extraction


-
Metric - AUROC Range, Silhouette Coefficient


-
Information Utilization Gap
