
Thumnail. Overall Framework
π‘
Summary
| Duration | 2025.12 - 2026.02 |
|---|---|
| Role | Co-First Author |
| Research Area | Computational Pathology |
| Keywords | VLM Β· MIL Β· Semantic Alignment Β· Digital Pathology |
| Status | Reject (MICCAI 2026 Submitted, Reviewed) |
π― What problem did I solve?
1. Background
Attention-based Multiple Instance Learning (MIL)μ Whole Slide Image λΆμμμ λ리 μ¬μ©λμ§λ§, attentionμ΄ λ³λ¦¬νμ μΌλ‘ μλ―Έ μλ μμμ΄ μλλΌ ν΅κ³μ μΌλ‘λ§ μμΈ‘λ ₯μ΄ λμ μμμ μ§μ€νλ λ¬Έμ κ° μμ΅λλ€. μ΄λ¬ν attentionμ λΆμμ μ±μ λ³μ κ° λ°μ΄ν° λΆν¬κ° λ¬λΌμ§ λ μΌλ°ν μ±λ₯ μ νλ‘ μ΄μ΄μ§ μ μμ΅λλ€.
β Research Questions
- Can semantic knowledge guide MIL attention toward clinically meaningful regions?
- Can semantic alignment improve cross-institution generalization?
- Which level of semantic alignment contributes most to robust MIL learning?
π― Objective
To develop a backbone-agnostic semantic regularization framework (CASA) that aligns MIL attention with pathology concepts derived from vision-language models while preserving the original MIL architecture.
π§ How did I solve it?
2. Dataset
- Internal Yonsei Cohort (489 WSIs)
- Institutional External Cohort
- BCR-Net External Cohort
μ΄ 3κ°μ λ 립 μ½νΈνΈλ₯Ό νμ©νμ¬ κΈ°κ΄ κ° μΌλ°ν μ±λ₯μ νκ°νμ΅λλ€.
3. Methods

Concept-guided Semantic Alignment
Vision-Language Modelμμ μμ±ν λ³λ¦¬ν κ°λ
μλ² λ©μ νμ©νμ¬ attentionμ μ§μ regularizationνλ νλ μμν¬μ
λλ€.
κΈ°μ‘΄ MIL backboneμ λ³κ²½νμ§ μκ³ λ€μ μΈ κ°μ§ μμ€μμ semantic alignmentλ₯Ό μννμ΅λλ€.
π‘
πͺ Method highlights
β Backbone-agnostic Framework
β Vision-Language Model Integration
β Concept-guided Attention Learning
β Multi-level Semantic Alignment
β Cross-institution Evaluation
β Component Ablation Study
Semantic Concept Modeling
To provide clinically meaningful supervision, pathology concepts were defined based on the WHO breast tumor classification and encoded into a shared imageβtext embedding space using a pretrained Vision-Language Model.
Each image patch was compared with these concept embeddings to estimate its semantic relevance, creating a concept-guided reference for attention learning.
π Multi-level Semantic Alignment
CASA introduces semantic constraints at three complementary levels.
- Distribution-level Alignment: Aligns the attention distribution with concept-derived semantic relevance, encouraging the model to focus on pathology-related regions instead of statistically correlated artifacts.
- Feature-level Alignment: Regularizes patch-level feature representations toward concept embeddings within the VLM semantic space, improving semantic consistency across instances.
- Global-level Alignment: Constrains the slide-level representation to preserve overall semantic coherence while maintaining compatibility with different MIL backbones.
Experimental Design
- 6 Attention-based MIL Backbones
- CLAM
- DSMIL
- ABMIL
- TransMIL
- DTFD-MIL
- CLAM-MB
Component Analysis
λ€μ μμλ€μ μν₯μ λ 립μ μΌλ‘ λΆμνμ΅λλ€.
- Warm-up Strategy
- Top-K Selection
- Feature Alignment
- Global Alignment
- Concept Tier Design
π What were the results?
- Finding 1: Semantic alignment generally improved robustness under cross-institution evaluation while preserving compatibility with existing MIL architectures.
- Finding 2: Feature-level semantic alignment (Lfeat) consistently contributed the largest performance gains across multiple MIL backbones.
- Finding 3: Global semantic alignment improved in-domain performance but required conservative tuning under severe domain shift.
- Finding 4: Different concept hierarchies exhibited backbone-dependent behavior, suggesting that semantic concept design should be optimized for each application rather than universally applied.

Discussion
- Semantic information from VLMs can serve as an effective inductive bias for attention-based MIL.
- Architecture modification is not necessary to incorporate semantic guidance.
- Multi-level alignment provides a flexible framework applicable to various MIL backbones.
Limitations
- Performance improvements remained modest across several backbones.
- Effectiveness depended on backbone architecture and domain characteristics.
- Concept quality and VLM representation significantly influenced alignment performance.
- Further validation on larger multi-institutional cohorts is required.
π©βπ» What was my contribution?
- Research idea formulation
- Literature review on MIL and Vision-Language Models
- CASA framework design
- Semantic concept engineering
- Experimental design
- Component ablation studies
- Cross-institution evaluation
- Manuscript writing
π‘ What did I learn?
Technical Insight
λͺ¨λΈ μ±λ₯ ν₯μμ λ¨μν μλ‘μ΄ μν€ν
μ²λ₯Ό μ€κ³νλ κ² λΏλ§ μλλΌ, μλ―Έμλ βκ·λ©μ νΈν₯βμ λμ
νλ κ²λ μ€μνλ€λ κ²μ 체κ°ν μ μμμ΅λλ€. VLMμ Semantic μ 보λ₯Ό μν€ν
μ²μ μμ μμ΄ κΈ°μ‘΄ MIL Frameworkμ ν΅ν©νλ λ°©λ²κ³Ό, μ΄λ€ κ΅¬μ± μμκ° μΌλ°νμ μ€μ λ‘ κΈ°μ¬νλμ§ μ΄ν΄νκΈ° μν΄ μ μ€ν ablation studyκ° νμμ μ΄λΌλ κ² λν 체λν μ μμμ΅λλ€.
Developing CASA taught me that improving model performance is not only about designing new architectures but also about introducing meaningful inductive biases. I learned how semantic information from vision-language models can be incorporated into existing MIL frameworks without architectural modifications, and how careful ablation studies are essential for understanding which components truly contribute to generalization.
Research Insight
Conferneceμ λ
Όλ¬Έμ μ μΆνκ³ , μ¬μ¬μμλ€μ μ μμ νΌλλ°±μ λ°μΌλ©΄μ βλ
μ°½μ±βκ³Ό βμ€μ¦μ κ·Όκ±°μ κ· νβμ μ€μμ±μ λ€μ ν λ² κΉ¨λ¬μμ΅λλ€. μ μλ νλ μμν¬λ ν΄μ κ°λ₯ν μλ―Έ μ λ ¬ μ λ΅μ λμ
νμ§λ§, λ€μν νκ²½μμ μΌκ΄λ μ±λ₯ ν₯μμ 보μ¬μ£Όκ³ , κ° μ€κ³ μ νμ μ€μ§μ μΈ μλ―Έλ₯Ό λͺ
ννκ² μ€λͺ
νλ κ²μ΄ μλ‘μ΄ λ°©λ²μ μ μνλ κ²λ§νΌ μ€μνλ€λ κ²μ μκ² λμμ΅λλ€.
λΉλ‘ λ―Έμ±νλ νλ‘μ νΈμμΌλ, μ μμ μ¬μ¬μμ λͺ¨λμ κ΄μ μμ μ°κ΅¬λ₯Ό λΉνμ μΌλ‘ μ€μ€λ‘ νκ°νλ κ΄μ μ 체λν μ μμμ΅λλ€.
Submitting this work to MICCAI and receiving reviewer feedback reinforced the importance of balancing novelty with empirical evidence. While the proposed framework introduced an interpretable semantic alignment strategy, I realized that demonstrating consistent improvements across diverse settings and clearly articulating the practical significance of each design choice are just as important as proposing a new method. This experience strengthened my ability to critically evaluate research from both the author's and reviewer's perspectives.
Appendix.
- Reviewer Feedback - Overall Score: 2 / 4 / 2
Major Comments
- Limited performance improvement across some backbones
- Need for stronger empirical validation
- Practical impact should be demonstrated more clearly
β Receiving reviewer feedback helped me recognize the gap between a technically interesting idea and a publication-ready contribution. Rather than viewing the rejection as a failure, I used the reviews to identify weaknesses in experimental validation and the clarity of the paper's narrative. This experience improved both my research design and scientific writing skills, and it continues to shape the next iteration of this work.