← Back to portfolio
Conference Β· MICCAI 2026

CASA: Concept-Attention Semantic Alignment Framework

Co-First AuthorVLM Β· MILComputational PathologyReject (Reviewed)

Thumnail. Overall Framework

Thumnail. Overall Framework

πŸ’‘

Summary

Duration 2025.12 - 2026.02
Role Co-First Author
Research Area Computational Pathology
Keywords VLM Β· MIL Β· Semantic Alignment Β· Digital Pathology
Status Reject (MICCAI 2026 Submitted, Reviewed)

🎯 What problem did I solve?

1. Background

Attention-based Multiple Instance Learning (MIL)은 Whole Slide Image λΆ„μ„μ—μ„œ 널리 μ‚¬μš©λ˜μ§€λ§Œ, attention이 λ³‘λ¦¬ν•™μ μœΌλ‘œ 의미 μžˆλŠ” μ˜μ—­μ΄ μ•„λ‹ˆλΌ ν†΅κ³„μ μœΌλ‘œλ§Œ 예츑λ ₯이 높은 μ˜μ—­μ— μ§‘μ€‘ν•˜λŠ” λ¬Έμ œκ°€ μžˆμŠ΅λ‹ˆλ‹€. μ΄λŸ¬ν•œ attention의 λΆˆμ•ˆμ •μ„±μ€ 병원 κ°„ 데이터 뢄포가 λ‹¬λΌμ§ˆ λ•Œ μΌλ°˜ν™” μ„±λŠ₯ μ €ν•˜λ‘œ μ΄μ–΄μ§ˆ 수 μžˆμŠ΅λ‹ˆλ‹€.

❓ Research Questions

🎯 Objective

To develop a backbone-agnostic semantic regularization framework (CASA) that aligns MIL attention with pathology concepts derived from vision-language models while preserving the original MIL architecture.

🧠 How did I solve it?

2. Dataset

총 3개의 독립 μ½”ν˜ΈνŠΈλ₯Ό ν™œμš©ν•˜μ—¬ κΈ°κ΄€ κ°„ μΌλ°˜ν™” μ„±λŠ₯을 ν‰κ°€ν–ˆμŠ΅λ‹ˆλ‹€.

3. Methods

fig_1_overview.png

Concept-guided Semantic Alignment

Vision-Language Modelμ—μ„œ μƒμ„±ν•œ 병리학 κ°œλ… μž„λ² λ”©μ„ ν™œμš©ν•˜μ—¬ attention을 직접 regularizationν•˜λŠ” ν”„λ ˆμž„μ›Œν¬μž…λ‹ˆλ‹€.
κΈ°μ‘΄ MIL backbone은 λ³€κ²½ν•˜μ§€ μ•Šκ³  λ‹€μŒ μ„Έ κ°€μ§€ μˆ˜μ€€μ—μ„œ semantic alignmentλ₯Ό μˆ˜ν–‰ν–ˆμŠ΅λ‹ˆλ‹€.


πŸ’‘

πŸͺ„ Method highlights

βœ” Backbone-agnostic Framework
βœ” Vision-Language Model Integration
βœ” Concept-guided Attention Learning
βœ” Multi-level Semantic Alignment
βœ” Cross-institution Evaluation
βœ” Component Ablation Study

Semantic Concept Modeling

To provide clinically meaningful supervision, pathology concepts were defined based on the WHO breast tumor classification and encoded into a shared image–text embedding space using a pretrained Vision-Language Model.

Each image patch was compared with these concept embeddings to estimate its semantic relevance, creating a concept-guided reference for attention learning.

πŸ“Œ Multi-level Semantic Alignment
CASA introduces semantic constraints at three complementary levels.


Experimental Design

Component Analysis

λ‹€μŒ μš”μ†Œλ“€μ˜ 영ν–₯을 λ…λ¦½μ μœΌλ‘œ λΆ„μ„ν–ˆμŠ΅λ‹ˆλ‹€.

πŸ“Š What were the results?

fig_2_topk.png

Discussion

Limitations

πŸ‘©β€πŸ’» What was my contribution?

πŸ’‘ What did I learn?

Technical Insight

λͺ¨λΈ μ„±λŠ₯ ν–₯상은 λ‹¨μˆœνžˆ μƒˆλ‘œμš΄ μ•„ν‚€ν…μ²˜λ₯Ό μ„€κ³„ν•˜λŠ” 것 뿐만 μ•„λ‹ˆλΌ, μ˜λ―ΈμžˆλŠ” β€œκ·€λ‚©μ  편ν–₯”을 λ„μž…ν•˜λŠ” 것도 μ€‘μš”ν•˜λ‹€λŠ” 것을 체감할 수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€. VLM의 Semantic 정보λ₯Ό μ•„ν‚€ν…μ²˜μ˜ μˆ˜μ •μ—†μ΄ κΈ°μ‘΄ MIL Framework에 ν†΅ν•©ν•˜λŠ” 방법과, μ–΄λ–€ ꡬ성 μš”μ†Œκ°€ μΌλ°˜ν™”μ— μ‹€μ œλ‘œ κΈ°μ—¬ν•˜λŠ”μ§€ μ΄ν•΄ν•˜κΈ° μœ„ν•΄ μ‹ μ€‘ν•œ ablation studyκ°€ ν•„μˆ˜μ μ΄λΌλŠ” 것 λ˜ν•œ 체득할 수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€.
Developing CASA taught me that improving model performance is not only about designing new architectures but also about introducing meaningful inductive biases. I learned how semantic information from vision-language models can be incorporated into existing MIL frameworks without architectural modifications, and how careful ablation studies are essential for understanding which components truly contribute to generalization.

Research Insight

Confernece에 논문을 μ œμΆœν•˜κ³ , μ‹¬μ‚¬μœ„μ›λ“€μ˜ μ μˆ˜μ™€ ν”Όλ“œλ°±μ„ λ°›μœΌλ©΄μ„œ β€œλ…μ°½μ„±β€κ³Ό β€œμ‹€μ¦μ  근거의 κ· ν˜•β€μ˜ μ€‘μš”μ„±μ„ λ‹€μ‹œ ν•œ 번 κΉ¨λ‹¬μ•˜μŠ΅λ‹ˆλ‹€. μ œμ•ˆλœ ν”„λ ˆμž„μ›Œν¬λŠ” 해석 κ°€λŠ₯ν•œ 의미 μ •λ ¬ μ „λž΅μ„ λ„μž…ν–ˆμ§€λ§Œ, λ‹€μ–‘ν•œ ν™˜κ²½μ—μ„œ μΌκ΄€λœ μ„±λŠ₯ ν–₯상을 보여주고, 각 섀계 μ„ νƒμ˜ μ‹€μ§ˆμ μΈ 의미λ₯Ό λͺ…ν™•ν•˜κ²Œ μ„€λͺ…ν•˜λŠ” 것이 μƒˆλ‘œμš΄ 방법을 μ œμ•ˆν•˜λŠ” κ²ƒλ§ŒνΌ μ€‘μš”ν•˜λ‹€λŠ” 것을 μ•Œκ²Œ λ˜μ—ˆμŠ΅λ‹ˆλ‹€.
비둝 λ―Έμ±„νƒλœ ν”„λ‘œμ νŠΈμ˜€μœΌλ‚˜, μ €μžμ™€ μ‹¬μ‚¬μœ„μ› λͺ¨λ‘μ˜ κ΄€μ μ—μ„œ 연ꡬλ₯Ό λΉ„νŒμ μœΌλ‘œ 슀슀둜 ν‰κ°€ν•˜λŠ” 관점을 체득할 수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€.
Submitting this work to MICCAI and receiving reviewer feedback reinforced the importance of balancing novelty with empirical evidence. While the proposed framework introduced an interpretable semantic alignment strategy, I realized that demonstrating consistent improvements across diverse settings and clearly articulating the practical significance of each design choice are just as important as proposing a new method. This experience strengthened my ability to critically evaluate research from both the author's and reviewer's perspectives.

Appendix.

Major Comments

β†’ Receiving reviewer feedback helped me recognize the gap between a technically interesting idea and a publication-ready contribution. Rather than viewing the rejection as a failure, I used the reviews to identify weaknesses in experimental validation and the clarity of the paper's narrative. This experience improved both my research design and scientific writing skills, and it continues to shape the next iteration of this work.