← Back to portfolio
Conference Β· NeurIPS 2026

Are Compact Rationales Free? Measuring Tile Selection Headroom in Frozen WSI-MIL

Co-AuthorFoundation Model Β· MILExplainable AIUnder Review

Thumnail. FOCI Architecture

Thumnail. FOCI Architecture

πŸ’‘

Summary: Link - arxiv

Duration 2026.02 - 2026.05
Role Co Author
Research Area Computational Pathology Β· Explainable AI
Keywords WSI Β· MIL Β· Interpretability Β· Foundation Model
Status NeurIPS 2026 Submission

🎯 What problem did we solve?

1. Background

Attention scores are widely used as explanations in Whole Slide Image (WSI) Multiple Instance Learning (MIL). However, high attention does not necessarily indicate that a small subset of image patches is sufficient to reproduce the model's prediction. Existing attention-based explanations therefore provide limited evidence for whether a prediction can truly be supported by a compact and faithful rationale.

❓ Research Questions

🎯 Objective

To develop a lightweight post-hoc rationale readout framework that identifies compact, output-consistent tile subsets from frozen WSI-MIL models and systematically measures how much explanation headroom exists across different MIL architectures.

🧠 How did we solve it?

2. Dataset

Three public pathology benchmarks were used to evaluate rationale compactness across diverse tissue types and prediction tasks.

3. Method

Fig 1. Overall Framework

Fig 1. Overall Framework

Proposed Framework
We proposed FOCI (Finding Optimal Contextual Instances), a lightweight rationale-readout module that operates on top of frozen MIL backbones without modifying their original inference pipeline. Instead of retraining the classifier, FOCI learns to identify compact tile subsets that preserve the original model prediction.

FOCI Architecture

Figure 2. FOCI Architecture

Figure 2. FOCI Architecture

Key Components

Evaluation Strategy

Rather than evaluating only slide-level classification performance, the study introduced new rationale evaluation metrics:

These metrics quantify how efficiently a model's prediction can be reconstructed from progressively revealed image tiles.

πŸ“Š What were the results?

4. Results

Table 1. Main Performance

Table 1. Main Performance

Figure 3. Main performance

Figure 3. Main performance

5. Discussion & Limitation

Discussion

Limitation

πŸ‘©β€πŸ’» What was my contribution?

Co-Author

πŸ’‘ What did I learn?

Technical Insight

이 연ꡬ에 κ³΅μ €μžλ‘œ μ°Έμ—¬ν•˜λ©΄μ„œ, β€œν•΄μ„κ°€λŠ₯성”에 λŒ€ν•œ 이해도가 ν™•μž₯λ˜μ—ˆμŠ΅λ‹ˆλ‹€. μ„€λͺ…μ˜ qualityλŠ” λΆ„λ₯˜ μ„±λŠ₯κ³ΌλŠ” λ…λ¦½μ μœΌλ‘œ ν‰κ°€λ˜μ–΄μ•Ό ν•˜κ³ , μƒˆλ‘œμš΄ 평가 ν”„λ‘œν† μ½œμ„ 톡해 κΈ°μ‘΄ μ§€ν‘œλ₯Ό λ„˜μ–΄ λͺ¨λΈμ΄ λ™μž‘ν•˜λŠ”λ° μžˆμ–΄ μˆ¨κ²¨μ§„ component듀을 λ°ν˜€λ‚Ό 수 μžˆλ‹€λŠ” 것을 μ•Œκ²Œ λ˜μ—ˆμŠ΅λ‹ˆλ‹€.
This project broadened my understanding of post-hoc interpretability in computational pathology. I learned that explanation quality should be evaluated independently from classification performance and that new evaluation protocols can reveal hidden properties of model behavior beyond conventional metrics.

Research Insight

κ³΅μ €μžλ‘œ 연ꡬ에 μ°Έμ—¬ν•˜λ©΄μ„œ ν˜‘μ—… 연ꡬλ₯Ό λ‹€λ₯Έ κ΄€μ μ—μ„œ κ²½ν—˜ν•  수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€. κ°œλ³„μ μΈ 기술적 κΈ°μ—¬κ°€ μ–΄λ–»κ²Œ 더 λ²”μœ„κ°€ 큰 ν”„λ‘œμ νŠΈμ— ν†΅ν•©λ˜λŠ”μ§€, 그리고 μ—„κ²©ν•œ μ‹€ν—˜ 섀계, μ‹¬μ‚¬μœ„μ› μ€‘μ‹¬μ˜ 평가, λͺ…ν™•ν•œ 과학적 μŠ€ν† λ¦¬ν‹Έλ μ΄ 영ν–₯λ ₯ μžˆλŠ” AI 연ꡬ λ…Όλ¬Έ 게재λ₯Ό μœ„ν•΄ μ–Όλ§ˆλ‚˜ μ€‘μš”ν•œμ§€ 체감할 수 μžˆμ—ˆμŠ΅λ‹ˆλ‹€.
Participating as a co-author allowed me to experience collaborative research from a different perspective. I learned how individual technical contributions integrate into a larger research project and how rigorous experimental design, reviewer-oriented evaluation, and clear scientific storytelling are essential for publishing high-impact AI research.