iFAN: Inference-Aware Learning for Plain Mask Transformers
Abstract
A training framework called iFAN improves mask transformers by aligning query ranking with mask quality and distilling stronger intermediate predictions to the final layer.
Query-based mask transformers assemble segmentation outputs through pixel-wise competition among query predictions of the final layer, yet this inference process is not explicitly optimized during training. We identify two key mismatches: the query with the highest probability-mask score does not necessarily produce the most accurate mask, and final-layer decoding may discard superior predictions from intermediate layers. To address these issues, we propose Inference-Aware Learning (iFAN), a general training framework for plain mask transformers. iFAN introduces Adjusted Probability-Mask Ranking (APMR), which aligns query competition with predicted mask quality and suppresses high-confidence but inaccurate competitors. We further employ Cross-Layer Self-Distillation (CLSD) to transfer stronger intermediate predictions to the final layer. The ranking and distillation objectives are training-only, while inference retains efficient final-layer decoding. Experiments on COCO, ADE20K, and Cityscapes demonstrate consistent improvements across panoptic, instance, and semantic segmentation, as well as across different architectures, backbone scales, and input resolutions. Overall, iFAN improves performance by an average of 1.20 PQ, 1.30 AP, and 0.63 mIoU, with negligible additional parameters, FLOPs and inference latency.
Community
iFAN: Inference-Aware Learning for Plain Mask Transformers
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Soft Mixture-of-Recursions: Going Deeper with Recursive Vision Transformers (2026)
- Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection (2026)
- Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation (2026)
- TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion (2026)
- Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception (2026)
- LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter (2026)
- Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.03216 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper