QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

This repository hosts the official model weights and evaluation benchmarks for QueenVIS.

🔗 Quick Links


📖 Model Description

QueenVIS rethinks the paradigm of image-only training for Video Instance Segmentation (VIS). By introducing training-only auxiliary heads for feature prediction and center prediction, QueenVIS embeds dense appearance and spatial priors directly into transformer object queries without requiring video clip supervision.

  • Key Feature: Adds zero parameters and zero extra GFLOPs at inference relative to image-only baselines like MinVIS.
  • Temporal Association: Employs a training-free confidence-guided query propagation mechanism paired with a non-parametric memory bank to maintain temporal identity across frames.

📊 Datasets & Evaluation Metrics

QueenVIS is evaluated across standard VIS benchmarks:

  • YouTube-VIS (2019 / 2021 / 2022 Long Split): Standard and long-sequence video benchmarks.
  • OVIS (Occluded Video Instance Segmentation): A challenging benchmark characterized by severe occlusion and complex object interactions.

Performance is measured using standard Video Instance Segmentation metrics:

  • AP (Average Precision): Primary metric across IoU thresholds.
  • AP₅₀ & AP₇₅: Average Precision at IoU thresholds of 0.50 and 0.75.
  • AR₁ & AR₁₀: Average Recall given 1 and 10 detections per video.

📈 Evaluation Results

1. Main Benchmarks (ResNet-50 Backbone)

Dataset AP AP₅₀ AP₇₅ AR₁ AR₁₀
YouTube-VIS 2019 51.8 75.2 58.3 49.4 62.1
YouTube-VIS 2021 50.9 72.1 55.8 42.8 57.3
YouTube-VIS 2022 (Long) 33.6 56.6 34.0 31.7 39.9
OVIS (Occluded VIS) 29.8 52.7 29.0 15.4 34.6

2. Scaling to Vision Transformers (Swin-L Backbone)

Dataset AP AP₅₀ AP₇₅ AR₁ AR₁₀
YouTube-VIS 2019 63.2 85.3 68.3 55.5 68.4
YouTube-VIS 2021 59.8 81.5 65.2 48.3 64.2
OVIS (Occluded VIS) 41.0 62.2 43.0 18.4 45.0

🚀 How to Use

For complete instructions on installation, environment setup, dataset preparation, and inference scripts, visit the official GitHub Repository.


📚 Citation

If you find QueenVIS useful in your research, please cite our work:

@misc{kheirandish2026queenvis,
      title={QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment}, 
      author={Arian Kheirandish and Fardin Ayar and Ehsan Javanmardi and Manabu Tsukada and Mahdi Javanmardi},
      year={2026},
      eprint={2607.24598},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={[https://doi.org/10.48550/arXiv.2607.24598](https://doi.org/10.48550/arXiv.2607.24598)}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ArianKheir/QueenVIS

Finetuned
(1)
this model

Space using ArianKheir/QueenVIS 1

Paper for ArianKheir/QueenVIS

Free AI Image Generator No sign-up. Instant results. Open Now