Awesome Referring Image Segmentation
:books: A collection of papers about Referring Image Segmentation.
Install / Use
npx skills add MarkMoHR/Awesome-Referring-Image-SegmentationInstalls into whichever agent you are using.
README
Awesome-Referring-Image-Segmentation
A collection of referring image segmentation papers and datasets.
Feel free to create a PR or an issue.

Outline
1. Datasets
| Short name | Paper | Source | Code/Project Link | | --- | --- | --- | --- | | MeViS | MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions | ICCV 2023 | [dataset] [project] | | gRefCOCO | GRES: Generalized Referring Expression Segmentation | CVPR 2023 | [dataset] [project] | | ClevrTex | ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation | NeurIPS Datasets and Benchmarks 2021 | [project] | | ScanRefer | ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language | ECCV 2020 | [project] | | VGPhraseCut | PhraseCut: Language-based Image Segmentation in the Wild | CVPR 2020 | [project] | | CLEVR-Ref+ | CLEVR-Ref+: Diagnosing Visual Reasoning with Referring Expressions | CVPR 2019 | [project] | | UNC | Modeling context in referring expressions | ECCV 2016 | [dataset] | | UNC+ | Modeling context in referring expressions | ECCV 2016 | [dataset] | | Google-Ref | Generation and comprehension of unambiguous object descriptions | CVPR 2016 | [dataset] | | ReferIt | Referit game: Referring to objects in photographs of natural scenes | EMNLP 2014 | [project] |
2. Challenges
| Name | Workshop | Date | Submission Link | | --- | --- | --- | --- | | 1st MeViS Challenge | CVPR 2024 Workshop: Pixel-level Video Understanding in the Wild | May 2024| [CodaLab] | | RVOS Challenge | ECCV 2024 Workshop: The 6th Large-scale Video Object Segmentation Challenge | Aug 2024| [CodaLab] |
3. Traditional Referring Image Segmentation
| Short name | Paper | Source | Code/Project Link | | --- | --- | --- | --- | | | Improving Target Presence and Plurality Recognition for Generalized Referring Image Segmentation | AAAI 2026 | [webpage] | | PixelRefer | PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity | arxiv 25.10 | [code] | | CoPatch | CoPatch: Zero-Shot Referring Image Segmentation by Leveraging Untapped Spatial Knowledge in CLIP | arxiv 25.09 | [code] | | SaFiRe | SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation | NeurIPS 2025 | | | UniPixel | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning | NeurIPS 2025 | [code] [webpage] | | RaAM | Region-aware Anchoring Mechanism for Efficient Referring Visual Grounding | ICCV 2025 | | | Latent-VG | Latent Expression Generation for Referring Image Segmentation and Grounding | ICCV 2025 | | | DeRIS | DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy | ICCV 2025 | [code] | | WeakMCN | WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation | CVPR 2025 | [code] | | HybridGL | Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation | CVPR 2025 | [code] | | IteRPrimE | IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis | AAAI 2025 | [code] | | DETRIS | Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation | AAAI 2025 | [code] | | VATEX | Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding | WACV 2025 | [code] [webpage] | | Shared-RIS | A Simple Baseline with Single-encoder for Referring Image Segmentation | arxiv 24.08 | [code] | | ASDA | Adaptive Selection based Referring Image Segmentation | ACM MM 2024 | code | | NeMo | Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation | ECCV 2024 | [webpage] [code] | | ReMamber | ReMamber: Referring Image Segmentation with Mamba Twister | ECCV 2024 | [code] | | GTMS | GTMS: A Gradient-driven Tree-guided Mask-free Referring Image Segmentation Method | ECCV 2024 | [code] | | SAM4MLLM | SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation | ECCV 2024 | [code] | | Pseudo-RIS | Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image Segmentation | ECCV 2024 | [code] | | SafaRi | SafaRi: Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation | ECCV 2024 | [webpage] | | CM-MaskSD | CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation | TMM 2024 | | | Prompt-RIS | Prompt-Driven Referring Image Segmentation with Instance Contrasting | CVPR 2024 | | | LQMFormer | LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation | CVPR 2024 | | | PPT | Curriculum Point Prompting for Weakly-Supervised Referring Image Segmentation | CVPR 2024 | | | GSVA | GSVA: Generalized Segmentation via Multimodal Large Language Models | CVPR 2024 | [[code]](
Related Skills
node-connect
385.5kDiagnose OpenClaw Android, iOS, or macOS node pairing, QR/setup code, route, auth, and connection failures.
blender-python-addon
40.5kBlender Python add-on rules for operators, panels, properties, registration, testing, and API-safe scripting
flutter-development-guidelines-cursorrules-prompt-file
40.5kCursor rules for Flutter development with MVVM architecture, Riverpod state management, Material widgets, and Dart style guidelines.
commit-push-pr
140.7kCommit, push, and open a PR
Security Score
Audited on Jul 7, 2026
