Awesome Unimodal Training
text-only training or language-free training for multimodal tasks (image/audio/video caption, retrieval, text2image)
Install / Use
npx skills add iOPENCap/awesome-unimodal-trainingInstalls into whichever agent you are using.
README
Awesome-Text-only-Training
The project is used to store text-only training, image-free training for multimodal tasks related papers.
Include text-only training for:
- zero-shot image/audio/video captioning
- zero-shot composed image retrieval
- visual storytelling, visual question answer...
papers
<br/>> 2024
-
[AAAI] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training [paper] [code][⭐8]<br/>
-
[ACM] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] TOMGPT: Reliable Text-Only Training Approach for Cost-Effective Multi-modal Large Language Model[paper]<br/>
-
[IJCV] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 5] Learning to Prompt with Text Only Supervision for Vision-Language Models[paper] [code][⭐80]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Text Data-Centric Image Captioning with Interactive Prompts[paper]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 2] MeaCap: Memory-Augmented Zero-shot Image Captioning[paper] [code][⭐27]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks [paper]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion [paper]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning [paper] [code][⭐4]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] From Unimodal to Multimodal: Scaling up Projectors to Align Modalities [paper]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning [paper] [code][⭐0]<br/>
-
[arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] DRCap: Decoding CLAP Latents with Retrieval-augmented Generation for Zero-shot Audio Captioning [paper] <br/>
<br/>> 2023
-
[arxiv] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] Improved Factorized Neural Transducer Model For text-only Domain Adaptation[paper] <br/>
-
[AAAI] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 2] Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning[paper] <br/>
-
[ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Text-Only Training for Visual Storytelling[paper] <br/>
-
[ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] VLIS: Unimodal Language Models Guide Multimodal Language Generation[paper] [code][⭐24] <br/>
-
[NeurlPS] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 7] LOVM:Language-Only Vision Model Selection[paper] [code][⭐18]<br/>
-
[IJCAI] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 4] From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping[paper] [code][⭐11]<br/>
-
[ICLR] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 47] Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training[paper] [code][⭐118]<br/>
-
[DCASE] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 4] Weakly-supervised Automated Audio Captioning via text only training[paper] [code][⭐1]<br/>
-
[ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 3] CgT-GAN: CLIP-guided Text GAN for Image Captioning[paper] [code][⭐16]<br/>
<br/>> 2022
-
[EMNLP] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 64] Text-Only Training for Image Captioning using Noise-Injected CLIP[paper] [code][⭐179]<br/>
-
[ICCV] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 11] I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision[paper][code][⭐55]<br/>
-
[arxiv] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 28] Multimodal Knowledge Alignment with Reinforcement Learning[paper][code][⭐22]<br/>
<br/>> 2021
Related Skills
qqbot-channel
385.5kQQ channel management skill. Use qqbot_channel_api for explicit QQ channel-management requests; confirm write, delete, and bulk actions before calling authenticated QQ Open Platform endpoints.
docs-writer
106.4kAlways use this skill when the task involves writing, reviewing, or editing files in the `/docs` directory or any `.md` files in the repository.
cpp
40.5kGuide Cursor to write modern C++ and CMake code with clear structure, RAII, const-correctness, and safe error handling.
gamemaker-gml
40.5kGameMaker Language (GML) rules for scripts, objects, events, rooms, data structures, and performance-minded game code
Security Score
Audited on Jul 28, 2026
