SkillAgentSearch skills...

Awesome Unimodal Training

text-only training or language-free training for multimodal tasks (image/audio/video caption, retrieval, text2image)

Install / Use

npx skills add iOPENCap/awesome-unimodal-training

Installs into whichever agent you are using.

README

Awesome-Text-only-Training

The project is used to store text-only training, image-free training for multimodal tasks related papers.

Include text-only training for:

  • zero-shot image/audio/video captioning
  • zero-shot composed image retrieval
  • visual storytelling, visual question answer...

papers

<br/>> 2024

  • [AAAI] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training [paper] [code][⭐8]<br/>

  • [ACM] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] TOMGPT: Reliable Text-Only Training Approach for Cost-Effective Multi-modal Large Language Model[paper]<br/>

  • [IJCV] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 5] Learning to Prompt with Text Only Supervision for Vision-Language Models[paper] [code][⭐80]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Text Data-Centric Image Captioning with Interactive Prompts[paper]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 2] MeaCap: Memory-Augmented Zero-shot Image Captioning[paper] [code][⭐27]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks [paper]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion [paper]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning [paper] [code][⭐4]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] From Unimodal to Multimodal: Scaling up Projectors to Align Modalities [paper]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning [paper] [code][⭐0]<br/>

  • [arxiv] | [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] DRCap: Decoding CLAP Latents with Retrieval-augmented Generation for Zero-shot Audio Captioning [paper] <br/>

<br/>> 2023

  • [arxiv] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] Improved Factorized Neural Transducer Model For text-only Domain Adaptation[paper] <br/>

  • [AAAI] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 2] Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning[paper] <br/>

  • [ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 0] Text-Only Training for Visual Storytelling[paper] <br/>

  • [ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 1] VLIS: Unimodal Language Models Guide Multimodal Language Generation[paper] [code][⭐24] <br/>

  • [NeurlPS] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 7] LOVM:Language-Only Vision Model Selection[paper] [code][⭐18]<br/>

  • [IJCAI] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 4] From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping[paper] [code][⭐11]<br/>

  • [ICLR] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 47] Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training[paper] [code][⭐118]<br/>

  • [DCASE] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 4] Weakly-supervised Automated Audio Captioning via text only training[paper] [code][⭐1]<br/>

  • [ACM] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 3] CgT-GAN: CLIP-guided Text GAN for Image Captioning[paper] [code][⭐16]<br/>

<br/>> 2022

  • [EMNLP] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 64] Text-Only Training for Image Captioning using Noise-Injected CLIP[paper] [code][⭐179]<br/>

  • [ICCV] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 11] I Can't Believe There's No Images! Learning Visual Tasks Using only Language Supervision[paper][code][⭐55]<br/>

  • [arxiv] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 28] Multimodal Knowledge Alignment with Reinforcement Learning[paper][code][⭐22]<br/>

<br/>> 2021

  • [CVPR] [<img src="https://github.com/user-attachments/assets/c30947ec-4d5a-424d-89eb-583d8efd2801" width="15"> 124] LAFITE: Towards Language-Free Training for Text-to-Image Generation[paper] [code][⭐180]<br/>

Related Skills

View on GitHub
GitHub Stars13
CategoryContent
Updated11d ago
Forks0

Security Score

80/100

Audited on Jul 28, 2026

No findings