Visage

Visage contains an image dataset of images with human annotations on whether or not certain attributes are present or depicted in the image. The attribute may either be stereotypical or non-stereotypical w.r.t. to the identity group in the image. It also contains a list of attributes in English along with annotations about whether they are visual.

Generate Convert Improve

Install / Use

/learn @google-research-datasets/Visage

About this skill

Quality Score

0/100

README

ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation

This repository contains data resources for the paper "ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation."

Overview

Recent studies have shown that Text-to-Image (T2I) model generations can reflect social stereotypes present in the real world. However, existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes. To address this gap, we introduce the ViSAGe (Visual Stereotypes Around the Globe) dataset to enable the evaluation of known nationality-based stereotypes in T2I models, across 135 nationalities. We enrich an existing textual stereotype resource by distinguishing between stereotypical associations that are more likely to have visual depictions, such as `sombrero', from those that are less visually concrete, such as 'attractive'. We demonstrate ViSAGe's utility through a multi-faceted evaluation of T2I generations. First, we show that stereotypical attributes in ViSAGe are thrice as likely to be present in generated images of corresponding identities as compared to other attributes and that the offensiveness of these depictions is especially higher for identities from Africa, South America, and Southeast Asia. Second, we assess the stereotypical pull of visual depictions of identity groups, which reveals how the 'default' representations of all identity groups in ViSAGe have a pull towards stereotypical depictions, and that this pull is even more prominent for identity groups from the Global South.

Dataset Description

The repo contains the data cards for the ViSAGe dataset, following the format proposed by Pushkarna et al.. The data card includes details of the dataset such as intended usage, field names and meanings, annotator recruitment, and payments. The file visual_attributes contains the annotations for the visual nature of attributes as denoted on a Likert Scale ranging from "Strongly Agree" to "Strongly Disagree". The file Image_Annotations contains annotations for the presence or absence of attributes in the images along with their co-ordinates.

Citation

@inproceedings{jha-etal-2024-visage,
    title = "{V}i{SAG}e: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation",
    author = "Jha, Akshita  and
      Prabhakaran, Vinodkumar  and
      Denton, Remi  and
      Laszlo, Sarah  and
      Dave, Shachi  and
      Qadri, Rida  and
      Reddy, Chandan  and
      Dev, Sunipa",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.667",
    pages = "12333--12347",
}

Related Skills

node-connect

351.2k

Diagnose OpenClaw node connection and pairing failures for Android, iOS, and macOS companion apps

frontend-design

110.6k

Create distinctive, production-grade frontend interfaces with high design quality. Use this skill when the user asks to build web components, pages, or applications. Generates creative, polished code that avoids generic AI aesthetics.

openai-whisper-api

351.2k

Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

qqbot-media

351.2k

QQBot 富媒体收发能力。使用 <qqmedia> 标签，系统根据文件扩展名自动识别类型（图片/语音/视频/文件）。