SkillAgentSearch skills...

MER Factory

πŸš€ Pre-process, annotate, evaluate, and train your Affect Computing (e.g., Multimodal Emotion Recognition, Sentiment Analysis) datasets ALL within MER-Factory! (LangGraph Based Agent Workflow)

Install / Use

npx skills add Lum1104/MER-Factory

Installs into whichever agent you are using.

README

πŸ‘‰πŸ» MER-Factory πŸ‘ˆπŸ»

<p align="left"> <a href="README_zh.md">δΈ­ζ–‡</a> &nbsp| &nbsp English&nbsp&nbsp </p> <br> <p align="center"> <a href="https://lum1104.github.io/MER-Factory/" target="_blank">πŸ“– Documentation</a> </p> <p align="center"> <img src="https://img.shields.io/badge/Task-Multimodal_Emotion_Reasoning-red"> <img src="https://img.shields.io/badge/Task-Multimodal_Emotion_Recognition-red"> <a href="https://zread.ai/Lum1104/MER-Factory" target="_blank"><img src="https://img.shields.io/badge/Ask_Zread-_.svg?style=plastic&color=00b0aa&labelColor=000000&logo=data%3Aimage%2Fsvg%2Bxml%3Bbase64%2CPHN2ZyB3aWR0aD0iMTYiIGhlaWdodD0iMTYiIHZpZXdCb3g9IjAgMCAxNiAxNiIgZmlsbD0ibm9uZSIgeG1sbnM9Imh0dHA6Ly93d3cudzMub3JnLzIwMDAvc3ZnIj4KPHBhdGggZD0iTTQuOTYxNTYgMS42MDAxSDIuMjQxNTZDMS44ODgxIDEuNjAwMSAxLjYwMTU2IDEuODg2NjQgMS42MDE1NiAyLjI0MDFWNC45NjAxQzEuNjAxNTYgNS4zMTM1NiAxLjg4ODEgNS42MDAxIDIuMjQxNTYgNS42MDAxSDQuOTYxNTZDNS4zMTUwMiA1LjYwMDEgNS42MDE1NiA1LjMxMzU2IDUuNjAxNTYgNC45NjAxVjIuMjQwMUM1LjYwMTU2IDEuODg2NjQgNS4zMTUwMiAxLjYwMDEgNC45NjE1NiAxLjYwMDFaIiBmaWxsPSIjZmZmIi8%2BCjxwYXRoIGQ9Ik00Ljk2MTU2IDEwLjM5OTlIMi4yNDE1NkMxLjg4ODEgMTAuMzk5OSAxLjYwMTU2IDEwLjY4NjQgMS42MDE1NiAxMS4wMzk5VjEzLjc1OTlDMS42MDE1NiAxNC4xMTM0IDEuODg4MSAxNC4zOTk5IDIuMjQxNTYgMTQuMzk5OUg0Ljk2MTU2QzUuMzE1MDIgMTQuMzk5OSA1LjYwMTU2IDE0LjExMzQgNS42MDE1NiAxMy43NTk5VjExLjAzOTlDNS42MDE1NiAxMC42ODY0IDUuMzE1MDIgMTAuMzk5OSA0Ljk2MTU2IDEwLjM5OTlaIiBmaWxsPSIjZmZmIi8%2BCjxwYXRoIGQ9Ik0xMy43NTg0IDEuNjAwMUgxMS4wMzg0QzEwLjY4NSAxLjYwMDEgMTAuMzk4NCAxLjg4NjY0IDEwLjM5ODQgMi4yNDAxVjQuOTYwMUMxMC4zOTg0IDUuMzEzNTYgMTAuNjg1IDUuNjAwMSAxMS4wMzg0IDUuNjAwMUgxMy43NTg0QzE0LjExMTkgNS42MDAxIDE0LjM5ODQgNS4zMTM1NiAxNC4zOTg0IDQuOTYwMVYyLjI0MDFDMTQuMzk4NCAxLjg4NjY0IDE0LjExMTkgMS42MDAxIDEzLjc1ODQgMS42MDAxWiIgZmlsbD0iI2ZmZiIvPgo8cGF0aCBkPSJNNCAxMkwxMiA0TDQgMTJaIiBmaWxsPSIjZmZmIi8%2BCjxwYXRoIGQ9Ik00IDEyTDEyIDQiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIxLjUiIHN0cm9rZS1saW5lY2FwPSJyb3VuZCIvPgo8L3N2Zz4K&logoColor=ffffff" alt="zread"/></a> <img src="https://zenodo.org/badge/1007639998.svg" alt="DOI"> </p> <p align="center"> <a href="https://lum1104.github.io/MER-Factory/"> <img src="docs/assets/logo.svg" width="700"> </a> </p>

[!IMPORTANT] ✍️ Challenge: Multimodal Affect Computing isn't just one step-it's the entire fragmented pipeline. The journey from raw files to a trained model is a gauntlet of tedious data preprocessing, slow and inconsistent manual annotation, and complex model training setups.

🏭 MER-Factory: Unifies this entire workflow into one seamless factory. We automate the heavy lifting of preprocessing and annotation to generate high-quality, reason-augmented datasets, and then bridge the gap directly to model training.

πŸš€ Stop juggling different tools: Let our factory handle the pipeline, so you can focus on what matters: your research.

<!-- <p align="center"> <a href="https://lum1104.github.io/MER-Factory/"> <img src="https://svg-banners.vercel.app/api?type=origin&text1=MER-Factory%20🧰&text2=✨%20Factory%20for%20Multimodal%20Emotion%20Recognition%20Reasoning%20(MERR)%20datasets&width=800&height=200" alt="MER-Factory Banner"> </a> </p> -->

πŸš€ Project Roadmap

MER-Factory is under active development with new features being added regularly - check our roadmap and welcome contributions!

<div style="text-align: center;"> <img src="docs/assets/mer-factory.jpeg" style="border: none; width: 100%; max-width: 1000px;"> <!-- the figure generate by gemini 3, many thanks! --> </div>

Table of Contents

Pipeline Structure

<details> <summary>Click here to expand/collapse</summary>

Remove for now, call the (print(app.get_graph().draw_mermaid())) graph.py to view

</details>

Features

  • Action Unit (AU) Pipeline: Extracts facial Action Units (AUs) and translates them into descriptive natural language.
  • Audio Analysis Pipeline: Extracts audio, transcribes speech, and performs detailed tonal analysis.
  • Video Analysis Pipeline: Generates comprehensive descriptions of video content and context.
  • Image Analysis Pipeline: Provides end-to-end emotion recognition for static images, complete with visual descriptions and emotional synthesis.
  • Full MER Pipeline: An end-to-end multimodal pipeline that identifies peak emotional moments, analyzes all modalities (visual, audio, facial), and synthesizes a holistic emotional reasoning summary.
  • Gate Agent (Experimental): An optional quality control layer that reviews intermediate analysis results. Following the "garbage in, garbage out" principle, it rejects low-quality or conflicting outputs and prompts sub-agents to refine their analysis before final synthesis. Enable with --use-gate-agent.

Check out example outputs here:

Installation

<p align="center"> πŸ“š Please visit <a href="https://lum1104.github.io/MER-Factory/" target="_blank">project documentation</a> for detailed installation and usage instructions. </p>

[!Note] For Windows users, simply download the pre-built ffmpeg and OpenFace and place them as requested.

We highly recommend serving the HF model/Ollama model on Linux and running MER-Factory on Windows to reduce installation time.

But, for those love the command line (e.g., me), a complete installation example for Linux environments (including Google Colab) can be found at:

Usage

Basic Command Structure

python main.py [INPUT_PATH] [OUTPUT_DIR] [OPTIONS]

Examples

# Show all supported args.
python main.py --help

# Full MER pipeline with Gemini (default)
python main.py path_to_video/ output/ --type MER --silent --threshold 0.8

# Using Sentiment Analysis task instead of MERR
python main.py path_to_video/ output/ --type MER --task "Sentiment Analysis" --silent

# Using ChatGPT models
python main.py path_to_video/ output/ --type MER --chatgpt-model gpt-4o --silent

# Using local Ollama models
python main.py path_to_video/ output/ --type MER --ollama-vision-model llava-llama3:latest --ollama-text-model llama3.2 --silent

# Using Hugging Face model
python main.py path_to_video/ output/ --type MER --huggingface-model google/gemma-3n-E4B-it --silent

# Process images instead of videos
python main.py ./images ./output --type MER

Note: Run ollama pull llama3.2 etc, if Ollama model is needed. Ollama does not support video analysis for now.

Hugging Face Client-Server Setup

When selecting a Hugging Face model with --huggingface-model, MER-Factory forwards all calls through a lightweight client that talks to a local/remote API server which actually hosts the HF model. This keeps your main environment clean and allows easy scaling.

  1. Start the HF API Server (in a separate terminal):
# Example: serve Whisper base on port 7860
python -m mer_factory.models.hf_api_server --model_id openai/whisper-base --host 0.0.0.0 --port 7860
  1. Run MER-Factory as usual and select the HF model by ID:
python main.py path_to_video/ output/ --type MER --huggingface-model openai/whisper-base --silent

Dashboard for Data Curation and Hyperparameter Tuning

We provide an interactive dashboard webpage to facilitate data curation and hyperparameter tuning. The dashboard allows you to test different prompts, save and run configurations, and rate the generated data.

To launch the dashboard, use the following command:

python dashboard.py

Command Line Options

| Option | Short | Description | Default | |--------|-------|-------------|---------| | --type | -t | Processing type (AU, audio, video, image, MER) | MER | | --task | -tk | Analysis task type (MERR, Sentiment Analysis) | MERR | | --label-file | -l | Path to a CSV file with 'name' and 'label' columns. Optional, for ground truth labels. | None | | --threshold | -th | Emotion detection threshold (0.0-5.0) | 0.8 | | --peak_dis | -pd | Steps between peak frame detection (min 8) | 15 | | --silent | -s | Run with minimal output | False | | --cache | -ca | Reuse existing audio/video/AU results from previous pipeline runs | False | | --concurrency | -c | Concurrent files for async processing (min 1) | 4 | | --ollama-vision-model | -ovm | Ollama vision model name | None | | --ollama-text-model | -otm | Ollama text model name | None | | --chatgpt-model | -cgm | ChatGPT model name (e.g., gpt-4o) | None | | --huggingface-model | -hfm | Hugging Face model ID | None | | --use-gate-agent | -uga | Enable Gate Agent for quality control (Dev Feature) | False |

Processing Types

1. Action Unit (AU) Extraction

Extracts facial Action Units and generates natural language descriptions:

python main.py video.mp4 output/ --type AU

2. Audio Analysis

Extracts audio, transcribes speech, and analyzes tone:

python main.py video.mp4 output/ --type audio

3. Video Analysis

Generates comprehensive video content descriptions:

python 

Related Skills

View on GitHub
GitHub Stars116
CategoryDevelopment
Updated7d ago
Forks14

Languages

Python

Security Score

100/100

Audited on Jul 31, 2026

No findings