21 skills found
stepfun-ai / Step Audio EditXA powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
david-yoon / Multimodal Speech EmotionTensorFlow implementation of "Multimodal Speech Emotion Recognition using Audio and Text," IEEE SLT-18
naver-ai / UsdmOfficial PyTorch implementation of "Paralinguistics-Aware Speech-Empowered LLMs for Natural Conversation" (NeurIPS 2024)
flybirdxx / ComfyUI SoulX PodcastSoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
ShawnPi233 / SynParaSpeechOfficial Repository of Paper: "SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding" (ICASSP 2026)
haoweilou / ParaStyleTTSThis is the official code for ACM CIKM 2025 Paper: ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
Alittleegg / Eureka AudioEureka-Audio: A 1.7B lightweight audio–language model that matches 7B–30B models on ASR, audio understanding, and paralinguistic reasoning.
NARUTO-2024 / WavBenchWavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models
david-yoon / Attentive Modality Hopping For SERTensorFlow implementation of "Attentive Modality Hopping for Speech Emotion Recognition," ICASSP-20
KeiKinn / ParaCLAPTowards a general language-audio model for computational paralinguistic tasks
smthemex / ComfyUI Step Audio EditX SMStep_Audio_EditX:the first open-source LLM-based audio model excelling at expressive and iterative audio editing—encompassing emotion, speaking style, and paralinguistics—alongside robust zero-shot text-to-speech (TTS) capabilities,try it in comfyUI
neosun100 / Soulx Podcast Docker🎙️ SoulX-Podcast Docker Deployment - Realistic Long-form Podcast Speech Synthesis with Dialectal and Paralinguistic Diversity | 播客语音合成 Docker 部署
haoweilou / ParaMETAOfficial Implementation for AAAI2026 Paper: ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations From Speech
hlt-cuhksz / EchoMindEchoMind is an interrelated multi‑level benchmark evaluating empathetic dialogue in speech language models by unifying linguistic and paralinguistic understanding in a context‑linked framework.
gizemsogancioglu / Elderly Emotion SCThis project contains the scripts for our entry in the Elderly Emotion Sub-Challenge which is part of INTERSPEECH 2020 Computational Paralinguistics Challenge (ComParE)
tltrogl / Diaremot2 OnDiaRemot2-ON: CPU-only audio intelligence pipeline (Faster-Whisper, ONNX, diarization, paralinguistics)
hwk06023 / SONATASONATA (SOund and Narrative Advanced Transcription Assistant): An advanced ASR system that captures human expressions including emotive sounds and non-verbal cues.
glam-imperial / ComParE2020 Breathing End2EndThis repository contains the code to replicate the End-to-End Deep Sequence Modelling baseline for the Breathing Challenge of the Interspeech 2020 Computational Paralinguistics Challenge (ComParE).
hhoangphuoc / SpeechLaughRecogniserAn ASR model for transcribing laughter and speech-laugh for spontaneous conversational speech
vantagewithai / Vantage Step Audio EditXThis project is a custom node implementation built on top of Step-Audio-EditX. It adapts and extends EditX capabilities to support multi‑speaker, long‑format, voice cloning, and emotion/style/speed editing, enabling you to feed in a script with multiple speakers, inline pauses, paralinguistic cues, and get a concatenated audio output in one pass.