829 skills found · Page 1 of 28
datawhalechina / Self Llm《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程
OpenBMB / MiniCPM VA Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
microsoft / UnilmLarge-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
modelscope / Ms SwiftUse PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
simular-ai / Agent SAgent S: an open agentic framework that uses computers like a human
X-PLUG / MobileAgentMobile-Agent: The Powerful GUI Agent Family
manycore-research / SpatialLM[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
microsoft / LMOpsGeneral technology for enabling AI capabilities w/ LLMs and MLLMs
ant-research / MagicQuill[CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System
atfortes / Awesome LLM ReasoningFrom Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓
NExT-GPT / NExT GPTCode and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
NVlabs / EagleEagle: Frontier Vision-Language Models with Data-Centric Strategies
timerring / Bilive极快的B站直播录制、自动切片、自动渲染弹幕以及字幕并投稿至B站,综合多种模态模型,兼容超低配置机器。Extremely fast live recording, automatic slicing, rendering, uploading and Integrating MLLMs. Compatible with low configurations machines.
InternLM / InternLM XComposerInternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
sherlockchou86 / VideoPipeA cross-platform video structuring (video analysis) framework based on CV models & mLLM.
X-PLUG / MPLUG DocOwlmPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
cambrian-mllm / CambrianCambrian-1 is a family of multimodal LLMs with a vision-centric design.
coderonion / Awesome Yolo Object Detection🚀🚀🚀 A collection of some awesome public YOLO object detection series projects and the related object detection datasets.
bytedance / Sa2VAOfficial Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
Osilly / Vision R1[ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incentivize reasoning capability.