Open models for speech, video, language, and embodied intelligence.
Models, datasets, benchmarks, and tools you can run, study, and build on.
Website · Model downloads · All repositories · Follow on X
English · 简体中文
Models, training frameworks, benchmarks, and research resources, organized by area.
Language · Vision · Speech · Audio · Embodied AI · Training · New architectures · Interpretability · Data · Benchmarks · Tools & resources
MOSS
Chinese and English dialogue with tool use
Models · Get started
MOSS-VL
11B models for long-form and real-time video understanding
Demo · Models
MOVA
Generate video and synchronized audio, with training and LoRA workflows
Samples · Models
OmniVAE
Aligned audio-video representations for reconstruction and joint generation
Project · Model
AnyGPT
A multimodal language model for text, speech, images, and music
Demos · Dataset
Also explore: MOSS-Video-Preview (earlier streaming video model)
MOSS-TTS
Model family for narration, dialogue, voice design, and sound effects
Model guide · Models
MOSS-TTS-Nano
100M-parameter multilingual voice cloning with CPU and ONNX inference
Demo · Run locally
MOSS-TTSD
Long-form, multi-speaker dialogue and podcast synthesis
Demo · Model
MOSS-Speech
End-to-end speech-to-speech dialogue without text guidance
Project · Paper
MOSS-Audio-Tokenizer
Streaming audio tokenization across speech, sound, and music
Model · Paper
Also explore: MOSS-TTS-Nano-Reader (in-browser reading) · SpeechGPT-2.0-preview (real-time spoken dialogue)
MOSS-Transcribe-Diarize
0.9B transcription in 50+ languages, with speaker labels and timestamps
Subtitle app · Model
MOSS-Audio
Captioning, question answering, and reasoning over real-world audio
Project · Models
MOSS-Music
Music captioning, lyrics transcription, structural analysis, and QA
Model
Also explore: MOSS-Audio-Tokenizer-Eval (codec reconstruction evaluation) · TTSD-eval (multi-speaker speech evaluation)
EasyWAM
Train, fine-tune, and evaluate World Action Models in a shared framework
Project · Models
OpenETA
An embodied agent linking perception, action, verification, and learning
Project
RoboOmni
Proactive robot manipulation in multimodal physical environments
Project · Model
FRoM-W1
Language-guided whole-body control for humanoid robots
Project
Also explore: Embodied-Planner-R1 (reinforcement learning for embodied planning)
CoLLiE
Efficient collaborative training of large language models
Docs
DiRL
Supervised fine-tuning and reinforcement learning for diffusion language models
Model · Paper
LongLLaDA
Extend the context length of diffusion language models
Paper
Sparse-dLLM
Accelerate diffusion LLMs with cache eviction and sparse attention
Paper
Also explore: rope_pp (rotary position embeddings for long context) · ReAttention (training-free context extension)
Llamascopium
Train and analyze sparse autoencoders, trace circuits, and visualize features
Docs
Also explore: Lorsa (low-rank sparse attention decomposition)
Ultra-Innerthought
Bilingual reasoning data
Dataset
SWE-bench-Science
Coding-agent evaluation on engineering tasks in scientific software
Leaderboard · Dataset
ContextWeave
Evaluate coding-agent memory through long-horizon worklog tasks
Paper
AgentHPOBench
Evaluate LLM agents as sequential hyperparameter optimizers
Paper
ABC-Bench
Test whether coding agents can build, deploy, and verify backend services
Project
FutureOmni
Forecast future events from audio and video context
Project · Dataset
VLABench
Evaluate vision-language-action models and embodied agents
Project
Also explore: LongSafety (long-context safety) · VehicleWorld (connected cockpit interaction) · HalluQA (Chinese hallucination evaluation) · GAOKAO-MM (Chinese multimodal evaluation) · Say-I-Dont-Know (knowing when to abstain)
Awesome-WAM
World Action Model survey, paper collection, and benchmark comparisons
Browse · Benchmarks
Thus-Spake-Long-Context-LLM
A survey of long-context architectures, infrastructure, training, and evaluation
Paper
Also explore: claude-codex-handoff (agent collaboration protocol) · OurClaw (multi-user OpenClaw deployment) · imclaw-skill (agent messaging)
Start with a project's README, then use its Issues or pull requests to report a reproducible bug, improve documentation, or contribute an integration. If you use a model or dataset in research, please cite the corresponding work linked in that repository.
Follow OpenMOSS on Hugging Face for models and datasets, and @Open_MOSS on X for release announcements.
OpenMOSS is led by Prof. Xipeng Qiu at the Shanghai Innovation Institute (SII), in collaboration with Fudan University and MOSI.AI. For research collaborations, PhD and internship inquiries, contact openmoss@sii.edu.cn.