Skip to content
@OpenMOSS

OpenMOSS (SII)

OpenMOSS Team is a research group under the Shanghai Innovation Institute (SII), working in close collaboration with Fudan University and MOSI Intelligence.
OpenMOSS

Open models for speech, video, language, and embodied intelligence.

Models, datasets, benchmarks, and tools you can run, study, and build on.

Website · Model downloads · All repositories · Follow on X

English · 简体中文

Explore our projects

Models, training frameworks, benchmarks, and research resources, organized by area.

Language · Vision · Speech · Audio · Embodied AI · Training · New architectures · Interpretability · Data · Benchmarks · Tools & resources

Language models

MOSS logo MOSS
Chinese and English dialogue with tool use
Models · Get started

Vision & multimodal generation

MOSS-VL logo MOSS-VL
11B models for long-form and real-time video understanding
Demo · Models

MOVA logo MOVA
Generate video and synchronized audio, with training and LoRA workflows
Samples · Models

OmniVAE
Aligned audio-video representations for reconstruction and joint generation
Project · Model

AnyGPT logo AnyGPT
A multimodal language model for text, speech, images, and music
Demos · Dataset

Also explore: MOSS-Video-Preview (earlier streaming video model)

Speech & audio generation

MOSS-TTS
Model family for narration, dialogue, voice design, and sound effects
Model guide · Models

MOSS-TTS-Nano
100M-parameter multilingual voice cloning with CPU and ONNX inference
Demo · Run locally

MOSS-TTSD logo MOSS-TTSD
Long-form, multi-speaker dialogue and podcast synthesis
Demo · Model

MOSS-Speech logo MOSS-Speech
End-to-end speech-to-speech dialogue without text guidance
Project · Paper

MOSS-Audio-Tokenizer
Streaming audio tokenization across speech, sound, and music
Model · Paper

Also explore: MOSS-TTS-Nano-Reader (in-browser reading) · SpeechGPT-2.0-preview (real-time spoken dialogue)

Speech, audio & music understanding

MOSS-Transcribe-Diarize logo MOSS-Transcribe-Diarize
0.9B transcription in 50+ languages, with speaker labels and timestamps
Subtitle app · Model

MOSS-Audio logo MOSS-Audio
Captioning, question answering, and reasoning over real-world audio
Project · Models

MOSS-Music logo MOSS-Music
Music captioning, lyrics transcription, structural analysis, and QA
Model

Also explore: MOSS-Audio-Tokenizer-Eval (codec reconstruction evaluation) · TTSD-eval (multi-speaker speech evaluation)

Embodied intelligence & robotics

EasyWAM logo EasyWAM
Train, fine-tune, and evaluate World Action Models in a shared framework
Project · Models

OpenETA logo OpenETA
An embodied agent linking perception, action, verification, and learning
Project

RoboOmni logo RoboOmni
Proactive robot manipulation in multimodal physical environments
Project · Model

FRoM-W1
Language-guided whole-body control for humanoid robots
Project

Also explore: Embodied-Planner-R1 (reinforcement learning for embodied planning)

Training & post-training

CoLLiE logo CoLLiE
Efficient collaborative training of large language models
Docs

DiRL logo DiRL
Supervised fine-tuning and reinforcement learning for diffusion language models
Model · Paper

New architectures

LongLLaDA
Extend the context length of diffusion language models
Paper

Sparse-dLLM
Accelerate diffusion LLMs with cache eviction and sparse attention
Paper

Also explore: rope_pp (rotary position embeddings for long context) · ReAttention (training-free context extension)

Interpretability

Llamascopium
Train and analyze sparse autoencoders, trace circuits, and visualize features
Docs

Also explore: Lorsa (low-rank sparse attention decomposition)

Data

Ultra-Innerthought
Bilingual reasoning data
Dataset

Benchmarks & agent evaluation

SWE-bench-Science
Coding-agent evaluation on engineering tasks in scientific software
Leaderboard · Dataset

ContextWeave
Evaluate coding-agent memory through long-horizon worklog tasks
Paper

AgentHPOBench logo AgentHPOBench
Evaluate LLM agents as sequential hyperparameter optimizers
Paper

ABC-Bench
Test whether coding agents can build, deploy, and verify backend services
Project

FutureOmni logo FutureOmni
Forecast future events from audio and video context
Project · Dataset

VLABench
Evaluate vision-language-action models and embodied agents
Project

Also explore: LongSafety (long-context safety) · VehicleWorld (connected cockpit interaction) · HalluQA (Chinese hallucination evaluation) · GAOKAO-MM (Chinese multimodal evaluation) · Say-I-Dont-Know (knowing when to abstain)

Surveys & developer tools

Awesome-WAM
World Action Model survey, paper collection, and benchmark comparisons
Browse · Benchmarks

Thus-Spake-Long-Context-LLM
A survey of long-context architectures, infrastructure, training, and evaluation
Paper

Also explore: claude-codex-handoff (agent collaboration protocol) · OurClaw (multi-user OpenClaw deployment) · imclaw-skill (agent messaging)

Build with us

Start with a project's README, then use its Issues or pull requests to report a reproducible bug, improve documentation, or contribute an integration. If you use a model or dataset in research, please cite the corresponding work linked in that repository.

Follow OpenMOSS on Hugging Face for models and datasets, and @Open_MOSS on X for release announcements.

OpenMOSS is led by Prof. Xipeng Qiu at the Shanghai Innovation Institute (SII), in collaboration with Fudan University and MOSI.AI. For research collaborations, PhD and internship inquiries, contact openmoss@sii.edu.cn.

Pinned Loading

  1. MOSS MOSS Public

    An open-source, tool-augmented conversational language model from Fudan University

    Python 12.2k 1.1k

  2. MOSS-VL MOSS-VL Public

    An open-weight 11B model series for long-form and real-time video understanding

    Python 596 20

  3. MOVA MOVA Public

    A foundation model that generates synchronized video and audio in a single model

    Python 1.1k 91

  4. MOSS-TTS-Nano MOSS-TTS-Nano Public

    A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation

    Python 4.3k 548

  5. Llamascopium Llamascopium Public

    A framework for training, analyzing, and visualizing sparse autoencoders and related interpretability methods

    Python 227 29

  6. MOSS-Transcribe-Diarize MOSS-Transcribe-Diarize Public

    A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness

    Python 1.9k 109

Repositories

Showing 10 of 63 repositories

Top languages

Loading…

Most used topics

Loading…