An open-source, tool-augmented conversational language model from Fudan University
A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning
An efficient framework for collaborative training of large language models
MOSS-Audio-Tokenizer is a Causal Transformer-based audio tokenizer built on the CAT architecture. Trained on 3M hours of diverse audio, it supports streaming and variable bitrates, delivering SOTA reconstruction and strong performance in generation and understanding—serving as a unified interface for next-generation native audio language models.
A unified framework for training, fine-tuning, and evaluating World Action Models
A Chinese human-level benchmark for evaluating multimodal large language models
A benchmark for testing whether coding agents can resolve engineering tasks in scientific software
A survey of long-context language models covering architecture, infrastructure, training, and evaluation