FunASR open-source speech recognition family: SenseVoice model achieves 17x real-time on CPU, 170x real-time on GPU, with half the Chinese error rate of Whisper, plus emotion and speaker recognition. Minimum requirements included.
There's an unwritten rule in the AI photo editing world: the larger the model, the better the results. FLUX.1-Fill-Dev has 11.9B parameters, SD3.5 Large pushes beyond 10B+, and running it once occupies a whole A100. The industry defaults to: if you want good results, stack compute first. Then HUST + VIVO AI Lab unveiled Moebius. 0.22B parameters. 226 million. Less than 2% of FLUX, 15x faster inference, 26ms per step on a single GPU. Across 6 standard benchmarks, it rivals FLUX.1-Fill-Dev—and even surpasses it in facial details and complex textures. The project has been accepted by ECCV 2026, with code and weights fully open-sourced under Apache-2.0. It ranks #1 on Hugging Face's daily leaderboard and #4 on the weekly leaderboard.