Back to List
TechnologyAIMobileMultimodal

MiniCPM-o: Gemini 2.5 Flash-Level MLLM for Mobile Devices with Vision, Speech, and Full-Duplex Multimodal Live Support

OpenBMB has introduced MiniCPM-o, a new multimodal large language model (MLLM) designed for mobile devices. This model is positioned as a Gemini 2.5 Flash-level equivalent, offering robust capabilities in vision, speech, and full-duplex multimodal live interactions. MiniCPM-o aims to bring advanced AI functionalities directly to users' smartphones, enabling a seamless and interactive experience across various modalities.

GitHub Trending

OpenBMB has unveiled MiniCPM-o, a cutting-edge multimodal large language model (MLLM) specifically engineered for mobile phone applications. This innovative model is touted as achieving a performance level comparable to Gemini 2.5 Flash, signifying its advanced capabilities in processing and understanding diverse data types. MiniCPM-o integrates support for vision, allowing it to interpret and respond to visual inputs, and speech, enabling voice-based interactions. A key feature is its full-duplex multimodal live support, which suggests the model can engage in real-time, continuous, and bidirectional communication across these different modalities. This development aims to enhance the user experience on mobile devices by providing sophisticated AI assistance that can understand and interact with users through a combination of visual and auditory cues in a live setting.

Related News

Superpowers: A Proven Agent Skill Framework and Software Development Methodology for Coding Agents
Technology

Superpowers: A Proven Agent Skill Framework and Software Development Methodology for Coding Agents

Superpowers is presented as an effective agent skill framework and a comprehensive software development methodology. It is designed for coding agents, built upon a foundation of composable 'skills' and a set of initial skills. This framework offers a complete workflow for developing agents, emphasizing a structured approach to agent-based software creation.

OpenViking: An Open-Source Context Database for AI Agents, Designed for Hierarchical Context Management and Self-Evolution
Technology

OpenViking: An Open-Source Context Database for AI Agents, Designed for Hierarchical Context Management and Self-Evolution

OpenViking, an open-source context database developed by volcengine, is specifically designed for AI agents like openclaw. It unifies the management of agent context, including memory, resources, and skills, through a file system paradigm. This innovative approach enables hierarchical context passing and supports the self-evolution of AI agents, streamlining how agents access and utilize necessary information for their operations and development.

dimos: A New Proxy Operating System Built on the Dimensional Framework Emerges on GitHub Trending
Technology

dimos: A New Proxy Operating System Built on the Dimensional Framework Emerges on GitHub Trending

dimos, described as a 'Proxy Operating System' and built upon a 'Dimensional Framework,' has recently appeared on GitHub Trending. Developed by dimensionalOS, this project was published on March 16, 2026. The limited information available suggests it is a foundational system, with its core components rooted in a dimensional architecture, aiming to provide a new approach to operating system design.