Back to List
TechnologyAIPerformanceAPI

Kimi K2 Thinking Achieves Top Performance on Vending-Bench, Outperforming Open-Source Models with Moonshot API Integration

Kimi.ai announced that its Kimi K2 Thinking model has become the leading open-source model on the Vending-Bench benchmark. This improved performance was observed after re-running the model using Moonshot's own API, a method suggested to enhance tool calling capabilities. The re-evaluation by Andon Labs confirmed that integrating with the Moonshot API significantly boosted Kimi K2's average net worth achieved on the benchmark, solidifying its position as the top performer among open-source alternatives.

twitter-Kimi.ai

Kimi.ai has highlighted the superior performance of its Kimi K2 Thinking model, which has now been recognized as the best open-source model on the Vending-Bench benchmark. This achievement follows a re-evaluation conducted by Andon Labs. The re-run of Kimi K2 Thinking on Vending-Bench utilized Moonshot’s proprietary API, a strategy that was suggested to improve the model's performance specifically in tool calling. Andon Labs confirmed the efficacy of this approach, stating that the integration with Moonshot's API indeed led to a significant improvement. Consequently, Kimi K2 Thinking has now secured the top position among open-source models on Vending-Bench, based on the average net worth achieved. Kimi.ai encourages users to review the Kimi K2 Thinking benchmark best practices and obtain an API key via their platform.

Related News

Technology

Hugging Face Introduces 'Skills' for AI/ML Task Definition, Compatible with Major Coding Agent Tools

Hugging Face has launched 'Skills,' a new framework designed to define AI/ML tasks such as dataset creation, model training, and evaluation. These 'Skills' are built to be compatible with leading coding agent tools, including OpenAI Codex, Anthropic's Claude Code, and Google De. This initiative aims to standardize and streamline the definition of various AI and machine learning tasks, facilitating integration across different development platforms.

Technology

Moonshine Voice: Fast and Accurate Automatic Speech Recognition (ASR) for Edge Devices Trends on GitHub

Moonshine Voice, a project by moonshine-ai, is gaining traction on GitHub Trending for its focus on delivering fast and accurate Automatic Speech Recognition (ASR) specifically designed for edge devices. Published on February 28, 2026, this initiative aims to optimize ASR capabilities for resource-constrained environments, making advanced speech recognition more accessible and efficient for a wide range of edge computing applications. The project's presence on GitHub Trending highlights its potential impact in the field of AI and edge device technology.

Technology

cc-switch: A Cross-Platform Desktop Assistant for Claude Code, Codex, OpenCode, and Gemini CLI Trending on GitHub

cc-switch is an innovative cross-platform desktop integrated assistant tool designed to streamline workflows for developers utilizing Claude Code, Codex, OpenCode, and Gemini CLI. Recently trending on GitHub, this tool aims to provide an all-in-one solution for managing these diverse coding and AI command-line interfaces, enhancing productivity and user experience across different operating systems. The project is authored by farion1231 and was published on February 28, 2026.