Back to List
TechnologyAIPerformanceAPI

Kimi K2 Thinking Achieves Top Performance on Vending-Bench, Outperforming Open-Source Models with Moonshot API Integration

Kimi.ai announced that its Kimi K2 Thinking model has become the leading open-source model on the Vending-Bench benchmark. This improved performance was observed after re-running the model using Moonshot's own API, a method suggested to enhance tool calling capabilities. The re-evaluation by Andon Labs confirmed that integrating with the Moonshot API significantly boosted Kimi K2's average net worth achieved on the benchmark, solidifying its position as the top performer among open-source alternatives.

twitter-Kimi.ai

Kimi.ai has highlighted the superior performance of its Kimi K2 Thinking model, which has now been recognized as the best open-source model on the Vending-Bench benchmark. This achievement follows a re-evaluation conducted by Andon Labs. The re-run of Kimi K2 Thinking on Vending-Bench utilized Moonshot’s proprietary API, a strategy that was suggested to improve the model's performance specifically in tool calling. Andon Labs confirmed the efficacy of this approach, stating that the integration with Moonshot's API indeed led to a significant improvement. Consequently, Kimi K2 Thinking has now secured the top position among open-source models on Vending-Bench, based on the average net worth achieved. Kimi.ai encourages users to review the Kimi K2 Thinking benchmark best practices and obtain an API key via their platform.

Related News

Technology

Google's New Python Library 'langextract' Leverages LLMs for Precise Structured Information Extraction from Unstructured Text with Interactive Visualization

Google has released 'langextract', a new Python library designed to extract structured information from unstructured text. This library utilizes Large Language Models (LLMs) to achieve high precision in source localization and offers interactive visualization capabilities. 'langextract' aims to streamline the process of converting free-form text into organized data, making it a valuable tool for developers and researchers working with large volumes of text.

Technology

GitButler: A New Version Control Client Powered by Git, Tauri, Rust, and Svelte Hits GitHub Trending

GitButler, a new version control client, has emerged on GitHub Trending. This client leverages the robust capabilities of Git for its core version control functionalities. It is built using a modern tech stack, driven by Tauri, Rust, and Svelte, indicating a focus on performance, cross-platform compatibility, and a responsive user interface. The project is developed by gitbutlerapp and was published on February 10, 2026.

Technology

Dexter: An Autonomous AI Agent for In-Depth Financial Research and Market Analysis

Dexter is an autonomous AI agent designed for deep financial research. It operates by thinking, planning, and learning throughout its tasks. The agent conducts analysis using task planning, self-reflection, and real-time market data, making it a sophisticated tool for financial insights.