Kimi K2 Thinking Achieves Top Performance on Vending-Bench, Outperforming Open-Source Models with Moonshot API Integration
Kimi.ai announced that its Kimi K2 Thinking model has become the leading open-source model on the Vending-Bench benchmark. This improved performance was observed after re-running the model using Moonshot's own API, a method suggested to enhance tool calling capabilities. The re-evaluation by Andon Labs confirmed that integrating with the Moonshot API significantly boosted Kimi K2's average net worth achieved on the benchmark, solidifying its position as the top performer among open-source alternatives.
Kimi.ai has highlighted the superior performance of its Kimi K2 Thinking model, which has now been recognized as the best open-source model on the Vending-Bench benchmark. This achievement follows a re-evaluation conducted by Andon Labs. The re-run of Kimi K2 Thinking on Vending-Bench utilized Moonshot’s proprietary API, a strategy that was suggested to improve the model's performance specifically in tool calling. Andon Labs confirmed the efficacy of this approach, stating that the integration with Moonshot's API indeed led to a significant improvement. Consequently, Kimi K2 Thinking has now secured the top position among open-source models on Vending-Bench, based on the average net worth achieved. Kimi.ai encourages users to review the Kimi K2 Thinking benchmark best practices and obtain an API key via their platform.