Back to list
Cerebras Inference Platform Achieves Record Speeds with Qwen 3.8 27B and OpenAI GPT OSS 120B
Industry NewsCerebrasQwenGPT OSS

Cerebras Inference Platform Achieves Record Speeds with Qwen 3.8 27B and OpenAI GPT OSS 120B

Cerebras Systems has announced a significant performance update to its inference platform, featuring the Qwen 3.8 27B and OpenAI GPT OSS 120B models. According to the latest documentation, the Qwen 3.8 27B model now operates at approximately 1500 tokens per second, while the GPT OSS 120B model reaches an impressive 3000 tokens per second. These models are available through various access tiers, including free trials and pay-as-you-go options, with context windows extending up to 131k. A key highlight of this release is Cerebras' commitment to model quality; all models served via public endpoints are unpruned versions. The platform utilizes selective weight-only quantization for storage to maintain high precision during operations, ensuring that quality-sensitive layers remain at full precision through on-the-fly dequantization.

Hacker News

Key Takeaways

  • Exceptional Inference Speeds: Qwen 3.8 27B achieves ~1500 tokens/s, and OpenAI GPT OSS 120B reaches ~3000 tokens/s on Cerebras infrastructure.
  • Unpruned Model Integrity: Cerebras hosts original, unpruned versions of open-source models to ensure maximum output quality and architectural transparency.
  • Advanced Context Support: Context windows are optimized for both free and paid tiers, ranging from 64k to 131k depending on the model and subscription level.
  • Sophisticated Quantization Strategy: The platform employs selective weight-only quantization (4-bit to 16-bit) for storage while maintaining high-precision execution for sensitive layers.
  • Flexible Access Tiers: Users can access these high-speed models via public endpoints (Free/Pay-as-you-go) or through Dedicated Endpoints for production-grade SLAs.

In-Depth Analysis

Breakthrough Performance in LLM Inference

The latest documentation from Cerebras Inference reveals a significant leap in processing speeds for large language models (LLMs). The Qwen 3.8 27B model, a 27-billion parameter powerhouse, is now clocked at approximately 1500 tokens per second. Even more striking is the performance of the OpenAI GPT OSS 120B model, which, despite its massive 120-billion parameter size, achieves a throughput of ~3000 tokens per second. These speeds represent a critical benchmark for real-time AI applications where latency and throughput are paramount.

To support these high speeds across various use cases, Cerebras has defined specific context window limits. For the Qwen 3.8 27B model, the context window is set at 64k for the free tier and 128k for the paid tier. The GPT OSS 120B model offers a slightly larger range, with 65k for free users and 131k for paid subscribers. This tiered approach allows developers to scale their applications from initial testing to large-scale document processing without switching platforms.

Model Quality and Compression Transparency

A central theme in the Cerebras documentation is the preservation of model quality through rigorous architectural standards. Unlike many platforms that utilize pruned models to save on computational costs, Cerebras explicitly states that it does not host pruned models on its public endpoints. All models available are the original, unpruned versions. While the company continues to conduct research into pruning techniques—such as REAP (Router-weighted Expert Activation Pruning)—these experimental models are reserved for the research community on Hugging Face and are not part of the standard API offering.

To balance the needs of storage efficiency and model fidelity, Cerebras utilizes selective weight-only quantization. This process involves storing weights in partial 16-bit, 8-bit, or 4-bit formats, which aligns with current industry standards. However, the platform distinguishes itself by ensuring that quality-sensitive layers are stored at full precision. By performing dequantization on the fly, Cerebras ensures that the actual operations—including activations and attention mechanisms—are conducted in high precision. This technical choice is designed to preserve the maximal quality of the original model while optimizing the underlying hardware storage.

Infrastructure and Deployment Options

Cerebras provides a structured path for developers to transition from experimentation to production. The public endpoints are designed for accessibility, offering a free trial and a pay-as-you-go tier. These are subject to rate limits and standard pricing models. For enterprises requiring more robust solutions, Cerebras offers Dedicated Endpoints. These reserved capacities provide higher throughput, additional model families, and production-level Service Level Agreements (SLAs).

The documentation also highlights a user-friendly onboarding process, including a Quickstart guide for making initial API calls and a model selection guide to help users choose the appropriate model based on their specific use case. This infrastructure is built to handle the complexities of modern LLMs while providing the transparency needed for developers to understand exactly how their models are being stored and executed.

Industry Impact

The availability of Qwen 3.8 27B and GPT OSS 120B at these speeds significantly lowers the barrier for high-performance AI integration. By delivering thousands of tokens per second, Cerebras is enabling a new class of responsive AI applications that were previously hindered by the slow inference speeds of large-scale models. Furthermore, the commitment to unpruned models and high-precision operations sets a high standard for quality in the inference-as-a-service market. This approach ensures that the intelligence of the original open-source models is not compromised for the sake of speed, providing a reliable foundation for developers who prioritize accuracy and architectural integrity.

Frequently Asked Questions

Question: What are the specific inference speeds for the new models on Cerebras?

Answer: The Qwen 3.8 27B model achieves approximately 1500 tokens per second, while the OpenAI GPT OSS 120B model reaches approximately 3000 tokens per second.

Question: Does Cerebras use pruned models for its public API endpoints?

Answer: No. Cerebras hosts only the original, unpruned versions of models on its public endpoints to preserve maximal quality. While they research pruning techniques like REAP, those models are not available through the shared API.

Question: How does Cerebras handle model quantization without losing quality?

Answer: Cerebras uses selective weight-only quantization during storage (16-bit, 8-bit, or 4-bit). Sensitive layers are kept at full precision, and dequantization is performed on the fly so that operations are executed in high precision.

Related News

Evaluating AI in Electronic Design: How GPT-6 Astra and EEBench Are Shaping Circuit Board Engineering
Industry News

Evaluating AI in Electronic Design: How GPT-6 Astra and EEBench Are Shaping Circuit Board Engineering

The recent demonstration of OpenAI's GPT-6 Astra working within KiCad has sparked a significant discussion regarding the current capabilities of AI in the field of electronics design. While modern AI models possess extensive theoretical knowledge derived from textbooks and datasheets, their practical application in traditional graphical CAD tools remains limited by interface complexities. EEBench introduces a shift toward declarative code using the "atopile" framework, allowing AI agents to interact directly with electrical constraints and components rather than navigating complex GUIs. This approach facilitates automated simulations and iterative design improvements, moving closer to functional hardware engineering. By focusing on code-based design, benchmarks like EEBench can more accurately measure an AI's engineering logic, as seen in tasks involving residential energy meters and hold-up circuits, highlighting the transition from simple visual drawing to robust electronic design automation.

OpenAI Unveils GPT-6 Astra and Proclaims the Commencement of the AGI Era
Industry News

OpenAI Unveils GPT-6 Astra and Proclaims the Commencement of the AGI Era

In a landmark announcement, OpenAI has introduced its latest flagship model, GPT-6 Astra, while simultaneously declaring that the world has officially entered the "AGI era." This development, featured on The Vergecast, marks a significant shift in the company's positioning of its technology. The announcement was accompanied by news of a strategic acquisition by Nvidia, highlighting the rapid evolution of the AI industry's infrastructure. Senior AI reporter Hayden Field and a panel of experts discussed the implications of these claims, focusing on the subjective definition of Artificial General Intelligence and what this transition means for the future of technology. The release of GPT-6 Astra is framed not just as a technical update, but as the realization of a long-held industry goal.

Microsoft Defends Copilot in Copyright Lawsuit Claiming Minimal Reproduction of New York Times Content
Industry News

Microsoft Defends Copilot in Copyright Lawsuit Claiming Minimal Reproduction of New York Times Content

Microsoft has filed new legal documents in its ongoing copyright battle against The New York Times and several book authors, asserting that its AI chatbot, Copilot, rarely reproduces full sentences or significant portions of copyrighted material. The tech giant argues that the tool does not serve as a substitute for original news articles or books. As part of the discovery process, Microsoft provided 8.2 million Copilot interaction records to demonstrate that users are not utilizing the AI to bypass original sources. This defense aims to undermine claims that AI models infringe on intellectual property by providing verbatim excerpts that could replace the need for the original content.