Back to List
Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost
Industry NewsAI AgentsOpen SourceDatabase Technology

Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost

A breakthrough collaboration between Castform and Neon has demonstrated that a 4B open-source model, post-trained with Reinforcement Learning (RL), can match the retrieval accuracy of frontier models like GPT-5.6 Sol. This approach addresses the primary challenges of modern AI agents: providing the right context and enabling intelligent search decisions. By leveraging Neon’s Lakebase Postgres and Castform’s specialized post-training, the solution achieves a 100x cost reduction compared to frontier models. While multi-turn search requests using GPT-5.6 Sol typically cost $0.03 and take over 10 seconds, the Castform-Neon synergy offers a significantly more efficient and scalable alternative for agentic retrieval, marking a shift from traditional RAG pipelines to complex, multi-hop search workflows.

Hacker News

Key Takeaways

  • Performance Parity: A 4B open-source model post-trained with Castform achieves search retrieval accuracy equal to GPT-5.6 Sol.
  • Massive Cost Efficiency: The specialized open-source approach is 100x cheaper than using frontier models for the same tasks.
  • Infrastructure Synergy: The solution combines Neon’s Lakebase Postgres for data context with Castform’s model logic for search decision-making.
  • Evolution of Search: The industry is moving from 2022-era one-shot RAG pipelines to 2025-era multi-hop agentic retrieval loops.
  • Latency Reduction: The new method addresses the >10s latency issues common in multi-turn search requests handled by frontier models.

In-Depth Analysis

The Evolution from RAG to Agentic Retrieval

The landscape of AI-driven data retrieval has undergone a significant transformation over the last few years. In approximately 2022, the industry focused heavily on embedding search, with database providers like Neon integrating extensions such as pgvector to support Retrieval-Augmented Generation (RAG). These early systems were largely "one-shot," where a single query was issued to find relevant context for a Large Language Model (LLM).

By 2025, the paradigm shifted toward agentic search. In this new model, developers create multi-hop workflows where agents decompose complex problems into smaller, manageable tasks. Instead of a single search, the model operates in a loop, planning and searching multiple times to refine results. However, this iterative process introduced a critical bottleneck: every loop iteration required a call to a frontier model, leading to exponential increases in both cost and latency. A typical multi-turn search with a model like GPT-5.6 Sol can cost roughly $0.03 and take more than 10 seconds, making it impractical for high-scale applications.

Bridging the Gap with RL Post-Training

While small open-weights models are inherently 100x cheaper than frontier models, their out-of-the-box capabilities often lag behind closed API models. The collaboration between Castform and Neon proves that this gap can be bridged through Reinforcement Learning (RL) post-training. By post-training a 4B model specifically for retrieval tasks, Castform has enabled it to match the accuracy of much larger, more expensive frontier models.

This strategy focuses on the two core requirements of a "good agent":

  1. Context: The ability to provide tools to find the right data, solved by Neon’s Lakebase Postgres and its new Search extensions.
  2. Model: The ability of the model to decide what to search for, solved by Castform’s post-training logic.

By pointing Castform at Neon, organizations can transform raw database data into usable intelligence without the need for advanced, expensive infrastructure or handcrafted RAG pipelines.

Infrastructure and Efficiency Gains

The primary advantage of the Castform and Neon approach lies in its ability to utilize existing database assets. As Ying Hang Seah, co-founder of Castform, noted, most teams have their best training data sitting idle in databases. The difficulty has always been turning that raw data into something usable for agents to read, search, and mutate at scale.

Neon’s Lakebase Postgres serves as the foundational infrastructure that allows agents to access data efficiently. When combined with a 4B model that has been optimized for decision-making, the system eliminates the prohibitive costs of frontier models. This allows for agentic retrieval that is not only accurate but also fast enough for real-time user requests, overcoming the 10-second latency barrier that currently plagues multi-turn frontier model workflows.

Industry Impact

The success of the Castform and Neon integration signals a major shift in how AI agents will be deployed in the enterprise. By proving that a 4B model can rival GPT-5.6 Sol in specific retrieval tasks, the industry may see a move away from general-purpose frontier models toward specialized, post-trained open-source models. This democratization of high-performance retrieval allows smaller teams to build sophisticated agents that were previously too expensive to operate. Furthermore, the emphasis on "agentic retrieval" over simple RAG suggests that the next generation of AI tools will be defined by their ability to perform complex, multi-step reasoning within a database environment at a fraction of current costs.

Frequently Asked Questions

Question: How does the cost of the 4B model compare to GPT-5.6 Sol?

According to the report, the 4B open-source model post-trained with Castform is 100x cheaper than GPT-5.6 Sol. While a typical multi-turn search with GPT-5.6 Sol costs approximately $0.03, the open-source alternative provides the same accuracy at a fraction of that price.

Question: What is the difference between traditional RAG and agentic retrieval?

Traditional RAG (common around 2022) usually involves a one-shot embedding similarity search to provide context to an LLM. Agentic retrieval (emerging in 2025) involves multi-hop workflows where the model plans and searches multiple times in a loop, decomposing large problems into smaller ones to achieve higher accuracy.

Question: What role does Neon play in this solution?

Neon provides the "Context" component through its Lakebase Postgres and Search extensions. This infrastructure allows agents to find, read, and search the right data efficiently, which is then processed by the Castform-optimized model to make search decisions.

Related News

Microsoft Reports $24.1 Billion in OpenAI-Linked Revenue, Dominating Over Half of Its AI Business
Industry News

Microsoft Reports $24.1 Billion in OpenAI-Linked Revenue, Dominating Over Half of Its AI Business

Microsoft has disclosed a significant financial milestone, reporting $24.1 billion in revenue directly linked to its partnership with OpenAI. This figure represents a pivotal shift in the company's financial structure, as OpenAI-related contributions now account for more than half of Microsoft's total AI-driven business for the period. The data underscores the immense commercial success of the Microsoft-OpenAI alliance and highlights the rapid enterprise adoption of generative AI technologies. As this partnership becomes the primary engine for Microsoft's AI growth, it sets a new benchmark for the industry regarding the monetization of advanced artificial intelligence models and the strategic value of deep-tech collaborations.

The Paradox of Typography: Analyzing the Emotional Impact of Fixed-Width Fonts and Blade Runner Title Cards
Industry News

The Paradox of Typography: Analyzing the Emotional Impact of Fixed-Width Fonts and Blade Runner Title Cards

This analysis explores the intricate relationship between functional design and emotional resonance in typography, as discussed in the context of developer environments and cinematic history. The article examines the 'typography paradox'—the idea that while well-designed type should be invisible to the reader, it inevitably conveys a specific 'feeling.' By looking at the author's experience using Claude Code within the Ghostty terminal on macOS, the piece highlights the importance of fixed-width typefaces like Apple's SF Mono. It further bridges the gap between technical utility and artistic expression by referencing the iconic title cards of Blade Runner, suggesting that even in data-heavy or structural environments, the visual form of letters builds a personal and emotional impression that transcends mere information delivery.

NVIDIA Vera Whitepaper Analysis: Examining the Olympus Core Architecture and Marketing Claims Against x86 Standards
Industry News

NVIDIA Vera Whitepaper Analysis: Examining the Olympus Core Architecture and Marketing Claims Against x86 Standards

NVIDIA has released a detailed 45-page whitepaper for Vera, its inaugural server CPU powered by the custom-designed Olympus core. The technical specifications reveal a formidable 88-core monolithic compute die utilizing the Arm v9.2 architecture, featuring a 10-wide decode front end, value prediction, and a substantial cache hierarchy. Despite the impressive hardware—which includes a 1.2 TB/s memory interface and a 3.4 TB/s coherency fabric—the whitepaper has drawn criticism for its marketing narrative. Analysts point out that NVIDIA's documentation mischaracterizes established x86 technologies, such as simultaneous multithreading and NUMA topologies, while employing unconventional metrics like "agentic benchmarks." This analysis explores the tension between Vera's genuine architectural innovations and the controversial storytelling used to promote it.