Back to list
Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost
Industry NewsAI AgentsOpen SourceDatabase Technology

Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost

A breakthrough collaboration between Castform and Neon has demonstrated that a 4B open-source model, post-trained with Reinforcement Learning (RL), can match the retrieval accuracy of frontier models like GPT-5.6 Sol. This approach addresses the primary challenges of modern AI agents: providing the right context and enabling intelligent search decisions. By leveraging Neon’s Lakebase Postgres and Castform’s specialized post-training, the solution achieves a 100x cost reduction compared to frontier models. While multi-turn search requests using GPT-5.6 Sol typically cost $0.03 and take over 10 seconds, the Castform-Neon synergy offers a significantly more efficient and scalable alternative for agentic retrieval, marking a shift from traditional RAG pipelines to complex, multi-hop search workflows.

Hacker News

Key Takeaways

  • Performance Parity: A 4B open-source model post-trained with Castform achieves search retrieval accuracy equal to GPT-5.6 Sol.
  • Massive Cost Efficiency: The specialized open-source approach is 100x cheaper than using frontier models for the same tasks.
  • Infrastructure Synergy: The solution combines Neon’s Lakebase Postgres for data context with Castform’s model logic for search decision-making.
  • Evolution of Search: The industry is moving from 2022-era one-shot RAG pipelines to 2025-era multi-hop agentic retrieval loops.
  • Latency Reduction: The new method addresses the >10s latency issues common in multi-turn search requests handled by frontier models.

In-Depth Analysis

The Evolution from RAG to Agentic Retrieval

The landscape of AI-driven data retrieval has undergone a significant transformation over the last few years. In approximately 2022, the industry focused heavily on embedding search, with database providers like Neon integrating extensions such as pgvector to support Retrieval-Augmented Generation (RAG). These early systems were largely "one-shot," where a single query was issued to find relevant context for a Large Language Model (LLM).

By 2025, the paradigm shifted toward agentic search. In this new model, developers create multi-hop workflows where agents decompose complex problems into smaller, manageable tasks. Instead of a single search, the model operates in a loop, planning and searching multiple times to refine results. However, this iterative process introduced a critical bottleneck: every loop iteration required a call to a frontier model, leading to exponential increases in both cost and latency. A typical multi-turn search with a model like GPT-5.6 Sol can cost roughly $0.03 and take more than 10 seconds, making it impractical for high-scale applications.

Bridging the Gap with RL Post-Training

While small open-weights models are inherently 100x cheaper than frontier models, their out-of-the-box capabilities often lag behind closed API models. The collaboration between Castform and Neon proves that this gap can be bridged through Reinforcement Learning (RL) post-training. By post-training a 4B model specifically for retrieval tasks, Castform has enabled it to match the accuracy of much larger, more expensive frontier models.

This strategy focuses on the two core requirements of a "good agent":

  1. Context: The ability to provide tools to find the right data, solved by Neon’s Lakebase Postgres and its new Search extensions.
  2. Model: The ability of the model to decide what to search for, solved by Castform’s post-training logic.

By pointing Castform at Neon, organizations can transform raw database data into usable intelligence without the need for advanced, expensive infrastructure or handcrafted RAG pipelines.

Infrastructure and Efficiency Gains

The primary advantage of the Castform and Neon approach lies in its ability to utilize existing database assets. As Ying Hang Seah, co-founder of Castform, noted, most teams have their best training data sitting idle in databases. The difficulty has always been turning that raw data into something usable for agents to read, search, and mutate at scale.

Neon’s Lakebase Postgres serves as the foundational infrastructure that allows agents to access data efficiently. When combined with a 4B model that has been optimized for decision-making, the system eliminates the prohibitive costs of frontier models. This allows for agentic retrieval that is not only accurate but also fast enough for real-time user requests, overcoming the 10-second latency barrier that currently plagues multi-turn frontier model workflows.

Industry Impact

The success of the Castform and Neon integration signals a major shift in how AI agents will be deployed in the enterprise. By proving that a 4B model can rival GPT-5.6 Sol in specific retrieval tasks, the industry may see a move away from general-purpose frontier models toward specialized, post-trained open-source models. This democratization of high-performance retrieval allows smaller teams to build sophisticated agents that were previously too expensive to operate. Furthermore, the emphasis on "agentic retrieval" over simple RAG suggests that the next generation of AI tools will be defined by their ability to perform complex, multi-step reasoning within a database environment at a fraction of current costs.

Frequently Asked Questions

Question: How does the cost of the 4B model compare to GPT-5.6 Sol?

According to the report, the 4B open-source model post-trained with Castform is 100x cheaper than GPT-5.6 Sol. While a typical multi-turn search with GPT-5.6 Sol costs approximately $0.03, the open-source alternative provides the same accuracy at a fraction of that price.

Question: What is the difference between traditional RAG and agentic retrieval?

Traditional RAG (common around 2022) usually involves a one-shot embedding similarity search to provide context to an LLM. Agentic retrieval (emerging in 2025) involves multi-hop workflows where the model plans and searches multiple times in a loop, decomposing large problems into smaller ones to achieve higher accuracy.

Question: What role does Neon play in this solution?

Neon provides the "Context" component through its Lakebase Postgres and Search extensions. This infrastructure allows agents to find, read, and search the right data efficiently, which is then processed by the Castform-optimized model to make search decisions.

Related News

Microsoft Sets October 7 Windows and Surface Event in San Francisco to Outline Local AI Future
Industry News

Microsoft Sets October 7 Windows and Surface Event in San Francisco to Outline Local AI Future

Microsoft has officially scheduled a major Windows and Surface event for October 7th in San Francisco, marking its first major Windows gathering in more than two years. According to an announcement reported by The Verge, the upcoming presentation will center on outlining the future trajectory of the Windows operating system alongside its Surface hardware lineup. A central theme highlighted by Microsoft is a dedicated conversation exploring how local artificial intelligence will shape the next chapter of computing devices and software platforms. Coming after a prolonged hiatus since the company's last major Windows showcase, this event represents a pivotal milestone for Microsoft as it connects its hardware roadmap directly with on-device artificial intelligence capabilities.

GoTo Adopts Pragmatic AI Strategy Focused on Conversion and Cost Efficiency Ahead of 2027 Rollout
Industry News

GoTo Adopts Pragmatic AI Strategy Focused on Conversion and Cost Efficiency Ahead of 2027 Rollout

GoTo, the parent company of Gojek, is pursuing a grounded and practical approach to artificial intelligence rather than aiming for grandiose, far-reaching initiatives. Characterizing its current posture as 'not trying to solve world hunger,' the Southeast Asian tech group is deliberately prioritizing pragmatic AI implementations capable of delivering tangible commercial results. Specifically, GoTo's immediate operational focus centers on deploying artificial intelligence solutions that directly enhance conversion rates or drive cost reductions across its business. This measured, ROI-driven strategy serves as the foundation leading up to an anticipated wider deployment of AI capabilities scheduled for 2027. By concentrating strictly on bottom-line efficiencies and revenue conversion ahead of broader expansion, GoTo highlights an industry trend toward financial discipline in enterprise artificial intelligence adoption.

Vietjet and Thales Partner on Aircraft Maintenance, Digital Aviation, Cybersecurity, and Artificial Intelligence Operations
Industry News

Vietjet and Thales Partner on Aircraft Maintenance, Digital Aviation, Cybersecurity, and Artificial Intelligence Operations

Vietjet and Thales have signed strategic cooperation agreements covering aircraft maintenance and modern operational technologies. The collaboration between the airline and the global technology group extends across several critical domains, including aircraft maintenance services, digital aviation, connectivity solutions, artificial intelligence, and cybersecurity designed for airline operations. By uniting foundational maintenance needs with advanced digital capabilities, the agreements reflect a multifaceted approach to modernizing airline operational infrastructure. The partnership establishes a collaborative framework centered on combining physical fleet reliability with intelligent digital technologies and resilient operational security.