Back to list
Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost
Industry NewsAI AgentsOpen SourceDatabase Technology

Castform and Neon Outperform GPT-5.6 Sol in Retrieval Accuracy at 100x Lower Cost

A breakthrough collaboration between Castform and Neon has demonstrated that a 4B open-source model, post-trained with Reinforcement Learning (RL), can match the retrieval accuracy of frontier models like GPT-5.6 Sol. This approach addresses the primary challenges of modern AI agents: providing the right context and enabling intelligent search decisions. By leveraging Neon’s Lakebase Postgres and Castform’s specialized post-training, the solution achieves a 100x cost reduction compared to frontier models. While multi-turn search requests using GPT-5.6 Sol typically cost $0.03 and take over 10 seconds, the Castform-Neon synergy offers a significantly more efficient and scalable alternative for agentic retrieval, marking a shift from traditional RAG pipelines to complex, multi-hop search workflows.

Hacker News

Key Takeaways

  • Performance Parity: A 4B open-source model post-trained with Castform achieves search retrieval accuracy equal to GPT-5.6 Sol.
  • Massive Cost Efficiency: The specialized open-source approach is 100x cheaper than using frontier models for the same tasks.
  • Infrastructure Synergy: The solution combines Neon’s Lakebase Postgres for data context with Castform’s model logic for search decision-making.
  • Evolution of Search: The industry is moving from 2022-era one-shot RAG pipelines to 2025-era multi-hop agentic retrieval loops.
  • Latency Reduction: The new method addresses the >10s latency issues common in multi-turn search requests handled by frontier models.

In-Depth Analysis

The Evolution from RAG to Agentic Retrieval

The landscape of AI-driven data retrieval has undergone a significant transformation over the last few years. In approximately 2022, the industry focused heavily on embedding search, with database providers like Neon integrating extensions such as pgvector to support Retrieval-Augmented Generation (RAG). These early systems were largely "one-shot," where a single query was issued to find relevant context for a Large Language Model (LLM).

By 2025, the paradigm shifted toward agentic search. In this new model, developers create multi-hop workflows where agents decompose complex problems into smaller, manageable tasks. Instead of a single search, the model operates in a loop, planning and searching multiple times to refine results. However, this iterative process introduced a critical bottleneck: every loop iteration required a call to a frontier model, leading to exponential increases in both cost and latency. A typical multi-turn search with a model like GPT-5.6 Sol can cost roughly $0.03 and take more than 10 seconds, making it impractical for high-scale applications.

Bridging the Gap with RL Post-Training

While small open-weights models are inherently 100x cheaper than frontier models, their out-of-the-box capabilities often lag behind closed API models. The collaboration between Castform and Neon proves that this gap can be bridged through Reinforcement Learning (RL) post-training. By post-training a 4B model specifically for retrieval tasks, Castform has enabled it to match the accuracy of much larger, more expensive frontier models.

This strategy focuses on the two core requirements of a "good agent":

  1. Context: The ability to provide tools to find the right data, solved by Neon’s Lakebase Postgres and its new Search extensions.
  2. Model: The ability of the model to decide what to search for, solved by Castform’s post-training logic.

By pointing Castform at Neon, organizations can transform raw database data into usable intelligence without the need for advanced, expensive infrastructure or handcrafted RAG pipelines.

Infrastructure and Efficiency Gains

The primary advantage of the Castform and Neon approach lies in its ability to utilize existing database assets. As Ying Hang Seah, co-founder of Castform, noted, most teams have their best training data sitting idle in databases. The difficulty has always been turning that raw data into something usable for agents to read, search, and mutate at scale.

Neon’s Lakebase Postgres serves as the foundational infrastructure that allows agents to access data efficiently. When combined with a 4B model that has been optimized for decision-making, the system eliminates the prohibitive costs of frontier models. This allows for agentic retrieval that is not only accurate but also fast enough for real-time user requests, overcoming the 10-second latency barrier that currently plagues multi-turn frontier model workflows.

Industry Impact

The success of the Castform and Neon integration signals a major shift in how AI agents will be deployed in the enterprise. By proving that a 4B model can rival GPT-5.6 Sol in specific retrieval tasks, the industry may see a move away from general-purpose frontier models toward specialized, post-trained open-source models. This democratization of high-performance retrieval allows smaller teams to build sophisticated agents that were previously too expensive to operate. Furthermore, the emphasis on "agentic retrieval" over simple RAG suggests that the next generation of AI tools will be defined by their ability to perform complex, multi-step reasoning within a database environment at a fraction of current costs.

Frequently Asked Questions

Question: How does the cost of the 4B model compare to GPT-5.6 Sol?

According to the report, the 4B open-source model post-trained with Castform is 100x cheaper than GPT-5.6 Sol. While a typical multi-turn search with GPT-5.6 Sol costs approximately $0.03, the open-source alternative provides the same accuracy at a fraction of that price.

Question: What is the difference between traditional RAG and agentic retrieval?

Traditional RAG (common around 2022) usually involves a one-shot embedding similarity search to provide context to an LLM. Agentic retrieval (emerging in 2025) involves multi-hop workflows where the model plans and searches multiple times in a loop, decomposing large problems into smaller ones to achieve higher accuracy.

Question: What role does Neon play in this solution?

Neon provides the "Context" component through its Lakebase Postgres and Search extensions. This infrastructure allows agents to find, read, and search the right data efficiently, which is then processed by the Castform-optimized model to make search decisions.

Related News

Industry News

Atlassian and OpenAI Expand Strategic Partnership to Turn Enterprise Knowledge into Action Across Team Workflows

Atlassian and OpenAI have announced an expansion of their strategic partnership, aimed at connecting frontier artificial intelligence models with enterprise knowledge to empower organizations across their operational lifecycles. By integrating cutting-edge frontier model capabilities directly with institutional context, the collaboration is designed to help teams seamlessly plan, build, and deliver work. The initiative addresses a critical gap in enterprise operations: moving beyond passive information retrieval to active, context-aware execution. Rather than treating organizational knowledge as static repositories, the joint effort seeks to transform institutional data into actionable workflows, enabling cross-functional teams to streamline project management, improve collaborative alignment, and accelerate delivery outcomes. This strategic move marks a meaningful step forward in embedding frontier AI into everyday enterprise tools and critical business processes.

Singapore Security Firm V-Key Takes Stake in CloudsineAI to Unify Cryptographic Identity and AI Defense
Industry News

Singapore Security Firm V-Key Takes Stake in CloudsineAI to Unify Cryptographic Identity and AI Defense

Singapore-based digital security firm V-Key has officially taken a stake in CloudsineAI, marking a significant strategic move aimed at unifying digital trust with artificial intelligence defenses. Under the agreement, the two technology companies announced plans to integrate V-Key's established identity verification and cryptographic tools directly with CloudsineAI's web integrity and dedicated AI security solutions. By joining forces, the organizations aim to deliver an integrated defense architecture capable of safeguarding both traditional web infrastructure and modern artificial intelligence environments. While specific transactional figures and financial valuations were not disclosed in the initial report, the collaboration highlights an intensifying industry focus on combining identity verification with AI-specific cybersecurity tools to mitigate emerging technological threats across mission-critical systems.

Industry News

How Jump Trading Scales Quantitative Research Using OpenAI ChatGPT and Long-Running Workflows

Jump Trading is leveraging OpenAI's ChatGPT technology to significantly scale and expand its quantitative research operations. According to an announcement from OpenAI, the initiative centers on deploying longer-running artificial intelligence workflows engineered to synthesize and analyze information across multiple diverse data sources. Crucially, these automated research pipelines are paired with human review to maintain high standards of precision and oversight. By integrating AI-driven workflows into quantitative research, Jump Trading illustrates how modern financial firms are augmenting analytical operations with advanced language models. The strategic development underscores a broader trend where autonomous, extended AI tasks operate in tandem with domain experts to process complex financial information effectively.