Home / Blog / Search agent beats GPT-6 Astra on benchmarks, just days…
Tech News

Search agent beats GPT-6 Astra on benchmarks, just days after release

Points: 1 # Comments: 0 Search agent beats GPT-6 Astra on benchmarks, just days after release Coverage based on HN AI Agents reporting.

By Dillip Chowdary • Sep 06, 2026 • Source: HN AI Agents

Search agent beats GPT-6 Astra on benchmarks, just days after release

What happened

A newly launched shopping agent called Findcheap scored 92.8% savings capture on PriceBench, a task-specific benchmark for agentic product search, while GPT-6 Astra running at high reasoning effort achieved only 43.9% on the same test. The difference is more than a doubled score: Findcheap completed each search in a median of 15.6 seconds and at a cost of $0.03 per query, against GPT-6 Astra's 87.5 seconds and $1.32. Perplexity, Claude Opus 5, and Gemini 3.1 Pro also ran in PriceBench and finished at 34.1%, 20.2%, and 10.7% savings capture respectively.

This piece walks through what Findcheap is, how PriceBench measures it, and what the results mean for anyone building or evaluating agentic search tools. Developers integrating shopping agents into products and researchers tracking the narrow-AI-versus-frontier-model tradeoff will find the most relevance here.

Findcheap published benchmark results showing its AI agent outperforming every other tested system on product-search savings capture. On the equivalent-product track — where the agent must find a functionally similar item for less — Findcheap reached 92.8% while GPT-6 Astra, the next best, reached 43.9%. On the exact-product track — find the same brand and model at the lowest available price — Findcheap scored 79.6% against GPT-6 Astra's 52.5%. Perplexity hit 20.7% on the exact track, Claude Opus 5 reached 24.9%, and Gemini 3.1 Pro remained at 10.7% across both tracks.

How it works

The company also announced two simultaneous product launches: a Chrome extension for consumers and an API for developers. The Chrome extension is free and installs directly from the Chrome Web Store. The API carries the per-query cost visible in PriceBench: $0.03, versus $1.32 for GPT-6 Astra at equivalent task completion. Findcheap claims its agent finds cheaper options for 94% of products on the internet and saves users 60.4% on average.

Search agent beats GPT-6 Astra on benchmarks, just days after release
Illustration · Pexels

Findcheap describes its approach as multimodal agentic search — the system interprets product pages visually and textually, identifies comparable or identical items across retailers, and returns a ranked result. The agent runs autonomously once triggered by a product page detection event in the browser. Rather than relying on a general-purpose frontier model to handle every possible task, Findcheap was designed to specialize in product search as a narrow subfield of agentic AI, trading breadth for depth and speed in that single domain.

Why it matters

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

The PriceBench evaluation used a 100-product human-verified test set. Tasks were sourced to reflect actual e-commerce distribution: Amazon accounts for roughly 42% of US e-commerce sales, so 42 of the 100 input products came from Amazon. A three-model judging panel composed of GPT-5.6 Sol, Claude Opus 5, and Gemini 3.5 Flash reviewed results autonomously to reduce human bias. PriceBench does not use a fixed task set, so overfitting to known inputs is structurally prevented.

The gap between Findcheap's 92.8% and GPT-6 Astra's 43.9% on the equivalent-product track — a 111.4% relative improvement, as Findcheap describes it — is large enough to change a product decision, not just a benchmark leaderboard ranking. The simultaneous latency gap, from 87.5 seconds to 15.6 seconds, is comparably significant for any user-facing integration where session continuity matters. Together, those two numbers make the case that task-specific agents can outperform general reasoning models on well-scoped search problems without needing the same compute budget.

The cost differential reinforces that point. At $0.03 per search versus $1.32 for GPT-6 Astra, a builder integrating Findcheap's API across high-query-volume use cases faces a roughly 44-times lower inference cost on the task PriceBench measures. Whether that efficiency holds across categories and retailers outside PriceBench's 100-product sample is a separate question, but the numbers are published, the benchmark is open source, and independent replication is possible.

Who is affected

Online shoppers who install the Chrome extension gain an always-on agent that detects product pages and surfaces cheaper alternatives without any search query required. Findcheap says it finds lower prices for 94% of products and saves users an average of 60.4%, so anyone buying consumer goods regularly on the web is the primary beneficiary. The tool works passively, which distinguishes it from search-style tools where a user must actively initiate a query.

Developers building commerce tooling are the second audience. The API's $0.03-per-search price point, combined with the PriceBench results, gives an engineering team concrete inputs for a build-versus-buy calculation. The Findcheap team also frames its bet in structural terms: they argue the $36 trillion global e-commerce industry is rebuildable from agentic-commerce primitives, positioning the API as infrastructure rather than a consumer feature dressed as a service.

What to watch next

PriceBench itself is the most important variable to track. Because it sources tasks at runtime rather than from a fixed list, leaderboard positions can shift as the benchmark runs more products and more agents. The benchmark is open source on GitHub under the llmbender organization, so independent researchers and competing teams can run their own evaluations. Builders considering the Findcheap API should run PriceBench themselves against their actual product categories before committing to the 94% and 60.4% figures, which come from Findcheap's own testing.

The company describes additional products as coming soon without specifying what they are. The two currently announced — the Chrome extension and the developer API — cover consumer and B2B access to the same underlying agent. Observing whether the savings-capture advantage holds on categories underrepresented in the current 100-product set, and whether latency remains near 15.6 seconds as query volume scales, will tell builders whether PriceBench results translate to production. The source article and PriceBench methodology are available at findcheap.ai and github.com/llmbender/pricebench.

Developer Action Items

  • Diff the official changelog for Claude / Gemini / Amazon 92.8 before you bump — APIs, defaults, and removed flags only.
  • Install through the vendor's documented channel in staging; keep a one-command rollback and time-box the canary.
  • Grep your repo for old flag names, lockfile pins, and plugin versions that the notes mark as breaking.
  • Prefer the first patch cut over the day-zero tag unless you have a reason to be on the leading edge.
  • If HN AI Agents did not name a region, plan, or SKU, screenshot the official availability line before you promise it to users.
Dillip Chowdary

Author

Dillip Chowdary

Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.

Related on Tech Bytes

Advertisement

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam · Unsubscribe anytime

Advertisement

✈️ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings — fit scores, job-specific resume optimization and email alerts.

Find matching jobs →

Free Tools

Browse all tools →