Onton's Neurosymbolic AI Ontology 1 Beats Google Shopping and Amazon on Product Search Accuracy While Indexing Just 1% of Their Catalogs
How Onton's Ontology 1 neurosymbolic product search model works, what its benchmark results against Google Shopping and Amazon mean, and why an inspectable product-knowledge layer matters for agentic commerce.
Key Takeaways
- Onton, a San Francisco search startup, has released Ontology 1, a neurosymbolic product search model. On a 90-query benchmark it scored P@10 0.630, beating Google Shopping (0.543) and Amazon (0.469)
- It achieved this accuracy while indexing only about 1% of those catalogs. Instead of matching labels, it decomposes vague intent such as "pet-friendly sectional" into verifiable attributes and reasons over them
- In an era where AI agents choose products, a trust layer for product understanding that is independent of labels and ads becomes a competitive asset. E-commerce operators should now audit whether their product data holds up as raw material for machine reasoning
Onton Releases Ontology 1 and Outscores Google and Amazon on a 90-Query Search Benchmark

Onton's Ontology 1 scores P@10 0.630 on Subtext-Decor-90, beating Google Shopping 0.543 and Amazon 0.469 on intent-heavy queries
www.marktechpost.comOnton, a San Francisco-based search and discovery company, released its neurosymbolic product search model Ontology 1 on July 29, 2026. On Subtext-Decor-90, a 90-query benchmark the company built itself, Ontology 1 reached a precision at 10 (P@10) of 0.630. On the same queries, Google Shopping scored 0.543 and Amazon 0.469. Counting per-query wins, Onton took 52 of 90 queries outright, against 19 for Google and 16 for Amazon.
The asymmetry of conditions is what stands out. Onton indexes a single vertical, home decor and furniture, at roughly 1% of the catalog size of Amazon or Google. According to the primary research post, Amazon carries around 600 million unique products, meaning Onton beat it on precision with about one hundredth of the inventory.
Onton announced a 7.5 million dollar seed round led by Footwork in November 2025, bringing total funding to roughly 10 million dollars, and reports more than 2 million monthly users. Ontology 1 is the company's first systematic disclosure of its technical foundation. The model is live for consumers at Onton.com, and partner access is granted case by case to teams building on the agentic web.
Why Keyword and Vector Search Break on Intent-Heavy Queries
The e-commerce search experience has barely changed in nearly 30 years: a category page, a product page, a checkout page. Users search a category, then narrow with filters for size, price, material, and brand. The mechanism works as long as intent maps cleanly onto attributes. The trouble is how abruptly it collapses the moment intent does not. There is no filter for "pet-friendly," and no query condition for "furniture that fits my room."
These intent-heavy queries are becoming the mainstream rather than the exception. A 2025 Bloomreach study found that 57% of shoppers have used AI to help them shop and 41% now search in natural language rather than keywords. A 2026 Klaviyo study likewise found that daily AI users routinely search with queries of eight or more words, sometimes entire paragraphs. The longer and vaguer the words thrown into the search box, the worse a stack built on keyword matching and vector similarity performs.
Subtext-Decor-90 targets exactly this territory. Its queries include "lighting that makes my apartment feel like a Tokyo cocktail bar at 11pm," "rug that hides cat puke but isn't beige," and "a laundry hamper I won't hate looking at for 10 years," blending aesthetic judgment, negation, cultural references, and emotional framing. Google ignored the negation and returned beige rugs; Amazon returned zero results for "a sofa my husband won't call feminine and I won't call a man cave." Onton scored 0.93, 0.97, and 0.60 respectively on these queries.
An Inspectable World Model: How Ontology 1 Works
The design philosophy of Ontology 1 comes down to one principle: do not trust the label. For "pet-friendly sectional," existing engines check whether the phrase appears in the title or description. But the seller may never have written it, and even if written, it may not be true. Ontology 1 instead reasons from attributes that are easier to verify objectively, such as fiber, weave, and construction, and flags claims the product data contradicts. The company says it also weighs source credibility, accounting for listings that game the algorithm and reviews that are bought.
Its way of learning also contrasts with large language models. An LLM absorbs patterns into weights that cannot be inspected from the outside, while Ontology 1 builds an explicit knowledge graph of how things relate and why products are the way they are. The advantage of this approach is that the system can recognize holes in its own world model. When it has no account of "pet-friendly," it registers the gap, decomposes the concept into cleanability and durability, and works out that polyester upholstery is a strong indicator. The learned relationship is reused on later queries such as "pet-friendly chair" or "cleanable blue couch," and the loop runs continuously so accuracy compounds with use.
The reasoning is powered by Ograph, a custom graph database. Onton reports that a single Ograph core beats SuiteSparse:GraphBLAS running on 14 cores, roughly 100 times the throughput per core. The GPU build runs 43 times faster than the CPU variant, with early runs touching 1,000 times as the implementation is tuned. This is infrastructure investment aimed at running multi-step reasoning at practical speed, and the consistency of the neurosymbolic design choice all the way down to the database layer deserves credit.
How Far Should You Trust the Benchmark
The numbers require caveats. First, Subtext-Decor-90 was designed and executed by Onton itself. Data and code are published on Hugging Face, so re-verification is possible, but the possibility that query selection favors Onton's strengths cannot be excluded. The e-commerce trade publication Shopifreaks, while relaying the company's claim of beating competitors on nearly every dimension, explicitly noted that the benchmarks are Onton's own.
Judge reliability has limits as well. Scoring was done by three multimodal LLMs, Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5, but Krippendorff's alpha across the judges is 0.465, meaning absolute values carry judge-dependent noise. All three judges did rank the engines in the same order, so the ordering is credible, but taking a headline multiplier like "2.7x more accurate" at face value is risky.
The company also disclosed where it loses. On "a lamp that won't wake my partner if I read at 3am," Amazon scored 0.9 to Onton's 0.4, and on "something to put on a weirdly deep windowsill," the gap was 0.67 to 0.07. On queries dominated by functional specifications, Amazon's category metadata and sheer catalog breadth win. Indexing 1% of the catalog looks like a strength in the precision framing, but it also means recall was never measured. As a single-vertical system, this is not yet a replacement for general-purpose e-commerce search.
What This Means for Agentic Commerce
The release still matters because it targets a weak point of the era in which AI agents shop on our behalf. Agents work through review roundups, influencer picks, and sponsored comparisons, and the concern is that they recommend products without being able to tell genuine information from synthetic, paid, or manipulated content. Onton itself argues that trust and authenticity gaps are part of why the first wave of agentic commerce in 2025 underdelivered, and positions Ontology 1 as a grounding layer for shopping agents.
For e-commerce operators, the implications come in two stages. In the short term, audit how badly your site search degrades on long, natural-language queries. Consumer search behavior has already shifted toward conversation, and a search box that cannot absorb intent-heavy queries quietly leaks revenue. In the medium term, the verifiability of product data becomes the differentiator. Because models like Ontology 1 reason from materials and construction rather than seller labels, merchants with accurate, structured attribute data gain an advantage in AI-mediated product discovery. Exaggerated titles and keyword-stuffed descriptions may not just fail to help discoverability in a reasoning-based search world; they risk being flagged as contradictions.
Pricing for external access to Ontology 1 and any timeline for a public API remain undisclosed. The only adoption path today is a negotiated partnership, and the model weights are not open. For teams evaluating it, the practical first step is to test queries close to your own category on Onton.com.
Conclusion
Ontology 1 presents an alternative to product search built on keywords and vector similarity: reasoning over a knowledge graph combined with continuous self-learning. With the caveats that the benchmark is self-produced and judge agreement is modest, beating Google Shopping and Amazon while indexing 1% of their catalogs is strong evidence that there are domains where depth of product understanding beats catalog scale.
For agentic commerce to work, agents need a layer of product information they can trust. That such a candidate has now appeared outside the LLM, in the form of an inspectable world model, suggests the menu of search infrastructure choices is about to diversify. Whichever way it plays out, the move that pays off for e-commerce operators is the same: invest in accurate, structured product data.


