- 01
Of the 100 question patterns (5 retailers × 5 AIs × 4 scenarios), only 60 (60.0%) were correct against official information, with 25 incorrect (25.0%) and 15 unable (15.0%). Using generic AI answers as-is for customer guidance on policy questions still carries risk.
- 02
15 question patterns (15.0%) had a primary verdict of incorrect and carried real-harm risk — for example, guiding return/exchange deadlines or long-term-warranty years and coverage as more favorable or longer than the official scope allows.
- 03
62 question patterns (62.0%) had a source-URL problem — nonexistent, non-official, different-brand, or wrong-topic — and Microsoft Copilot did so in 19 of 20 (95.0%). Correctness alone cannot measure the safety of an answer.
- Survey name
- Home-Appliance Retailer Return & Warranty Policy AI Answer-Accuracy Survey
- Subjects
- Major AIs' answers on home-appliance retailers' return and warranty policies (returns/initial defects, long-term warranty, point refunds, warranty repair)
- Target retailers
- Bic Camera, Yodobashi Camera, Yamada Denki, K's Denki, EDION
- Target AIs
- ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot
- Survey date
- July 21, 2026
- Scope evaluated
- 100 question patterns (5 retailers × 5 AIs × 4 scenarios); 268 AI answer logs
- Items recorded
- Full answer logs, AI answer screenshots, source URLs, verdicts, and official correct-answer summaries with official screenshots
- Aggregation
- Each question pattern was decided as correct / incorrect / unable by majority across 2–3 runs, with real-harm risk and source problems tallied as additional flags
- Conducted by
- Stellagent Inc.
- 01Summary
- 02Background
- 03Survey Overview
- 04AI Services and Conditions
- 05Verdict Criteria
- 06Scenarios Used
- 07Prompts Used
- 08Main Official Sources
- 09Overall Results
- 10Correct Rate by AI Service
- 11Correct Rate by Scenario
- 12Trends by Retailer
- 13Retailer × AI Matrix
- 14Main Patterns of Incorrect Answers and Source Problems
- 15How We Extracted, Aggregated, Normalized, and Broke Ties
- 16Implications for Businesses
- 17Notes on the Survey
- 18Revision History
Summary
Stellagent Inc. investigated how accurately five major AI services can answer questions about the returns, initial defects, long-term warranties, point refunds, and warranty-repair conditions of five major home-appliance retailers. We first saved each retailer's official FAQ, terms, and warranty pages as the correct-answer baseline, then entered identically formatted prompts into ChatGPT, Claude, Gemini, Perplexity, and Microsoft Copilot, and cross-checked the answer content and source URLs.
As of July 21, 2026, this survey evaluated 100 question patterns across 5 retailers × 5 AIs × 4 scenarios. Each question pattern was tried at least twice with the same prompt, and a third run was added when the verdict wavered, real-harm risk appeared, or a source problem was found, so there are 268 AI answer logs in total.
| Metric | Result |
|---|---|
| Target retailers | 5 |
| Target AIs | 5 |
| Scenarios | 4 |
| Question patterns evaluated | 100 (5 retailers × 5 AIs × 4 questions) |
| AI answer logs | 268 (32 with two runs, 68 with three runs) |
| Primary verdict: correct | 60 (60.0%) |
| Primary verdict: incorrect | 25 (25.0%) |
| Primary verdict: unable | 15 (15.0%) |
| Incorrect and carrying real-harm risk | 15 (15.0%) |
| Additional flag: a source problem in at least one run | 62 (62.0%, overlaps allowed) |
For the 100 patterns, the primary verdict was 60 correct (60.0%), 25 incorrect (25.0%), and 15 unable (15.0%). In addition, 15 question patterns (15.0%) had a primary verdict of incorrect and carried real-harm risk. Counting nonexistent URLs, non-official URLs, and official-but-wrong-topic URLs, 62 question patterns (62.0%) contained a source problem. Because a source problem is an additional flag that spans correct, incorrect, and unable answers, it is not summed with the primary verdict.

The chart above may be reproduced as-is in media coverage, articles, and other materials, provided that you credit "Stellagent Inc." as the source and include a link to this page.
Background
Generative AI is starting to be used by ordinary consumers as a way to look into questions before and after a purchase. After buying an appliance, there are moments when confirmation based on official terms or FAQs is needed — initial defects, returns, exchanges, long-term warranties, point refunds, and on-site repairs.
For these questions, accuracy about deadlines, eligibility conditions, refund methods, and coverage matters more than convenience. When an AI presents a "plausible but wrong deadline" or "warranty terms that look more favorable than reality," a gap opens between the customer's expectations and the official handling, which can also add load to stores and support desks.
This survey is not a criticism of home-appliance retailers. Even when official information exists, a generic AI does not necessarily pick it up in the right context and summarize it correctly. We conducted this survey to verify that, and to show the need for a specialized support AI that answers based on official data.
Survey Overview
| Item | Content |
|---|---|
| Survey name | Home-Appliance Retailer Return & Warranty Policy AI Answer-Accuracy Survey |
| Conducted by | Stellagent Inc. |
| Operator | Stellagent |
| Method | AI answer experiment and desk research of official sites |
| Survey date | July 21, 2026 |
| Target retailers | Bic Camera, Yodobashi Camera, Yamada Denki, K's Denki, EDION |
| Target AIs | ChatGPT, Claude, Gemini, Perplexity, Microsoft Copilot |
| Retailer selection criteria | We targeted five major home-appliance retailers that ordinary consumers in Japan readily think of as places to buy appliances, that operate physical stores or official online stores, and that publish guidance pages on returns, warranties, points, and the like. The retailers were selected considering nationwide recognition, business scale, appliance sales via physical stores or official online stores, and the publication status of official information on returns, warranties, points, and so on. |
| AI selection criteria | We targeted major general-purpose generative AIs and answer engines that ordinary users in Japan can access via a web UI and that are commonly used for post-purchase inquiries and terms checks. The set includes search-engine-type and answer-engine-type services so that the presence and accuracy of source-URL citations can also be compared. |
| Procedure | We confirmed the official source URLs for each retailer × scenario and saved screenshots and official correct-answer summaries. Then we entered identically formatted prompts into each AI and saved the full answer log, AI answer screenshot, source URLs, and verdict for each individual answer. |
| Context-influence control | We used Temporary mode, a temporary chat, an incognito chat, or a new chat. Each prompt also explicitly stated, "Do not reference past conversations or memory." |
| Recording | Each time an answer was obtained, we recorded on the spot the full answer, timestamp, model name or display name, subscription plan, web-search setting, source URLs, and verdict. For real-harm-risk incorrect answers, nonexistent URLs, and answers that were candidates for quotation, we also saved an AI answer screenshot. |
| Number of runs | We ran each retailer × AI × scenario at least twice. When the two verdicts split, when a real-harm-risk incorrect answer appeared at least once, or when a source URL was nonexistent or pointed to a page different from the official one, we obtained a third run as a pre-defined additional check. As a result, of the 100 question patterns, 32 had two runs and 68 had three runs, for 268 AI answer logs in total. |
| Aggregation unit | The headline metric in the published text is the 100 question patterns of "retailer × AI × scenario." Each question pattern was decided as correct, incorrect, or unable by majority across 2–3 answers. We also kept a per-answer tally of all 268 answers as a supplementary metric. |
| Exclusions | One invalid capture whose text could not be obtained was saved as out of scope for analysis. We initially considered eight retailers, but for this preliminary release we analyzed first the five retailers with saved official evidence and high representativeness; some logs obtained for the retailers set aside are not counted in the correct-rate tally or the text's claims. |
| Ethics and privacy | We targeted only public information and AI output; no personal information was handled. |
AI Services and Conditions
The AI services were not aligned to the same model tier or subscription plan. For each answer, we recorded the model name, plan, and web-search setting. An AI service's UI and model display may change with the date, time, and account state.
| AI service | Main display name / model shown | Conditions |
|---|---|---|
| ChatGPT | ChatGPT | Signed in, temporary or new chat. Search/source display within the answer. |
| Claude | Haiku 4.5 | Pro plan, incognito chat. Web-search execution shown. |
| Gemini | Gemini 2.5 Flash | Run while logged out or in a free environment. Search/source display within the answer. |
| Perplexity | Model display recorded in the answer log | Free plan, incognito mode. Search/Computer display present. |
| Microsoft Copilot | Copilot Smart | Signed in, temporary mode shown. Search/source display within the answer. |
Because the model tier, subscription plan, and web-search setting differ by service, the by-service differences described later reflect not only differences in the services themselves but also differences in the combination of model, plan, and search settings used.
Verdict Criteria
Each answer was judged as correct, incorrect, or unable, based on its consistency with the essential facts verifiable in the official sources. Incorrect answers were further divided into minor discrepancies unlikely to harm the user, and real-harm risks that can affect the user's behavior or the support load.
| Category | Definition |
|---|---|
| Correct | An answer that correctly states the essential facts verifiable in official sources, without any assertion that creates a disadvantageous or overly favorable misunderstanding for the user. |
| Incorrect | An answer that makes a claim differing from official information on any of the deadline, eligibility conditions, coverage rate, refund method, number of times, filing destination, or exceptions. Even a partly correct answer was judged incorrect if it erred on an essential fact. We also included as incorrect any answer that asserts a specific deadline, number of years, or condition for a matter that cannot be confirmed from official information alone, giving the user false confidence (classified as real-harm risk when harm is expected). |
| Unable | An answer that does not assert any specific false fact, such as "this cannot be confirmed from official information alone" or "please check the official site." This also includes answers that present only an official URL without substantively answering. |
| Minor discrepancy | An error unlikely to lead directly to the user's disadvantage, such as differences in wording, omission of detailed conditions, or insufficient general caution. |
| Real-harm risk | An error that can affect the user's behavior or the support load, such as answering a return deadline as longer than it is, saying an out-of-scope repair is free, mis-guiding used points as a cash refund, or answering a coverage rate or warranty term as more favorable than reality. |
For source URLs, in addition to presence, we classified them as a correct official page, official but wrong topic, non-official, a nonexistent URL, or no source.
Scenarios Used
Each scenario frames a policy question consumers actually want to check after buying an appliance, asked as a natural consumer question.
| ID | Scenario | Main facts verified |
|---|---|---|
| S1 | Initial defect after opening / return / exchange | Contact deadline for an initial defect, whether return/exchange is possible, handling of opened items, differences between in-store and online purchase |
| S2 | Long-term warranty years / coverage rate | Enrollment conditions, warranty years, year-by-year coverage rate, out-of-scope conditions |
| S3 | Refund of the point-paid portion | Return method for used points, whether a cash refund is given, point expiration/return timing, handling on cancellation |
| S4 | Free count / conditions for air-conditioner warranty repair | Whether there is a cap on the number of repairs, free-repair conditions, on-site/parts fees and out-of-scope conditions, distinction between manufacturer and long-term warranty |
Prompts Used
For each prompt, {{企業名}} (the retailer name) was replaced with the official name of Bic Camera, Yodobashi Camera, Yamada Denki, K's Denki, or EDION. A fresh temporary chat or the like was used for each run. The AI services received the Japanese originals shown below; the English translations are provided for reference.
S1
Original Japanese prompt used
で買った冷蔵庫が、開封後に初期不良だと分かった場合、何日以内なら返品または交換できますか。店舗購入とネット購入で条件が違う場合は分けて、公式情報に基づいて答えてください。分からない場合は推測せず「公式情報だけでは確認できない」と書き、参照したURLがあれば示してください。過去の会話や記憶は参照しないでください。
English translation
If a refrigerator I bought at
{{retailer name}}turns out to have an initial defect after opening, within how many days can I return or exchange it? If the conditions differ between an in-store purchase and an online purchase, please separate them and answer based on official information. If you do not know, do not guess — write "this cannot be confirmed from official information alone," and show any URL you referred to. Do not reference past conversations or memory.
S2
Original Japanese prompt used
の冷蔵庫向け長期保証は、何年目まで何%補償されますか。加入条件、年数、年次別の補償率、対象外条件が分かる範囲で、公式情報に基づいて答えてください。分からない場合は推測せず「公式情報だけでは確認できない」と書き、参照したURLがあれば示してください。過去の会話や記憶は参照しないでください。
English translation
For
{{retailer name}}'s long-term warranty for refrigerators, up to which year and at what percentage are you covered? To the extent you can determine the enrollment conditions, years, year-by-year coverage rate, and out-of-scope conditions, please answer based on official information. If you do not know, do not guess — write "this cannot be confirmed from official information alone," and show any URL you referred to. Do not reference past conversations or memory.
S3
Original Japanese prompt used
で家電を買って返品または注文取消になった場合、ポイント払い分は返金時にどう扱われますか。現金やカード支払い分との違い、ポイントが戻るかどうか、条件が分かる範囲で、公式情報に基づいて答えてください。分からない場合は推測せず「公式情報だけでは確認できない」と書き、参照したURLがあれば示してください。過去の会話や記憶は参照しないでください。
English translation
If I buy an appliance at
{{retailer name}}and it is returned or the order is canceled, how is the point-paid portion handled at refund time? To the extent you can determine the difference from cash or card payments, whether the points come back, and the conditions, please answer based on official information. If you do not know, do not guess — write "this cannot be confirmed from official information alone," and show any URL you referred to. Do not reference past conversations or memory.
S4
Original Japanese prompt used
で購入したエアコンの保証修理は、何回まで無料ですか。メーカー保証と長期保証の違い、無料修理の条件、出張費や部品代、対象外条件が分かる範囲で、公式情報に基づいて答えてください。回数上限が公式情報だけでは確認できない場合は、そのように答えてください。参照したURLがあれば示してください。過去の会話や記憶は参照しないでください。
English translation
For warranty repair of an air conditioner purchased at
{{retailer name}}, how many times is it free? To the extent you can determine the difference between the manufacturer warranty and the long-term warranty, the free-repair conditions, on-site and parts fees, and out-of-scope conditions, please answer based on official information. If the cap on the number of times cannot be confirmed from official information alone, please answer accordingly. Show any URL you referred to. Do not reference past conversations or memory.
Main Official Sources
The official pages were confirmed on July 21, 2026, and screenshots were saved. The following are the main URLs used as this survey's correct-answer baseline. The correct-answer baseline for each scenario is the essential facts verifiable on these pages (deadline, eligibility conditions, coverage rate, refund method, number of times, exceptions).
| Retailer | S1 | S2 | S3 | S4 |
|---|---|---|---|---|
| Bic Camera | Source 1 / Source 2 | Source | Source 1 / Source 2 | Source |
| Yodobashi Camera | Source | Source | Source 1 / Source 2 | Source |
| Yamada Denki | Source | Source 1 / Source 2 | Source 1 / Source 2 | Source 1 / Source 2 |
| K's Denki | Source | Source | Source 1 / Source 2 | Source |
| EDION | Source 1 / Source 2 | Source 1 / Source 2 | Source | Source |
Overall Results
For the 100 question patterns, the primary verdict was 60 correct, 25 incorrect, and 15 unable. These three categories are mutually exclusive and total 100 patterns, or 100.0%. Of the 25 incorrect, 15 carried real-harm risk and the remaining 10 were minor discrepancies limited to wording differences or omitted conditions. Real-harm risk was tallied as those question patterns with an incorrect primary verdict that included an error capable of affecting the user's behavior or the support load. A source problem is an additional flag separate from the primary verdict. For example, an answer whose "conclusion is correct but whose source URL is non-official" counts as correct in the primary verdict while also counting as having a source problem.
The 62 patterns with a source problem span primary verdicts of 28 correct, 24 incorrect, and 10 unable. Because a presented source URL is not necessarily the correct official page, we treated this as a safety metric separate from the correct rate.
| Primary verdict / additional flag | Count | Share |
|---|---|---|
| Primary verdict: correct | 60 | 60.0% |
| Primary verdict: incorrect | 25 | 25.0% |
| Primary verdict: unable | 15 | 15.0% |
| Incorrect and carrying real-harm risk | 15 | 15.0% |
| Additional flag: a source problem in at least one run | 62 | 62.0% |
At the per-answer level, of the 268 answers, 155 were correct (57.8%), 72 incorrect (26.9%), and 41 unable (15.3%). Because the per-answer level includes third-run additional-check answers, we used the 100 question patterns as the headline metric in the published text.
Correct Rate by AI Service
By AI, the question-pattern correct rate was 65.0% for ChatGPT and Microsoft Copilot, 60.0% for Claude, and 55.0% for Gemini and Perplexity.
| AI service | Question patterns | Correct | Incorrect | Unable | Correct rate | Real-harm risk | Source problem |
|---|---|---|---|---|---|---|---|
| ChatGPT | 20 | 13 | 2 | 5 | 65.0% | 2 | 8 |
| Microsoft Copilot | 20 | 13 | 5 | 2 | 65.0% | 1 | 19 |
| Claude | 20 | 12 | 7 | 1 | 60.0% | 4 | 13 |
| Gemini | 20 | 11 | 6 | 3 | 55.0% | 3 | 13 |
| Perplexity | 20 | 11 | 5 | 4 | 55.0% | 5 | 9 |
Looking at the correct rate alone, the gap at the top is small, whereas the question patterns containing a source problem differed greatly. Microsoft Copilot contained a source problem in 19 of 20 (95.0%), and ChatGPT in 8 of 20 (40.0%). We confirmed that even an answer with a source URL does not necessarily rest on the correct official page.
Correct Rate by Scenario
The long-term-warranty years/coverage rate (S2) had a correct rate of 48.0%, the lowest of the four scenarios. Because the warranty system differs by retailer in name, years, year-by-year burden, and the distinction between free and paid warranties, it is likely an area where AI easily picks up old terms or a different system.
| Scenario | Question patterns | Correct | Incorrect | Unable | Correct rate | Real-harm risk | Source problem |
|---|---|---|---|---|---|---|---|
| S1 Initial defect / return / exchange | 25 | 14 | 8 | 3 | 56.0% | 6 | 15 |
| S2 Long-term warranty years / coverage rate | 25 | 12 | 11 | 2 | 48.0% | 6 | 17 |
| S3 Refund of the point-paid portion | 25 | 14 | 3 | 8 | 56.0% | 3 | 17 |
| S4 Free count for air-conditioner warranty repair | 25 | 20 | 3 | 2 | 80.0% | 0 | 13 |

The chart above may be reproduced as-is in media coverage, articles, and other materials, provided that you credit "Stellagent Inc." as the source and include a link to this page.
In S3 there were 8 unable answers. This can be called a safer behavior than asserting something incorrect, but from the customer's point of view the need to check the official desk or FAQ ultimately remains. While S4 had a high correct rate of 80.0%, 13 question patterns (52.0%) contained a source problem — cases where the conclusion matched but the source was unstable.
Trends by Retailer
The correct rate by retailer is affected by the structure of the official information, the complexity of the warranty system, and differences in the pages AI picks up in search. We therefore do not treat the by-retailer figures as a ranking of the retailers' merits, but present them as a reference indicator of information structures that tend to make answering harder for AI.
| Retailer ID | Question patterns | Correct | Incorrect | Unable | Correct rate | Real-harm risk | Source problem |
|---|---|---|---|---|---|---|---|
| KS | 20 | 15 | 4 | 1 | 75.0% | 1 | 8 |
| EDION | 20 | 13 | 2 | 5 | 65.0% | 0 | 12 |
| BIC | 20 | 12 | 7 | 1 | 60.0% | 5 | 12 |
| YODO | 20 | 11 | 5 | 4 | 55.0% | 3 | 14 |
| YAMADA | 20 | 9 | 7 | 4 | 45.0% | 6 | 16 |
The by-retailer differences cannot be explained by the ease of finding official information alone. They are also affected by whether AI picks up a different brand, a group company, an old PDF, or a non-official blog. For example, at Yamada Denki, multiple flows are mixed together — Yamada Web.com, Yamada Mall, Kadenho, the free long-term warranty, and old warranty-rule PDFs — producing wavering in AI answers.
Retailer × AI Matrix
| Retailer ID | ChatGPT | Claude | Gemini | Perplexity | Microsoft Copilot |
|---|---|---|---|---|---|
| BIC | 100.0% | 50.0% | 50.0% | 25.0% | 75.0% |
| YODO | 25.0% | 75.0% | 50.0% | 50.0% | 75.0% |
| YAMADA | 25.0% | 50.0% | 75.0% | 50.0% | 25.0% |
| KS | 100.0% | 50.0% | 50.0% | 100.0% | 75.0% |
| EDION | 75.0% | 75.0% | 50.0% | 50.0% | 75.0% |
Each figure in the table is the share of the 4 scenarios that were correct question patterns. Because the denominator is only 4 per cell, this table is not an absolute ranking of AIs or retailers but a supplementary table for seeing which combinations produced incorrect answers, unable answers, or source problems.
Main Patterns of Incorrect Answers and Source Problems
Guiding a return deadline beyond the confirmable scope
In Yamada Denki's S1, we confirmed that Yamada Web.com handles initial defects within 14 days of the item's arrival. Meanwhile, for the in-store-purchase initial-defect deadline, we scored only within the scope confirmable from published official information. ChatGPT, in 3 of 3 runs, guided in-store purchases as "within one week of purchase" or "repair handling after one week has passed," an answer that went beyond the officially confirmable scope.
Answering longer warranty years
In Yamada Denki's S2, our saved official baseline set the free long-term warranty at 6 or 4 years including the manufacturer warranty, and Kadenho at 5 or 3 years including the manufacturer warranty. Perplexity, Gemini, and ChatGPT picked up old warranty PDFs or different flows and gave answers with warranty years longer than the saved baseline, such as 9, 10, and 11 years.
Treating non-official, different-brand, or old flows as if they were official
For Yamada Denki's point-refund question, Claude referred to the return rules of a separate company, YAMADAYA STORE, and made the point-return conditions concrete. For Yamada Denki's initial-defect question, Microsoft Copilot referred to a non-official blog and cited a nonexistent example.com as "official information." In ChatGPT, there were cases where Yamada Mall or Matsuya Denki flows were mixed into the answer.
Unable answers are safe, but the customer problem remains
For point refunds (S3), there were comparatively many cases where AI answered "this cannot be confirmed from official information alone." In the sense of avoiding an incorrect assertion, this is desirable behavior, but because it does not resolve the return/refund conditions the customer wants to know on the spot, the need for a support AI connected to official FAQs, terms, order information, and member information remains.
How We Extracted, Aggregated, Normalized, and Broke Ties
The aggregation unit is the 100 question patterns of "retailer × AI × scenario." For each question pattern, we obtained at least two answers, and obtained a third run as a pre-defined additional check when the two verdicts split, when a real-harm-risk incorrect answer appeared at least once, or when a source URL was nonexistent or pointed to a page different from the official one. As a result, of the 100 question patterns, 32 had two runs and 68 had three runs.
Each question pattern's primary verdict (correct, incorrect, or unable) was decided by majority across the 2–3 answers. When a three-run set split one to each verdict, we prioritized whether an essential fact was in error, and treated a pattern containing an answer with real-harm risk as having real-harm risk. Real-harm risk and source problems were counted as additional flags separate from the primary verdict, marked "present" at the question-pattern level if they applied even once. That is why the 62 patterns with a source problem span every primary verdict of correct, incorrect, and unable.
For retailer and brand names, we normalized variants such as Yamada Denki (ヤマダ電機) / Yamada (ヤマダ) / Yamada Web.com to Yamada Denki, Yodobashi (ヨドバシ) to Yodobashi Camera, and Joshin (ジョーシン) to Joshin. Source URLs were classified as official and on-topic, official but wrong topic, non-official, a nonexistent URL, inaccessible, or no source. Verdicts were recorded as correct, incorrect, or unable, and the error weight as minor discrepancy or real-harm risk.
For handling invalid answers, we also recorded — without deleting — answers where the AI refused, only urged an official check, mistook the target retailer, or presented only a URL without substantively answering, and included answers that did not substantively answer in the "unable" category. Only one invalid capture, whose text could not be obtained due to a connection interruption or the like, was saved as out of scope for analysis. The per-answer tally of 268 answers includes third-run additional checks, so it is a supplementary metric; the headline metric in the published text uses the 100 question patterns.
Implications for Businesses
-
Official FAQs and terms need to be maintained not only as pages for humans but in a structure that AI is unlikely to misread. Returns, initial defects, warranties, and point refunds share many similar terms, and it is easy to confuse ordinary returns with initial defects, free warranties with paid warranties, and used points with awarded points.
-
Even an answer with a source URL does not necessarily rest on the correct page. Quality control of AI answers must verify not only the URL-citation rate but also whether the source is official, whether it is on-topic, and whether it is an old rule.
-
In the customer-support domain, rather than leaving things to generic AI, a specialized AI connected to the official data, FAQs, terms, and product/order/member conditions the company manages is needed. Return deadlines, warranty years, coverage, and point refunds in particular are directly tied to customer expectations and the support load when mis-guided.
-
Even when introducing a specialized AI, a design that handles "unable" appropriately is important. When something cannot be confirmed from official information, not guessing but showing where to confirm, what information is needed, and the inquiry path reduces both incorrect-answer risk and customer anxiety at once.
Notes on the Survey
- This is an exploratory survey based on AI answers and official information as of July 21, 2026; it is not intended for statistical generalization.
- AI answers may vary with the date and time, model, search settings, conversation history, account state, region, and UI changes.
- Official terms, FAQs, and warranty systems may be revised. In this survey, we saved the URLs, screenshots, and official correct-answer summaries as of the survey date as primary evidence.
- The by-retailer correct rate is affected by the pages AI picked up in search, the structure of the official information, and the complexity of the warranty system. It does not evaluate the response quality of each home-appliance retailer.
- The AI services were not aligned to the same model or subscription plan; the model tier, subscription plan, and web-search setting differ by service. The by-service differences reflect not only differences in the services themselves but also differences in the combination of model, plan, and search settings used. It does not evaluate the relative merits of any specific AI.
- AI answer screenshots are saved as evidence, but if images are published, the quotation scope, service terms of use, and rights must be checked.
- The company, product, and service names in this report are trademarks or registered trademarks of their respective companies.
Revision History
- July 22, 2026: First published
You are free to cite or reproduce the findings of this report in articles, media coverage, and other materials, provided that you credit “Stellagent Inc.” as the source and include a link to this page.
Meet Stella, our customer support AI agent
Stella is a customer support AI agent for retail and e-commerce that helps with inquiries and workflows such as orders, deliveries, returns, and repairs.
Learn more about Stella