Do not only ask whether AI mentions your product. Test whether it understands your product correctly.
A company can be highly visible in AI answers and still have a serious problem.
The answers are wrong.
That makes product understanding testing one of the most practical GEO exercises a team can run.
You do not need sophisticated software to begin.
You need a disciplined question set.
Step 1: Build a product truth set
Choose 20 to 50 questions whose answers you know.
Include several categories.
Capability
- Does Acme support regional forecasting?
- Can Smart Routing use CRM fields?
- Does the product generate automated summaries?
Availability
- Which plan includes Smart Routing?
- Is Feature X generally available?
- Is it available in Europe?
Integration
- Does Acme integrate with Salesforce?
- Which SSO providers are supported?
Comparison
- What is the difference between Forecasting and Pipeline Analytics?
- When should a customer use Feature A instead of B?
Freshness
- When was Feature X introduced?
- What changed in the latest release?
Limitations
- What does Feature X not support?
- Are there usage limits?
Step 2: Write the canonical answer
Before testing models, document the correct answer.
Include the exact fact, the authoritative source URL, and the last verified date.
This prevents the team from judging answers based on memory.
Step 3: Test several systems
Run the same questions across systems your customers are likely to use.
Record:
- the answer,
- visible sources,
- the date,
- the model or product,
- whether live retrieval appeared to occur.
Do not assume every system behaves identically.
Test sheet, one row per question
Step 4: Score accuracy
A simple rubric:
| Score | Meaning |
|---|---|
| 3 | Correct and complete |
| 2 | Mostly correct, missing material context |
| 1 | Substantially incomplete or outdated |
| 0 | Wrong or hallucinated |
Score source quality separately.
Step 5: Diagnose the failure
If an answer is wrong, ask why.
- No source exists. Create a canonical source.
- The source is vague. Make the fact explicit.
- The source is outdated. Update or redirect it.
- Conflicting sources exist. Clarify current truth and canonicalize.
- A desired crawler is blocked. Review your access policy.
- Third party misinformation dominates. Strengthen first party evidence and consider outreach where appropriate.
Step 6: Retest
Changes in AI answers may not be immediate.
Retest on a consistent cadence.
Monthly is enough for many teams. Faster moving products may want additional checks for high value capabilities.
Step 7: Add customer language
Internal teams often test using product names customers do not know.
Include natural language:
Can Acme automatically route enterprise leads to different teams by country?
not only:
Does Smart Routing support geographic conditions?
Both matter.
Track understanding over time
A useful internal metric is:
Product Answer Accuracy. The percentage of test questions answered correctly or mostly correctly.
Another:
Authoritative Source Rate. The percentage of answers that rely on current first party or approved sources when sources are visible.
These are not industry standards.
They are operational tools.
Turn errors into an editorial backlog
The most valuable output is not the score.
It is the list of missing knowledge.
If several AI systems cannot explain a feature correctly, there is a decent chance a human visitor cannot either.
That is a content problem worth fixing regardless of GEO.
Questions people ask
- Should we automate these tests?
- Automation can help at scale, but manual review remains valuable. Reading full answers reveals nuance, framing, and partial errors that a pass or fail script can miss.
- Should we test only our brand name?
- No. Include problem oriented and capability oriented questions, because many buyers describe the problem they have rather than the product name you use.
- How often should we retest?
- Monthly is a reasonable starting point, plus checks after major product or website changes.
- What if different AI systems disagree?
- Record the disagreement. It can reveal differences in retrieval, freshness, or model knowledge, and it tells you which sources each system is leaning on.