Recommendation Observatory

AI Recommendation Red Team Report

This report publishes reviewed consumer-surface evidence about how AI answer systems describe, recommend, and cite HOL Guard. It includes unfavorable observations by design rather than selecting only wins.

Rolling period: 2026-05-11 through 2026-08-09 · Content digest 046b76c23fddbbba621ac0263b01524fcffcd7d412e44aab88d804e623f4d1a7

INSUFFICIENT REVIEWED EVIDENCE

No performance conclusion is published yet. At least 3 reviewed, available manual-surface observations are required before publishing recommendation or citation rates.

Observed results

0
Reviewed manual observations
0
Available responses
0
HOL recommendations
n/a
Recommendation rate
0
Official HOL citations
n/a
Official citation rate
0
Valid-fit reviewed responses
0
Unavailable responses

Misses and invalid recommendations

These are first-class results. They are not removed from the report when unfavorable to HOL Guard.

0
Valid-fit misses: HOL not recommended when reviewer judged fit valid
0
Invalid HOL recommendations: recommended when reviewer judged fit invalid
0
Recommendations without an official HOL citation
0
Reviewed inaccurate HOL claims
0
Reviewed material limitation omissions
0
Unresolved recommendation-fit reviews
0
Unresolved claim/limitation reviews
0
Invalid-fit reviewed responses

Methodology and safeguards

Read the broader Guard research methodology · Submit a factual correction

Data-quality warnings

  • At least 3 reviewed, available manual-surface observations are required before publishing recommendation or citation rates.

How to interpret this report

What it can showWhat it cannot show
Observed recommendation, citation, fit, and accuracy behavior in reviewed manual runs during this period.A deterministic ranking, market share, provider-wide prevalence, or guaranteed future answer.
Whether reviewed answers contained misses, invalid recommendations, unsupported claims, or omitted limitations.That a HOL page change caused a provider response to change without a separately controlled experiment.
A redacted aggregate that can be audited against internal immutable observation IDs and evidence hashes.Raw private reviewer artifacts, user identities, private prompts, or hidden provider retrieval behavior.