← Back to Research AI Safety · Recommendation Agents · EMNLP 2026 Accepted

SafeGEO

Paid search once steered patients toward dubious hospitals; with AI assistants as the new gateway, will the same story repeat? SafeGEO offers the first systematic measurement.

arXiv
600recommendation cases
22GEO attack variants
40,800materialized samples
83.2%max harmful promotion increase
39.2%max Target@3 mitigation reduction

More and more purchase decisions now start with a question to an AI assistant: “Which air purifier is worth buying?” Sellers have noticed — web copy is increasingly written for AI readers, a practice with a name: Generative Engine Optimization (GEO), the SEO of the AI era. SafeGEO measures the question that follows: when a genuinely flawed product dresses its page up as an “independent buyer guide”, can the agent recommending on the user’s behalf hold its judgment?

User query“Which air purifieris worth buying?”Retrieved web evidenceGEO-rewritten seller sourceRecommendationagentwithout the rewritewith the rewriteTop 3 · no GEOTop 3 · with GEO1Product A1Product B · flawed2Product C2Product A3Product D3Product Cflawed product B never enters the top 3product D is pushed out
Redrawn from the paper. The agent reads the same candidate set either way; rewriting one seller-controlled source is enough to move a flawed product into the top three and push a good one out.

When product pages are written for AI readers

Recommendation agents (AI systems that search, compare, and advise on the user’s behalf) are becoming a new entry point for commerce and content platforms. Their judgment rests on what they can read online — product pages, reviews, FAQs; and seller-controlled sources such as product pages (material written by the sellers themselves) carry a built-in incentive to be optimized, or manipulated.

The story in the headline deserves unpacking. In the era of paid search ranking, the script — platform optimized by sellers, price paid by users — played out once already: patients in urgent need were steered toward the highest bidder, not the best doctor. With the entry point shifting from the search box to AI assistants, the same contest moves to a new arena — what decides the ranking is no longer keyword bids, but the “evidence” an AI reads.

GEO is not inherently harmful: a clearer page helps humans and AI alike. The problem is the boundary — once a rewrite starts hiding flaws, fabricating reputation, or impersonating an “independent review”, the evidence the agent reads is systematically polluted. SafeGEO is the first controlled measurement of this risk: how far attacks can push recommendations away from the user’s interest, and how much existing defenses recover.

A controlled testbed

SafeGEO covers six evidence-grounded product verticals: AI meeting transcription tools, baby monitors, carry-on backpacks, home air purifiers, noise-canceling headphones, and office chairs. Each case keeps candidate products, canonical attributes, non-target evidence, and hidden utility labels (ground-truth quality annotations known to the evaluator but never shown to the model) fixed while rewriting exactly one seller-controlled source — so any change in the recommendation can be cleanly attributed to that single rewrite.

Recommendation cases600
Avg. candidates per case19.96
GEO targets per case3
Attack variants22
Total samples40,800
Evaluation metricsTarget@3, HCV@1, GT@3, uNDCG@5

Benchmark statistics. Each base case expands into control and attack conditions so recommendation changes can be attributed to one rewritten seller source.

Evaluation uses four metrics: Target@3 (how often the attacked, flawed product enters the top three recommendations), HCV@1 (how often it takes the top slot), GT@3 (how often genuinely good products remain in the top three), and uNDCG@5 (how well the top five match the user’s true utility). The 40,800 materialized samples (concrete evaluation instances expanded from each case under different attack conditions) keep every condition pairwise comparable.

Base case · fixedGEO attack constructionMaterialize and evaluateCandidate products19.96 per case on averageSeller-controlled sourcesproduct pages · reviews · FAQsHidden utility labelsknown to the evaluator onlyNon-target evidenceleft untouched600 recommendation cases① Pick 3 candidatesthen one GEO target② Pick 1 of 22 variantsatomic · block · cross · realistic③ Rewrite one seller sourcethe only thing that changesOne rewrite per sample, so any shift is attributableMaterialize instances600 cases × 68 conditions40,800 samplesEvaluationTarget@3HCV@1GT@3uNDCG@5utility labels stay hidden
Redrawn from the paper. Candidate sets, hidden utility labels and non-target evidence stay fixed; GEO rewrites exactly one target product source at a time, so any change in the recommendation is attributable to that rewrite.
ConditionAverage source-text length
No GEO3,911 [3,901, 3,921]
Truthful-rewrite3,905 [3,895, 3,915]
Avg. GEO, 22 variants3,925 [3,924, 3,926]

Source-length control. GEO and control contexts are closely matched in length, so the observed harm is not explained by simply giving the model more text.

How far attacks go

Experiments show that GEO attacks can substantially promote flawed target products. Realistic seller-facing variants are especially strong: they package false fit, evidence padding, and salience manipulation into one plausible-looking seller document, rather than mechanically stacking keywords. On DeepSeek-V4-Flash, the flawed product enters the top three only 6.2% of the time with no attack; under the realistic “selective comparison note” variant, that rises to 82.3%.

Target@3 uplift vs. the truthful-rewrite control (pp)20406080A-only43.9U-only41.3C-only65.4R-only38.7E-only48.7S-only59.1M-only15.7Content bundle45.4Epistemic bundle49.1Model-facing bundle17.0Content + epistemic49.1Content + model-facing21.5Epistemic + model-facing25.9Full-stack diagnostic26.3Caveat-buried FAQ77.2Popularity-heavy profile75.7Citation-padded note77.2Independent buyer guide77.2False-fit checklist71.9Selective comparison note80.6AI-directed source text72.8Full-stack realistic76.3AtomicBlockCross-blockRealisticThe eight realistic seller-page variants add 72–81 pp on their own.
Redrawn from the paper, values approximate. Target@3 uplift over the truthful-rewrite control for all 22 attack variants. The realistic seller-facing variants dominate, suggesting that a coherent source document matters more than mechanically stacking manipulation primitives.
Representative realistic GEO attack results (DeepSeek-V4-Flash)
SettingTarget@3ΔHCV@1ΔGT@3ΔuNDCG@5Δ
No GEO6.2--24.5--66.7--77.0--
Truthful-rewrite control4.6--23.0--67.7--78.8--
Caveat-buried FAQ77.5+72.976.2+53.257.7-10.066.3-12.5
Popularity-heavy profile71.2+66.671.4+48.457.6-10.167.3-11.5
Citation-padded note78.7+74.178.4+55.458.1-9.666.2-12.7
Independent buyer guide77.9+73.377.3+54.356.5-11.266.0-12.9
False-fit checklist79.1+74.678.4+55.457.7-9.966.1-12.7
Selective comparison note82.3+77.781.8+58.856.9-10.865.4-13.5
Avg. realistic72.6+68.073.4+50.457.7-10.066.9-11.9
Realistic variants raise Target@3 and HCV@1 while degrading utility-quality metrics.

Mechanistically, an attack succeeds almost exactly to the extent that it hijacks the agent’s citations: the more the model’s citations are steered toward misleading lines, the higher the flawed product ranks — a correlation of r=0.91.

20203030404050506060707080809090Misleading GEO-line citation rate (%)Attacked product in top 3 (%)M-onlyC-onlythe 8 realistic variantsr = 0.91
Redrawn from the paper, values approximate. Each point is one attack variant. Variants that redirect the agent’s citations toward misleading GEO lines also achieve higher target placement; the paper reports r = 0.91.

How much defenses recover

Simple mitigations help, but they are not enough. Defensive prompting (explicitly instructing the agent to watch for marketing manipulation) reduces harmful promotion; evidence breakdown (requiring the agent to list supported, missing, and conflicting evidence for each candidate before ranking) is strongest, cutting Target@3 by 39.2 percentage points on Qwen3.6 27B. Even the strongest defense, though, does not restore no-attack behavior.

Reduction in Target@3 (pp) · Gemma 4 31B ITbroadest and strongestL1PromptL2RationaleL3Evidence sheetL4Context balanceL5Instruction filterCaveat-buried FAQ7.115.324.412.41.5Popularity-heavy profile8.615.326.216.91.6Citation-padded note6.414.921.68.70.2Independent buyer guide6.115.023.98.40.3False-fit checklist12.013.235.2-1.81.4Selective comparison note6.015.920.67.10.8AI-directed source text37.914.938.931.34.7Full-stack realistic36.815.446.79.27.2Its best cell, 46.7 pp, still gives back only part of a 76 pp attack.
Redrawn from the paper. Variant-level mitigation effects on Gemma 4 31B IT: the L3 evidence sheet is the broadest and strongest layer, prompt-only defenses help unevenly, and one layer even helps a single attack.
Mitigation results on the same attacked instances (excerpt)
ModelMitigationTarget@3ΔHCV@1ΔGT@3ΔuNDCG@5Δ
Gemma 4 31B ITNo mitigation79.6--75.6--67.9--68.6--
Gemma 4 31B ITDefensive prompt64.5-15.160.8-14.869.3+1.372.6+4.0
Gemma 4 31B ITEvidence breakdown49.9-29.746.6-29.169.5+1.674.4+5.7
Qwen3.6 27BNo mitigation78.3--83.7--60.8--63.6--
Qwen3.6 27BDefensive prompt67.3-11.066.2-17.568.5+7.673.4+9.8
Qwen3.6 27BEvidence breakdown39.1-39.242.1-41.669.7+8.877.4+13.9
Devstral Small 2No mitigation90.9--90.7--47.9--59.2--
Devstral Small 2Evidence breakdown73.2-17.778.9-11.843.4-4.556.3-2.8
Evidence breakdown is usually the strongest mitigation, but it still does not restore no-GEO behavior.

A caution for agent safety

Makes GEO risk measurable.The work grounds visibility optimization in concrete recommendation choices.
Focuses on seller-controlled evidence.It studies information sources that can realistically be optimized or manipulated.
Shows simple defenses are incomplete.Prompting and evidence checks help, but GEO remains a serious agent-safety risk.

Research resources

SafeGEO has been accepted to EMNLP 2026 and is available on arXiv. Contact the lab for evaluation details or research collaboration.

arXiv Contact the research team