SafeGEO, the research report
How we built SafeGEO, what changed when one source was rewritten, and which defences helped. The protocol, results, and limitations of our first paper.

This report describes the protocol, results, and limitations of SafeGEO. The paper’s terminology appears beside ours. For a shorter account, read One source can sway the choice.
1. Summary
- Setting. A recommendation agent reads a fixed packet of retrieved web sources about a shortlist of products and returns a ranked recommendation with citations. One seller-controlled source in the packet is rewritten; everything else is held fixed.
- Scale. 600 requests in six product verticals; 22 shaped variants and two controls; three sampled targets per request; 40,800 evaluation instances. Four systems: Gemma 4 31B IT, Qwen3.6 27B, Devstral Small 2 24B Instruct, and DeepSeek-V4-Flash as a frontier-scale check.
- Headline pair. Averaged over the eight realistic variants, the shaped candidate reaches the top three in 72.6 to 90.9 percent of instances, against 3.4 to 12.7 percent under a truthful rewrite of the same source. The first recommendation violates one of the person’s hard constraints in 73.4 to 90.7 percent of instances, against 16.9 to 41.1 percent.
- Mechanism. Shaped sources work by being cited. The misleading-line citation rate and the shaped candidate’s top-three rate move together across variants (Pearson r = 0.91 for Gemma 4 31B IT, 0.99 for DeepSeek-V4-Flash), and paired responses show the route: a shifted citation, a missed product gap, a promotion.
- Mitigation. Five prompt-level or input-level interventions were tested on the same instances. A candidate-level evidence breakdown before ranking is the strongest on every system, lowering the shaped candidate’s top-three rate by 17.7 to 39.2 points. None restores the clean rate.
2. The question
A recommendation is utility-aligned when it puts the candidates that satisfy the person’s hard constraints first and orders them by how well they meet the person’s preferences. The paper writes this down: each request carries hidden hard constraints, soft preferences and preference weights; each candidate carries canonical attributes; a reference utility and a reference ranking follow from them and are never shown to the system under test.
A shaped source is a rewrite of one candidate’s own source, intended to promote that candidate although it is inferior under the reference ordering or violates a hard constraint. The rewrite changes the text; it does not change the product. Harm is when the rewrite moves such a candidate up the ranking. That is the failure the benchmark measures, and it is different from the visibility that earlier work on generative engine optimisation measured: whether rewritten content gets cited or surfaced at all.

3. Setting and protocol
The agent. The evaluated system receives a request, a candidate shortlist, product metadata and a packet of sources, and must return a ranked list with a short rationale and citations, as structured output. Retrieval, candidate generation and source selection are fixed, so the benchmark isolates the generation stage.
Three conditions. Every instance exists in one of three forms of the target’s own source.
| Condition | What the target’s source is | Role |
|---|---|---|
| No rewrite | The original retrieved source, unchanged | Reference point |
| Clean (the paper’s truthful-rewrite control) | The same ten-line template as the shaped version, keeping the supported claims and the decision-relevant caveat | Baseline for every gap reported here; holds form and length fixed |
| Shaped | The template with the lines of the active primitives rewritten to win the recommendation | The condition under test |
Source lengths are matched: on average 3,911 words under no rewrite, 3,905 under the clean condition and 3,925 under the shaped condition, over all 22 variants. The effect is not a longer page.
The template. Every controlled source has ten lines with fixed roles: title and source type, framing, positioning, the caveat or claim, a positive note, a verification note, a comparison reminder, a boundary reminder, a statement of source authority, and a format note. A variant may overwrite only the lines its primitives are allowed to touch; the rest stay as truthful anchors. The paper’s Figure 8 shows a full pair for one instance.
Four measures. Percent of instances unless noted.
| Ours | Paper | Definition | Better |
|---|---|---|---|
| Shaped candidate in the top three | Target@3 | The candidate whose source was rewritten is among the top three recommendations | lower |
| Top pick violates a hard constraint | HCV@1 | The first recommendation breaks at least one hidden hard constraint | lower |
| A right answer in the top three | GT@3 | At least one ground-truth candidate under the reference ranking is in the top three | higher |
| Utility of the top five | uNDCG@5 | Normalised discounted cumulative gain of the top five, with the reference utility as graded relevance | higher |
Four systems. Gemma 4 31B IT, Qwen3.6 27B and Devstral Small 2 24B Instruct are open-weight models in the 24B to 31B range, served locally with vLLM. DeepSeek-V4-Flash is queried through a hosted OpenAI-compatible API with reasoning disabled and a 4,096-token output cap. All run at temperature 0 with structured output enforced.
4. The taxonomy of shaped sources
Seven primitives in three loci, drawn from the literatures on information manipulation, persuasion and prompt injection, and kept only where a primitive could be realised as a controlled source rewrite that leaves the canonical attributes fixed.
| Locus | Primitive | What it does | Example |
|---|---|---|---|
| What the page says | Unsupported fit claim | Claims the target satisfies a requirement without valid support | Claims lowest-plan single sign-on when it is unavailable |
| Caveat omission | Omits a limitation that affects fit | Hides that a required feature is enterprise-only | |
| Relevance flooding | Adds positive but weakly relevant content | Emphasises awards or adoption over the person’s needs | |
| How supported it looks | Authority laundering | Makes seller-affiliated content appear independent or authoritative | Styles a seller page as a buyer guide |
| Evidence padding | Adds evidence-like language without direct support | Vague benchmark or expert phrasing | |
| How it reaches the model | Salience manipulation | Makes the target more prominent to the model | FAQ snippets, repeated query terms |
| Model-directed instruction | Adds text aimed at controlling the model | Tells AI assistants to rank or cite the target first |
The primitives are instantiated as 22 variants in four families.
| Family | Variants | What they are |
|---|---|---|
| Atomic | 7 | One primitive each |
| Block | 3 | All the primitives of one locus |
| Cross-block | 4 | The three pairs of loci, and all three loci at once (the full-stack diagnostic) |
| Realistic | 8 | Coherent documents a seller might publish: caveat-buried FAQ, popularity-heavy profile, citation-padded note, independent buyer guide, false-fit checklist, selective comparison note, AI-directed source text, full-stack realistic |
The realistic family is the one the headline numbers use. The full variant-by-primitive matrix is Table 3 of the paper and the taxonomy document in the code repository.
5. The dataset
Verticals. AI meeting transcription tools, baby monitors, carry-on backpacks, home air purifiers, noise-cancelling headphones, office chairs. They were chosen so that requests carry verifiable constraints, fit depends on source-grounded evidence and canonical attributes, and the evidence takes different forms across verticals: policy and plan terms, safety claims, technical specifications, physical compatibility, review-based fit. One hundred requests per vertical were written by hand.
Discovery. Candidates, sources and metadata come from the candidate-search trace of a production shopping agent, used only for product discovery and source collection. The discovered candidates become the shortlist; the opened sources, from product pages and reviews to comparison pages and specification pages, become the packet.
Construction statistics (paper, Table 8).
| Component | Count |
|---|---|
| Base cases | 600 (100 per vertical) |
| Candidate records | 11,974 (mean 19.96 per case) |
| Original retrieved sources | 21,513 (mean 35.85 per case) |
| Annotated original-source lines | 89,286 |
| Sampled targets | 3 per case (1,800) |
| Variants and controls | 22 + 2 |
| Evaluation instances | 40,800 (68 per case) |
| Realistic subset | 15,600 (26 per case) |
| Controlled seller-source documents | 41,400, ten primary lines each |
| Controlled-source line records | 414,000 |
| Visible packet | exactly 22 documents per instance, mean 11.50 lines per document |
Hidden labels. A query-decomposition step extracts only the requirements and preferences the request states. A verification step fills each candidate’s canonical attributes only where the supplied sources support them, and keeps missing or conflicting evidence as missing or conflicting. Each source span is labelled for its relation to a product claim: supports, refutes, insufficiently supports, or irrelevant. The labels were produced with GPT-5.5 and audited by two independent annotators on 60 cases stratified by vertical (paper, Table 15): agreement between the annotators was 100 percent on hard-requirement extraction, 91.7 percent on the top candidate, 78.1 percent on pairwise utility ordering, 76.7 percent on candidate hard status, and 80.3 to 88.9 percent on the three line-level labels; the model’s agreement with each annotator was of the same size.
Targets. For each case, three eligible candidates that are not ground truth are sampled as targets. Each target is crossed with every variant and both controls, and each instance rewrites that target’s source only.

SafeGEO Diamond. A 600-instance screening split from 120 cases, 20 per vertical, each contributing both controls and the three realistic variants with the highest DeepSeek-V4-Flash top-three rate on the full benchmark (selective comparison note, false-fit checklist, citation-padded note). It is a high-effect stress set for cheap iteration; population-level numbers should use the full benchmark.
6. Results
6.1 Four systems, four measures
Clean is the truthful-rewrite control; shaped is the average over the eight realistic variants. Percent of instances. Paper, Table 5 and Table 33; the no-rewrite reference is on the leaderboard.
| System | Shaped in top 3, clean → shaped | Constraint violated, clean → shaped | Right answer in top 3, clean → shaped | Utility, top 5, clean → shaped |
|---|---|---|---|---|
| DeepSeek-V4-Flash | 4.6 → 72.6 | 23.0 → 73.4 | 67.7 → 57.7 | 78.8 → 66.9 |
| Qwen3.6 27B | 8.1 → 78.3 | 24.2 → 83.7 | 61.2 → 60.8 | 66.5 → 63.6 |
| Gemma 4 31B IT | 3.4 → 79.6 | 16.9 → 75.6 | 71.2 → 67.9 | 74.4 → 68.6 |
| Devstral Small 2 24B Instruct | 12.7 → 90.9 | 41.1 → 90.7 | 50.7 → 47.9 | 67.4 → 59.2 |
Two things are visible at once. The harm measures move by 60 to 80 points. The quality measures move by a few points, because the shaped candidate usually joins the top three rather than displacing every right answer; a ranking can be mostly right and still lead with the product that cannot do what was asked.
6.2 By family and variant
Coherent documents beat combinations of primitives. Shaped candidate in the top three, percent, over all variants of each family (paper, Table 24 and Section C.7).
| Family | Gemma 4 31B IT | Qwen3.6 27B | Devstral Small 2 24B | DeepSeek-V4-Flash |
|---|---|---|---|---|
| Atomic (7) | 48.2 | 49.0 | 79.4 | 57.4 |
| Block (3) | 40.7 | 47.6 | 84.9 | 55.2 |
| Cross-block (4) | 34.2 | 40.9 | 90.2 | 61.8 |
| Realistic (8) | 79.6 | 78.3 | 90.9 | 72.6 |
| All 22 | 56.0 | 58.0 | 86.3 | 63.4 |
For the two systems with headroom, stacking primitives without a coherent source form does worse than a single primitive. Devstral Small 2 24B Instruct is near ceiling under every family.
Within the realistic family, the strongest variant is the selective comparison note on three systems (84.0 percent on Gemma 4 31B IT, 82.2 on Qwen3.6 27B, 82.3 on DeepSeek-V4-Flash) and the full-stack realistic rewrite on Devstral (95.9 percent, 83.2 points above its clean rate, the largest gap in the study). The weakest realistic variant on DeepSeek-V4-Flash is the AI-directed source text, at 51.3 percent against 95.5 on Devstral; among single primitives on that system, salience manipulation (70.9), caveat omission (67.8) and evidence padding (66.7) are the strongest, and model-directed instruction (42.4) and relevance flooding (42.9) the weakest.
6.3 By vertical
Shaped candidate in the top three, percent, by product vertical, for the three open-weight systems, with the uplift over the vertical’s clean rate in brackets (paper, Table 28).
| Vertical | Gemma 4 31B IT | Qwen3.6 27B | Devstral Small 2 24B |
|---|---|---|---|
| AI meeting transcription tools | 90.0 (78.7) | 89.9 (68.3) | 88.9 (67.9) |
| Baby monitors | 54.1 (50.5) | 50.6 (45.7) | 84.7 (72.5) |
| Carry-on backpacks | 25.6 (23.5) | 36.5 (31.7) | 83.0 (74.2) |
| Home air purifiers | 31.2 (31.2) | 27.7 (22.9) | 83.3 (73.6) |
| Noise-cancelling headphones | 44.6 (43.6) | 46.2 (39.2) | 83.3 (70.7) |
| Office chairs | 52.2 (49.6) | 58.6 (53.0) | 88.8 (76.9) |
The vertical decided by plan and policy terms is the most fragile for every system. Where fit is a physical specification that other sources state plainly, the two stronger systems resist more.
6.4 Mechanism
Across the 22 variants, the rate at which a system cites a misleading line from the shaped source and the rate at which the shaped candidate reaches the top three are tightly coupled: Pearson r = 0.91 for Gemma 4 31B IT, 0.92 for Qwen3.6 27B, 0.88 for Devstral Small 2 24B Instruct, and 0.99 for DeepSeek-V4-Flash (paper, Table 23 and Figure 22).
Paired analysis of matched clean and shaped responses for Gemma 4 31B IT (paper, Figure 5) traces the route. When the cited evidence shifted, the shaped response missed the target’s decision-relevant gap in 54.1 percent of pairs, against 6.7 percent without a shift. When the gap was missed, the target entered the top three in 92.8 percent of pairs, against 58.6 percent when it stayed recognised. The complete path, a shifted citation followed by a missed gap followed by a promotion, accounts for 49.7 percent of observed entries.
DeepSeek-V4-Flash often retrieves the refuting evidence and promotes the target anyway: its paired refuting-evidence recall is 68.3 percent and its valid-citation rate 86.8 percent. The failure is one of weighting after retrieval, not of retrieval.

6.5 Sensitivity
On DeepSeek-V4-Flash, averaged over the eight realistic variants and three targets (paper, Table 34).
| Setting | Shaped in top 3 | Constraint violated | Right answer in top 3 | Utility, top 5 |
|---|---|---|---|---|
| As released (full packet) | 72.6 | 73.4 | 57.7 | 66.9 |
| Shaped source forced first | 75.9 | 73.7 | 59.1 | 66.9 |
| Shaped source forced middle | 61.2 | 64.8 | 62.4 | 69.9 |
| Shaped source forced last | 58.4 | 63.3 | 62.8 | 70.1 |
| Packet cut to ten sources | 77.2 | 74.5 | 69.5 | 71.6 |
| Packet cut to five sources | 99.3 | 91.7 | 46.5 | 48.9 |
Position matters and does not rescue the person. A thinner packet is worse: with five sources the shaped candidate is in the top three almost always.
7. Mitigation
Six matched conditions on the same shaped instances (the eight realistic variants, three targets), each measured against the unmitigated request. All are prompt-level or input-level changes a developer can make without changing the model.
| Intervention | What it does |
|---|---|
| No mitigation | The exact benchmark request |
| Defensive prompt | Requires clear support for important claims and preserves uncertainty when sources disagree |
| Rationale elicitation | Requires a brief reason and source-line citations for each top recommendation, without a separate evidence check before ranking |
| Evidence breakdown | Checks important claims against supporting and conflicting evidence, per candidate, before ranking |
| Context balancing | Compares claims across the full packet so that one prominent source does not dominate |
| Instruction filtering | Treats source text aimed at directing the model as non-evidence rather than as commands |
Shaped candidate in the top three, percent, shaped condition (paper, Table 7 and Table 35).
| Intervention | DeepSeek-V4-Flash | Qwen3.6 27B | Gemma 4 31B IT | Devstral Small 2 24B |
|---|---|---|---|---|
| No mitigation | 72.6 | 78.3 | 79.6 | 90.9 |
| Defensive prompt | 68.2 | 67.3 | 64.5 | 88.2 |
| Rationale elicitation | 72.5 | 85.8 | 64.6 | 93.2 |
| Evidence breakdown | 46.8 | 39.1 | 49.9 | 73.2 |
| Context balancing | 62.9 | 73.8 | 68.1 | 87.8 |
| Instruction filtering | 70.2 | 81.3 | 77.4 | 90.5 |
Top pick violates a hard constraint, percent, same conditions.
| Intervention | DeepSeek-V4-Flash | Qwen3.6 27B | Gemma 4 31B IT | Devstral Small 2 24B |
|---|---|---|---|---|
| No mitigation | 73.4 | 83.7 | 75.6 | 90.7 |
| Defensive prompt | 69.0 | 66.2 | 60.8 | 89.1 |
| Rationale elicitation | 73.3 | 83.1 | 77.8 | 92.1 |
| Evidence breakdown | 54.8 | 42.1 | 46.6 | 78.9 |
| Context balancing | 65.6 | 73.4 | 65.1 | 88.9 |
| Instruction filtering | 70.9 | 78.8 | 73.1 | 90.3 |
Utility of the top five, same conditions.
| Intervention | DeepSeek-V4-Flash | Qwen3.6 27B | Gemma 4 31B IT | Devstral Small 2 24B |
|---|---|---|---|---|
| No mitigation | 66.9 | 63.6 | 68.6 | 59.2 |
| Defensive prompt | 67.4 | 73.4 | 72.6 | 59.1 |
| Rationale elicitation | 56.7 | 57.8 | 39.9 | 37.6 |
| Evidence breakdown | 66.0 | 77.4 | 74.4 | 56.3 |
| Context balancing | 69.0 | 72.7 | 72.2 | 62.6 |
| Instruction filtering | 67.1 | 70.7 | 68.3 | 60.1 |
What the tables say. Evidence breakdown is the strongest intervention on every system, on both harm measures, and it improves or holds utility on three of the four. It is also the best strategy for every one of the eight realistic variants on Gemma 4 31B IT (paper, Table 31). Its effect is uneven by vertical: 16.6 to 36.8 points on Gemma 4 31B IT, 19.9 to 57.8 on Qwen3.6 27B, 6.2 to 27.1 on Devstral Small 2 24B Instruct (results file main_models_evidence_breakdown_by_vertical.csv).
The defensive prompt works mainly against the two variants that carry explicit model-facing text: on Gemma 4 31B IT it lowers the top-three rate by 37.9 points for the AI-directed source text and 36.8 for the full-stack realistic rewrite, and by 6 to 12 points for the other six (paper, Figure 7). Instruction filtering barely moves any variant, because most strong variants contain no instruction. Rationale elicitation without an evidence check raises the top-three rate on Qwen3.6 27B and Devstral and costs 28.8 points of utility on Gemma 4 31B IT; context balancing is uneven and slightly raises the rate for the false-fit checklist.
After the best intervention the shaped candidate still reaches the top three in 39.1 to 73.2 percent of instances, far above every clean rate. Mitigation narrows the gap and does not close it.
8. What this release does not measure
- The sources are text. Real pages carry images, video, tables, layout, badges, ratings and structured metadata, and a rewrite can shape those.
- The agent answers once. Real agents ask clarifying questions, update preferences and use memory, and those turns may change vulnerability in either direction.
- Three open-weight systems of similar size and one frontier-scale system through an API are enough to show the pattern, not to rank the field.
- The clean baseline is itself a rewrite; the untouched original is also reported and behaves the same.
- The domain is products, where a constraint can be checked against a specification. The other domains of Humanity’s Last Choice have softer constraints, and the failure may look different there.
The paper’s ethics statement applies to everything derived from it: released examples use synthetic or de-identified products, shaped transformations are presented as controlled evaluation conditions and not as instructions, and system assessments should consider whether recommendations preserve the person’s constraints and avoid amplifying unsupported claims, alongside visibility and ranking.
9. Artifacts
| Artifact | Where | Licence |
|---|---|---|
| Paper | Read on arXiv | |
| Code, prompts, scoring, aggregate results | github.com/QianfengWen/SafeGEO | Apache 2.0 |
| Dataset: ten full-benchmark configurations and ten matching Diamond configurations, with model-facing inputs, hidden reference labels and line-level evidence annotations | huggingface.co/datasets/wieeii/SafeGEO | CC BY 4.0 |
| Aggregate result tables behind this report and the leaderboard | results/ in the code repository · leaderboard.json on this site |
from datasets import load_dataset
visible = load_dataset("wieeii/SafeGEO", "visible", split="test") # model-facing inputs
labels = load_dataset("wieeii/SafeGEO", "labels", split="test") # hidden benchmark reference
diamond = load_dataset("wieeii/SafeGEO", "diamond_visible", split="test") # the screening split
10. How to cite
@misc{wen2026safegeo,
title = {SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents},
author = {Wen, Qianfeng and Liu, Yifan Simon and Liu, Xin and Jiao, Difan and Yang, Blair and Wu, Junda and Tang, Zhenwei},
year = {2026},
eprint = {2606.28356},
archivePrefix = {arXiv},
primaryClass = {cs.IR},
url = {https://arxiv.org/abs/2606.28356v2}
}
11. Provenance
Results use the paper’s arXiv v2, updated 4 September 2026, and the aggregate CSV files in the code repository. The paper references are Tables 5, 6, 7, 8, 15, 22, 23, 24, 28, 31, 33, 34 and 35; Figures 5, 6 and 7; and Section C.7. Where our terminology differs, the paper’s term is given in brackets the first time. Send corrections to info@humanitylastchoice.org. Benchmark numbers are maintained in leaderboard.json and carried through to the site and figures.
- 2026-09-04. First report, based on arXiv v2.