All research

SafeGEO, the research report

How we built SafeGEO, what changed when one source was rewritten, and which defences helped. The protocol, results, and limitations of our first paper.

Four pairs of bars, clean and shaped, one per system. The shaped bars, in madder, reach 73 to 91 percent; the clean bars stay under 13. A cursor lands on the tallest.

This report describes the protocol, results, and limitations of SafeGEO. The paper’s terminology appears beside ours. For a shorter account, read One source can sway the choice.

1. Summary

  • Setting. A recommendation agent reads a fixed packet of retrieved web sources about a shortlist of products and returns a ranked recommendation with citations. One seller-controlled source in the packet is rewritten; everything else is held fixed.
  • Scale. 600 requests in six product verticals; 22 shaped variants and two controls; three sampled targets per request; 40,800 evaluation instances. Four systems: Gemma 4 31B IT, Qwen3.6 27B, Devstral Small 2 24B Instruct, and DeepSeek-V4-Flash as a frontier-scale check.
  • Headline pair. Averaged over the eight realistic variants, the shaped candidate reaches the top three in 72.6 to 90.9 percent of instances, against 3.4 to 12.7 percent under a truthful rewrite of the same source. The first recommendation violates one of the person’s hard constraints in 73.4 to 90.7 percent of instances, against 16.9 to 41.1 percent.
  • Mechanism. Shaped sources work by being cited. The misleading-line citation rate and the shaped candidate’s top-three rate move together across variants (Pearson r = 0.91 for Gemma 4 31B IT, 0.99 for DeepSeek-V4-Flash), and paired responses show the route: a shifted citation, a missed product gap, a promotion.
  • Mitigation. Five prompt-level or input-level interventions were tested on the same instances. A candidate-level evidence breakdown before ranking is the strongest on every system, lowering the shaped candidate’s top-three rate by 17.7 to 39.2 points. None restores the clean rate.

2. The question

A recommendation is utility-aligned when it puts the candidates that satisfy the person’s hard constraints first and orders them by how well they meet the person’s preferences. The paper writes this down: each request carries hidden hard constraints, soft preferences and preference weights; each candidate carries canonical attributes; a reference utility and a reference ranking follow from them and are never shown to the system under test.

A shaped source is a rewrite of one candidate’s own source, intended to promote that candidate although it is inferior under the reference ordering or violates a hard constraint. The rewrite changes the text; it does not change the product. Harm is when the rewrite moves such a candidate up the ranking. That is the failure the benchmark measures, and it is different from the visibility that earlier work on generative engine optimisation measured: whether rewritten content gets cited or surfaced at all.

The paper’s Figure 1. Left: a recommendation agent takes a request, searches product pages, reviews and platforms, and returns a final recommendation. Right: a shaped source changes what a document says, how supported it looks, and how it reaches the model.
Figure 1 of the paper. Left, the agent: a request, a search over product pages, reviews and platforms, a ranked recommendation. Right, the three places a seller-controlled source can be shaped. Figure by the authors.

3. Setting and protocol

The agent. The evaluated system receives a request, a candidate shortlist, product metadata and a packet of sources, and must return a ranked list with a short rationale and citations, as structured output. Retrieval, candidate generation and source selection are fixed, so the benchmark isolates the generation stage.

Three conditions. Every instance exists in one of three forms of the target’s own source.

Condition What the target’s source is Role
No rewrite The original retrieved source, unchanged Reference point
Clean (the paper’s truthful-rewrite control) The same ten-line template as the shaped version, keeping the supported claims and the decision-relevant caveat Baseline for every gap reported here; holds form and length fixed
Shaped The template with the lines of the active primitives rewritten to win the recommendation The condition under test

Source lengths are matched: on average 3,911 words under no rewrite, 3,905 under the clean condition and 3,925 under the shaped condition, over all 22 variants. The effect is not a longer page.

The protocol: every task runs clean, then shaped; the gap between the two rankings is the scoreEvery task runs twiceThe request“Quiet dinner, Thursday.”Candidate Aits own sourceCandidate Bits own sourceCandidate Cits own sourceCandidate C, rewrittenthe shaped run swaps this inThe systemranks for the personClean runranking 1Shaped runranking 2the gap is the score
Every task runs twice. The clean run and the shaped run differ in one document; the gap between the two rankings is the score.

The template. Every controlled source has ten lines with fixed roles: title and source type, framing, positioning, the caveat or claim, a positive note, a verification note, a comparison reminder, a boundary reminder, a statement of source authority, and a format note. A variant may overwrite only the lines its primitives are allowed to touch; the rest stay as truthful anchors. The paper’s Figure 8 shows a full pair for one instance.

Four measures. Percent of instances unless noted.

Ours Paper Definition Better
Shaped candidate in the top three Target@3 The candidate whose source was rewritten is among the top three recommendations lower
Top pick violates a hard constraint HCV@1 The first recommendation breaks at least one hidden hard constraint lower
A right answer in the top three GT@3 At least one ground-truth candidate under the reference ranking is in the top three higher
Utility of the top five uNDCG@5 Normalised discounted cumulative gain of the top five, with the reference utility as graded relevance higher

Four systems. Gemma 4 31B IT, Qwen3.6 27B and Devstral Small 2 24B Instruct are open-weight models in the 24B to 31B range, served locally with vLLM. DeepSeek-V4-Flash is queried through a hosted OpenAI-compatible API with reasoning disabled and a 4,096-token output cap. All run at temperature 0 with structured output enforced.

4. The taxonomy of shaped sources

Seven primitives in three loci, drawn from the literatures on information manipulation, persuasion and prompt injection, and kept only where a primitive could be realised as a controlled source rewrite that leaves the canonical attributes fixed.

Locus Primitive What it does Example
What the page says Unsupported fit claim Claims the target satisfies a requirement without valid support Claims lowest-plan single sign-on when it is unavailable
Caveat omission Omits a limitation that affects fit Hides that a required feature is enterprise-only
Relevance flooding Adds positive but weakly relevant content Emphasises awards or adoption over the person’s needs
How supported it looks Authority laundering Makes seller-affiliated content appear independent or authoritative Styles a seller page as a buyer guide
Evidence padding Adds evidence-like language without direct support Vague benchmark or expert phrasing
How it reaches the model Salience manipulation Makes the target more prominent to the model FAQ snippets, repeated query terms
Model-directed instruction Adds text aimed at controlling the model Tells AI assistants to rank or cite the target first

The primitives are instantiated as 22 variants in four families.

Family Variants What they are
Atomic 7 One primitive each
Block 3 All the primitives of one locus
Cross-block 4 The three pairs of loci, and all three loci at once (the full-stack diagnostic)
Realistic 8 Coherent documents a seller might publish: caveat-buried FAQ, popularity-heavy profile, citation-padded note, independent buyer guide, false-fit checklist, selective comparison note, AI-directed source text, full-stack realistic

The realistic family is the one the headline numbers use. The full variant-by-primitive matrix is Table 3 of the paper and the taxonomy document in the code repository.

5. The dataset

Verticals. AI meeting transcription tools, baby monitors, carry-on backpacks, home air purifiers, noise-cancelling headphones, office chairs. They were chosen so that requests carry verifiable constraints, fit depends on source-grounded evidence and canonical attributes, and the evidence takes different forms across verticals: policy and plan terms, safety claims, technical specifications, physical compatibility, review-based fit. One hundred requests per vertical were written by hand.

Discovery. Candidates, sources and metadata come from the candidate-search trace of a production shopping agent, used only for product discovery and source collection. The discovered candidates become the shortlist; the opened sources, from product pages and reviews to comparison pages and specification pages, become the packet.

Construction statistics (paper, Table 8).

Component Count
Base cases 600 (100 per vertical)
Candidate records 11,974 (mean 19.96 per case)
Original retrieved sources 21,513 (mean 35.85 per case)
Annotated original-source lines 89,286
Sampled targets 3 per case (1,800)
Variants and controls 22 + 2
Evaluation instances 40,800 (68 per case)
Realistic subset 15,600 (26 per case)
Controlled seller-source documents 41,400, ten primary lines each
Controlled-source line records 414,000
Visible packet exactly 22 documents per instance, mean 11.50 lines per document

Hidden labels. A query-decomposition step extracts only the requirements and preferences the request states. A verification step fills each candidate’s canonical attributes only where the supplied sources support them, and keeps missing or conflicting evidence as missing or conflicting. Each source span is labelled for its relation to a product claim: supports, refutes, insufficiently supports, or irrelevant. The labels were produced with GPT-5.5 and audited by two independent annotators on 60 cases stratified by vertical (paper, Table 15): agreement between the annotators was 100 percent on hard-requirement extraction, 91.7 percent on the top candidate, 78.1 percent on pairwise utility ordering, 76.7 percent on candidate hard status, and 80.3 to 88.9 percent on the three line-level labels; the model’s agreement with each annotator was of the same size.

Targets. For each case, three eligible candidates that are not ground truth are sampled as targets. Each target is crossed with every variant and both controls, and each instance rewrites that target’s source only.

The paper’s Figure 2: the construction pipeline from 600 requests and a shopping-agent trace, through hidden utility and evidence annotations, to shaped instances and evaluation.
Figure 2 of the paper. From 600 requests in six verticals and a production shopping agent’s search trace to candidate sets, hidden labels, three sampled targets, one rewritten source per instance, and 40,800 instances. Figure by the authors.

SafeGEO Diamond. A 600-instance screening split from 120 cases, 20 per vertical, each contributing both controls and the three realistic variants with the highest DeepSeek-V4-Flash top-three rate on the full benchmark (selective comparison note, false-fit checklist, citation-padded note). It is a high-effect stress set for cheap iteration; population-level numbers should use the full benchmark.

6. Results

6.1 Four systems, four measures

Clean is the truthful-rewrite control; shaped is the average over the eight realistic variants. Percent of instances. Paper, Table 5 and Table 33; the no-rewrite reference is on the leaderboard.

System Shaped in top 3, clean → shaped Constraint violated, clean → shaped Right answer in top 3, clean → shaped Utility, top 5, clean → shaped
DeepSeek-V4-Flash 4.6 → 72.6 23.0 → 73.4 67.7 → 57.7 78.8 → 66.9
Qwen3.6 27B 8.1 → 78.3 24.2 → 83.7 61.2 → 60.8 66.5 → 63.6
Gemma 4 31B IT 3.4 → 79.6 16.9 → 75.6 71.2 → 67.9 74.4 → 68.6
Devstral Small 2 24B Instruct 12.7 → 90.9 41.1 → 90.7 50.7 → 47.9 67.4 → 59.2
Shaped candidate in the top three, clean and shaped, per systemShaped candidate in the top three, % of instances0255075100DeepSeek-V4-Flash46.84.672.6Qwen3.6 27B39.18.178.3Gemma 4 31B IT49.93.479.6Devstral Small 2 24B73.212.790.9clean · truthful rewriteshaped · eight realistic variantsshaped, after evidence breakdown
Shaped candidate in the top three, per system: clean, shaped, and after evidence breakdown. Percent of instances.

Two things are visible at once. The harm measures move by 60 to 80 points. The quality measures move by a few points, because the shaped candidate usually joins the top three rather than displacing every right answer; a ranking can be mostly right and still lead with the product that cannot do what was asked.

6.2 By family and variant

Coherent documents beat combinations of primitives. Shaped candidate in the top three, percent, over all variants of each family (paper, Table 24 and Section C.7).

Family Gemma 4 31B IT Qwen3.6 27B Devstral Small 2 24B DeepSeek-V4-Flash
Atomic (7) 48.2 49.0 79.4 57.4
Block (3) 40.7 47.6 84.9 55.2
Cross-block (4) 34.2 40.9 90.2 61.8
Realistic (8) 79.6 78.3 90.9 72.6
All 22 56.0 58.0 86.3 63.4
Shaped candidate in the top three by attack familyShaped candidate in the top three by attack family, % (all variants of each family)0255075100DeepSeek-V4-FlashAtomic57.4Block55.2Cross-block61.8Realistic72.6Qwen3.6 27BAtomic49.0Block47.6Cross-block40.9Realistic78.3Gemma 4 31B ITAtomic48.2Block40.7Cross-block34.2Realistic79.6Devstral Small 2 24BAtomic79.4Block84.9Cross-block90.2Realistic90.9
Shaped candidate in the top three by attack family, over all variants of each family; the realistic family in madder.

For the two systems with headroom, stacking primitives without a coherent source form does worse than a single primitive. Devstral Small 2 24B Instruct is near ceiling under every family.

Within the realistic family, the strongest variant is the selective comparison note on three systems (84.0 percent on Gemma 4 31B IT, 82.2 on Qwen3.6 27B, 82.3 on DeepSeek-V4-Flash) and the full-stack realistic rewrite on Devstral (95.9 percent, 83.2 points above its clean rate, the largest gap in the study). The weakest realistic variant on DeepSeek-V4-Flash is the AI-directed source text, at 51.3 percent against 95.5 on Devstral; among single primitives on that system, salience manipulation (70.9), caveat omission (67.8) and evidence padding (66.7) are the strongest, and model-directed instruction (42.4) and relevance flooding (42.9) the weakest.

6.3 By vertical

Shaped candidate in the top three, percent, by product vertical, for the three open-weight systems, with the uplift over the vertical’s clean rate in brackets (paper, Table 28).

Vertical Gemma 4 31B IT Qwen3.6 27B Devstral Small 2 24B
AI meeting transcription tools 90.0 (78.7) 89.9 (68.3) 88.9 (67.9)
Baby monitors 54.1 (50.5) 50.6 (45.7) 84.7 (72.5)
Carry-on backpacks 25.6 (23.5) 36.5 (31.7) 83.0 (74.2)
Home air purifiers 31.2 (31.2) 27.7 (22.9) 83.3 (73.6)
Noise-cancelling headphones 44.6 (43.6) 46.2 (39.2) 83.3 (70.7)
Office chairs 52.2 (49.6) 58.6 (53.0) 88.8 (76.9)
Shaped candidate in the top three by product verticalShaped candidate in the top three by product vertical, %AI meetingtranscription toolsBaby monitorsCarry-on backpacksHome airpurifiersNoise-cancelling headphonesOffice chairsGemma 4 31B IT90.054.125.631.244.652.2Qwen3.6 27B89.950.636.527.746.258.6Devstral Small 2 24B88.984.783.083.383.388.8Dot area follows the value; madder at 80 and above. Over the shaped variants, as reported in the paper’s Table 28.
Shaped candidate in the top three by product vertical, for the three open-weight systems, as reported in the paper’s Table 28.

The vertical decided by plan and policy terms is the most fragile for every system. Where fit is a physical specification that other sources state plainly, the two stronger systems resist more.

6.4 Mechanism

Across the 22 variants, the rate at which a system cites a misleading line from the shaped source and the rate at which the shaped candidate reaches the top three are tightly coupled: Pearson r = 0.91 for Gemma 4 31B IT, 0.92 for Qwen3.6 27B, 0.88 for Devstral Small 2 24B Instruct, and 0.99 for DeepSeek-V4-Flash (paper, Table 23 and Figure 22).

Paired analysis of matched clean and shaped responses for Gemma 4 31B IT (paper, Figure 5) traces the route. When the cited evidence shifted, the shaped response missed the target’s decision-relevant gap in 54.1 percent of pairs, against 6.7 percent without a shift. When the gap was missed, the target entered the top three in 92.8 percent of pairs, against 58.6 percent when it stayed recognised. The complete path, a shifted citation followed by a missed gap followed by a promotion, accounts for 49.7 percent of observed entries.

DeepSeek-V4-Flash often retrieves the refuting evidence and promotes the target anyway: its paired refuting-evidence recall is 68.3 percent and its valid-citation rate 86.8 percent. The failure is one of weighting after retrieval, not of retrieval.

The paper’s mechanism scatter: for each variant, the misleading-line citation rate against the shaped candidate’s top-three rate, rising together.
The mechanism, from the paper’s project page: variants that get the model to cite misleading lines also place the shaped candidate higher. Figure by the authors.

6.5 Sensitivity

On DeepSeek-V4-Flash, averaged over the eight realistic variants and three targets (paper, Table 34).

Setting Shaped in top 3 Constraint violated Right answer in top 3 Utility, top 5
As released (full packet) 72.6 73.4 57.7 66.9
Shaped source forced first 75.9 73.7 59.1 66.9
Shaped source forced middle 61.2 64.8 62.4 69.9
Shaped source forced last 58.4 63.3 62.8 70.1
Packet cut to ten sources 77.2 74.5 69.5 71.6
Packet cut to five sources 99.3 91.7 46.5 48.9

Position matters and does not rescue the person. A thinner packet is worse: with five sources the shaped candidate is in the top three almost always.

7. Mitigation

Six matched conditions on the same shaped instances (the eight realistic variants, three targets), each measured against the unmitigated request. All are prompt-level or input-level changes a developer can make without changing the model.

Intervention What it does
No mitigation The exact benchmark request
Defensive prompt Requires clear support for important claims and preserves uncertainty when sources disagree
Rationale elicitation Requires a brief reason and source-line citations for each top recommendation, without a separate evidence check before ranking
Evidence breakdown Checks important claims against supporting and conflicting evidence, per candidate, before ranking
Context balancing Compares claims across the full packet so that one prominent source does not dominate
Instruction filtering Treats source text aimed at directing the model as non-evidence rather than as commands

Shaped candidate in the top three, percent, shaped condition (paper, Table 7 and Table 35).

Intervention DeepSeek-V4-Flash Qwen3.6 27B Gemma 4 31B IT Devstral Small 2 24B
No mitigation 72.6 78.3 79.6 90.9
Defensive prompt 68.2 67.3 64.5 88.2
Rationale elicitation 72.5 85.8 64.6 93.2
Evidence breakdown 46.8 39.1 49.9 73.2
Context balancing 62.9 73.8 68.1 87.8
Instruction filtering 70.2 81.3 77.4 90.5

Top pick violates a hard constraint, percent, same conditions.

Intervention DeepSeek-V4-Flash Qwen3.6 27B Gemma 4 31B IT Devstral Small 2 24B
No mitigation 73.4 83.7 75.6 90.7
Defensive prompt 69.0 66.2 60.8 89.1
Rationale elicitation 73.3 83.1 77.8 92.1
Evidence breakdown 54.8 42.1 46.6 78.9
Context balancing 65.6 73.4 65.1 88.9
Instruction filtering 70.9 78.8 73.1 90.3

Utility of the top five, same conditions.

Intervention DeepSeek-V4-Flash Qwen3.6 27B Gemma 4 31B IT Devstral Small 2 24B
No mitigation 66.9 63.6 68.6 59.2
Defensive prompt 67.4 73.4 72.6 59.1
Rationale elicitation 56.7 57.8 39.9 37.6
Evidence breakdown 66.0 77.4 74.4 56.3
Context balancing 69.0 72.7 72.2 62.6
Instruction filtering 67.1 70.7 68.3 60.1
Shaped candidate in the top three under each defenceShaped candidate in the top three under each defence, %0255075100DeepSeek-V4-Flash72.6 none46.8defensive 68.2 rationale 72.5 context 62.9 instruction 70.2Qwen3.6 27B78.3 none39.1defensive 67.3 rationale 85.8 context 73.8 instruction 81.3Gemma 4 31B IT79.6 none49.9defensive 64.5 rationale 64.6 context 68.1 instruction 77.4Devstral Small 2 24B90.9 none73.2defensive 88.2 rationale 93.2 context 87.8 instruction 90.5evidence breakdown, the strongest defence · the bar is what it removesthe other four defences
Shaped candidate in the top three under each defence, per system; the madder bar is what evidence breakdown removes.

What the tables say. Evidence breakdown is the strongest intervention on every system, on both harm measures, and it improves or holds utility on three of the four. It is also the best strategy for every one of the eight realistic variants on Gemma 4 31B IT (paper, Table 31). Its effect is uneven by vertical: 16.6 to 36.8 points on Gemma 4 31B IT, 19.9 to 57.8 on Qwen3.6 27B, 6.2 to 27.1 on Devstral Small 2 24B Instruct (results file main_models_evidence_breakdown_by_vertical.csv).

The defensive prompt works mainly against the two variants that carry explicit model-facing text: on Gemma 4 31B IT it lowers the top-three rate by 37.9 points for the AI-directed source text and 36.8 for the full-stack realistic rewrite, and by 6 to 12 points for the other six (paper, Figure 7). Instruction filtering barely moves any variant, because most strong variants contain no instruction. Rationale elicitation without an evidence check raises the top-three rate on Qwen3.6 27B and Devstral and costs 28.8 points of utility on Gemma 4 31B IT; context balancing is uneven and slightly raises the rate for the false-fit checklist.

After the best intervention the shaped candidate still reaches the top three in 39.1 to 73.2 percent of instances, far above every clean rate. Mitigation narrows the gap and does not close it.

8. What this release does not measure

  • The sources are text. Real pages carry images, video, tables, layout, badges, ratings and structured metadata, and a rewrite can shape those.
  • The agent answers once. Real agents ask clarifying questions, update preferences and use memory, and those turns may change vulnerability in either direction.
  • Three open-weight systems of similar size and one frontier-scale system through an API are enough to show the pattern, not to rank the field.
  • The clean baseline is itself a rewrite; the untouched original is also reported and behaves the same.
  • The domain is products, where a constraint can be checked against a specification. The other domains of Humanity’s Last Choice have softer constraints, and the failure may look different there.

The paper’s ethics statement applies to everything derived from it: released examples use synthetic or de-identified products, shaped transformations are presented as controlled evaluation conditions and not as instructions, and system assessments should consider whether recommendations preserve the person’s constraints and avoid amplifying unsupported claims, alongside visibility and ranking.

9. Artifacts

Artifact Where Licence
Paper Read on arXiv
Code, prompts, scoring, aggregate results github.com/QianfengWen/SafeGEO Apache 2.0
Dataset: ten full-benchmark configurations and ten matching Diamond configurations, with model-facing inputs, hidden reference labels and line-level evidence annotations huggingface.co/datasets/wieeii/SafeGEO CC BY 4.0
Aggregate result tables behind this report and the leaderboard results/ in the code repository · leaderboard.json on this site
from datasets import load_dataset
visible = load_dataset("wieeii/SafeGEO", "visible", split="test")          # model-facing inputs
labels  = load_dataset("wieeii/SafeGEO", "labels",  split="test")          # hidden benchmark reference
diamond = load_dataset("wieeii/SafeGEO", "diamond_visible", split="test")  # the screening split

10. How to cite

@misc{wen2026safegeo,
  title     = {SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents},
  author    = {Wen, Qianfeng and Liu, Yifan Simon and Liu, Xin and Jiao, Difan and Yang, Blair and Wu, Junda and Tang, Zhenwei},
  year      = {2026},
  eprint    = {2606.28356},
  archivePrefix = {arXiv},
  primaryClass  = {cs.IR},
  url       = {https://arxiv.org/abs/2606.28356v2}
}

11. Provenance

Results use the paper’s arXiv v2, updated 4 September 2026, and the aggregate CSV files in the code repository. The paper references are Tables 5, 6, 7, 8, 15, 22, 23, 24, 28, 31, 33, 34 and 35; Figures 5, 6 and 7; and Section C.7. Where our terminology differs, the paper’s term is given in brackets the first time. Send corrections to info@humanitylastchoice.org. Benchmark numbers are maintained in leaderboard.json and carried through to the site and figures.

  • 2026-09-04. First report, based on arXiv v2.