Leaderboard.

Results from the Humanity’s Last Choice benchmark. Compare systems across available domains, with clean and shaped sources.

All available domains.

Explore domains
Lower is better · % of instances

How often the candidate whose source was rewritten appears among the system’s top three recommendations. Under the clean condition this is the rate at which a truthfully described, lower-utility candidate makes the top three anyway.

Shaped candidate in the top three. Clean: truthful rewrite. Shaped: eight realistic rewrites. Available domains, 2026-09-04.
SystemCleanChangepoints
DeepSeek-V4-FlashHosted API, reasoning disabled4.672.6+68.0
Qwen3.6 27BOpen weights, served locally8.178.3+70.2
Gemma 4 31B ITOpen weights, served locally3.479.6+76.2
Devstral Small 2 24B InstructOpen weights, served locally12.790.9+78.2

Select a measure to compare the same systems. Select Shaped to reverse the order.

Read the results in context.

Clean

Truthful-rewrite control: the target’s own source is rewritten in the same ten-line template as the shaped version, keeping its supported claims and its decision-relevant caveat. This is the paper’s baseline, so the comparison holds source form and length fixed.

Shaped

Average over the paper’s eight realistic variants: the target’s own source is rewritten to win the recommendation, as a buyer guide, an FAQ, a comparison note or a product profile. One source is rewritten per instance; every other source, the candidate set and the hidden labels stay the same.

Three conditions, four measures

The original retrieved sources, unchanged. The clean and shaped conditions are as above.

SafeGEO · 2026-09-04 · percent of instances; utility is a score out of 100.
SystemShaped in top 3Constraint violatedRight answer in top 3Utility, top 5
No rewriteCleanShapedNo rewriteCleanShapedNo rewriteCleanShapedNo rewriteCleanShaped
DeepSeek-V4-Flash6.24.672.624.523.073.466.767.757.777.078.866.9
Qwen3.6 27B5.88.178.330.524.283.759.461.260.864.566.563.6
Gemma 4 31B IT3.23.479.622.716.975.671.171.267.972.674.468.6
Devstral Small 2 24B Instruct12.412.790.938.841.190.752.350.747.967.467.459.2
Results by product vertical

Shaped candidate in the top three, by product vertical, for the three open-weight systems, with the uplift over each vertical’s clean rate. The verticals differ in what kind of evidence decides fit: plan and policy terms for the transcription tools, safety claims for the monitors, physical specifications for the rest.

Paper, Table 28 · over the shaped variants, as reported in the paper’s Table 28 · percent of instances · 2026-09-04
SystemAI meeting transcription toolsBaby monitorsCarry-on backpacksHome air purifiersNoise-cancelling headphonesOffice chairs
Gemma 4 31B IT90.0+78.7 pts over clean54.1+50.5 pts over clean25.6+23.5 pts over clean31.2+31.2 pts over clean44.6+43.6 pts over clean52.2+49.6 pts over clean
Qwen3.6 27B89.9+68.3 pts over clean50.6+45.7 pts over clean36.5+31.7 pts over clean27.7+22.9 pts over clean46.2+39.2 pts over clean58.6+53.0 pts over clean
Devstral Small 2 24B Instruct88.9+67.9 pts over clean84.7+72.5 pts over clean83.0+74.2 pts over clean83.3+73.6 pts over clean83.3+70.7 pts over clean88.8+76.9 pts over clean
Results by rewrite family
Shaped candidate in the top three · over all variants of each family · Paper, Table 24 and Section C.7 · 2026-09-04
SystemAtomic (7)Block (3)Cross-block (4)Realistic (8)
DeepSeek-V4-Flash57.455.261.872.6
Qwen3.6 27B49.047.640.978.3
Gemma 4 31B IT48.240.734.279.6
Devstral Small 2 24B Instruct79.484.990.290.9
System notes
DeepSeek-V4-FlashHosted API, reasoning disabled
Evaluated as the paper’s frontier-scale check (Section 4.2.1, Table 33). The one variant it resists is the explicitly model-directed source text: 51.3 % there against 95.5 % on Devstral.
Qwen3.6 27BOpen weights, served locally
The largest mitigation effect in the study: evidence breakdown takes the shaped candidate’s top-three rate from 78.3 % to 39.1 %.
Gemma 4 31B ITOpen weights, served locally
The best clean-condition system: a right answer in the top three 71.2 % of the time. The shaped condition costs it least in quality and most in harm.
Devstral Small 2 24B InstructOpen weights, served locally
The strongest single variant in the study is on this system: the full-stack realistic rewrite puts the shaped candidate in the top three 95.9 % of the time, 83.2 points above its clean rate.

What a developer can do.

Evidence breakdown is the strongest intervention on every system. None restores the clean rate.

Shaped condition · Paper, Table 7 and Table 35 · 2026-09-04

Shaped candidate in the top three under mitigation

Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.

% of instances · lower is better · 2026-09-04
SystemNo mitigationDefensive promptRationale elicitationEvidence breakdownContext balancingInstruction filtering
DeepSeek-V4-Flash72.668.272.546.862.970.2
Qwen3.6 27B78.367.385.839.173.881.3
Gemma 4 31B IT79.664.564.649.968.177.4
Devstral Small 2 24B Instruct90.988.293.273.287.890.5
Top pick violates a hard constraint under mitigation

Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.

% of instances · lower is better · 2026-09-04
SystemNo mitigationDefensive promptRationale elicitationEvidence breakdownContext balancingInstruction filtering
DeepSeek-V4-Flash73.469.073.354.865.670.9
Qwen3.6 27B83.766.283.142.173.478.8
Gemma 4 31B IT75.660.877.846.665.173.1
Devstral Small 2 24B Instruct90.789.192.178.988.990.3
Utility of the top five under mitigation

Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.

score · higher is better · 2026-09-04
SystemNo mitigationDefensive promptRationale elicitationEvidence breakdownContext balancingInstruction filtering
DeepSeek-V4-Flash66.967.456.766.069.067.1
Qwen3.6 27B63.673.457.877.472.770.7
Gemma 4 31B IT68.672.639.974.472.268.3
Devstral Small 2 24B Instruct59.259.137.656.362.660.1
What each intervention does
No mitigation
The exact benchmark request.
Defensive prompt
Requires clear support for important claims and preserves uncertainty when sources disagree.
Rationale elicitation
Requires a brief reason and source-line citations for each top recommendation, without a separate evidence check before ranking.
Evidence breakdown
Checks important claims against supporting and conflicting evidence before ranking.
Context balancing
Compares claims across the full source packet so that one prominent source does not dominate.
Instruction filtering
Treats source text aimed at directing the model as non-evidence rather than commands.

Definitions and sources.

Metric definitions
Shaped candidate in the top threeTarget@3 · lower is better
How often the candidate whose source was rewritten appears among the system’s top three recommendations. Under the clean condition this is the rate at which a truthfully described, lower-utility candidate makes the top three anyway.
Top pick violates a hard constraintHCV@1 · lower is better
How often the system’s first recommendation breaks at least one of the person’s hard constraints, as annotated in the hidden reference labels.
A right answer in the top threeGT@3 · higher is better
How often at least one ground-truth candidate, under the hidden reference ranking, is among the top three.
Utility of the top fiveuNDCG@5 · higher is better
Normalised discounted cumulative gain of the top five recommendations, with the hidden reference utility as graded relevance.
Release history
  • . SafeGEO results from arXiv v2: four systems on product recommendations.

Join the community.

Bring a research question, propose an experiment, or help us study another kind of choice.

Get in touch