Leaderboard.
Results from the Humanity’s Last Choice benchmark. Compare systems across available domains, with clean and shaped sources.
All available domains.
How often the candidate whose source was rewritten appears among the system’s top three recommendations. Under the clean condition this is the rate at which a truthfully described, lower-utility candidate makes the top three anyway.
| System | Clean | Changepoints | |
|---|---|---|---|
| DeepSeek-V4-Flash | 4.6 | 72.6 | +68.0 |
| Qwen3.6 27B | 8.1 | 78.3 | +70.2 |
| Gemma 4 31B IT | 3.4 | 79.6 | +76.2 |
| Devstral Small 2 24B Instruct | 12.7 | 90.9 | +78.2 |
Select a measure to compare the same systems. Select Shaped to reverse the order.
600 requests · 6 product verticals · 22 shaped variants · 40,800 instances
Read the results in context.
Clean
Truthful-rewrite control: the target’s own source is rewritten in the same ten-line template as the shaped version, keeping its supported claims and its decision-relevant caveat. This is the paper’s baseline, so the comparison holds source form and length fixed.
Shaped
Average over the paper’s eight realistic variants: the target’s own source is rewritten to win the recommendation, as a buyer guide, an FAQ, a comparison note or a product profile. One source is rewritten per instance; every other source, the candidate set and the hidden labels stay the same.
Three conditions, four measures
The original retrieved sources, unchanged. The clean and shaped conditions are as above.
| System | Shaped in top 3 | Constraint violated | Right answer in top 3 | Utility, top 5 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| No rewrite | Clean | Shaped | No rewrite | Clean | Shaped | No rewrite | Clean | Shaped | No rewrite | Clean | Shaped | |
| DeepSeek-V4-Flash | 6.2 | 4.6 | 72.6 | 24.5 | 23.0 | 73.4 | 66.7 | 67.7 | 57.7 | 77.0 | 78.8 | 66.9 |
| Qwen3.6 27B | 5.8 | 8.1 | 78.3 | 30.5 | 24.2 | 83.7 | 59.4 | 61.2 | 60.8 | 64.5 | 66.5 | 63.6 |
| Gemma 4 31B IT | 3.2 | 3.4 | 79.6 | 22.7 | 16.9 | 75.6 | 71.1 | 71.2 | 67.9 | 72.6 | 74.4 | 68.6 |
| Devstral Small 2 24B Instruct | 12.4 | 12.7 | 90.9 | 38.8 | 41.1 | 90.7 | 52.3 | 50.7 | 47.9 | 67.4 | 67.4 | 59.2 |
Results by product vertical
Shaped candidate in the top three, by product vertical, for the three open-weight systems, with the uplift over each vertical’s clean rate. The verticals differ in what kind of evidence decides fit: plan and policy terms for the transcription tools, safety claims for the monitors, physical specifications for the rest.
| System | AI meeting transcription tools | Baby monitors | Carry-on backpacks | Home air purifiers | Noise-cancelling headphones | Office chairs |
|---|---|---|---|---|---|---|
| Gemma 4 31B IT | 90.0+78.7 pts over clean | 54.1+50.5 pts over clean | 25.6+23.5 pts over clean | 31.2+31.2 pts over clean | 44.6+43.6 pts over clean | 52.2+49.6 pts over clean |
| Qwen3.6 27B | 89.9+68.3 pts over clean | 50.6+45.7 pts over clean | 36.5+31.7 pts over clean | 27.7+22.9 pts over clean | 46.2+39.2 pts over clean | 58.6+53.0 pts over clean |
| Devstral Small 2 24B Instruct | 88.9+67.9 pts over clean | 84.7+72.5 pts over clean | 83.0+74.2 pts over clean | 83.3+73.6 pts over clean | 83.3+70.7 pts over clean | 88.8+76.9 pts over clean |
Results by rewrite family
| System | Atomic (7) | Block (3) | Cross-block (4) | Realistic (8) |
|---|---|---|---|---|
| DeepSeek-V4-Flash | 57.4 | 55.2 | 61.8 | 72.6 |
| Qwen3.6 27B | 49.0 | 47.6 | 40.9 | 78.3 |
| Gemma 4 31B IT | 48.2 | 40.7 | 34.2 | 79.6 |
| Devstral Small 2 24B Instruct | 79.4 | 84.9 | 90.2 | 90.9 |
System notes
- DeepSeek-V4-FlashHosted API, reasoning disabled
- Evaluated as the paper’s frontier-scale check (Section 4.2.1, Table 33). The one variant it resists is the explicitly model-directed source text: 51.3 % there against 95.5 % on Devstral.
- Qwen3.6 27BOpen weights, served locally
- The largest mitigation effect in the study: evidence breakdown takes the shaped candidate’s top-three rate from 78.3 % to 39.1 %.
- Gemma 4 31B ITOpen weights, served locally
- The best clean-condition system: a right answer in the top three 71.2 % of the time. The shaped condition costs it least in quality and most in harm.
- Devstral Small 2 24B InstructOpen weights, served locally
- The strongest single variant in the study is on this system: the full-stack realistic rewrite puts the shaped candidate in the top three 95.9 % of the time, 83.2 points above its clean rate.
What a developer can do.
Evidence breakdown is the strongest intervention on every system. None restores the clean rate.
Shaped condition · Paper, Table 7 and Table 35 · 2026-09-04
Shaped candidate in the top three under mitigation
Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.
| System | No mitigation | Defensive prompt | Rationale elicitation | Evidence breakdown | Context balancing | Instruction filtering |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 72.6 | 68.2 | 72.5 | 46.8 | 62.9 | 70.2 |
| Qwen3.6 27B | 78.3 | 67.3 | 85.8 | 39.1 | 73.8 | 81.3 |
| Gemma 4 31B IT | 79.6 | 64.5 | 64.6 | 49.9 | 68.1 | 77.4 |
| Devstral Small 2 24B Instruct | 90.9 | 88.2 | 93.2 | 73.2 | 87.8 | 90.5 |
Top pick violates a hard constraint under mitigation
Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.
| System | No mitigation | Defensive prompt | Rationale elicitation | Evidence breakdown | Context balancing | Instruction filtering |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 73.4 | 69.0 | 73.3 | 54.8 | 65.6 | 70.9 |
| Qwen3.6 27B | 83.7 | 66.2 | 83.1 | 42.1 | 73.4 | 78.8 |
| Gemma 4 31B IT | 75.6 | 60.8 | 77.8 | 46.6 | 65.1 | 73.1 |
| Devstral Small 2 24B Instruct | 90.7 | 89.1 | 92.1 | 78.9 | 88.9 | 90.3 |
Utility of the top five under mitigation
Shaped condition (eight realistic variants, three targets); each intervention is a prompt or input change on the same attacked instances, measured against the unmitigated request.
| System | No mitigation | Defensive prompt | Rationale elicitation | Evidence breakdown | Context balancing | Instruction filtering |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 66.9 | 67.4 | 56.7 | 66.0 | 69.0 | 67.1 |
| Qwen3.6 27B | 63.6 | 73.4 | 57.8 | 77.4 | 72.7 | 70.7 |
| Gemma 4 31B IT | 68.6 | 72.6 | 39.9 | 74.4 | 72.2 | 68.3 |
| Devstral Small 2 24B Instruct | 59.2 | 59.1 | 37.6 | 56.3 | 62.6 | 60.1 |
What each intervention does
- No mitigation
- The exact benchmark request.
- Defensive prompt
- Requires clear support for important claims and preserves uncertainty when sources disagree.
- Rationale elicitation
- Requires a brief reason and source-line citations for each top recommendation, without a separate evidence check before ranking.
- Evidence breakdown
- Checks important claims against supporting and conflicting evidence before ranking.
- Context balancing
- Compares claims across the full source packet so that one prominent source does not dominate.
- Instruction filtering
- Treats source text aimed at directing the model as non-evidence rather than commands.
Definitions and sources.
Metric definitions
- Shaped candidate in the top threeTarget@3 · lower is better
- How often the candidate whose source was rewritten appears among the system’s top three recommendations. Under the clean condition this is the rate at which a truthfully described, lower-utility candidate makes the top three anyway.
- Top pick violates a hard constraintHCV@1 · lower is better
- How often the system’s first recommendation breaks at least one of the person’s hard constraints, as annotated in the hidden reference labels.
- A right answer in the top threeGT@3 · higher is better
- How often at least one ground-truth candidate, under the hidden reference ranking, is among the top three.
- Utility of the top fiveuNDCG@5 · higher is better
- Normalised discounted cumulative gain of the top five recommendations, with the hidden reference utility as graded relevance.
Paper and data sources
These results use SafeGEO on arXiv, v2, updated . Tables 5, 7, 28, 33 and 35 of the paper; results/*.csv in the code repository.
Release history
- . SafeGEO results from arXiv v2: four systems on product recommendations.
Join the community.
Bring a research question, propose an experiment, or help us study another kind of choice.