How well do decision models know geography?
Decision models answer typed questions instead of generating text: a yes/no question returns a probability, a multiple-choice question returns a probability for every option. Jev (TypeSafe), Clef and Clef-flash (Cloudflare) and Laya (open weights) all work this way.
A popular eval for language models asks "Land or Water?" at every 2° of the globe and plots the answers as a map (original post; results for 12 Claude models). Decision models are a natural fit for it: the probability comes back directly. Jev maps have been shown before, without scores. This post scores four decision models on the same grid, with seven ways of asking, and checks what they know when given place names instead of coordinates.
Summary
- Land or water from coordinates: Jev 84.1%, Clef 83.8%, Clef-flash 80.7%, Laya 29.0%. Answering "water" everywhere scores 71.0%. Jev is between Opus 4.5 and Sonnet 4.6 on the Claude chart.
- Names work better than numbers. Given a city name, Jev and Clef name the continent 98–99% of the time. Given a coordinate, they get the continent right 92.2% (Jev) and 73.4% (Clef) of the time.
- Laya has almost no geography, from coordinates or from names.
- Jev ranks best but its calibration depends on the question: 77%–86% at a 0.5 cutoff across seven question types. Clef stays at 80%–84%.
- Countries: Jev 75.3%, Clef 63.3%, Clef-flash 49.6% of land points. Keeping Jev's most confident half gives 96.5%.
- Jev is not deterministic: 17 of 648 answers changed when the same requests were sent twice. Clef returned identical results.
Setup
Every model got the same 16,200 points (cell centres of a 2° grid) and the same questions. The coordinate is written as 35°S, 175°W; this format beat signed decimals in a tuning run on a 10° grid. Up to 64 points go into one request, each as its own question; questions in a request are answered independently. The alternative is one point per request, with the coordinate as the whole state. Jev does equally well either way; both Clef models do better with one point per request. Results below use each model's better setting: Jev with the coordinate in the question, Clef and Clef-flash with the coordinate as the state. The one exception is the seven-ways table, where every model has the coordinate in the question.
Ground truth: the 1-km GLOBE land mask, ETOPO1 elevation, Natural Earth 1:50m countries. Accuracy is weighted by cos(latitude), as in the Claude chart. Laya ran on a Modal L4 GPU through its Python package; the others through their hosted APIs.
Land or water

Land/water accuracy at 2°, area-weighted. Grey: Claude models, read from the published chart. Coloured: decision models, same prompt and grid, each in its better setting.
Accuracy at a fixed cutoff mixes two things: whether a model ranks land above water, and whether its probabilities are centred. AUC measures the ranking alone (0.5 = no signal, 1 = perfect): Jev 0.93, Clef 0.90, Clef-flash 0.84, Laya 0.69.




P(Land) at every point. White = land, black = water, grey = unsure.
Jev's map has the clearest continents. Clef and Clef-flash find the large landmasses but blur them, with vertical streaks where single longitudes shift up or down. Both Clef maps show a faint X through 0°N 0°E, along the points where latitude and longitude have the same value; Jev shows the same X on a 1° grid. Laya returns about the same probability everywhere.
Where the coordinate goes
| “Land or Water?” | In question: accuracy | AUC | ECE | As state: accuracy | AUC | ECE |
|---|---|---|---|---|---|---|
| Jev 1.13 | 84.1% | 0.93 | 0.094 | 83.3% | 0.93 | 0.133 |
| Clef | 82.9% | 0.87 | 0.076 | 83.8% | 0.90 | 0.030 |
| Clef-flash | 73.3% | 0.75 | 0.088 | 80.7% | 0.84 | 0.031 |
| Laya typed-decisions | 29.0% | 0.69 | 0.334 | 51.3% | 0.63 | 0.221 |
Clef-flash moves from 73.3% to 80.7%. Prompt format matters more for the Clef models than for Jev.
Mirror images
Clef-flash's map looks symmetric around the Greenwich meridian: land patches at 100°W and at 100°E. One way to measure this: correlate each map with its own mirror image. A model that ignores the E/W or N/S letter would score high.
| Correlation with own mirror image | East–west | North–south |
|---|---|---|
| Real world | 0.30 | -0.18 |
| Jev 1.13 | 0.35 | 0.00 |
| Clef | 0.42 | -0.01 |
| Clef-flash | 0.47 | 0.49 |
All three maps are more symmetric than the real world. Clef-flash is close to symmetric north–south as well. This matches the names test below, where Clef-flash gets "east of Greenwich" right only 76% of the time.
Seven ways to ask
Every Choice answer includes a probability for each option, so P(land) can be read from questions that are not about land: add up the land options. Each model got all seven.
| “Land or Water?” (Choice) | “…is over land” (yes/no) | “…is on land, not in an ocean, sea or lake” (yes/no) | “On land or in water?” (Choice) | Which continent or ocean? | Terrain / sea depth | Colour on a physical atlas | |
|---|---|---|---|---|---|---|---|
| Jev 1.13 | 84.1%AUC 0.93 | 79.8%AUC 0.90 | 78.5%AUC 0.91 | 84.4%AUC 0.92 | 77.2%AUC 0.93 | 86.5%AUC 0.93 | 84.6%AUC 0.92 |
| Clef | 82.9%AUC 0.87 | 84.0%AUC 0.88 | 83.6%AUC 0.88 | 82.3%AUC 0.87 | 83.3%AUC 0.88 | 83.6%AUC 0.89 | 79.9%AUC 0.85 |
| Clef-flash | 73.3%AUC 0.75 | 75.7%AUC 0.77 | 74.4%AUC 0.78 | 72.5%AUC 0.76 | 61.1%AUC 0.74 | 61.7%AUC 0.71 | 63.2%AUC 0.73 |
| Laya typed-decisions | 29.0%AUC 0.69 | 71.0%AUC 0.64 | 29.3%AUC 0.59 | 54.6%AUC 0.64 | 71.0%AUC 0.59 | 29.0%AUC 0.56 | 71.0%AUC 0.42 |
Accuracy at P(land) ≥ 0.5 and AUC, coordinate in the question for every model. Shading starts at the all-water baseline (71.0%).
Jev's ranking (AUC) barely moves between questions; its accuracy at 0.5 does, from 77%–86%. The yes/no phrasings lean towards land and need a cutoff near 0.7. The terrain question is best calibrated and gives Jev's best land map. Clef's accuracy stays within 80%–84% for every question.
Calibration
| “Land or Water?” | ECE | Brier | Best cutoff | Accuracy there |
|---|---|---|---|---|
| Jev 1.13 | 0.094 | 0.111 | 0.68 | 85.9% |
| Clef | 0.030 | 0.114 | 0.45 | 83.9% |
| Clef-flash | 0.031 | 0.139 | 0.51 | 80.7% |
| Laya typed-decisions | 0.334 | 0.314 | 0.64 | 74.6% |
Dashed line: perfect calibration. ECE: expected calibration error, 10 bins, area-weighted. Lower is better for ECE and Brier.
Clef and Clef-flash are the best calibrated on this question. Jev's curve sits below the diagonal in the middle range: when it says 0.6, the point is land less than half the time.
Names, not numbers
A wrong map can mean the model lacks the knowledge, or that it cannot read the coordinate. To separate the two, the same models got the 500 largest cities in Natural Earth by name (state: Lagos), and every country, with five questions: continent, northern hemisphere, east of Greenwich, and which of six latitude and six longitude bands.
Continent accuracy from a city name (filled) and from a coordinate on land (hollow).
| City → continent | Country → continent | North? | East of Greenwich? | Latitude band | Longitude band | |
|---|---|---|---|---|---|---|
| Jev 1.13 | 99% | 99% | 98% | 81% | 94% | 81% |
| Clef | 98% | 99% | 98% | 98% | 95% | 89% |
| Clef-flash | 97% | 97% | 94% | 76% | 78% | 64% |
| Laya typed-decisions | 45% | 33% | 17% | 52% | 3% | 28% |
| laya-english | 44% | 25% | 18% | 59% | 8% | 29% |
| laya-multilingual | 61% | 68% | 52% | 51% | 1% | 9% |
| Most common answer | 47% | – | 86% | 70% | 49% | 32% |
Jev and Clef know where cities are. They name the continent of 98.8% and 98.2% of cities and place most in the right latitude band. The weak step is reading a coordinate, and it costs Clef more than Jev: 98% by name, 73.4% by coordinate. Jev has one gap by name: east or west of Greenwich, 81% (Clef 98%).
Laya's typed-decisions and English checkpoints are at or below the most-common-answer baseline on every task. The multilingual checkpoint gets continents right 61–68% of the time (baseline 47%) and nothing else. The encoders (322–421M parameters) are trained to classify text, and hold little knowledge about the world.
To rule out a setup error, Laya also got its own README example and questions whose answer is in the text. All three checkpoints route the README's double-billing ticket to "billing" (0.81–1.00). "The point is in the Sahara desert, in southern Algeria" comes back as land in Africa; "…in the middle of the Pacific Ocean…" as water; "Our office is in Lima, Peru" as South America. Laya reads the text correctly. It cannot supply facts the text does not contain.
Continents
"Which continent or ocean is this location in?", with seven continents and five oceans as options.




| Right continent | Jev | Clef | Clef-flash | Laya |
|---|---|---|---|---|
| Africa | 93% | 76% | 97% | 0% |
| Antarctica | 100% | 32% | 99% | 0% |
| Asia | 93% | 85% | 91% | 0% |
| Europe | 72% | 64% | 69% | 0% |
| North America | 88% | 82% | 83% | 0% |
| Oceania | 94% | 53% | 68% | 0% |
| South America | 98% | 72% | 97% | 0% |
| All land | 92.2% | 73.4% | 89.5% | 0.0% |
| Ocean points called an ocean | 69.6% | 90.0% | 63.8% | 74.5% |
All three draw continents as blocks. Clef-flash gains the most from one point per request: 70.4% with the coordinate in the question, 89.5% as the state. Part of that comes from oversized continents that also cover much of the ocean; Clef is the opposite, cautious on land and right about the ocean 90% of the time. Laya answers "Pacific Ocean", "Oceania" or "Atlantic Ocean" regardless of the coordinate.
Countries
One Choice over all 242 Natural Earth countries and territories plus "Ocean". Not run on Laya: its question budget truncates options beyond about 20.




Answers drawn on true land points only.
| Right country | In top 3 | Most confident 50% | Ocean called ocean | Countries never chosen | |
|---|---|---|---|---|---|
| Jev 1.13 | 75.3% | 92.6% | 96.5% | 33.1% | 27 |
| Clef | 63.3% | 89.5% | 82.1% | 79.0% | 55 |
| Clef-flash | 49.6% | 70.4% | 70.5% | 23.9% | 82 |
Left: accuracy when only the most confident answers are kept.
All three rank their own answers usefully: accuracy rises as low-confidence answers are dropped. Jev's confidence separates much better. Its most confident half is 96.5% right, Clef's 82.1%. Clef's confidence values run much lower (median 0.12 against Jev's 0.62), so a cutoff tuned on Jev does not transfer. On open sea, Jev usually names the nearest coastal country; Clef answers "Ocean" more often (79% against 33%). Some countries are never the answer at any point, even when they are often second or third: 27 for Jev, including Mali and Chad.
| Right country, by country | Jev | Clef | Clef-flash |
|---|---|---|---|
| Russia | 63% | 81% | 41% |
| Canada | 93% | 91% | 81% |
| USA | 95% | 83% | 69% |
| China | 76% | 82% | 66% |
| Brazil | 76% | 52% | 58% |
| Australia | 99% | 79% | 74% |
| India | 96% | 76% | 76% |
| Argentina | 81% | 88% | 90% |
| Kazakhstan | 76% | 24% | 50% |
| Algeria | 52% | 39% | 0% |
| DR Congo | 59% | 41% | 0% |
| Saudi Arabia | 68% | 28% | 16% |
Physical map
Eight elevation and depth bands, asked two ways: by terrain ("Mountains: land 1,500 to 3,000 m") and by the colour a physical atlas uses ("brown").




Atlas-colour version.
| Terrain: exact | within one band | Colour: exact | within one band | |
|---|---|---|---|---|
| Jev 1.13 | 49.8% | 83.5% | 56.7% | 77.4% |
| Clef | 58.8% | 77.6% | 54.7% | 76.8% |
| Clef-flash | 50.4% | 70.3% | 39.2% | 60.0% |
| Laya typed-decisions | 8.2% | 21.3% | 12.2% | 71.0% |
Exact bands are hard for every model. Clef gets the most exact terrain bands (58.8%); Jev is closest within one band (83.5%) and places the Himalaya, Andes and Rockies.
Cost, speed, determinism
| Access | $ / M input tokens | Tokens per question | One 2° map | Latency | Changed on repeat (mean |Δp|) | |
|---|---|---|---|---|---|---|
| Jev 1.13 | Hosted API | $0.042 | 59 | $0.04 | 0.33 s / 100 q | 17 of 648 (0.022) |
| Clef | Hosted API | $0.24 | 87 | $0.34 | 1.46 s / 64 q | 0 of 648 (0.000) |
| Clef-flash | Hosted API | $0.24 | 87 | $0.34 | 0.89 s / 64 q | 0 of 648 (0.000) |
| Laya typed-decisions | Open weights (Apache-2.0) | own GPU | – | < $0.01 GPU | 16,200 q in 6–30 s on one L4 | – |
The country question costs the most: 243 options in every question. One 2° country map cost $1.33 on Jev and $14.62 on each Clef model, whose tokenizer counts the option list at twice the tokens. Repeat runs used the 648 points of the 10° grid; Jev's largest change was 0.20.
Method notes
- The land mask counts large lakes (Caspian Sea, Lake Victoria) as land.
- Coordinate format was chosen on Jev and then checked on Clef and Clef-flash on the 10° grid: the same format wins for both. Where the coordinate goes does differ by model, and each section uses each model's better setting.
- Claude scores are read from the published chart; that eval read P(Land) from token log-probabilities.
- Laya was also run with the coordinate as the state instead of in the question. Results were the same: near chance.
- OpenAI has announced a decision API but has not published it, so it is not included.
Code, raw answers and journal: github.com/move38studios/decision-map-bench. Total API spend for the study: $80.