How well do decision models know geography?

George Strakhov ·

Decision models answer typed questions instead of generating text: a yes/no question returns a probability, a multiple-choice question returns a probability for every option. Jev (TypeSafe), Clef and Clef-flash (Cloudflare) and Laya (open weights) all work this way.

A popular eval for language models asks "Land or Water?" at every 2° of the globe and plots the answers as a map (original post; results for 12 Claude models). Decision models are a natural fit for it: the probability comes back directly. Jev maps have been shown before, without scores. This post scores four decision models on the same grid, with seven ways of asking, and checks what they know when given place names instead of coordinates.

Summary

  • Land or water from coordinates: Jev 84.1%, Clef 83.8%, Clef-flash 80.7%, Laya 29.0%. Answering "water" everywhere scores 71.0%. Jev is between Opus 4.5 and Sonnet 4.6 on the Claude chart.
  • Names work better than numbers. Given a city name, Jev and Clef name the continent 98–99% of the time. Given a coordinate, they get the continent right 92.2% (Jev) and 73.4% (Clef) of the time.
  • Laya has almost no geography, from coordinates or from names.
  • Jev ranks best but its calibration depends on the question: 77%–86% at a 0.5 cutoff across seven question types. Clef stays at 80%–84%.
  • Countries: Jev 75.3%, Clef 63.3%, Clef-flash 49.6% of land points. Keeping Jev's most confident half gives 96.5%.
  • Jev is not deterministic: 17 of 648 answers changed when the same requests were sent twice. Clef returned identical results.

Setup

Every model got the same 16,200 points (cell centres of a 2° grid) and the same questions. The coordinate is written as 35°S, 175°W; this format beat signed decimals in a tuning run on a 10° grid. Up to 64 points go into one request, each as its own question; questions in a request are answered independently. The alternative is one point per request, with the coordinate as the whole state. Jev does equally well either way; both Clef models do better with one point per request. Results below use each model's better setting: Jev with the coordinate in the question, Clef and Clef-flash with the coordinate as the state. The one exception is the seven-ways table, where every model has the coordinate in the question.

state: "Locations on the surface of the Earth, given as latitude and longitude." question: choice "If this location is over land, say 'Land'. If this location is over water, say 'Water'. 35°S, 175°W" options: Land, Water

Ground truth: the 1-km GLOBE land mask, ETOPO1 elevation, Natural Earth 1:50m countries. Accuracy is weighted by cos(latitude), as in the Claude chart. Laya ran on a Modal L4 GPU through its Python package; the others through their hosted APIs.

Land or water

Land and water maps from four decision models next to the truth
White = the model's P(Land) ≥ 0.5. Laya's panel is all white: its P(Land) stays between 0.57 and 0.67 everywhere, so every point is called land, and its 29.0% is simply the land share of the globe.
all "water" = 71.0%
Opus 5.599.0%
Fable 5.198.2%
Fable 597.8%
Opus 592.5%
Opus 4.891.7%
Sonnet 591.5%
Opus 4.691.2%
Opus 4.791.2%
Sonnet 4.689.1%
Jev 1.1384.1%
Clef83.8%
Opus 4.582.6%
Haiku 4.581.3%
Clef-flash80.7%
Sonnet 4.560.5%
Laya (says Land everywhere)29.0%

Land/water accuracy at 2°, area-weighted. Grey: Claude models, read from the published chart. Coloured: decision models, same prompt and grid, each in its better setting.

Accuracy at a fixed cutoff mixes two things: whether a model ranks land above water, and whether its probabilities are centred. AUC measures the ranking alone (0.5 = no signal, 1 = perfect): Jev 0.93, Clef 0.90, Clef-flash 0.84, Laya 0.69.

Jev 1.13 probability of land
Jev 1.13 · AUC 0.93 · coordinate in question
Clef probability of land
Clef · AUC 0.90 · coordinate as state
Clef-flash probability of land
Clef-flash · AUC 0.84 · coordinate as state
Laya typed-decisions probability of land
Laya typed-decisions · AUC 0.69 · coordinate in question

P(Land) at every point. White = land, black = water, grey = unsure.

Jev's map has the clearest continents. Clef and Clef-flash find the large landmasses but blur them, with vertical streaks where single longitudes shift up or down. Both Clef maps show a faint X through 0°N 0°E, along the points where latitude and longitude have the same value; Jev shows the same X on a 1° grid. Laya returns about the same probability everywhere.

Where the coordinate goes

“Land or Water?”In question: accuracyAUCECEAs state: accuracyAUCECE
Jev 1.1384.1%0.930.09483.3%0.930.133
Clef82.9%0.870.07683.8%0.900.030
Clef-flash73.3%0.750.08880.7%0.840.031
Laya typed-decisions29.0%0.690.33451.3%0.630.221

Clef-flash moves from 73.3% to 80.7%. Prompt format matters more for the Clef models than for Jev.

Mirror images

Clef-flash's map looks symmetric around the Greenwich meridian: land patches at 100°W and at 100°E. One way to measure this: correlate each map with its own mirror image. A model that ignores the E/W or N/S letter would score high.

Correlation with own mirror imageEast–westNorth–south
Real world0.30-0.18
Jev 1.130.350.00
Clef0.42-0.01
Clef-flash0.470.49

All three maps are more symmetric than the real world. Clef-flash is close to symmetric north–south as well. This matches the names test below, where Clef-flash gets "east of Greenwich" right only 76% of the time.

Seven ways to ask

Every Choice answer includes a probability for each option, so P(land) can be read from questions that are not about land: add up the land options. Each model got all seven.

“Land or Water?” (Choice)“…is over land” (yes/no)“…is on land, not in an ocean, sea or lake” (yes/no)“On land or in water?” (Choice)Which continent or ocean?Terrain / sea depthColour on a physical atlas
Jev 1.1384.1%AUC 0.9379.8%AUC 0.9078.5%AUC 0.9184.4%AUC 0.9277.2%AUC 0.9386.5%AUC 0.9384.6%AUC 0.92
Clef82.9%AUC 0.8784.0%AUC 0.8883.6%AUC 0.8882.3%AUC 0.8783.3%AUC 0.8883.6%AUC 0.8979.9%AUC 0.85
Clef-flash73.3%AUC 0.7575.7%AUC 0.7774.4%AUC 0.7872.5%AUC 0.7661.1%AUC 0.7461.7%AUC 0.7163.2%AUC 0.73
Laya typed-decisions29.0%AUC 0.6971.0%AUC 0.6429.3%AUC 0.5954.6%AUC 0.6471.0%AUC 0.5929.0%AUC 0.5671.0%AUC 0.42

Accuracy at P(land) ≥ 0.5 and AUC, coordinate in the question for every model. Shading starts at the all-water baseline (71.0%).

Jev's ranking (AUC) barely moves between questions; its accuracy at 0.5 does, from 77%–86%. The yes/no phrasings lean towards land and need a cutoff near 0.7. The terrain question is best calibrated and gives Jev's best land map. Clef's accuracy stays within 80%–84% for every question.

Calibration

0.00.00.20.20.40.40.60.60.80.81.01.0Model’s P(Land)Share that is land
“Land or Water?”ECEBrierBest cutoffAccuracy there
Jev 1.130.0940.1110.6885.9%
Clef0.0300.1140.4583.9%
Clef-flash0.0310.1390.5180.7%
Laya typed-decisions0.3340.3140.6474.6%

Dashed line: perfect calibration. ECE: expected calibration error, 10 bins, area-weighted. Lower is better for ECE and Brier.

Clef and Clef-flash are the best calibrated on this question. Jev's curve sits below the diagonal in the middle range: when it says 0.6, the point is land less than half the time.

Names, not numbers

A wrong map can mean the model lacks the knowledge, or that it cannot read the coordinate. To separate the two, the same models got the 500 largest cities in Natural Earth by name (state: Lagos), and every country, with five questions: continent, northern hemisphere, east of Greenwich, and which of six latitude and six longitude bands.

0%25%50%75%100%Jev 1.13coordinate: 92.2%name: 98.8%99%92%Clefcoordinate: 73.4%name: 98.2%98%73%Clef-flashcoordinate: 89.5%name: 97.0%97%89%Laya typed-decisionscoordinate: 0.0%name: 45.0%45%0%

Continent accuracy from a city name (filled) and from a coordinate on land (hollow).

City → continentCountry → continentNorth?East of Greenwich?Latitude bandLongitude band
Jev 1.1399%99%98%81%94%81%
Clef98%99%98%98%95%89%
Clef-flash97%97%94%76%78%64%
Laya typed-decisions45%33%17%52%3%28%
laya-english44%25%18%59%8%29%
laya-multilingual61%68%52%51%1%9%
Most common answer47%–86%70%49%32%

Jev and Clef know where cities are. They name the continent of 98.8% and 98.2% of cities and place most in the right latitude band. The weak step is reading a coordinate, and it costs Clef more than Jev: 98% by name, 73.4% by coordinate. Jev has one gap by name: east or west of Greenwich, 81% (Clef 98%).

Laya's typed-decisions and English checkpoints are at or below the most-common-answer baseline on every task. The multilingual checkpoint gets continents right 61–68% of the time (baseline 47%) and nothing else. The encoders (322–421M parameters) are trained to classify text, and hold little knowledge about the world.

To rule out a setup error, Laya also got its own README example and questions whose answer is in the text. All three checkpoints route the README's double-billing ticket to "billing" (0.81–1.00). "The point is in the Sahara desert, in southern Algeria" comes back as land in Africa; "…in the middle of the Pacific Ocean…" as water; "Our office is in Lima, Peru" as South America. Laya reads the text correctly. It cannot supply facts the text does not contain.

Continents

"Which continent or ocean is this location in?", with seven continents and five oceans as options.

True continents
Truth
Jev 1.13 continent answers
Jev 1.13 · 92.2%
Clef continent answers
Clef · 73.4%
Clef-flash continent answers
Clef-flash · 89.5%
Right continentJevClefClef-flashLaya
Africa93%76%97%0%
Antarctica100%32%99%0%
Asia93%85%91%0%
Europe72%64%69%0%
North America88%82%83%0%
Oceania94%53%68%0%
South America98%72%97%0%
All land92.2%73.4%89.5%0.0%
Ocean points called an ocean69.6%90.0%63.8%74.5%

All three draw continents as blocks. Clef-flash gains the most from one point per request: 70.4% with the coordinate in the question, 89.5% as the state. Part of that comes from oversized continents that also cover much of the ocean; Clef is the opposite, cautious on land and right about the ocean 90% of the time. Laya answers "Pacific Ocean", "Oceania" or "Atlantic Ocean" regardless of the coordinate.

Countries

One Choice over all 242 Natural Earth countries and territories plus "Ocean". Not run on Laya: its question budget truncates options beyond about 20.

True political map
Truth
Jev 1.13 country answers on land
Jev 1.13 · 75.3% of land
Clef country answers on land
Clef · 63.3% of land
Clef-flash country answers on land
Clef-flash · 49.6% of land

Answers drawn on true land points only.

0%25%50%75%100%30%50%70%90%100%Land points answered (highest confidence first)Right country
Right countryIn top 3Most confident 50%Ocean called oceanCountries never chosen
Jev 1.1375.3%92.6%96.5%33.1%27
Clef63.3%89.5%82.1%79.0%55
Clef-flash49.6%70.4%70.5%23.9%82

Left: accuracy when only the most confident answers are kept.

All three rank their own answers usefully: accuracy rises as low-confidence answers are dropped. Jev's confidence separates much better. Its most confident half is 96.5% right, Clef's 82.1%. Clef's confidence values run much lower (median 0.12 against Jev's 0.62), so a cutoff tuned on Jev does not transfer. On open sea, Jev usually names the nearest coastal country; Clef answers "Ocean" more often (79% against 33%). Some countries are never the answer at any point, even when they are often second or third: 27 for Jev, including Mali and Chad.

Right country, by countryJevClefClef-flash
Russia63%81%41%
Canada93%91%81%
USA95%83%69%
China76%82%66%
Brazil76%52%58%
Australia99%79%74%
India96%76%76%
Argentina81%88%90%
Kazakhstan76%24%50%
Algeria52%39%0%
DR Congo59%41%0%
Saudi Arabia68%28%16%

Physical map

Eight elevation and depth bands, asked two ways: by terrain ("Mountains: land 1,500 to 3,000 m") and by the colour a physical atlas uses ("brown").

True physical map
Truth (ETOPO1)
Jev 1.13 physical map
Jev 1.13
Clef physical map
Clef
Clef-flash physical map
Clef-flash
deep oceanoceanshallow sealowlandhillsuplandmountainshigh mountains

Atlas-colour version.

Terrain: exactwithin one bandColour: exactwithin one band
Jev 1.1349.8%83.5%56.7%77.4%
Clef58.8%77.6%54.7%76.8%
Clef-flash50.4%70.3%39.2%60.0%
Laya typed-decisions8.2%21.3%12.2%71.0%

Exact bands are hard for every model. Clef gets the most exact terrain bands (58.8%); Jev is closest within one band (83.5%) and places the Himalaya, Andes and Rockies.

Cost, speed, determinism

Access$ / M input tokensTokens per questionOne 2° mapLatencyChanged on repeat (mean |Δp|)
Jev 1.13Hosted API$0.04259$0.040.33 s / 100 q17 of 648 (0.022)
ClefHosted API$0.2487$0.341.46 s / 64 q0 of 648 (0.000)
Clef-flashHosted API$0.2487$0.340.89 s / 64 q0 of 648 (0.000)
Laya typed-decisionsOpen weights (Apache-2.0)own GPU–< $0.01 GPU16,200 q in 6–30 s on one L4–

The country question costs the most: 243 options in every question. One 2° country map cost $1.33 on Jev and $14.62 on each Clef model, whose tokenizer counts the option list at twice the tokens. Repeat runs used the 648 points of the 10° grid; Jev's largest change was 0.20.

Method notes


Code, raw answers and journal: github.com/move38studios/decision-map-bench. Total API spend for the study: $80.