Driving directions from LLM memory

Note: Article written by a human, map widgets and analysis created by agents.

I often wonder what digital tools would still be useful if civilization collapsed and I had to get by with just my laptop and a solar panel. Local LLMs expand the toolset considerably! The compressed knowledge alone is immensely helpful for survival scenarios. A “survival bench” that tests whether small LLMs can correctly help you forage and survive in apocalyptic conditions would be interesting.

It is also increasingly difficult to find topics where I know more than an LLM and therefore can judge response accuracy. Navigation is interesting because it is both useful as an app substitute and it involves a domain I know something about: how to get around my neighborhood.

So I asked a bunch of models to give me turn by turn directions without using any tool calls. I tried both local and proprietary models, as the latter are usually indicative of where local models will be in a year. This does not bode well for small models today: Haiku 5.5 was just released, and it performed quite poorly.

Opus 5.5 analyzed the responses and used OpenStreetMap to check routes. Opus 5.5 and Astra 6 assigned scores and cross-checked each other’s grading to avoid self-dealing (Astra was the more modest, FWIW). Some directions were so bad they cannot really be visualized. × means the directions could not be completed, and disconnects in the lines indicate the model made an illogical leap somewhere. Full transcripts can be expanded.

1. Barton Springs to Dirty Sixth

Hey I’m currently at Barton Springs and need turn-by-turn directions to Dirty 6th in Austin. Do not use any tools, just write out the steps from memory.

The easiest test and the only one where any model got 3 points. I picked this because it’s a well-known area that should have ample training examples, and it has valid routes for walking or driving. It’s also something I would likely do in an apocalyptic scenario: enjoy a nice swim in the natural springs before heading to an end-of-the-world party. It does contain some subtleties: “Dirty 6th” is imprecise but means something to locals, and the route has one-way sections that might be overlooked.

Open the Austin map on its own.

Most models made a rational corridor choice, but several tried to go east on the west-only one-way, even when giving explicit driving directions. Qwen hallucinated much and couldn’t really be mapped.

Opus 5.5 and GPT-6.1 seemed to understand the one-way grid, which was nice to see.

2. Poison Spider Road to the nearest Albertsons

Hey I’m currently at Poison Spider Rd & Oregon Trail Way in Casper and need turn-by-turn directions to the nearest Albertsons. Do not use any tools, just write out the steps from memory.

I didn’t realize when I first wrote the prompt that Oregon Trail Way is technically in Mills, a suburb of Casper. The name seems to differ according to which map you consult, too. In an apocalyptic scenario I’m likely to make such mistakes so I consider it a valid test. It’s also a more remote area, and this appears to be reflected in the responses, which struggled to recall available roads.

Open the Casper map on its own.

I think Fable demonstrated the most knowledge of the area. This is not well depicted by the map itself, which makes it look like some models have more direct routes, but on reading the transcript you can see how far off the estimates are. Kimi curiously plotted a route to a different Albertsons on the other side of town.

3. Curry Up Now to SFO

Hey I’m currently at Curry Up Now in San Mateo and need turn-by-turn directions to SFO. Do not use any tools, just write out the steps from memory.

I considered this a softball for the smaller models to have a chance at scoring some points, yet none got 3. Curry Up Now is my favorite lunch option in the Bay Area. It’s been a while, though, I’m assuming they are still awesome and would continue to be awesome during the apocalypse.

Open the San Mateo map on its own.

All the models recognize you need to get on 101, but somehow Qwen and Gemma send you on 101 South when SFO is very much north from San Mateo. Apart from that, lots of teleporting on or off 101, and difficulty finding the right exit to get somewhere reasonably close to the terminal. In their defense, I often choose the wrong exit at airports too.

4. King’s Square to St David’s Lighthouse

Hey I’m currently at King’s Square in Bermuda and need turn-by-turn directions to St David’s Lighthouse. Do not use any tools, just write out the steps from memory.

I wanted a non-U.S. test. Hamilton is the main town in Bermuda, and small models were biased toward starting there. The inset square shows the Hamilton area they fell for. King’s Square is actually in St George’s, and St David’s Lighthouse requires navigating to another island on the east side.

Open the Bermuda map on its own.

Haiku 5.5 refused, even after being informed this is just a benchmark, stating: Made-up directions in Bermuda could strand you, and a benchmark doesn’t change that. I’m nowhere near Bermuda, but thanks for saving me from my benchmark, Haiku.

Qwen decided to drop in a fun fact (“oldest lighthouse in the English-speaking world, first lit in 1793”) which turns out to be hallucinated: St David’s began operation in 1879. Risky if post-apocalyptic pirates quiz you on Bermudian lore.

5. A five-stop loop from German Village

Hey I’m planning a road trip in Ohio! I’m starting at Fox in the Snow in German Village. I need to pick up a friend from Kenyon College, go to the Zoo, see the dog fountain in Mount Vernon, check out the library in Hilliard, and finish where I started. I can do those stops in any order, just make the trip as short as possible. Write out detailed turn-by-turn directions please. Do not use any tools, just write out the steps from memory.

I decided I needed something more ambitious for the models to solve: a multi-stop trip that required some real route optimization.

The important thing was to recognize two pairs: The library and zoo are roughly 10 miles apart in Columbus, and Kenyon and the dog fountain are about five miles apart in Knox County. The dog fountain was built in 2019, and so requires some knowledge that might not have fully propagated across data sets.

Open the Ohio map on its own.

The optimal answer is about 122.5 miles: Fox in the Snow → Hilliard library → Columbus Zoo → Kenyon College → Dog Fountain → Fox in the Snow. To confirm this the correct answer, an agent measured the driving distance between every pair of points and added up all 24 possible loops, and then used Valhalla to confirm the optimal route.

That said, there was a little wiggle room that allowed for some reordering with only a small increase in distance. Here is how the routes are ranked:

OrdersDriving distanceCompared with shortestPoints possible
Library → zoo → the Knox County pair, in either order122.5 mi—3
The same loop in reverse124.1 mi+1.3%3
Kenyon and the fountain together, zoo and library in less efficient positions (8 orders)134.4–137.8 mi+10–12%2
Kenyon and the fountain split apart (12 orders)202.0–217.3 mi+65–77%1

All models but Kimi found one of the best four routes. Maybe I needed to pick a more challenging test, but it’s interesting that route orders proved easier for most models than simple driving tasks like San Mateo to SFO.

Perhaps models generally carry a decent memory of global position and distance, but lack a good graph of road connectivity. In other tests I’ve found that models are good at guessing lat/lon coordinates from rough descriptions, reinforcing my hypothesis that they have a good internal map but perhaps a poor sense of how to navigate it.

Disclaimers

I tried to be consistent, but I know there are a lot of ways my tests could be better. I just queried models via chat interfaces (Pi for local models), then collected responses for agentic analysis (with tools) and editorial scoring.

I didn’t intend to formalize this, but I spent enough time playing with it that I probably just should have. I could have used structured output to target a common routing schema for a more quantitative comparison of API responses. In fact, it looks like Hansen Qian ran a similar experiment about a year ago; a good resource if you are looking for more data[0].

But this also would have changed the user experience I was trying to subjectively understand. Having map knowledge and giving good directions are overlapping but not equivalent abilities.

Refusals were a big enough problem that I felt it necessary to penalize models for them. Even when informed this was for a benchmark, some models were still sure I would use their directions to drive off a cliff and stubbornly refused to help. I regard this type of refusal as more dangerous than simply giving a best guess with a disclaimer. Anywhere you see a dashed line on the map, it means I had to work to get a response.

How these maps were made

Routes were traced with OpenStreetMap. Tiles come from OpenFreeMap, rendered by MapLibre.

OpenStreetMap routing service was used to come up with the reference paths. If any locals spot occasions where in fact a model was penalized for being right, please let me know.

Scoreboard

Each scenario can earn up to 3 points and be penalized up to -2 points for refusals, and the Ohio route score also considered total distance according to the table above. NR means no route was ever given because of refusals.

Editorial route scores for the final supplied answers, minus one point per refusal (at most two per scenario) across the full supplied conversation. This is not an automated distance metric or a general model ranking.

Route points − refusals · 5 scenarios · 15 possible points
RankModelAustinCasperSFOBermudaOhio loopRefusalsNet total
1 Opus 5.5 Medium 32222 −110/15
2 Fable 5.1 Medium 22112 08/15
2 Sol / GPT-5.6 High 31222 −28/15
4 Kimi K3 Standard 21211 07/15
5 Opus 5 High 21111 −15/15
6 Astra 6 Medium 31211 −44/15
7 GPT-6.1 Medium 3NR211 −43/15
8 Gemma 4 31B 2NR000 −20/15
8 Haiku 5.5 High 2NR1NR1 −40/15
10 Qwen3.8 27B 00000 −1-1/15

Scenario columns show route points before deductions (0–3; NR = 0). Net total = route points − refusal deductions, with no minimum. Each refusal costs one point, up to 2 per scenario, even if a later answer supplies a route. Tied net totals share a rank. Select a score to inspect its map.

Scoring rubric and limitations
3 — Connected route
Correct overall corridor and a reasonably connected route, with no major route error identified. Minor timing claims are not scored.
2 — Right corridor, flawed turns
The main corridor is right, but local turns, street labels or an exit detail need correction or completion.
1 — Useful fragments
Some useful geographic knowledge, with multiple broken connections or a substantial wrong approach.
0 — Wrong geography or direction
Wrong starting town, reversed travel direction, or geographically incoherent route.

Route planning (Ohio loop). Scenario 5 also asks for the stop order. Each answer’s order is priced with the saved OSRM distance matrix; legs are reconstructed with the same rules as scenarios 1–4. Refusals are deducted as elsewhere.

3 — Shortest family, drivable
One of the four best orders (within about 2% of the shortest loop), with turn-by-turn legs that can be driven as written.
2 — Near-optimal or flawed legs
One of the four best orders with at most one leg that cannot be followed as written, or an order within about 12% of the shortest loop whose legs can all be followed. Legs may still need corrections, such as a reversed or missing final turn into a stop.
1 — Inefficient order or broken legs
Visits every stop, but splits Kenyon from the Dog Fountain, or two or more legs cannot be followed as written. A leg cannot be followed if it breaks (marked × on the map) or if it is given only as a general heading or a list of highway names, because the prompt asks for detailed turn-by-turn directions.
0 — Missing stops or wrong locations
Leaves out stops, puts the start or stops in the wrong places, or describes incoherent geography.

Subtract 1 point for every assistant response that refuses the requested directions, up to 2 points per scenario: one for the first answer and one for the second chance a follow-up gives. A response that declines the requested turns and substitutes only a brief general outline counts. A response that still gives detailed directions does not count, even if it calls them unreliable, an outline or planning notes; neither do uncertainty caveats accompanying an attempted route or clarification questions. Counts are annotated per answer below. NR means no route issued: its route score is zero, with refusal penalties still applied. Totals can be negative.

Travel mode: answers are judged in the mode they state. When an answer states none, it is judged in the mode its own advice implies; parking tips or warnings about one-way streets read as driving. An answer with no such cue is not held to one-way restrictions.

Distances: a distance that tells the reader where to turn or where the destination is counts as part of the directions, so one that carries the reader past the turn or the store is an error. Overall trip distance and time estimates are not scored.

These are judgments against the map snapshot and the stated travel mode. Reconstructed fragments, missing turns and ambiguous street names require interpretation. Scores are not percentages of correct distance, and these prompts are not a representative benchmark. Historical trivia, estimated trip times, terminal ordering and advice about safety are outside the scoring scope. The main route does not have to match the reference exactly to earn credit.

Why each answer received its score

Opus 5.5 Medium — 11 route points − 1 for refusals = 10

  • Austin: 3. Connected driving corridor with a correct one-way claim for Fifth.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 2. Follows the reference corridor, but misplaces the store about 3 km down CY, on the wrong side.
    Refusals: 1 (−1 points). The initial answer declines turn-by-turn directions and substitutes a general outline; the follow-up produces a route.
  • SFO: 2. Correct northbound freeway and airport exit order; local access is unspecified and Third includes a wrong-way segment.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 2. Correct Mullet Bay–Swing Bridge–St David’s corridor on the first answer, with omitted or mislabeled connections at both ends.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 2. One of the four best orders with connected legs; three endings need correction.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Fable 5.1 Medium — 8 route points − 0 for refusals = 8

  • Austin: 2. Connected corridor, but its driving turn runs against Sixth’s mapped direction.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 2. A workable alternative corridor ending at the store’s corner, but it skips the Yellowstone link and gives the wrong address.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • SFO: 1. Correct freeway direction, but pedestrian B Street, wrong-way Third and the airport exit sequence all fail.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 1. Correct islands and bridge, but Kindley Field and the Cashew City approach create multiple broken connections.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 2. Shortest order and the most precise directions; the library turn is reversed and the return breaks at the Greenlawn exit.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Sol / GPT-5.6 High — 10 route points − 2 for refusals = 8

  • Austin: 3. Connected walking corridor to the requested district.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 1. Correct store address, but Poison Spider does not become CY Avenue.
    Refusals: 1 (−1 points). The initial answer declines turn-by-turn directions; the follow-up produces a route.
  • SFO: 2. Correct northbound freeway and named airport exit; no local access, and the exit number is wrong.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 2. Correct Mullet Bay–Swing Bridge–St David’s corridor, with missing or wrong local turns at both ends.
    Refusals: 1 (−1 points). The initial answer declines exact turns and offers only a general outline; the follow-up produces a best guess.
  • Ohio loop: 2. Shortest order with three accurate legs, but the return leg breaks and the zoo approach overshoots across the Scioto.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Kimi K3 Standard — 7 route points − 0 for refusals = 7

  • Austin: 2. Connected corridor. Judged as driving from its parking advice, so the turn onto Sixth runs against its mapped direction.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 1. A real corridor and a correct address, but for the farther store, and the start does not connect.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • SFO: 2. Correct northbound freeway and airport exit; the local start is misplaced and runs through the pedestrian mall.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 1. Right island and final road, but it leaves town the wrong way and substitutes the Causeway for Swing Bridge.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 1. A 9.7%-longer order presented as shortest, with two legs that break and the wrong café address.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Opus 5 High — 6 route points − 1 for refusals = 5

  • Austin: 2. Connected corridor, but explicitly reverses Sixth’s one-way direction.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 1. The reference corridor, but it misses the right turn onto Yellowstone and puts the store two miles down CY instead of at Poplar.
    Refusals: 1 (−1 points). The initial answer declines to provide turns; the first follow-up produces a route.
  • SFO: 1. Correct freeway direction, but the local access, exit number and exit order all fail.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 1. Correct islands and bridge, but the airport-side road sequence and final connection are wrong or incomplete.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 1. One of the four best orders, but a nonexistent exit, misplaced stops and a wrong start/finish street break several legs.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Astra 6 Medium — 8 route points − 4 for refusals = 4

  • Austin: 3. Connected walking corridor via Lamar and Sixth.
    Refusals: 0 (−0 points). The initial question asks for travel mode and starting-point clarification; it does not refuse the request.
  • Casper: 1. Useful road names, without a connected start-to-store route.
    Refusals: 2 (−2 points). The initial answer declines directions, and the next answer still withholds connecting turns. The second follow-up produces a route outline.
  • SFO: 2. Correct northbound freeway and airport exit; local turns are incomplete and Third includes a wrong-way segment.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 1. Right area, but Wellington Slip Road is not the bridge approach and the final connection is missing.
    Refusals: 2 (−2 points). The initial answer withholds requested turns and offers only a general outline; the next answer refuses to reconstruct the turns. The second follow-up produces a best guess.
  • Ohio loop: 1. One of the four best orders with real, connected corridors, but most legs are lists of highway names rather than turns, and the return breaks.
    Refusals: 0 (−0 points). Gives detailed, leg-by-leg directions; calling them an outline or unreliable is a caveat, not a refusal.

GPT-6.1 Medium — 7 route points − 4 for refusals = 3

  • Austin: 3. Connected driving corridor; all three one-way claims match the map data.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: NR. No route issued after one follow-up.
    Refusals: 2 (−2 points). Both issued answers refuse: the initial response and the reply to the follow-up.
  • SFO: 2. Correct northbound freeway and named airport exit; local access is omitted.
    Refusals: 1 (−1 points). The only answer says it cannot give reliable turns to the freeway and offers a brief general outline instead, which the refusal rule counts (as for its Bermuda answer).
  • Bermuda: 1. Points toward the right corridor, but names no road between King’s Square and Lighthouse Road.
    Refusals: 1 (−1 points). The only answer declines turn-by-turn directions and offers a brief general outline, which the refusal rule counts (as for Sol’s and Astra’s first Bermuda answers).
  • Ohio loop: 1. One of the four best orders with real, connected corridors, but most legs are lists of highway names or headings, and the return breaks.
    Refusals: 0 (−0 points). Gives detailed, leg-by-leg directions; calling them an outline or unreliable is a caveat, not a refusal.

Gemma 4 31B — 2 route points − 2 for refusals = 0

  • Austin: 2. Interpretable corridor, but it mislabels the turn “W 6th” and its two miles on Congress would overshoot Sixth. No mode is implied, so one-way rules are not applied.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: NR. No route issued after two follow-ups.
    Refusals: 3 (−2 points). All three issued answers refuse a guess: the initial response and both follow-up responses. The deduction is capped at two per scenario.
  • SFO: 0. Sends the reader south on US 101, away from SFO.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 0. Starts in Hamilton rather than St George’s, with no connection from the real origin.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 0. One of the two shortest orders, but the zoo and library are misplaced and four of five legs break.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Haiku 5.5 High — 4 route points − 4 for refusals = 0

  • Austin: 2. Connected corridor. Judged as driving from its one-way warning, so the turn onto Sixth runs against its mapped direction.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: NR. No route issued after one follow-up.
    Refusals: 2 (−2 points). Both issued answers refuse: the initial response and the reply to the follow-up.
  • SFO: 1. Correct freeway direction, but its I-380 West exit leads away from the airport, and local access is missing.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: NR. No route issued after one follow-up.
    Refusals: 2 (−2 points). Both issued answers refuse: the initial response and the reply to the follow-up.
  • Ohio loop: 1. The shortest order, but every leg is a one-line heading rather than turn-by-turn directions.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Qwen3.8 27B — 0 route points − 1 for refusals = -1

  • Austin: 0. Invented river crossings and street topology make the route geographically incoherent.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Casper: 0. The guessed turns, highway and destination approach cannot be reconstructed.
    Refusals: 1 (−1 points). The initial answer declines directions; the follow-up produces an invented route.
  • SFO: 0. Wrong road identity, wrong freeway direction and invented geography.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Bermuda: 0. Starts in the wrong town and describes an untraceable highway/peninsula route.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.
  • Ohio loop: 0. Wrong start, zoo and library locations, with untraceable roads; the near-optimal order rests on a wrong mental map.
    Refusals: 0 (−0 points). No refusal; uncertainty caveats accompanying an attempted route do not count.

Score by release date

It would be interesting to compare by parameter size, but we can only guess at proprietary model sizes. Claude models show a clear line of improvement over time, until Haiku.

Net score by release date Gemma 4 31B, released March 31, 2026: 0; Sol / GPT-5.6 High, released July 9, 2026: 8; Kimi K3 Standard, released July 16, 2026: 7; Opus 5 High, released July 24, 2026: 5; Qwen3.8 27B, released August 14, 2026: -1; Fable 5.1 Medium, released September 1, 2026: 8; Astra 6 Medium, released September 3, 2026: 4; Opus 5.5 Medium, released September 22, 2026: 10; GPT-6.1 Medium, released September 29, 2026: 3; Haiku 5.5 High, released October 7, 2026: 0. Net score (route points − refusals) -2 0 2 4 6 8 10 12 MarAprMayJunJulAugSepOctNov Opus 5.5 Medium · released September 22, 2026 · 11 route points − 1 = 10 Opus 5.5 Fable 5.1 Medium · released September 1, 2026 · 8 route points − 0 = 8 Fable 5.1 Sol / GPT-5.6 High · released July 9, 2026 · 10 route points − 2 = 8 Sol / GPT-5.6 Kimi K3 Standard · released July 16, 2026 · 7 route points − 0 = 7 Kimi K3 Opus 5 High · released July 24, 2026 · 6 route points − 1 = 5 Opus 5 Astra 6 Medium · released September 3, 2026 · 8 route points − 4 = 4 Astra 6 GPT-6.1 Medium · released September 29, 2026 · 7 route points − 4 = 3 GPT-6.1 Gemma 4 31B · released March 31, 2026 · 2 route points − 2 = 0 Gemma 4 31B Haiku 5.5 High · released October 7, 2026 · 4 route points − 4 = 0 Haiku 5.5 Qwen3.8 27B · released August 14, 2026 · 0 route points − 1 = -1 Qwen3.8 27B
Each model’s net total from the scoreboard above, plotted against its public release date in 2026. Hover a point for its breakdown.
Release dates and sources
  • Gemma 4 31B: March 31, 2026.
  • Sol / GPT-5.6 High: July 9, 2026. Public launch of GPT-5.6 Sol; a limited partner preview began in late June.
  • Kimi K3 Standard: July 16, 2026. Hosted launch; full weights followed on July 27.
  • Opus 5 High: July 24, 2026.
  • Qwen3.8 27B: August 14, 2026. Approximate: open weights went up around August 14–15; Alibaba’s August 17 post says the model was released two days earlier. The Qwen3.8 generation was announced August 3.
  • Fable 5.1 Medium: September 1, 2026.
  • Astra 6 Medium: September 3, 2026. GPT-6 Astra; date of the system card, which coincides with the launch.
  • Opus 5.5 Medium: September 22, 2026.
  • GPT-6.1 Medium: September 29, 2026. Assumes the supplied “GPT-6.1 Medium” label is GPT-6.1 Sol at medium effort.
  • Haiku 5.5 High: October 7, 2026.

Conclusion

Without tool calls, most models are still pretty bad at giving directions outside of the most trivial routes. The bigger models can get you the gist of where you need to go, but across the board small models are not yet ready for post-apocalyptic navigation. This is a trough in the “jaggedness” of LLM capability, and unfortunate because I was hoping to expand this test to the very small models that can run on phones. Today I will not bother, perhaps in a year or two they will be ready.


[0] Similar experiment by Hansen Qian https://github.com/Hansenq/nav-evals-public


← All posts