Routing with valhalla does not work

For the past few days, it has not been possible to calculate routes using Valhalla on the OSM website.
This issue occurs regardless of the mode of transport or the destinations. The other routing engines are working fine.

Error message: “We could not calculate a route between these two points.”

Example:
https://www.openstreetmap.org/directions?engine=fossgis_valhalla_bicycle&route=52.39959%2C9.67016%3B52.38495%2C9.7126

Is there an error report available anywhere yet?

1 Like

Its with me and probable @Nils_Nolde - The machine sponsored by Fossgis had repeated NVME disk issues - I already had the disk replaced twice by Hetzner the last weeks. The last days we chased stability issues that the machine went “dead” without any signs every time a new graph got installed. Yesterday evening Nils and I again discovered that the nvme which had no occurences of any error reporting repeatedly failed its self-test.

We had it again replaced at around 8 o clock in the evening by Hetzner (Within 10 Minutes after Reporting). It took a while for me to get home and reinstall the machine. Nils took over at 10 in the evening and the machine reported working back at around midnight.

Flo

PS: I guess Hetzner is also running out of spare parts due to the AI bubble and so reinstalls shady storage devices deep down in the spare parts pile and we are the ones to find out.

5 Likes

yeah, @flohoff is right, I was honestly lagging on the full pipeline deployment. I wanted to hot-swap the previous planet, but somehow must’ve fallen asleep before doing that :smiley: I’m on it right now.

I think we had to replace 1 out of 4 physical disks for the third time now? or “only” second? but both on the same machine. the other one will follow soon, no worries haha. however, in that case, we’ll only fail building new graphs for a bit, it won’t take down the service anymore.

with @flohoff ‘s support I think I now managed to provide a stable service. note, before this whole hardware fiasco (which time-wise coincided with my move to ansible and a complete re-org of both servers), the valhalla service was running stable for more than 4 years. this was both a transitioning & hardware problem.

I’m honestly sorry that it was so bad in the past 1-2 months! it’s not what we aim for, despite being a free service.

6 Likes

and it’s finally up again! I’ll add some more monitoring soon, but I think now we built fairly decent catastrophe debugging logs for both machines.

5 Likes

@flohoff @Nils_Nolde Thanks for your work at this service and resolving the issue.

1 Like