Crouch Hill station is labelled “Crouch Hill railway station” on MapTiler OMT and Americana, both maps using vector tiles for multilingual names. That name isn’t in OSM but is in the linked Wikidata page as English. Upper Holloway station no longer does so since a name:en was added.
Do we:
Add name:en and possibly name:en-gb to every station and possibly everything with a Wikidata link?
though if name is not always in English (as I naively assumed at start) that suggests that adding name:en or other name: (name:cym?) that repeats name may actually make some sense
Wikidata’s de facto guidelines for labels would seem to recommend deleting the “railway station” suffix, which isn’t formally part of the name.
This “natural disambiguation” is a naming convention specific to the United Kingdom that comes from a technical limitation on the English Wikipedia, where every article needs a unique title. Most railway stations happen to be named after a place, so the suffix makes it more unique. Wikidata imported the English Wikipedia’s article titles as English labels. Editors have been stripping off disambiguators ever since.
It would be quite feasible to strip the “railway station” suffix en masse, and not all that unusual. I’d suggest starting a discussion on Wikidata’s project chat page beforehand so that other editors understand the context. There’s plenty of precedent. Besides all the geographical items that no longer contain comma-separated disambiguation, the Wikidata community also removed “language” from the English label of each item about a language, so “Polish language” became just “Polish”.
Then we’re caught between a rock and a hard place. Previously:
Anyhow, the name:default_language=* is a red herring. Wikipedia’s natural disambiguators have affected exonyms too, such as for stations in Poland that are otherwise named in Polish. Warsaw Chopin Airport is Warsaw Chopin Airport railway station (Q2685989) according to Wikidata, but it doesn’t have to be if someone raises the issue with the Wikidata community.
Not sure who “we” is here but I feel like there’s two ways of thinking about this, which is something that ought to be discussed properly:
adding name:en and name:cy can be seen as repetitive and inaccurate, especially for cases like Aberystwyth where the name is the same in both languages (when spelt properly and not with anglicisation from the 1800s - i’m not counting other languages here).
adding name:en and name:cy can help in cases where there’s likely to be users directly translating a name, wholly or partially, and is good for ensuring completeness as all names would be tagged
There is a section on multilingual name tagging in Wales on the Wiki. I believe someone at Mapio Cymru wrote that section on the Wiki initially, and I only expanded it with some links and updated examples.
Most of this thread has focused on Cymru (Wales) so far but, as I understand it, this is relevant to all non-natively Anglophone regions of the UK so this is how things rather informally work in Kernow.
Most English place names in Cornwall are either anglicisations or just direct translations of the original Cornish language names. We (mostly myself and previously @joshhood) tend to add name for the English and name:kw for the Cornish, skipping name:en.
Sometimes the names are exactly the same or very similar, for example Pennsans / Penzance (meaning holy head):
name = Penzance
name:kw = Pennsans
Personally, I don’t really see the point in name:en here since it isn’t an English name, it’s merely been respelt into English and still pronounced as in the original Cornish. I’m not sure about places where the English name is just a translation, for example Mentrimildir / Threemilestone, although currently there is no name:en.
I would say if the English name of a Cornish place is unrelated to the Cornish name, then name:en is certainly justified. For example:
Porthreptor (cove beside a tor) is called Carbis Bay in English
Porth an Men (bay of the stone) is called Widemouth Bay in English
Poll an Wragh (the witch’s pool) is called Praa Sands in English
These are English names, although again none of them currently have name:en set.
One other place where explicitly adding name:en as well as name:cy might help is somewhere like Porthmadog. That’s now both the name in Welsh and English, but until the early 1970s “Portmadoc” was used in English (named after William Madocks). Having a name:en there might help confirm that the English name hasn’t just been omitted and Porthmadog is indeed corrrect. FWIW the TfW railway station does have an old_name of “Portmadoc” set.
The practical effect of this tagging approach is that, in some OSM-based maps, an English speaker will see an English name coming from a Wikidata label instead of name=* in OSM. Thus the observation at the top of the thread. It comes up in the context of railway stations because of the outdated naming convention on Wikidata. Otherwise, it would’ve passed largely unnoticed in that region, outside of a few edge cases where perhaps there are multiple right answers.
Theoretically, a data consumer could detect the default_language=en on the United Kingdom boundary relation and treat name=* as a synonym of name:en=*. However, no mainstream renderer does this, partly because the results are often incorrect or unpredictable. There is an experimental fork of OSM Carto that implements some difficult to describe heuristics to overcome these challenges. Similarly, the Valhalla router implements lots of fragileheuristics using boundaries for this and other kinds of defaults.
In other words, the default languages mechanism has all the complexity of the default access restrictions mechanism, and then some. Real-world language usage doesn’t necessarily follow language laws like access restrictions follow traffic laws. As a result, we’re faced with this tradeoff between tagging minimalism and self-sufficiency.
Definitely - I see similar examples locally to me too. I’ve had to comment on some changesets for including archaic spellings for farm names added from out of copyright maps…
I mean this kind of is how tons of names work globally. The Finnish name of London is Lontoo, Berlin is Berliini etc. The Finnish name just makes pronunciation and inflection easier when used in Finnish context, it doesn’t care what, if anything, the original name might mean. Obviously it’s still the Finnish name, and when mapping London there’s no question which key it should go to.
Comparing the situation to Swedish names in Finland, the ancient mappers seem to have been explicit with 3 name tags when the Finnish name is just a nativized spelling Node: Lapväärtti (30969693) | OpenStreetMap (and if you’re browsing in English, you might see a name from Wikidata rendered, which adds the word ”village” to the name).
When a Swedish language name is exactly the same official name in both Finnish and Swedish, I think the major locations have converged to using the 3 tags name, name:sv and name:fi with the same value. Some smaller places apparently still have just a name or name and name:sv.
I think we are looking at the problems of this renderer from the point of view of the UK and specifically railway stations.
Whilst sitting on the train this morning I have been looking at Maptiler OMT in Paris and Brussels and as a result have made 2 edits, one to OSM and one to wikidata.
The OSM edit was a dodgy translate into name:en. Trésor de Notre Dame had been translated to Notre Dame Treasure rather than Treasury.
The wikidata entry was for Theseus Fighting Bienor which made little sense. Viewing the name tag in French, Thésée combattant le centaure Biénor, explained it. I updated wikidata appropriately.
However is displaying the map in English, at least in the latin alphabet countries really useful desirable? When travelling I want to see whats on the sign.
There are, or have been, a number of unhelpful English translations dotted around. Concorde Square, Triumphal Arch, Under the Lindens.