Hello United Kingdom forum / talk-gb mailing list, as previously discussed on talk-gb
( [Talk-GB] Revisiting OS trigpoints - viability of importing? ), I am proposing to import a subset of the Ordnance Survey (OS) triangulation station dataset, sourced from the OS website.
(as a new user I can only post 3 links - so please consult the wiki page and github pages for further details - I’ve had to remove a number of links from this message…)
You can also find the OSC XML file(s) intended to be used for the import there, and also a link to a slippy map view that can be helpful in evaluating the data.
License
I have checked with the LWG that this data is compatible with the ODbL.
This data is distributed under OGL-UK-3.0.
Abstract
The import will add any Ordnance Survey trigpoint pillar that currently has no representation in OSM within 15m. Note, it will not edit/merge/delete any existing OSM nodes.
Current expected import size is ~4437 new nodes.
**Please** consult the wiki and github pages and read the existing talk-gb thread before asking questions that may already have been answered!
I think ref:os is too generic, I would prefer if it had the country code in it and was clearer what it referred to, e.g. ref:GB:ordnance_survey or ref:GB:os_trigpoint.
I disagree with adding a note telling people not to move the point, unless the points in the OS database are always 100% correct.
Is it worth adding an operator and operator:wikidata tag? Do the OS “operate” these? As you mentioned man_made=survey_point is used for a few different things. Are there any trig point pillars in the UK that aren’t operated by OS?
I’ve looked at the wiki page you’ve created - it seems very comprehensive - good job! However, I have a few comments / suggestions…
First, I couldn’t see where OS have stated that the data can be used under the OGL. Are you able to demonstrate that, or explain how you know it’s true? (And how other people can verify it as well?) Could you also add more detail about what was checked with LWG, as most people can’t access the tickert? If this was just that OGL v3 data can be used in OSM, then that’s already well-established, so doesn’t need much more commentary. If it was something more complicated, then it would be good to spell this out.
It would also be good to have some commentary on the wiki page about the positional accuracy and completeness of the dataset. Do we believe the locations are always 100% accurate (e.g. to within 1m)? Do we believe that the data you’ve extracted is 100% complete and doesn’t contain any trig points that no longer exist?
I had a look at the slippymap, and was a bit confused by what was happening at OSM OK trig pillar import status . Am I right in thinking that the big red marker means you would be adding a new OSM node for S5275 there? But there’s an OSM node within 25m, which is almost certainly the trig point in question, even if it’s slightly misplaced. I don’t think we’d want to be duplicating data like this.
One option would be to manually review all unmatched OSM survey_point nodes prior to doing a new conflation before the bulk import. (I assume there may be a lot of flush-bracket benchmarks in the OSM data. Is there an established way to tag them differently from trig points?) Otherwise, I think you would need to significantly increase the radius for an OSM node preventing the import of a new node.
That node also raises another issue. While the trig points in the dataset might have very accurate positions, there is an issue if the existing mapping in their neighbourhood isn’t well aligned. You could end up importing trigpoints inside buildings or (in the example above) in the middle of roads, or on the wrong side of existing mapped roads/paths. Arguably, this isn’t really a problem with the import itself, but it would be good if this could somehow be addressed, by attempting to flag the alignment of any nearby objects for checking. Otherwise, it’s likely that well-meaning mappers might adjust the trig point location later on to make it’s position relative to the mis-aligned objects correct.
I don’t think the wiki page describes what you’re planning to do with the different refs very well. Am I right in thinking that there is a flush bracket number (that’s displayed on the pillar) and also an internal OS ref (the “new name” column in the trig-point dataset)? What are you proposing to add to what OSM key? Only ref:os is mentioned in the table on the wiki page. If both refs are to be added, they need to go in different keys, but in different places, the wiki page suggests both would be added to the same ref:os=* tag. I agree with the previous poster, that if we’re having a custom ref value, then we should use the now standard ref:GB:* form. I think the final part should reflect the scope of the dataset not just the assigning organisation (which may have several different datasets).
Also, could you add a bit more on the wiki page about what you’re doing to remove various lines from the main dataset (e.g. for trigpoints marked as destroyed). This filtering isn’t mentioned at all in the wiki (I think it should be). There’s also not much discussion in the github notes to explain how you decide on destroyed pillars, or where to find the code to review this. (The notes column is note entirely consistent with how these are marked, so presumably needs quite a bit of effort to parse.)
Sure, it makes sense. I only chose ref:os because that was a tag that was already in mild use with trigpoints. The only other similar tag in use is ref:GB:trigpointinguk, but that is specific to that trigpointing website, and not related to the trigpoint data itself. I’ll expand in Roberts reply…
I think the general premise should be that the points in the OS trig database are correct - possibly even fairly definitive. Sure, there are some still in the database which will since have been moved or destroyed for one reason or another and are not marked as such, but I think we can take it that almost all of the points will be accurate.
I would moot that they will be more accurate than anything measured with a consumer GPS or found via visual map alignment. In fact one key potential benefit of adding these trigpoints is that they could then be used for image alignment?
I really would like a note hinting that the trigpoint should not be moved without good reason - so I’m not saying they cannot be moved, but it should be a very concious and fairly proven decision. What I’m trying to avoid here is them being moved because they do not line up with somebodies GPS waymark or visual checking of an aerial image - as nominally the trigpoint data should be the definitive source and be used to correct the waymarks or images, no?
Ah, indeed, that is absolutely something we could and probably should add. I don’t think we could say the trigpoints are ‘operated’ any more by the OS, as they are no longer maintained and updated, but I believe they are still owned by the OS.
I’ve added a github issue #15 to track this so I don’t forget ;-)
afaik there are no other authorities who have or maintain, specifically, pillars that cover the same area as these trigpoints. There are other groups, such as trigpoints in Ireland, I believe, but this import does not cover those. But that doesn’t mean we should not be OS specific in our tagging.
I’m also hesitant on this note. I accept the reasoning, but all it takes is someone moving it (with good, but misguided intent) either in an application which doesn’t show the tag or just not paying attention, for the node to then be in the wrong position and the note still saying not to move it.
Having the note will probably reduce, but not eliminate the number of accidental moves, but we might need some form of a period re-run to fix those nodes that have wandered.
Sure, np. Creating the wiki isn’t so bad - at least there is a template supplied which makes it easier…
Sure. The OGL details are listed on the OS web pages, but annoyingly you have to click the ‘Read More’ to show the information. There it is stated to be OGL, with a link off to the OGL docs.
For LWG, I followed the guidelines around OGL checking requirements (inspired by CC0 requirements I believe), whereby OGL itself does not necessarily cover all the data, and you have to get each individual data field type you are using checked (I can’t find the relevant wiki page right now - there are quite a few with sometimes contradictory information :-( )… so I emailed LWG with links to the datafiles, to their licences and with information around which fields I would be using and what they contained. The came back with ‘OK’… Indeed, I can’t see the ticket either :-) - I would link to the LWG meeting notes, but they seem to be lagging somewhat, so are not visible yet. I guess I could upload the emails to the wiki as user files and reference them - hmm, this being legal, I almost feel like I should check with Legal that they are OK with that first. Or, we can just wait and see if the LWG notes get updated…
Again, there is some info on this on the OS pages - once you’ve clicked ‘Read More’… I’ll post a quote here:
“The archive coordinates supplied for these triangulation stations have not been realised via ETRS89 and OSTN15. They are the original archive coordinates and can therefore no longer be considered as true OSGB36 National Grid Coordinates of the station. It is expected that agreement between ETRS89/OSTN15 derived coordinates and the original archive coordinates of triangulation stations (down to third order) will be at the 0.10 metre root mean square (r.m.s) level.”
I then use Proj, from within the Rsf library to do the transform to ETRS89 and force it to use the best transform it can (OSTN15 NTv2), and here is what my log output says
Transforming 7405 (OSGB36+ODN) -> 4937 (ETRS89) with accuracy 0.038m
As to accuracy of actual trigpoints included. The database has a marker for all trigpoints it knows have been damaged or destroyed, and I remove them from the dataset. So, we only process trigpoints that were known to exist the last time the database was updated. Now, I know there are trigpoints in that remaining data that have since been moved or destroyed - but I believe they will be relatively few. Personally I think it is going to be unavoidable to have a few anomalies in the import.
Yes, the red marker would be a new trigpoint. If you click on the marker you will get a popup with a bunch more information - like the fact that the nearest OSM node is 23.19m away :-)
You’ll find on the github page where I looked at the distance data between the OS and OSM points, and decided that 15m would be the cutoff point for ‘snapping’ an OS point to its nearest OSM neighbour. We have to make a cutoff choice somewhere. For reference, I ran the data process with the limit set to 25m and we get 4389 New nodes (as opposed to the 4437 with 15m cutoff) with the difference being split between Editable (mergeable) and Review (human intervention) nodes.
I could split out matched and unmatched OSM nodes on the map probably reasonably easily. I don’t collect that data at present. I don’t think I’d be doing the manual review myself though…
So, we should clarify Pillars and FlushBrackets and the databases.
The trigpoint database contains the Pillars - it has no flush bracket data in it
The flush bracket data is held in the benchmark database - which has half a million entries. Only ~4220 of those entries are Pillars, and only ~4114 of those have flush bracket data. We do not have a flush bracket number for every pillar!
Pillars tend to have a flush bracket attached to them
I’m not sure that all Flush Brackets are attached to Pillars (I suspect not)
I’m not sure if any Flush Brackets have their own OSM entries. The ones I see are held in tags (generally ref) on Pillar nodes.
Bottom line - we could expand the ‘snap’ distance. Looking at the graphs on github, the locality really seems to drop off after ~50m. The consequence of expanding the distance will be more nodes will end up ‘mismatched’, and fall into the Review or Edit categories, and not be part of the import.
So, I re-ran the code with a 50m cutoff - we get 4333 New Nodes - to summarise
cutoff
new nodes
15m
4437
25m
4389
50m
4333
Indeed. And I’d lean towards the trigpoints being accurate, so they should be placed where they really are, and would then be an indicator and aid to improving alignments, no?
The idea of the note field is to try and stop people moving trigpoints without having understood what they are and why they are where they are… but yes, I feel it is inevitable some will get moved around incorrectly in the future.
I don’t have the skills to automatically assess if a trigpoint falls within an existing OSM item though - I don’t think I can (or not easily) do that in my R code for instance.
Oh :-( I followed the import wiki template - I was hoping it was clear enough. I’ll have a stare and see if it can be made clearer…
There are two ref fields:
ref:os : would contain the unique OS Pillar reference taken from the Pillar DB New Name field.
ref : would optionally contain the Flush Bracket number for the Flush Bracket on that Pillar, if we have managed to extract that from the Benchmark database.
ref is in a little table of its own below the main table in the wiki page as it is optional.
But, yes, I agree we need to work out a ref:GB:* tag - or two - to hold these numbers… I’ve only proposed ref and ref:os as they are what already existed. I guess I should note that ref containing the Flush Bracket number is used a lot already in OSM - so we might want to stick with that one for historical reasons. ref:os is only currently used 5 times, so is an easy hand-fix.
There are details on the github of the process - but I can copy a bunch of that over to the wiki if need be. You can find the R code on github as well - if you are happy reading >2kloc of R?
Parsing destroyed Pillars from the pillar database is OK - there is a separate field. Parsing information out of the Benchmark database is a pain - a lot of the data is held in a single text field, and much of it is inconsistent. Probably most of the effort and processing is in trying to extract and match flush brackets to pillars.
So, I’ll see if I can embelish the wiki some more, and I’ll contemplate if I need to ask LWG if it’s OK to upload the emails for reference. We can then discuss ‘snap’ distances and tagging some more I think.
I think it is inevitable that some of the nodes will be moved in error in the future. Anything we can do to try and reduce that is good, no?
Sure, if the node is moved in error and the note tag left, and/or no good explanation left in a changeset note… what to do?
I am contemplating adding an ‘errata’ file to the code so it can be informed of known good changes and take them into account - I’ve already had 2 pm’s noting one pillar that has since been destroyed and one that has moved due to a new housing estate. Let me at least open a github issue for that later on as a start…
The code/slippy map does have ‘strict checking’ code, that will then qualify nodes that ‘match perfectly’ with the generated data into the ‘Good’ category. Right now that code is set to a more ‘lenient’ mode, as otherwise we would have zero good nodes on the map ;-) The plan is once we have the first node(s) imported I will enable strict mode and re-generate the map data. That will then show which nodes in OSM are matching the OS extracted data.
Now - as long as I’m importing, and up to the end of my import, I intend to re-generate the map data. But, once I’m done importing, I have no plans (nor automation in place - it’s currently a manual process) to do any periodic updates. Ideally the code and data would end up in some automated system (hint hint anybody ;-) ), but I’m not set up for that.
Sure, I’m going to be willing to do periodic updates on request in the future, but with all the best will in the world, motivation, time and knowledge of how to do that is all going to fade in the future.
This is great, and long overdue! Thank you for putting the effort in to get this happening.
I definitely agree there should be a note on the trigpoints to say not to move them, as long as we’re happy the underlying data is spatially accurate and the coordinate transforms it’s gone through in the import process are correct.
Regarding destroyed pillars, these could be tagged using a lifecycle prefix, e.g. destroyed:man_made=survey_point. I couldn’t see any discussion of that, but apologies if you’ve already thought of/covered it somewhere.
If it’s straightforward to do this, it would be worthwhile, as having a node in OSM with an appropriate lifecycle prefix reduces the possibility of a well-meaning mapper in future adding a trig point where one has previously been destroyed, without tagging it appropriately. e.g. If they’re working off a historic map and haven’t checked the current status of the trig.
Leaving them as an overlay is also a way of making sure OSM users don’t benefit from them.
If the licence is acceptable it would be good to get them added to the cadastral overlay layer though, that way any trig point node that isn’t in the △ can be investigated.
np. I quite like the challenge of massaging the data That’s the fun bit…
afaict the data should be accurate - but, if somebody wanted to check that over that’d be great.
That sounds good! No, that idea or discussion has not come up before.
Looking at that lifecycle page I think both/either demolished or destroyed would be appropriate - but I think it would be hard (or at least potentially hard work) to figure if a pillar had been deliberately removed or not. I’d probably err on the side of demolished.
I’ll add a ticket to the github repo noting this could be done. Nominally the coding should not be too hard - all the code is there already and we could just ‘flip’ the ‘not destroyed’ filtering and do a run for all the destroyed pillars, but I think there will probably be a few warts to iron out or some code refactoring if we wanted to have the code do separate data generation for existing and destroyed, and that would take a bit more effort. So, I’ll open the ticket for now and might just have a tinker to see how hard it might be to do.
But, yes, I agree having all the destroyed pillars in OSM might help improve the plethora of historic map markers that are spread across man_made=survey_point. I currently have no feel for if the OS data on destroyed pillars aligns with what is on the old maps or not.
I think you guys are talking ‘above my paygrade’ If there is anything I can do to help here - data generation etc. - let me know. You can get the data off github in OSC XML changefile format already if that helps. It should be reasonably easy to generate in a different form or shape if need be - it’s just an iteration over a dataset and R has a lot of support for different data formats.
The license is OGL. The LWG were happy with the data fields I’m using from the dataset.
I use https://trigpointing.uk/ a bit to check ones I’m heading to and correct ones I know, adding flash details or type to OSM as I check which data is needed here, especially as their members often report whether one has been moved or destroyed completely. I wonder if their database is usable, allowable be used, and something could help?
Hi @fredgolightly,
yes, I’m aware of the site, but I’m not sure we can use the data from there in OSM. I did some digging around on that topic whilst working on the trigpoint input data and found:
there was a thread on Talk-GB back in 2019 discussing it - it seemed inconclusive, but not positive.
Having a quick look at trigpointing.uk, there terms of use don’t look very OSM ODbL compatible to me - in particular mentions of non-commercial strike me (there was a thread on the forum recently regarding some other data with a non-commercial license).
and, afaik, nobody has done due diligence and taken that dataset through an approval with the LWG.
Personally, I don’t fancy trying to wrangle that through LWG and then do the database access and integration into my R code - so am not planning to do anything with the trigpointing.uk data - sorry!
Yes, it seems more trouble than it’s worth to try to use the TPUK data. But it may be worth cleaning up the tpuk_ref values. These seem mostly very consistent, but there are, as always, some rogue values (mostly just missing initial TP).
Tried to improve the tagging section by adding a column noting which dataset the information comes from, and pasting an OSC XML example of how a new node ‘looks’
Improved the hotlinks to the code (now links to the R file) and the dataflow (now links to the dataflow section on the github homepage).
Noted you need to click the ‘Read more’ links on the OS pages to see details of the license
Added a new section about positional accuracy. I’d be happy if somebody wanted to ponder the data there and confirm/deny/improve. Oh, I’ll note, the trigpoint data is supplied in OSGB36, which is tied to the UK landmass, so tectonic drift does not affect their accuracy afaict (it’s something somebody brought up on the mailing list).
Note - I’ve not copied over chunks of information from the github pages into the wiki page, as that’s only going to duplicate the information and cause me a bit of a maintenance headache trying to keep both in sync and accurate.
So, to move forwards, I’d like to make a few proposals …
update the ref:os tag to the modern style. Looking at the OS page, they seem to call the database the ‘Complete Trig Archive’, so I propose we move to ref:GB:complete_trig_archive, and store the unique New Name reference in there. I’m open to suggestions here…
Add an operator tag. There are already ~66 "operator"="Ordnance Survey" tags in the OSM trig data, so I propose we go with that.
Add a wikidata tag. There are already 46 instances of operator:wikidata=Q548721 in the OSM trig data - so I propose we go with that.
Make the note field a little less commanding, so change it to Please not move this node without reading https://wiki.openstreetmap.org/wiki/OS_pillar_trigpoint_import first
Right - but, not as part of this import I’ll leave TPUK stuff to others…
Inconsistency seems to be a way of life in the little man_made=survey_point world. As part of the (pre)import it’s likely I’ll do a few cleanups, like the few points that have their heights stored in ft and some other little niggles I keep bumping into. I’ll do those by hand though - also not part of the import.
@gurglypipe - I’ve added the destroyed pillars to the slippy map as an off-by-default layer of markers, as that was relatively easy. fyi, there are 1030 of them. From what I see some of them were then later re-instated (so have multiple entries in the db).
I’m not convinced it would be a good idea to import the destroyed trig points. OSM is supposed to be about what is present today. Fine if there’s an already mapped trig point that the data says is not gone, then we could leave the node mapped and mark it as destroyed. But I don’t really see any benefit to adding a whole bunch of no-longer-existing stuff from a third-party database.
One other thing to emphasis from my previous comments: I really think the import needs a better way of preventing adding duplicate nodes for the same trig point when the existing mapping does not have an accurate location. Sure, we probably don’t want to snap existing mapping to the locations from the dataset automatically if the existing mapping is too far away, but equally we should avoid adding new nodes that are likely to be duplicates. It may well be that the best option is to do a manual review of all the not-matched OSM trig point objects prior to the import, to either classify them at not-trigpoints, match them to a destroyed trigpoint in the dataset, or correct any obviously wrong locations.