Address Interpolation., or "Source Reliability and Damage if Wrong"

In a different topic I indirectly mentioned I had a negative view of address Interpolation for individual objects (there is a tag addr:interpolation for a different use), and said I would start a new topic to discuss it (to @LordGarySugar). So here we go.

To be clear, I am talking about assuming an address for an object, based on nearby addresses. The most common example being reliably sourced residential addresses along residential streets. The situation where we have house numbers
21, 23, ?, ?, ?, ?, ?, 35
A mapper fills in those gaps using interpolation. This I view as bad practice, as you can not have the appropriate reliability for the additions. Any one who has surveyed addresses on the ground must know that wrong addresses will be introduced by this method, due to issues such as 25a, no 27, 29 is 107 on road behind, and a 34a above 33 and 35.

This leads onto my view on

Source Reliability and Damage if Wrong

I’d say that generally OSM Source Reliability could be summed up with

  1. Guess — No meaningful evidence
  2. Assumption – Backed by some evidence, where incorrect outcome is not surprising
  3. Presumption – Backed by strong indirect evidence, where an incorrect outcome is clearly surprising
  4. Third Party fact – Up to date data supplied by map/database/organisation.
  5. Ground Verified Survey – Person(s) have ground surveyed and observed it.

Then I suggest you need to take into account Damage if Wrong, scaled from, Very Low to Very High. Below are some quick examples of impact if wrong eg

  • Very Low Impact — Route of 0.5m wide waterway culvert under a field.
  • Low Impact — Unattached Building outline shape from imagery
  • Medium — Footpath route across field
  • High — Interpolated (inferred) house numbers
  • Very High - road access restrictions for emergency vehicles.

The two checks above show why I disagree with interpolated (inferred) house numbers. The evidence, or process, that produces them is not reliable, it’s an assumption that will introduce wrong information. Then separately, and importantly, the impact of the wrong data is high or very high

If OSM allows adding addresses that are Assumed and will have mistakes, and the data looses value it had.

A suggestion is to add a tag source:addr:housenumber=interpolation when house numbers have been assumed via interpolation. This could be used on the ground for checking house numbers. But it already can be seen a survey is needed by the current lack of a house number. I believe the appearance of these house numbers on main maps would mean the addresses are less likely to be surveyed

1 Like

We have dedicated tagging schema already for interpolation and we can use it

Why invent a new one?

4 Likes

Impact on who or what if a datum is wrong?

The occasional wrong house number is not a big impact IMO. Hard to spot, yes, but nothing new. I’m sure most of us are capable of checking addresses either side or nearby if we’re ever directed to the wrong building by OSM.

What is your case for it to be a high impact?

3 Likes

OSM is iterative. The map gets more detailed and more complete as time goes on.

Address interpolation is just another example of this - a first pass that can be refined with individual addresses in due course. It’s no different from landuse polygons before we’ve mapped every building, river polylines before we’ve mapped the riverbank, simple shop nodes before we’ve mapped the opening hours.

11 Likes

Do you know of data being added due to interpolation without declaring it, or are you making a case against something that doesn’t happen?

As has been said, there is a method for saying “you can probably intetpolate addresses here” without mapping all the house numbers. See Addresses - OpenStreetMap Wiki

For your example of footpaths. I have mapped some that I wasn’t able to get a GPS trace for the full length, but I know the ends. I don’t add extra nodes (so it’s very straight) and use the “note” or “source” keys to explain how much of it has been accurately surveyed.

OpenStreetMap is editable and thus can be improved (some say this is part of the “wiki” definition). It can be good to put some data in, knowing that it could be improved later on. In the early days of OpenStreetMap lots of stuff was added as nodes (pubs, airports, parks) and when there was aerial imagery available or more mapper time these nodes would get replaced with areas (building outlines, landuse, the added detail of paths on the site, etc). OpenStreetMap would never had got where it was if the early volunteers had to map every detail first time.

7 Likes

Can you clarify, but what you mean, or give an example, of what you mean by Interpolated addresses and how they can be “refined with individual addresses”?

We are trying to create an appropriately sourced database of UK addresses. An address is a fact, there should be nothing subjective in it’s recording. It’s should be either incomplete, or complete and correct.

A complete address is not comparable to the subjective recording of a land-use polygon or path routes which will always be open for change. There is room for iterative improvement by adding detail, eg Add city, then postcode, then later street name, the later still house number. But adding the complete address by assumption is far from what I consider an iterative improvement.

My view of High Impact is done with regard of intended use, and impact of wrong information. If OSM data is used to provide someone with an address, and the address supplied is a full address, they’ll expect that address to exist. Assumed addresses in the data will have a High impact on the decision of data users to use OSM UK addresses. A situation where someone provides a full address from OSM, only to be told that address does not exist is harmful. eg filling in a web address form.

I feel that “most of us are capable of checking addresses either side or nearby if we’re ever directed to the wrong building by OSM” is very true, but the map isn’t for just for our use, we have to think of everyone including Evri drivers. In many situations providing a wrong address can simply end with a blank faced computer or person stating “Computer says no”.

That may be true in some cases, but it’s certainly not true in every case.

If those houses were on a new-build estate, you have good up-to-date aerial imagery, land registry parcels, and UPRNs, all showing you that there are exactly the right number of dwellings to fit the natural numbering (with no other possible houses on the ground in between, and no possibility any of the houses you have belong to a different street), there’s no number 13 to worry about, and you are confident in the numbers you already have, then I think you can be as confident in filling in that gap. (I mention a new-build estate, since that removes the possibility of historical anomalies, new-build houses tend to be more uniform, and new estates will have to comply with Local Authority numbering policies, which tend to be quite strict.)

7 Likes

So the person is filling out the form to say their address is #25, because that’s what OSM says, despite their letterbox out the front saying it’s #23A? :thinking:

1 Like

I can see how infilling missing addresses gives rise to a potential for incorrect data that can’t easily be separated from missing address data. The suggestions on adding a “check address” feature to Street Complete looking for a tag like source:addr:housenumber=interpolation seems to be a sensible way to avoid this issue. It would also avoid the issue of users of Street Complete not being prompted for a potentially missing house name, if there is already a number present.

A lot of information can be inferred from the postcodes linked to UPRNs, and more often than not, the order of the UPRNs themselves.

UPRNs typically follow the order of addresses on a street, so additional UPRNs that are out of this sequence are a good indication that there may be a new building (potentially with A, B in the number), worthy of further investigation.

In all but the most sparsely populated areas, postcodes map to at most one street. This allows edge cases such as houses at the end of terraces to be linked to the right street with high confidence.

In some cases I have filled in missing addresses through interpolation to discover that not all of the existing manually surveyed addresses could be correct without duplicate or missing numbers. This has then prompted further investigation, and in some cases has discovered errors in the initial (apparently manually surveyed) addresses.

3 Likes

That would be a bit weird, and it was not what I suggesting

OK, so what were you trying to suggest with that scenario?

1 Like

Who is ‘we’? I doubt all mappers are mapping for the same reasons. I’m not aware that there’s a separate OSM database for addresses.

An address can be good enough. It can have different forms and still be correct. Whether or not suburbs need to be included, they can be superfluous, but still correct. Errors are normal and, depending on how you define correct (spelling, 0 or O etc), often not a problem.

I’m pretty sure I did not campare to land use.

Assumption can be reasonable. I’ve no doubt that many correct addresses have been added by this method. If method limits acceptability, should correct addresses added by this method also be deleted? If you could identify them - probably (ha!) those on either side of an identified error.

An assumed address might have an impact, people have coped with errors in existing sources (likely PAF) very well over the last 30 years or so. There’s a reason good form programming allows for free form address collection and not just pick from a list.

A town of 30k addresses missing from the map, IMO, is of higher impact than 30k incorrect addresses across the UK.

That brings me back to my previous point of good form programming.

What are your definitions for a ‘marketable’ product for your database of UK Addresses? AT what stage would you ‘release’ it? Quotes because anyone can use it now, but I hope my meaning makes sense.

2 Likes

A person who takes an address from OSM and tries to use it in situation where it’s compared to the PAF database.

Your scenario is based on a person knowing their correct address filling in a form using incorrect OSM data which doesn’t make sense, which was your point.

To “market” or promote the data for use it must be provided with some confidence in the quality of the data. My view if a description of the data shows it contains assumed addresses, in reality assumed house numbers, then the data is of little value due to the need for end users expecting this type of data to be correct.

This can be used to sum up the different view I have to the majority.
IMO, 30K incorrect addresses across the UK is higher impact than 30K addresses of one town missing from the map.

There are clearly examples where it will be perfectly safe to interpolate addresses based off a survey of some numbers, standard pattens and other sources to confirm the right number of properties in the gap. There are also clearly situations where it would not be safe to do the interpolation.

So it’s a judgement call from the mapper in each case. The things to be weighed up include how confident can I be that the addresses to be added will be correct, what’s the benefit of adding the addresses now, how long might it be before someone gets round to surveying them to add manually.

There are clearly benefits to having a greater number of addresses mapped now in OSM. I think if the chance of an error is estimated to be below that from manual surveying then it’s fine to add the addresses. You can probably be a bit more relaxed if leaving an explicit fixme="check house number" or similar on the object. It’s also worth noting that different mappers will have different levels of experience of mapping addresses, so may be better at estimating the chance of making a mistake, and also better at using and interpreting other sources to increase confidence in their interpolations.

7 Likes

I’ve no idea how you determine if a correct address is assumed data. Would you delete it on the basis that it is assumed data?

If you were marketing a database of UK addresses, I doubt you’d sell any if it was known that a whole town was missing. 30k across the UK is likely acceptable as 0.1% assumed error.

What would be your acceptable assumed error value?

I would argue for deleting all assumed address data.
I’ve not linked “assumed” with an error value. I use assume to describe confidence in data.

How about adding a suitable source tag, where you are confident that it is assumed (perhaps after asking the person who added it)? Would that not achieve everyone’s aims?

1 Like

The appropriate and standardized tag for this is source:addr:housenumber=interpolation.

4 Likes