A series of changesets added the boundaries of dozens of core-based statistical areas (MSAs and μSAs). Each relation is tagged boundary=statistical border_type=metropolitan/micropolitan with a name that matches the Wikipedia article title or Wikidata label rather than the official name. For example, here’s the Chicago metropolitan area (Chicago–Naperville–Elgin). Should we delete the 77 MSAs and 9 μSAs or complete the full set of 393 MSAs, 542 μSAs, and 184 CSAs?
If we decide to keep and complete coverage of CBSAs, I would prefer this tag over the border_type=metropolitan that @aighes used. Yes, “statistical area” sounds redundant to boundary=statistical, but the redundancy is necessary because metropolitan would be ambiguous on its own. The OMB subdivides the largest MSAs into smaller metropolitan divisions.
To be honest, I don’t mind. I mainly went by taginfo and followed the character of other values. There are also micropolitan_statistical_area (on the same level as MSA) and then combined_statistical_area above them. The official name is recorded in the official_name tag.
In my opinion, the official name is the only name for these CBSAs. A moniker like “Chicago metropolitan area” refers to any of several definitions of the metropolitan area. “Chicago–Naperville–Elgin Metropolitan Area” refers to the specific OMB definition that you mapped, which applies to (most of) the federal government. However, state and local authorities often have their own definitions of the “Chicago” metro area, not to mention private industry.
The English Wikipedia intentionally conflates these concepts, causing Wikidata to often conflate them as well. Some other Wikipedias have been better about it. The worst offenders are the articles about the largest metro areas, like San Francisco and Los Angeles. I’ve been slowly splitting out Wikidata items to distinguish CBSAs where this has happened.
A couple weeks late to Minh’s follow up, catching up when the SLC MSA was added and digging for any conversation that happened. I agree with those from the original part of this thread, that these boundaries don’t belong in OSM.
I agree as well, MSAs/CSAs are just collections of counties that can be determined from the boundaries already in OSM.
The original Chicagoland boundary was deleted a year ago, and nobody complained, so I’d say the SLC one can be summarily deleted also.
What about the 85 other CBSA boundary relations that were introduced recently?
Oh, I missed your note above. I feel like in the context of this thread that these should not have been added and need to go. Lost work, but that’s the risk you take when doing a wide scale edit without community buy-in
Exactly, some of these statistical areas are just a single county (the Spearfish, SD µSA doesn’t even have its own Wikipedia page…).
Though we have ATYL, so there is no need for a community buy-in. As mentioned above, I’m open to change the tags I used, if the community prefers a different tagging. Though there was not so much feedback so far.
The tags are fine, we just think the features should be deleted.
I’ve split this conversation out to a separate topic to give the discussion of all MSAs (not just Chicago.) better visibility with a more accurate title.
I’m in favor of not having these objects in OSM. They are creations of the Census Bureau government for statistical convenience. They don’t represent anything in particular in the real world.
I don’t follow your argument.
Like @SD_Mapman wrote, those objects represent usually collections of counties. So they can be verified. So based on our ATYL, they can be added to the database.
Are you saying, the relations should be removed and instead I should add tags to the actual county boundary relation, like msa=Chicago...?
No. I am saying you should remove the boundaries without replacement. The statistical relationships can be modeled outside of OSM, e.g. in wikidata.
I retitled the new thread because the core-based statistical areas are defined and delineated by the White House’s Office of Management and Budget based on demographic data compiled by the Census Bureau. New changes are announced in an OMB bulletin that takes effect immediately.
Granted, the boundary=statistical tag is upfront that the boundary is a statistical convenience. My main concern at this point is that our coverage of these boundaries is misleading:
-
It isn’t enough to relegate “Chicago–Naperville–Elgin Metropolitan Statistical Area” to an
official_name=*. No one who calls it anything else has any business using that particular boundary in their statistical analysis. This isn’t Wikipedia, which likes to hang a wide-ranging discussion on a specific official designation just to satisfy notability criteria. -
If we have any CBSA boundary at all, we must have all of them. Otherwise, we’ll mislead an analyst to produce incomplete statistics. Are we prepared to maintain this expansion in scope? Boundaries are already somewhat outside of OSM’s focus, let alone statistical boundaries that exist purely on paper, that are adequately modeled as collections, and that laypeople rarely use specifically by name.
Looking abroad, we have coverage of analogous boundaries in Canada, but maybe that community has mappers who are on top of the inevitable boundary breakage. I could see OpenHistoricalMap adding coverage of CBSA boundaries. That community accepts the burden of keeping extra boundaries intact, as long as there’s something interesting to say about them. The CBSA boundaries have dramatically evolved over time, not only when county boundaries change.
Forgive me for not retaining the details. It will likely not be the last time I forget. ![]()
Yes, I would agree. Might as well have all the Combined Statistical Areas too, if we have any statistical areas.
If we were talking about collections of counties with a distinct identity used in the real world by everyday people, I’d be open to the argument that these belong in OSM. For example, three counties in Vermont are collectively known as the Northeast Kindgom. This name is well known and can be seen on various signs around the region. It’s clearly a regional place name used in the real world.
The Burlington-South Burlington Metropolitan Statistical Area is also a collection of three counties in Vermont, but it is not a regional place name used in the real world. US mappers are fairly open to supporting verifiability claims with government data, but we generally want to see at least some on the ground evidence in the real world as well. Otherwise we’d be including religious administration boundaries, school districts, congressional districts, and many more boundaries that exist for all sorts of different reasons, but on paper only.
That wasn’t the point I was trying to make. My point was that, similarly to US Catholic diocesan boundaries (see @ezekielf’s link in the previous post), these are collections of items we already have in OSM so there is no need to include them.
Boy am I glad that’s out of scope. SD School Districts would be a real pain to map.
CBSAs and CSAs are even more straightforward than diocesan boundaries: by law, they have to be collections of counties. It isn’t legally possible to split counties in the manner that a minority of dioceses do.