Hi all,
I’m Gopi, a PhD student at New York University. I’m starting a research study that focuses on how OSM contributors discuss and debate the use of AI for mapping, tagging, and validation, and what those discussions reveal about the values and practices the community wants to preserve. Here is a poster I presented about the study at the State of the Map US 2026 conference last week in Madison, WI, USA.
For this research, I’m considering fetching publicly visible threads and posts from this forum (titles, post text, dates, tags, and engagement counts like views/likes/reactions) using the Discourse API.
My plan is to make read-only requests, include a clearly identified User-Agent with my contact, and rate-limit to roughly one request every couple of seconds. Discovery will be limited to the “ai” tag plus a defined keyword search list including names of AI-assisted tools/products (like Rapid or Microsoft Building Footprints), process-related terms (like AI-assisted tagging), and governance/policies (like Automated Edits code of conduct). Any published results will quote sparingly with attribution and pseudonymize usernames.
Does this seem like a respectful way to use the forum’s shared infrastructure? Is there anything I should do differently, or anyone I should check with first?
(Side request - if you know threads I should make sure to read to better understand discussions about AI in OSM, I’d be grateful for those suggestions).
Thank you so much in advance!
