address parsing and normalization

Designs and implements systems that convert raw address text into structured, canonical representations and map locations: building parsers to extract address components, normalization pipelines that apply locale-specific rules, validation and deduplication logic, and geocoding modules that resolve addresses to geographic coordinates or map features. Also develops planning and evaluation tools for address-to-map workflows, including ambiguity resolution, record linkage/matching algorithms, and quality metrics for parsing, normalization, and geocoding.

addressparsingandnormalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Coordinates from Context: Using LLMs to Ground Complex Location References

Oct 09, 2025
TM
Tessa Masis
🏛️ University of Massachusetts Amherst

This paper addresses the geocoding challenge posed by compositional location references (e.g., “the Starbucks northwest of Zhongguancun Metro Station”). We propose a novel two-stage LLM-based approach: first decoupling spatial knowledge from logical reasoning capabilities, then jointly optimizing both via prompt engineering and supervised fine-tuning. Our key innovation is a knowledge-reasoning separation architecture, enabling lightweight fine-tuned models (e.g., 7B-parameter variants) to achieve accuracy on par with billion-parameter general-purpose LLMs on compositional referencing tasks. Experiments across multiple real-world datasets demonstrate an average 12.6% improvement in geocoding accuracy over baselines—including rule-based systems and end-to-end fine-tuned models—validating the efficacy and feasibility of domain-specialized small models for geographic semantic parsing.

Developing improved LLM-based geocoding for complex locationsEvaluating LLMs' geospatial knowledge versus reasoning skillsGeocoding compositional location references in text

Pingmark: A Textual Protocol for Universal Spatial Mentions

Oct 08, 2025
KD
Kalin Dimitrov
🏛️ Independent Researcher

This paper addresses the lack of lightweight, privacy-preserving, and decentralized spatial referencing mechanisms on the Internet. We propose Pingmark—a plain-text-based, universal spatial semantics protocol. Pingmark uses the trigger symbol “!@” to encode physical locations as coordinate-free, user-identifier-free short textual strings (e.g., `!@Zhongguancun Tower, Beijing`). Clients resolve these strings locally and instantaneously into standardized URLs by invoking open map APIs, requiring no registration or reliance on proprietary mapping services. Its key innovation lies in abstracting geographic context into human-readable, shareable, and machine-parsable text primitives—extended with optional timestamps for spatiotemporal expressiveness. We have completed the Pingmark Protocol Specification (PPS) v0.1 and implemented a reference parser, validating its usability across diverse scenarios. This work lays the foundation for an open, interoperable spatial semantics infrastructure.

Creating a universal textual protocol for spatial mentionsEstablishing an open standard for location expression in textUsing semantic triggers instead of coordinates or proprietary links

Data Processing for the OpenGPT-X Model Family

Oct 11, 2024
NB
Nicolo’ Brandizzi
🏛️ Fraunhofer IAIS | Fraunhofer IIS | DFKI

This work addresses data quality, multilingual coverage, and regulatory compliance challenges in training large language models (LLMs) for the OpenGPT-X initiative. Methodologically, it introduces a novel “dual-track” data processing paradigm: lightweight filtering for curated datasets and aggressive filtering combined with MinHash/LSH-based deduplication for large-scale web corpora—fully aligned with EU regulations such as the GDPR. The pipeline integrates fastText-based language identification, hybrid rule-and-statistics filtering, a learned quality scoring model, and end-to-end metadata provenance tracking. Its primary contribution is the construction of the first high-quality, EU-compliant multilingual corpus for LLM training, explicitly designed for public-sector applications. Empirical evaluation demonstrates substantial improvements in model robustness, transparency, and auditability—particularly in government and public service use cases—while ensuring legal and ethical adherence across 24 official EU languages.

Develop data pipeline for multilingual OpenGPT-X LLMsEnsure compliance with European data regulationsHandle curated and web data with distinct processing methods

MapQaTor: A System for Efficient Annotation of Map Query Datasets

Dec 30, 2024
ML
Mahir Labib Dihan
🏛️ Bangladesh University of Engineering and Technology (BUET) | Qatar Computing Research Institute (QCRI)

To address low efficiency, unstable ground truth, and poor cross-platform reproducibility in geospatial question answering (GeoQA) dataset construction, this paper introduces MapQator: an end-to-end, reproducible, and traceable map-based QA dataset construction platform. Methodologically, it adopts a plug-and-play architecture to seamlessly integrate multiple map APIs (e.g., Google Maps, OpenStreetMap); employs HTTP response caching to ensure ground-truth consistency amid dynamic geospatial data; and incorporates a structured annotation interface with integrated visualization and analysis modules. Key contributions include: (1) the first fully closed-loop framework—from map querying and response capture to natural language QA generation and quality validation; (2) over 30× improvement in annotation efficiency; and (3) support for generating high-fidelity, complex-reasoning GeoQA datasets. The platform is open-sourced and publicly deployed at mapqator.github.io.

Data Set ConstructionGeospatial Question AnsweringModel Performance Enhancement

Synthesizing JSON Schema Transformers

May 27, 2024
JS
Jack Stanek
🏛️ University of Wisconsin - Madison

To address the error-prone and inefficient manual rewriting of data transformation logic upon JSON Schema evolution, this paper proposes a type-directed, top-down program synthesis approach for automatically generating semantics-preserving JSON Schema converters. Our method integrates type inference, semantic constraint modeling, a rewrite system, and intermediate representation (IR)-driven code generation to guarantee lossless data transformation and formal verifiability. It natively supports complex nested schemas and synthesizes correct, efficient, and human-readable Python and JavaScript conversion code. We evaluate our approach on real-world API configuration schemas and healthcare data integration scenarios, demonstrating its safety—via formal guarantees and empirical validation—its practical utility in industrial settings, and its generalizability across diverse schema evolution patterns. Experimental results confirm high accuracy, robustness to structural changes (e.g., field additions, type refinements, nested object restructuring), and scalability to large, deeply nested schemas.

Automating transformation between different JSON Schema versionsGenerating programs to convert JSON data between schemasPreventing data loss during JSON Schema evolution

Latest Papers

What's happening recently
View more

This work proposes the first end-to-end geocoding framework based on large language models, reframing coordinate prediction as a text generation task. In contrast to traditional multi-stage approaches—which suffer from complex pipelines, error propagation, and heavy reliance on structured geographic knowledge bases—the proposed method eliminates the need for external databases by integrating chain-of-thought reasoning to enhance spatial relationship inference. Furthermore, it employs a reinforcement learning mechanism with a distance-bias-aware reward function to refine coordinate prediction accuracy. The framework not only achieves high precision in mapping explicit addresses to point coordinates but also effectively interprets ambiguous relative location descriptions and demonstrates strong generalization capabilities for non-point, region-based queries.

error propagationgeocodingmulti-stage approaches

This study addresses the limited native understanding of GPS coordinates and geospatial relationships in large language models (LLMs) within real-world applications. To systematically evaluate such capabilities without external tools, the authors introduce GPSBench, a novel benchmark comprising 57,800 samples across 17 diverse tasks. Through zero-shot and fine-tuned evaluations, noise robustness tests, and downstream transfer experiments, they find that LLMs exhibit relatively strong country-level geolocation accuracy but weaker performance at the city level, alongside generally poor geometric reasoning abilities. While coordinate-aware input augmentation enhances downstream task performance, fine-tuning risks degrading the model’s pre-existing world knowledge. This work establishes a new foundation for assessing and advancing geospatial intelligence in language models.

coordinate understandinggeographic knowledgegeospatial reasoning

Current large language models often generate GIS code that violates spatial rules—such as geographic semantics, topological relationships, coordinate reference systems (CRS), and units—leading to unreliable outputs. This work proposes GeoContra, a novel framework that formalizes geographic constraints into executable geographic contracts and integrates static checking, runtime verification, and semantic validation to establish a geography-aware, closed-loop repair mechanism. By embedding natural language understanding, CRS metadata, spatial predicates, and topological rules directly into the LLM generation pipeline, GeoContra significantly enhances spatial correctness across 7,079 real-world tasks: achieving up to 81.5% accuracy with proprietary models and yielding an average improvement of 26.6% across eleven open-source models.

coordinate semanticsgeographic plausibilitygeospatial analysis

This work addresses the prevalent issue in large language models (LLMs) of introducing control-flow, type, or I/O errors during code translation due to neglect of program intent. To mitigate this, the paper proposes the first systematic use of a language-agnostic, structured intermediate specification that preserves semantic fidelity through an intermediate representation, structured generation, and automated test-based validation. Evaluated on the Avatar and CodeNet datasets with five state-of-the-art LLMs, the approach significantly improves translation accuracy, raising the micro-averaged accuracy from 67.7% to 78.5%. It completely eliminates lexical errors and substantially reduces errors related to structure, declarations, and runtime dependencies.

code translationcross-language programmingLarge Language Models

Hot Scholars

CJ

Claudio J. Tessone

Professor for Blockchain & Distributed Ledger Technologies, Universität Zürich
BlockchainCryptoeconomicsDeFiBlockchain Analytics
GS

Guangyu Sun

School of Integrated Circuits, Peking University
Computer ArchitectureDesign AutomationEmerging Memory
LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
CF

Christof Ferreira Torres

Assistant Professor in Computer Science, INESC-ID / Instituto Superior Técnico, University of Lisbon
Information SecurityBlockchainWeb Privacy
JL

Jingwen Leng

Professor, Shanghai Jiao Tong University
Computer Architecture