Skip to content

fieldcheck: add uk/ca/de postal_code semantic type - #516

Open
wufangyong973 wants to merge 2 commits into
FreshCode-Org:mainfrom
wufangyong973:test/postal-code-validators
Open

wufangyong973 wants to merge 2 commits into
FreshCode-Org:mainfrom
wufangyong973:test/postal-code-validators

Conversation

@wufangyong973

Copy link
Copy Markdown

closes #512

each country gets its own compact regex: UK is 1-2 letters plus a district plus the inward, Canada alternates letter/digit, and the German PLZ is five digits. normalisation (uppercase + strip) happens in the per-cell step, the vectorised pre-screen only matches the space-free canonical form, so spaced or lowercase values fall through to the per-cell check and get re-decided there, same rhythm as email/url.

Signed-off-by: wufangyong973 <wufangyong973@users.noreply.github.com>
Signed-off-by: wufangyong973 <wufangyong973@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 519a12e9-b787-413c-b0d1-85b3f48b1032

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@JohnnyWilson16

Copy link
Copy Markdown
Contributor

Thanks for working on this, @wufangyong973! The helper functions, normalization, and integration into _suspect_rows and validate_fields are well structured.

However, there is a functional bug in the UK postcode regex that causes valid two-digit districts to fail:

Issue: Two-digit UK districts fail validation

In src/freshdata/fieldcheck.py:93:

_UK_POSTAL_RE = re.compile(r"^[A-Z]{1,2}[0-9][A-Z]{0,1}[0-9][A-Z]{2}$")

Notice that [A-Z]{0,1} only permits a second letter in the outward code (e.g. W1A, SW1A). It rejects postcodes with a second digit (standard formats A99 and AA99), such as:

  • B33 8TH (Birmingham) -> False
  • SN10 1AA (Swindon) -> False
  • M60 1NW (Manchester) -> False
  • OX14 4SE (Oxford) -> False

Interestingly, your docstring on line 91 explicitly lists B33 and SN10 as examples!

Fix:

Update _UK_POSTAL_RE and _POSTAL_COMPACT_RE to allow an optional second digit or letter:

_UK_POSTAL_RE = re.compile(r"^[A-Z]{1,2}[0-9][0-9A-Z]?[0-9][A-Z]{2}$")

_POSTAL_COMPACT_RE = re.compile(
    r"^([A-Z]{1,2}[0-9][0-9A-Z]?[0-9][A-Z]{2}|[A-Z][0-9][A-Z][0-9][A-Z][0-9]|[0-9]{5})$"
)

Please also add test cases in tests/test_fieldcheck.py covering two-digit districts like B33 8TH and SN10 1AA. Once updated, this will be good to go!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(validation): add context-aware postal code validators for UK, Canada, and Germany

2 participants