Skip to content

Dotted initials after a two-word surname read as a credential: García Márquez, G.J. has no given name #563

Description

@derek73

Since #516 (Policy.unlisted_dotted_suffixes, default on), a dotted token after a comma that follows two or more words reads as a credential. The parser cannot tell a given name plus a family name (John Smith) from a two-word surname (García Márquez, Lloyd Webber, De La Cruz, van der Berg). In the second case the token is the person's initials, and the reading drops or splits the name:

text 2.3.0 master
García Márquez, G.J. given G.J., family García Márquez given García, family Márquez, suffix G.J.
Lloyd Webber, A.L. given A.L., family Lloyd Webber given Lloyd, family Webber, suffix A.L.
De La Cruz, M.J. given M.J., family De La Cruz family De La Cruz, suffix M.J., no given name
van der Berg, A.J. given A.J., family van der Berg given van, family der Berg, suffix A.J.
Le Guin, U.K. given U.K., family Le Guin given Le, family Guin, suffix U.K.

The #516 decision (decisions.md, 2026-09-14) weighed Smith, A.B. (one word before the comma) and John Smith, A.B. (two words), but not a two-word part made entirely of surname. Checking for a particle chain would fix only De La Cruz and van der Berg; double surnames without particles look exactly like John Smith.

The opposite direction is over-reported. Nobody hesitates over these, yet each carries suffix-or-name:

>>> parse("John Smith, X.Y.Z. P.D.Q.").ambiguities   # one report over the whole run
>>> parse("John Smith, PhD P.D.Q.").ambiguities      # even beside an unambiguous credential

Proposed rule

Rationale: the token's own shape is the evidence the parser can see. Two dotted single letters are ordinary initials. Three or more in one unspaced token, or two dotted tokens in a row, are not how initials are written; the conventional form separates initials with spaces (Smith, J. R. R.), which this rule does not touch.

Statement: after a comma that follows two or more words, an unlisted dotted token

  • of exactly two single-letter chunks reads as the given name (initials), and the fork is reported SUFFIX_OR_NAME;
  • of three or more single-letter chunks reads as a credential and is not reported;
  • in a run of two or more dotted tokens, or in a run holding a listed unambiguous credential, reads as a credential and is not reported;
  • with any multi-letter chunk (B.Tech., Msc.Ed.) reads as a credential, as today.

Examples (target readings):

"García Márquez, G.J."        →  given="G.J."  family="García Márquez"   ambiguity: SUFFIX_OR_NAME
"John Smith, X.Y."            →  given="X.Y."  family="John Smith"       ambiguity: SUFFIX_OR_NAME
"John Smith, X.Y.Z."          →  suffix="X.Y.Z."                         no ambiguity
"John Smith, X.Y.Z. P.D.Q."   →  suffix="X.Y.Z. P.D.Q."                  no ambiguity
"John Smith, PhD P.D.Q."      →  suffix="PhD P.D.Q."                     no ambiguity
"John Smith, B.Tech."         →  suffix="B.Tech."                        · unchanged
"Tolkien, J.R.R."             →  given="J.R.R."                          · unchanged: one word before the comma
"Smith, J. R. R."             →  given="J."  middle="R. R."              · unchanged: spaced initials

Accepted consequences: García Márquez, G.J.R. (a double surname followed by three unspaced initials) reads as a credential. Accepted because spaced initials are the conventional form (Derek, 2026-09-30).

Also fix

On the no-comma path, a token taken into the class by shape alone gets the listed-member message, which is wrong on both counts:

>>> parse("John Smith X.Y.Z. P.D.Q.").ambiguities
# "'X.Y.Z.' written without periods is both a post-nominal and an ordinary name; ..."

X.Y.Z. is written with periods and is in no list.

Open questions

Related: #490, #516.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions