From 967dc1ae862fad04a8c90f933b43559ea02f1ccb Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 17:10:43 -0700 Subject: [PATCH 1/4] Part a trailing suffix part at a declared delimiter as a comma would (#549) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A delimiter declared through extra_suffix_delimiters / suffix_delimiter now separates a tail segment exactly as a comma typed in its place: group cuts the segment at its cores before grouping, groups each part as the comma twin's segment would be grouped, and drops the cores. No join, maiden walk or link search can reach across a core, so the #538 stepping machinery goes: the cores parameter on four functions, _maiden_take's seen index list, the skip arguments of peel_walk and trailing_start, and the post-join #206 drop block. 'Smith, John, PhD - and MD' reads suffix 'PhD, and MD' (1.4.0's answer; 2.0-2.3 gave 'PhD - and MD'), and 'Smith, John, PhD née Puig - i Soler' reads maiden 'Puig', suffix 'PhD, i Soler' again, as 2.3.0 did -- the boundary reading #538 declined, taken now. The default policy declares no delimiter; the differential gate exits 0 at all five baselines. One call per parse fewer. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 8 + docs/design/rules.md | 18 +- docs/release_log.rst | 4 +- nameparser/_pipeline/_group.py | 261 ++++++++++---------------- nameparser/_pipeline/_pieces.py | 12 +- nameparser/_pipeline/_post_rules.py | 2 +- tests/v2/cases.py | 11 ++ tests/v2/pipeline/test_group.py | 123 ++++++------ tests/v2/test_ledger_guards.py | 6 +- tests/v2/test_properties.py | 76 ++++---- tools/differential/corpus_rules.jsonl | 2 + 11 files changed, 242 insertions(+), 281 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 91021239..52464407 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -228,6 +228,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q - 2026-09-22 #397 follow-up — A DELIMITER CORE PAST THE CLAUSE'S FIRST WORD PASSES FOR THE NAME WORD BESIDE A LINK, RECORDED AS A DEVIATION RATHER THAN REPAIRED. The link exception above wants "a connective standing between two name words of the clause", and a separator the caller declared through `Policy.extra_suffix_delimiters` is structure, not a name word — so a link with one beside it joins nothing and should end the clause like any other suffix word. Between the marker and the clause's first word that already holds, the bound refusing the core before either piece test is asked. PAST that first word it does not: the core is an ordinary index to the run walk, which steps over connectives and nothing else, so it stands in for the name word on the link's left and the clause runs on past a title it would otherwise stop at. MEASURED 2026-09-22 under `extra_suffix_delimiters=(" - ",)`: `Smith, John, PhD née Puig Mr. - i Soler` reads maiden 'Puig Mr. i Soler' where its separator-less twin `Smith, John, PhD née Puig Mr. i Soler` stops at 'Puig Mr.'. THE POPULATION is the branch's own sweep, recorded with the code it describes (2026-09-21, corpus ∪ cases.py ∪ the property grids ∪ a 50,925-name generated set with cores, under thirteen core-bearing policies): the predicate is asked about a core in 51,072 of 900,023 calls, the answer differs from a core-skipping reading in 8,094 parses over 1,278 texts, and 1,824 of those move `maiden` on 288 texts — none of the 288 reachable at the default policy, `extra_suffix_delimiters` being empty there. NOT REPAIRED HERE: the fix threads the core set through three call sites into the run walk and moves the parent's reading as well, which makes it its own change rather than a rider on a review round. Open: #538. PINNED TWICE MEANWHILE. rules.md#M2 carries the shape as a `deviates: #538` example under an `extra_suffix_delimiters-dash` annotation — the first entry `tests/v2/rules_doc.py`'s registry has had for that field, named after the Policy field and carrying the delimiter in the suffix because the field's value is a set rather than a flag. And `tests/v2/pipeline/test_group.py::test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538` holds the pair at the piece level, named so nobody reads it as the contract. ONE COST OF THE DOC EXAMPLE, worth knowing before the repair lands: its string enters `corpus_rules.jsonl`, where the differential gate parses it with the DEFAULT facade — no delimiter declared, so the dash is an ordinary name word and #538's reading is off the path entirely. It moves there for the 2026-09-20 link fix instead, and is classified as that at all five baselines: added to the `fix(#397) a link inside a maiden clause stays in the birth name` alternation at the four 2.x ledgers (suffix 'PhD i Soler' → 'PhD', maiden 'Puig Mr. -' → 'Puig Mr. - i Soler', identical at each), and given its own rule at 1.4.0, where v1 had no maiden markers and read the whole suffix-comma tail as one suffix. - 2026-09-26 #538 — A DELIMITER CORE IS STEPPED OVER BY THE LINK'S NEIGHBOUR SEARCH, AS A CONNECTIVE IS (Derek, 2026-09-26). The 2026-09-22 follow-up above recorded the deviation; this resolves it. `_run_neighbours` treats a lone core like a connective, so the word on a link's side is the one past the core and a maiden clause reads exactly as the same text written without the core: `Smith, John, PhD née Puig Mr. - i Soler` under `extra_suffix_delimiters=(" - ",)` now stops at maiden 'Puig Mr.' as its separator-less twin does. INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_reads_as_if_it_were_not_written` over 81 generated texts, maiden field: 6 disagreed at e0f1a2fa, 0 after. Declined: THE BOUNDARY READING, which the 2026-09-22 entry's wording implied ("a link with one beside it joins nothing") — a core ending the neighbour search on its side. It fixes the reported shape equally, and it moves `Smith, John, PhD née Puig - i Soler` and `… Puig i - Soler` from maiden 'Puig i Soler' to maiden 'Puig', suffix 'PhD i Soler' — a birth-name link pushed into the credentials, where the skip reading leaves both as they read. Default-unreachable either way. THE SAME STEPPING APPLIES OUTSIDE A CLAUSE: the `frozen` loop's own `_run_neighbours` call also takes `cores`, so a connective beside a core looks past it; where a credential or nothing stands beyond, it stays a lone suffix word rather than joining, and the core it stood beside -- now a lone piece with nothing joined to it -- is dropped as #206 drops any lone core, the same way it already drops one between two ordinary post-nominals: `Smith, John, PhD - i Soler` 'PhD - i Soler' -> 'PhD, i Soler', `Smith, John, PhD - i - MD` 'PhD - i - MD' -> 'PhD, i, MD', `Smith, John, - i Puig` '- i Puig' -> 'i Puig', `Smith, John, Puig i -` 'Puig i -' -> 'Puig i'. MEASURED 2026-09-26: 191 of 8,097 generated core-bearing tail texts (the three heads `Smith, John,`, `Smith, John, Jr.,` and `John Smith,`, each followed by every 2-4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr.} containing a core) move `suffix` and no other field, none reachable at the default policy; the comparison is each text parsed under `extra_suffix_delimiters=(" - ",)` with `_run_neighbours` handed its cores and with it handed none. Not repaired here: a core the join MERGES into a joined piece -- interior (`… Puig Dr. i - y Soler`, still 'PhD i - y Soler') or at the edge a link joins across, a name word standing beyond it rather than a credential or nothing (`Smith, John, Puig - i Soler` 'Puig - i Soler', its separator-less twin 'Puig i Soler'; `Smith, John, Puig i - Soler` 'Puig i - Soler', measured 2026-09-26) -- is not dropped from the suffix text; the join is right, the surviving separator is the gap. Also in that family: an ORDINARY connective's join still takes a core as its neighbour, the stepping above being for generational vocabulary alone (rules.md#P3) -- `Smith, John, PhD - and MD` keeps suffix 'PhD - and MD' (measured 2026-09-26), pre-existing and unchanged here. Open: #549, which carries both the surviving core and this one. +- 2026-10-03 #549 — SUPERSEDED: the skip reading the #538 entry above settles is reversed, and a declared delimiter in a trailing part now parts it as a comma would, so `Smith, John, PhD née Puig - i Soler` reads maiden 'Puig' with suffix 'PhD, i Soler', the reading #538 declined and the one 2.3.0 gave. The decision, its measurement and the invariant that replaces #538's are under C1's 2026-10-03 entry. - 2026-09-26 #535 — THE MAIDEN WALK READS THE TRAILING TITLE CHAIN, AND EVERY STOP ASKS ONE RELEASE QUESTION OF THE SPAN IT GIVES UP (Derek, 2026-09-26). This resolves the NOT TRANSPARENT HERE clause of the 2026-09-19 #533 entry above, which recorded `Jane Doe nee Smith MA Prof.` and `Jane Doe nee Smith Prof. MA` disagreeing and deferred the pair as one decision about what "trailing" means inside a clause. BOTH HALVES WERE TAKEN TOGETHER: a trailing title ends the clause (`Jane Doe nee Smith Prof.` reads title 'Prof.', maiden 'Smith', where 2.3.0 read maiden 'Smith Prof.'), and a credential or numeral standing in front of that title gets the stop it gets with the title absent (`… Smith MA Prof.` gives suffix 'MA', `… Smith V Prof.` suffix 'V'). Only the first half would have left the pair disagreeing in the other direction; H5's transparency is a statement about the peel and the chain read together to their fixed point, so the walk now reads the end of the clause the way assign reads the end of the name — `tail_reading`, the shared predicate mechanisms.md#ONE-PREDICATE-PER-QUESTION names — and over the forms the transparency property test below exercises, the two spellings give one answer. They split where something in or ahead of the clause stands to take the title — for example a particle chain (ahead of the clause or inside it), a bound-given join, a name left with no name word, or, after a family comma, a title word in front of a class member with a title behind it, which the given part's chain stops short of exactly as it does with no marker (`Doe, Jane nee Smith Rev. MA Prof.` reads maiden 'Smith Rev.', title 'Prof.', suffix 'MA', where `… Rev. Prof. MA` reads title 'Rev. Prof.', maiden 'Smith', suffix 'MA'; bare `Doe, Jane Rev. MA Prof.` reads middle 'Rev.') — and the Accepted pairs recorded here and in rules.md#M2 are examples of that, not a complete list. `tests/v2/test_properties.py::test_a_trailing_title_is_transparent_to_the_maiden_clause` holds that as an invariant over the two inputs, over heads and runs where the credential is given up and nothing ahead would take the title, with its negative control at e0f1a2fa in its docstring. ACCEPTED, THE TRANSPARENCY BOUNDARY: where the clause KEEPS the credential the spellings differ, because a clause is one contiguous run and a title inside the kept text cannot leave without the words behind it — `Doe nee Smith ba Prof.` reads title 'Prof.', maiden 'Smith ba', and `Doe nee Smith Prof. ba` maiden 'Smith Prof. ba'; `abdul nee Smith MA Prof.` reads title 'Prof.', family 'abdul', maiden 'Smith MA', and `abdul nee Smith Prof. MA` given 'abdul', maiden 'Smith Prof. MA' (2.3.0 kept every word in all four). The report differs with them: `Doe, Dr. nee Smith MA` and `… Prof. MA` report `suffix-or-name` on the kept MA, `… MA Prof.` reports nothing, the kept member no longer ENDING the clause, which is what M2's report is asked of. rules.md#M2 carries the first pair as an Accepted example. ACCEPTED, FURTHER SPLITS, measured 2026-09-26. Where something ahead of the clause would take the title, the spellings differ even with the credential given up in both, because the first-suffix-word stop ends the clause before the title check runs and the title check alone declines: `Jane van der Berg nee Smith PhD Prof.` reads title 'Prof.', maiden 'Smith', suffix 'PhD', and `… Prof. PhD` maiden 'Smith Prof.', suffix 'PhD' — H5's accepted particle-chain boundary inherited, the bare `Jane van der Berg Smith PhD Prof.` / `… Prof. PhD` splitting the same way — and `Berg, abdul nee Smith PhD Prof.` (bound-given join) and `Doe, Dr. nee Smith PhD Prof.` (no name word left) split likewise; the parent d9d80492 and 2.3.0 read all of these as the tree does, so only the claim is new. And a released particle with a title behind it is withdrawn: `Jane Doe nee Smith DO Prof.` reads title 'Prof.', maiden 'Smith DO', where `… Prof. DO` reads title 'Prof.', maiden 'Smith', suffix 'DO' (2.3.0 maiden 'Smith DO Prof.' and 'Smith Prof. DO'); `do` likewise. A particle INSIDE the clause does it from the other side, its chain able to take the credential behind it: `Jane Doe nee Smith do MA Prof.` reads title 'Prof.', maiden 'Smith do MA', and `… do Prof. MA` title 'Prof.', maiden 'Smith do', suffix 'MA' (2.3.0 kept every word in both). rules.md#M2 carries a pair of each shape as Accepted examples. ACCEPTED, H5'S REACH: `Jane Doe nee Smith King.` reads title 'King.', maiden 'Smith' (2.3.0 maiden 'Smith King.'), H5's accepted `Mary Jane King.` cost now reaching the last word of a birth name as it reaches the last word of a current one; rules.md#M2 carries it as an Accepted example. READER SCOPE IS UNCHANGED: the chain is consulted only where a trailing rule reads the clause's words (no comma, the part before a suffix comma, the given part after a family comma); before a family comma and in a tail segment the walk reads the peel alone and the clause keeps the title (`Doe nee Smith Prof., Jane` keeps maiden 'Smith Prof.'). THE FIRST-WORD FLOOR MOVED INTO THE CHAIN. `trailing_titles` and `tail_reading` take a `floor` (1 everywhere else), and the walk passes the position just past the marker's first word, so the chain never takes that word: `Jane Doe nee King.` keeps its maiden name (TITLES holds borne surnames and the marker announced one), and `Jane Doe nee Prof. Dr.` keeps 'Prof.' and gives up 'Dr.'. The first draft applied the floor as a clamp on the title stop after the chain had run, and that is measurably wrong rather than merely inelegant: a chain allowed to take the first word has already spliced it out of the count the re-peel reads, so `Jane Doe nee King. ba` kept 'ba' in the clause where `Jane Doe nee Smith ba` gives it up. `test_a_title_first_word_counts_as_a_word` is the invariant (a title-vocabulary first word against an ordinary one, the credential behind it read alike) and its docstring records the clamp's failure count. ONE RELEASE CHECK FOR THREE STOPS. The numeral, the credential and the title stop each ask `_release_reads_off` whether the name the take would leave reads what they give up as titles or suffixes with no join below the take absorbing it — rules.md#M2's invariant, "a word the clause gives up reads as a post-nominal or the clause keeps it", which the title stop now answers to as well. It is asked of the whole SPAN a stop gives up, not of the stop's own word: a first draft that checked the word alone broke the invariant on 513 of a 5,198-parse sweep (measured 2026-09-26 on that draft, which is not in the tree to recompute), because a stop gives up every word behind it — `DOE NEE SMITH PROF. MA` stopped at the title and handed the MA to the family, where the name left standing reads MA as a name; it keeps maiden 'SMITH PROF. MA' now. The question is asked per reader, the way that reader reads: the TRAILING reader is assign's own `tail_reading` over the view; the GIVEN_SLOT reader has its count settled by the comma, so it is the writing alone with a name word ahead. The name-word-ahead half is what keeps `Dr. nee Jones Smith Prof.` whole — the take would leave `Dr. Prof.`, whose Prof. would be the family name. THE JOINS. A released span holding a title with a particle ahead of it declines, because P2's chain runs on over a trailing title (rules.md#H5's Accepted `John van der Berg Prof.`): `Jane van der Berg nee Smith Prof.` keeps maiden 'Smith Prof.', and a released particle with a title behind it is withdrawn for the same reason (`Jane Doe nee Smith MA do Prof.` keeps maiden 'Smith MA do'). The bound-given (P5) half of `_join_takes_the_member` is asked for the GIVEN_SLOT reader alone: P5 is `BoundJoin.LENIENT` only after a family comma, and before one its STRICT reserve already refuses a join that would change a suffix reading, so asking there over-declined — `abdul nee Smith V` kept 'V' in the clause, where the scoped check lets it go to suffix as 2.3.0 did. The consequence is recorded by a case row rather than prevented: `abdul nee Smith Dr.` reads title 'Dr.', family 'abdul', maiden 'Smith', which is how `abdul Dr.` reads bare (2.3.0 read given 'abdul', maiden 'Smith Dr.'). THREE NUMERAL-STOP READINGS MOVE WITH IT, decided in rather than deferred, since each is the same invariant broken at the stop this change was already rewriting. The numeral stop never asked the join question: `Berg, abdul nee Smith V` read given 'abdul V' at 2.2.0 and 2.3.0 (given 'abdul nee', suffix 'V' at 2.0.0 and 2.1.0, #411's reserve differing between those pairs before the walk runs) and reads maiden 'Smith V' now. After a family comma the given slot reads a lone numeral as a suffix only where the given part is the LAST comma part — assign's own two-segment condition from #144 — and the walk now asks it too (`tail_follows`): `Doe, Jane nee Smith V, PhD` read middle 'V' at 2.2.0, 2.3.0 and e0f1a2fa, an M2 violation predating this change that its first draft had extended to `… V Prof., PhD`, and it reads maiden 'Smith V' now, as 2.0.0 read it. The withdrawal reaches through a title behind the numeral, so `Doe, Jane nee Smith Prof. V, PhD` keeps maiden 'Smith Prof. V' where `Doe, Jane nee Smith V Prof., PhD` gives up the title — the same asymmetry the bare given slot has with no marker in it (`Doe, Jane Prof. V, PhD` against `Doe, Jane V Prof., PhD`). And the numeral stop reads FROM the marker, unlike the other two, so a numeral straight after the marker is not held to the first-word floor, and it may decline the clause; examples, not a rule over every shape: `Dr. nee V` and `Doe, J. nee V` keep maiden 'V' as 2.3.0 did, though `Doe, J. V` reads suffix 'V', and `Doe nee V, Jane` and `Smith, John, PhD nee V` read no clause at all (family 'Doe nee V', suffix 'PhD nee V', as at 2.3.0) — `Jane Doe nee V Prof.` read maiden 'V Prof.' at 2.3.0 and reads family 'nee', suffix 'V', title 'Prof.' now, as `Jane Doe nee V` reads plus the title. After a family comma with another comma part behind the given one, the same #144 condition that keeps `Doe, Jane nee Smith V, PhD` whole now keeps a lone numeral in the clause as well: `Doe, Jane nee V, PhD` read given 'Jane', middle 'nee V', suffix 'PhD' at 2.3.0 and at the parent d9d80492 and reads maiden 'V', suffix 'PhD' now, as 2.0.0 and 2.1.0 read it, `Doe, Jane nee V, Jr.` likewise — the marker had been read as a name word there at 2.2.0 and 2.3.0, and a case row records it. A LINK THE WALK STOPS AT ASKS THE RELEASE QUESTION TOO. Reading the end of the clause through the title chain makes the link exception refuse a link it used to join — in `Jane Doe nee Smith i DO Prof.` the DO is the peel's once the title is chained — and a stop at a link gives up the words behind it. Unchecked, that put 'Doe i' in the middle name and 'DO Prof.' in the family (e0f1a2fa read maiden 'Smith i DO Prof.'), and `Berg, abdul nee Smith i V Prof., MD` read middle 'V'. So a link the exception refuses asks `_release_reads_off` of the run it would give up — only where the exception, bounded by the peel over the words as WRITTEN, would have joined it (the refusal is the title chain's), where a trailing rule reads the clause, and where the link is not the first word after the marker (a first-word stop declines the clause and gives nothing up); where that fails the clause keeps the link and walks on. Scoped that narrowly because a first version asked it of every refused link and kept links the name left standing reads off: `Doe, Jane nee Smith i V` kept maiden 'Smith i' where every release gives up suffix 'i V' (23,472 of a 411,936-parse link grid moved clean readings, measured 2026-09-26). With it, the given-part model reads the lenient numeral in both of assign's passes — the literal last piece first, then the last one standing once the chain has taken the titles behind it — and only where no comma part follows (`Doe, Jane nee Smith i V Prof.` gives up 'i V' and the title, as bare `Doe, Jane i V Prof.` reads them; `… i V Prof., PhD` keeps maiden 'Smith i V'). The same model moves the credential stop where a lenient numeral follows the credential: `Doe, Jane nee Smith MA V` now gives up 'MA V' (suffix 'MA V', maiden 'Smith') as bare `Doe, Jane MA V` reads it, where d9d80492 and 2.3.0 read maiden 'Smith MA', suffix 'V'; `… MA V, PhD` keeps maiden 'Smith MA V'. Over that link grid (links i, y, e, and, i y, y i, and i × 8 heads × bodies × `, MD` tail × 3 cases × 3 name orders), against the tree before the link check: 0 new violations, 6,195 fixed, and the clean readings that move now match the bare given part's (`DOE, JANE NEE SMITH MA I` gives up 'MA I', as `DOE, JANE MA I` reads suffix 'MA I'). `Jane Doe nee Smith i DO Prof.` now reads title 'Prof.', maiden 'Smith i DO' and reports the kept DO, and `Berg, abdul nee Smith i V Prof., MD` maiden 'Smith i V Prof.', suffix 'MD'. A RELEASED TITLE THAT IS ALSO A PARTICLE IS KEPT AFTER A FAMILY COMMA, because P6 attaches a particle trailing the given part to the family: released by the title stop, 'St.' in `Doe, Jane nee Smith St.` reads family 'St. Doe', so `_join_takes_the_member` declines it and the name reads maiden 'Smith St.', as e0f1a2fa and 2.3.0 read it, and so does `Doe, Jane nee Smith MA St.`. Titles only — a credential that is also a particle ('DO') is the given slot's own lean, which the credential stop has already asked (#533). A FOURTH MEASUREMENT, against e0f1a2fa rather than d9d80492: the M2 invariant over a fuzz of twelve heads ({`Jane Doe`, `Doe, Jane`, `Doe, Prof.`, `Jane Doe, PhD`, `Doe, J.`, `Berg, Jane van der`, `Jane van der Berg`, `Berg, abdul`, `abdul Berg`, `J. Doe`, `Prof. Jane Doe`, `Dr.`}), ` nee Smith ` and every one-to-three-word sequence over {MA, Ma, V, Prof., St., King., PhD, ba, do, DO, Jr., i, van, y} holding at least one of Prof., St., King., then nothing or `, MD`, as written, upper- and lower-cased (107,352 parses, measured 2026-09-26 on the narrowed tree; the intermediate tree above read 2,228): 588 texts that violated it at e0f1a2fa no longer do, and 12 newly do, all `Dr. nee Smith i MA|V` followed by `Prof.`, `King.` or `St.`, with or without `, MD`. Each is the title-carrying twin of `Dr. nee Smith i MA` / `Dr. nee Smith i V`, which read family 'i' at e0f1a2fa already: the bare-title head leaves no name word, the unguarded stop's class (#548), which the title now reaches because it leaves the clause. ACCEPTED: `Jane Doe (nee Smith Prof.)` reads title 'Prof.', maiden 'Smith' — bracket content ending in a period is suffix-shaped (S1), so the brackets are dropped and the clause is read bare, outside the reach of the delimiter precedence; rules.md#M2 carries it as an Accepted example. MEASURED 2026-09-26, the tree against its parent d9d80492, over three generated grids, and these are dated snapshots. Recipe: grid A is each head in {`Doe, Jane`, `Doe, J.`, `Doe, Prof.`, `Berg, abdul`, `Doe, Jane van`, `Jane Doe`, `J. Doe`, `Dr.`, `abdul`, `Jane van der Berg`, `John`}, then ` nee `, then every sequence of one to three words (repetition allowed) from {Smith, Jones, Prof., Dr., Sir, MA, DO, PhD, V, III, i, do, van, Ma, M.A., King., Rev., ba}, then either nothing or `, PhD`, each text as written, upper-cased and lower-cased, duplicates removed — 327,936 texts; grid B, aimed at the bound-given join, is the same construction over heads {`Dr. abdul`, `abdul rahman`, `Dr. abdul rahman`, `Berg, Dr. abdul`, `Berg, abdul rahman`, `abu`, `Berg, abu`, `Mr. abu`, `Dr.`, `Berg, Dr.`, `abdul`, `Berg, abdul`}, words {Smith, Prof., MA, V, do, van, PhD, Ma, ba, III, Jr., M.A.} and tails {nothing, `, PhD`, `, Jr., MD`}, plus these fourteen texts, as written only: `Jane Doe nee Ph. D. Prof.`, `Jane Doe nee Ph. D. Smith Prof.`, `Jane Doe z domu King. ba`, `Jane Doe z domu Smith Prof.`, `Jane Doe nee King. Prof. ba`, `Jane Doe nee King. Prof.`, `Doe, Jane nee Smith V,`, `Doe, Jane nee Smith V, ` (with the trailing space), `Doe, Jane, nee Smith V`, `Doe, Jane nee Smith V Prof., PhD`, `Doe, Jane nee Smith Prof. V, PhD`, `Doe, Jane nee Smith MA, PhD`, `Doe, Jane nee Smith V (Jr.)`, `Doe, Jane nee Smith V "Bo"` — 173,057 texts. The invariant tested is M2's: in a parse with a non-empty maiden field, every token written after the ` nee ` marker is roled maiden, title or suffix (so the two `z domu` texts are parsed and not checked). Grid C, a four-word probe of the title chain after a family comma, is each head in {`Doe, Jane`, `Jane Doe`, `Doe, J.`, `John Smith`, `Doe, Jane Mary`}, then ` nee `, then every sequence of four words (repetition allowed) from {Smith, Jones, Prof., Dr., King., MA, PhD, V, Jr., ba, do}, then either nothing or `, PhD`, as written only — 146,410 texts. Per grid, the counts are: texts that newly violate the invariant at the tree, texts that violated it at the parent and no longer do, texts that violate it at both. Grid A: 0, 2,012 and 17,854; grid C: 0, 1,368 and 15,534; grid B: 0, 1,222 and 11,149 (all re-measured 2026-09-26 on the tree with the link and particle-title checks below, as narrowed; an intermediate tree whose link check fired on every refused link read grid A 0, 3,772 and 16,094, its extra 1,760 fixes being links the as-written reading refuses too, which are #548's class), one of those last being the check itself counting the quoted nickname in `Doe, Jane nee Smith V "Bo"`. In grids A and B, every other remaining violation has the clause ending immediately before a word of the unambiguous suffix vocabulary written the way the plain suffix-word stop takes it (PhD, III, M.A., Jr., and a lower-case i or v; never a bare capital, which reads as an initial), which is where that stop ends a clause without asking a release question. DEFERRED: the walk's plain suffix-word stop — the one ending the clause at the first suffix word, as distinct from the trailing numeral, credential and title stops — asks no release question at all, and where the words behind it cannot read as post-nominals they land in a name part: on degenerate heads the stop word itself does (`Dr. nee Smith PhD Prof.` reads family 'PhD', `Doe nee Smith Jr. Prof., Jane` family 'Doe Prof.'), and with an ordinary head the words behind it do (`Doe, Jane nee Smith PhD Smith` reads middle 'Smith'). Unchanged here, and what "the clause keeps it" should mean where the kept words would follow a credential the clause itself ended at is its own question; rules.md#M2 carries `Doe nee Smith Jr. Prof., Jane` as an Accepted example meanwhile. The trailing numeral's stop has the same gap before a family comma, where it is made over the peel alone with no release question: `Doe nee Smith V, Jane` reads family 'Doe V', maiden 'Smith' (so did 2.2.0, 2.3.0 and the parent d9d80492; 2.0.0 and 2.1.0 read maiden 'Smith V'), and rules.md#M2 carries it beside the other. Open: #548. COST, measured 2026-09-26 on CPython 3.11.16 as profiler call events in one `Parser.parse` after a warm-up parse, parent → tree: `John Smith` 171 unchanged, as are `Jane Doe Prof.` and `John Smith MA`; `Jane Doe nee Smith` 244 → 246; `Jane Doe nee Smith MA` 377 → 388; `Jane Doe nee Smith Prof.` 276 → 372, the one shape that now runs the chain and a release check it never ran. Recompute: count `sys.setprofile` call events around the second of two `Parser().parse` calls, with the tree and d9d80492 each first on `sys.path`. - 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'. @@ -404,6 +405,7 @@ The reconciled v1-style banks (`tests/test_*.py`) carried eight `@pytest.mark.xf - 2026-09-24 #478 — CROSS-REFERENCE, no change to this rule: case repair's hyphen clause is the one place a marked letter in a name written wholly in one case is read as the connective rather than an initial — `maria silva-e-sousa` repairs to `Maria Silva-e-Sousa` beside the spaced `Maria Silva E Sousa`, and `conjunctions_ambiguous` does not reach the hyphenated form. Decided there, with its cost (`J-E-P DUPONT` → `J-e-P Dupont`): decisions.md#R4's 2026-09-24 #478 bullet. - 2026-09-26 #538 — CROSS-REFERENCE: rules.md#P3's separator sentence (a declared delimiter is read past as a connective is) is decided and measured in the 2026-09-26 #538 entry under `### M2`. +- 2026-10-03 #549 — CROSS-REFERENCE: that separator sentence is reversed; it now says a declared separator ends the search as a comma would, a connective never joining across one, and the decision is C1's 2026-10-03 entry. - Provenance: the single-letter-connective guard is v1's fix for Google Code issue 11 ("john e smith", 2013, commit 33676c9) — the "#11" citations that circulated pointed at a GitHub accident, not the real source. Recorded so the archaeology stays done. @@ -910,6 +912,12 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): BLAST RADIUS, measured 2026-10-01: the gate exits 0 at all five baselines, and of the 1459 names in master's corpora exactly one moves, `De La Cruz, M.J. K.L.` (above, decided). Comparator: `parse(n).as_dict()`, the ambiguity kinds and `initials()` for every name in every `tools/differential/corpus*.jsonl` at `git archive origin/master`, under all three name orders, on master's tree against this one. The population that could move is small and the corpus is evidence about itself, not the rule: two corpus names have a part before the comma that is one particle surname of several words, with the part after the comma holding a credential — that one and `De La Cruz, M.J. PhD`, which keeps given 'M.J.'. Recompute: corpus names whose part before the first comma has more than one token and `_vocab.surname_unit_count` 1 (24 at master, nearly all `de la Vega, Juan`-type listings), then keep those whose second segment holds a word `_vocab.ambiguous_class_candidate` admits (a listed or dotted member of the credential class) — 2 at master. A filter on any suffix word keeps 17, the `de la Vega, Juan … III` listings among them. Against 2.3.0, `fix(#575)` classifies one name, `van der Berg, PhD`; the other particle-surname example names diff there only by this cycle's #289 count and report. Outside the corpora the move against 2.3.0 is a class, not a list: a leading ambiguous particle and one word, before a comma followed by an unambiguous credential or a title alone, now reads as one surname (`Abu Bakar, PhD`, `bin Laden, PhD`, `Mac Donald, PhD`, `van Gogh, Jr.`, `Van Johnson, Dr.`), where 2.3.0 read the particle as the given name; `Freiherr von Berg, Ed` and `Abu Bakar, Ed` move only against this cycle's master. The boundary case rows added in review carry no shape tag on purpose: their diffs against the older baselines come from earlier changes (#296's positional read, #289's count), so admitting them to the contract corpus would have stretched unrelated ledger rules over them; the case table asserts them either way. COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. - 2026-10-01 (Derek), #564 — AN UNLISTED ALL-CAPS WORD IS A CREDENTIAL IN THE PART AFTER A COMMA BEHIND TWO NAME WORDS, BY DEFAULT. The decision, its two-letter rule and the mixed-run membership are recorded under decisions.md#S2 (#564), where the caps shape's switch has always been decided; this bullet is C1's pointer to it. +- 2026-10-03 (Derek), #549 — A DECLARED DELIMITER PARTS A TRAILING SUFFIX PART AS A COMMA WOULD, AND THIS REVERSES #538'S SKIP READING (M2, 2026-09-26). The question #549 asked was whether a core a connective join swallows should be skipped (the text read as if the core were absent) or should end the run it stands in; Derek's ground for the second was the one a caller has when declaring one: a delimiter set to ` - ` is meant to behave like a comma around suffixes, and the simplest parser that honours that is one where it IS a comma there. So `group` cuts a tail segment at its cores before grouping anything, groups each part as the comma twin's own segment would be grouped, and drops the cores, which post_rules' entry pass still reads back as boundaries. No join, maiden walk or link search can reach across a core because none of them ever sees one, and the machinery that let them step over it is gone: the `cores` parameter of `_group_segment`, `_maiden_take`, `_run_neighbours` and `_link_joins_inside_the_clause`, `_maiden_take`'s `seen` index list, the `skip` arguments of `peel_walk` and `trailing_start`, and the #206 drop block that ran after the joins. One call per parse fewer (`tools/perf/call_count.py`, py3.11, 2026-10-03: parse 369 against 370 at 061f02da). + WHAT MOVES, against 1.4.0 and against 2.3.0. A connective beside a core no longer joins across it: `HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix` is 'PhD, and MD', and `... Puig - y Soler` 'Puig, y Soler' -- 1.4.0's answers (its `expand_suffix_delimiter` split the part's text before anything joined), where 2.0.0, 2.1.0, 2.2.0 and 2.3.0 each gave 'PhD - and MD' and 'Puig - y Soler' (all measured 2026-10-03 with the released wheels). The generational link moves back too: 2.3.0 read `Smith, John, Puig - i Soler` as 'Puig, i Soler' and `Smith, John, PhD née Puig - i Soler` as maiden 'Puig', suffix 'PhD, i Soler', `i` being no connective there, and #397/#538 had moved them to 'Puig - i Soler' and maiden 'Puig i Soler' before either shipped; both read as 2.3.0 did again. An interior core still parts: `... PhD née Puig Dr. i - y Soler` gives suffix 'PhD i, y Soler'. At the default policy nothing moves, `extra_suffix_delimiters` being empty: the gate exits 0 at all five baselines. + THE DECLINED READING, TAKEN. #538 declined the boundary reading because it moves `Smith, John, PhD née Puig - i Soler` from maiden 'Puig i Soler' to maiden 'Puig' with suffix 'PhD, i Soler', "a birth-name link pushed into the credentials". That is what the same text with a comma typed in the core's place reads, and it is what 2.3.0 read; the comma reading is accepted as the cost of one rule for a declared separator, where the skip reading kept two (a lone core dropped like a comma, a core beside a connective stepped over like a word) and #549 was the seam between them. + INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin`: every field of a name with a ` - ` core in a part after the suffix comma equals that of the same text with a comma in its place, over 129 generated texts; 102 disagreed at 061f02da, 0 after. It replaces #538's `test_a_delimiter_core_reads_as_if_it_were_not_written`, whose invariant this decision reverses. Scoped to TRAILING parts on purpose: before the first comma, or in the given part of the listing form, the core is a word (C1's Accepted entry, v1 parity), and a comma typed there changes the name's structure rather than parting a suffix run, so the twin is a different parse. + MEASURED 2026-10-03, branch against 061f02da. POPULATION: each head in {`Smith, John,`, `Smith, John, Jr.,`, `John Smith,`, `John Smith`} followed by every 2-, 3- and 4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr., née, and} that contains `-`, 19,972 texts, plus the 1,505 distinct names of `tools/differential/corpus*.jsonl` at 061f02da. Under `extra_suffix_delimiters=(" - ",)` 2,624 of the generated parses move (suffix alone 2,384, suffix and maiden 240); under `(" / ",)`, with each ` - ` rewritten to ` / `, 2,204 (1,924 and 280); and under `(" and ",)` over the texts as written, where `and` is now the core and the dash a word, 2,528 (2,468 and 60). No other field and no ambiguity report moves under any of the three. In the corpus, one name moves under each of two policies: `Smith, John, PhD née Puig - i Soler` under ` - ` (above), and `Doe, Jane, and Jr.` under ` and `, suffix 'and Jr.' to 'Jr.', a leading core dropped as any lone core is. COMMA-TWIN AGREEMENT over the generated texts under ` - `, comparing each against the text with ` - ` replaced by `, ` under the default policy and keeping only texts whose core is neither first nor last after the head nor beside another core: `Smith, John,` 752 of 2,100 disagreed before and 0 after, `Smith, John, Jr.,` 973 of 3,310 and 0; `John Smith,` 1,956 of 2,100 both before and after, the head where the typed comma moves the name's structure. RECOMPUTE: parse each text under each policy at both trees (pin `nameparser.__file__` on each side, AGENTS.md's two-tree gotcha), compare the seven role fields and the sorted ambiguity kinds, and count the texts whose record differs. + SUPERSEDES the 2026-09-26 #538 entry under M2 (its stepping, its invariant and its Not-repaired list, which this closes) and the 2026-09-22 #397 follow-up's bound argument there (a core can no longer stand between a marker and the clause's first word, the segment being cut at it first). Not changed: a core OUTSIDE a tail segment is still a word (C1's Accepted entry), and a core inside a token (`RN/CRNA` under `/`) still reads by `splits_into_suffixes` and keeps the token whole. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 932c4e46..ee14dda2 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -547,11 +547,8 @@ P3. Rationale: connective words ("y", "of the") bind name words into nothing, and a word of that vocabulary ending a name, or standing before the credential a name ends with, is the generation it also spells. - A separator the caller declared is not a word of the name, and - the search reads past it as it reads past a connective: the word - on the connective's side is the first one beyond the separator, - so a connective of that vocabulary beside one joins exactly where - it would with the separator absent. + A separator the caller declared ends the search as a comma would, + a connective never joining across one (C1). Both questions this rule asks of a name — how many words it has, and whether it is written in one case — are asked of the name's OWN words: a maiden marker taken as one, and the words it takes @@ -1521,7 +1518,7 @@ M2. Rationale: a maiden marker announces that what follows it is the "Jane Doe nee Smith i DO Prof." → maiden="Smith i DO" "Doe, Jane nee Smith St." → maiden="Smith St." "Smith, John, PhD née Puig Mr. - i Soler" extra_suffix_delimiters-dash → maiden="Puig Mr." - "Smith, John, PhD née Puig - i Soler" extra_suffix_delimiters-dash → maiden="Puig i Soler" + "Smith, John, PhD née Puig - i Soler" extra_suffix_delimiters-dash → maiden="Puig" "Jane Doe nee Smith Prof." → maiden="Smith" "Jane Doe nee Smith Prof." → title="Prof." "Jane Doe nee Smith MA Prof." → suffix="MA" @@ -1858,6 +1855,13 @@ C1. Rationale: a credential run after the comma means the name is in further comma. Longer suffix words are not in question either way, and the strict knob above still vetoes the initial-shaped ones, so the run ends at them there. + A delimiter the policy declares parts a trailing suffix part as a + comma would: the words on each side of it read exactly as the + parts of the same text written with a comma in its place, so no + join, maiden clause or connective reaches across it, and the + delimiter itself is dropped. Only a trailing part is parted: in + the part before the first comma, or in the given part of the + listing form, the delimiter is a word (the Accepted entry below). "Smith, John" → family="Smith" "سلمان، محمد" → family="سلمان" "田中、太郎" → family="" @@ -1940,6 +1944,8 @@ C1. Rationale: a credential run after the comma means the name is in "García Márquez, Ms G.J." → given="G.J." · boundary "García Márquez, Ed G.J." → given="Ed" · boundary "John Smith, A.B. Ph.D." → given="A.B." · boundary + "Smith, John, Puig - i Soler" extra_suffix_delimiters-dash → suffix="Puig, i Soler" + "Smith, John, PhD - and MD" extra_suffix_delimiters-dash → suffix="PhD, and MD" Accepted: a title in front of a particle surname, before the credential class's word, is read into the family. The surname counts once and the title is no name word, so the count reads the diff --git a/docs/release_log.rst b/docs/release_log.rst index 27fa6d0b..a325b4ed 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -38,7 +38,7 @@ Release Log - **Fix a roman numeral ending a maiden clause landing in the first or middle name.** ``HumanName("Berg, abdul nee Smith V")`` gives first ``abdul`` with maiden ``Smith V``, where 2.2 and 2.3 gave first ``abdul V`` -- a word of the birth name joined into the current given name -- and 2.0 and 2.1 gave first ``abdul nee``. ``Doe, Jane nee Smith V, PhD`` gives maiden ``Smith V``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``V``: after a family comma a lone numeral ending the given part is a suffix only when no further comma part follows it, and the clause now asks that as the given part itself does. Without the credential tail the numeral still goes to the suffix (``Doe, Jane nee Smith V`` gives maiden ``Smith``, suffix ``V``). A numeral straight after the marker can leave the marker an ordinary word, as ``Jane Smith née V`` does, and where it does a title behind the numeral no longer changes that: ``Jane Doe nee V Prof.`` gives title ``Prof.``, suffix ``V``, middle ``Doe``, last ``nee``, where 2.3.0 gave maiden ``V Prof.``. Elsewhere the clause keeps it -- ``Dr. nee V`` keeps maiden ``V`` as before -- and ``Doe, Jane nee V, PhD`` now gives maiden ``V``, suffix ``PhD``, as 2.0 and 2.1 did, where 2.2 and 2.3 gave middle ``nee V`` -- the same further-comma-part condition as ``Doe, Jane nee Smith V, PhD`` above -- with ``Doe, Jane nee V, Jr.`` moving the same way. See the ``M2`` entry of ``docs/design/decisions.md`` (#535) - - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Doe, Jane nee Puig i Soler`` and ``Jane Doe née Kowalska i Nowak`` move with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` is read past the way a connective is, so a link beside one reads as it would with the delimiter absent: under ``(" - ",)``, ``Smith, John, PhD née Puig Mr. - i Soler`` keeps maiden ``Puig Mr.`` and ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, both as 2.3.0 read them, where the link had first let the delimiter pass for a name word and given maiden ``Puig Mr. i Soler`` and suffix ``PhD - i Soler``. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) + - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Doe, Jane nee Puig i Soler`` and ``Jane Doe née Kowalska i Nowak`` move with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under ``(" - ",)``, ``Smith, John, PhD née Puig Mr. - i Soler`` keeps maiden ``Puig Mr.``, ``Smith, John, PhD née Puig - i Soler`` keeps maiden ``Puig`` with suffix ``PhD, i Soler``, and ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, all as 2.3.0 read them. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) - **Fix a connective contributing no initial even where it is joining nothing.** ``parse("Juan de y").initials()`` gives ``J. y.``, where every release gave ``J.`` while ``family_base`` said ``y`` -- two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, so ``John and Jane Smith`` gives ``J. J. S.`` where 2.0 through 2.3 gave ``J. a. J. S.`` and 1.4.0 the run-together ``J a J. S.``, ``Duke of Edinburgh`` gives ``D. E.`` where 2.0 through 2.3 gave ``D. o. E.`` and 1.4.0 ``D o E.``, and ``John & Jane`` gives ``J. J.``. The question is asked of the whole part and never of a word count, so ``Jon Dough and`` has base ``Dough and`` and keeps ``J. D.``, and ``Juan Velasquez y Garcia`` keeps ``J. V. G.``. ``HumanName.initials()`` moves with the core -- over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it: ``JUAN Y GARCIA`` and ``محمد و علي`` both give the answer 1.4.0 gave. Parsing got cheaper by the same change -- the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, so ``initials()`` and ``capitalize()`` disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #461) @@ -70,6 +70,8 @@ Release Log - **Fix a katakana name typed with a separate voicing mark being read family-first.** ``HumanName("ア゙イ タロウ")`` gives first ``ア゙イ``, last ``タロウ``, as ``アイ タロウ`` does, where 2.1 through 2.3 gave last ``ア゙イ``, first ``タロウ``. A dakuten or handakuten after a kana that has no precomposed voiced form (``ア゙``, ``ン゙``), or the spacing ``゛`` and ``゜``, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: ``マイケル ア゙イ`` gives first ``マイケル`` where it gave last ``マイケル``. Hiragana names and kanji-and-kana names with such a mark read as before. See the ``W4`` entry of ``docs/design/decisions.md`` (closes #596) + - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part exactly as a comma typed in its place would: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. Only a trailing suffix part -- one after the name's commas -- is affected; elsewhere the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) + - **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597) **Additions** diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 5669405f..ff9738d0 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -8,8 +8,8 @@ tail tokens get role=MAIDEN; marker tokens land in dropped. Reads: token tags (from classify), Lexicon.given_name_titles (the P5 licence, #369) and Policy.extra_suffix_delimiters, whose -delimiter-core tokens tail segments drop (v1 suffix_delimiter parity) --- no other Policy field. Policy.lenient_comma_suffixes left this list +delimiter-core tokens part a tail segment as a comma would and are +dropped (v1 suffix_delimiter parity, #549) -- no other Policy field. Policy.lenient_comma_suffixes left this list with #436: it reached here only through segment_suffix_reading, whose render consumer was this stage's one-entry join and now lives in post_rules. The v1 "derived titles/prefixes" @@ -197,23 +197,24 @@ def marker_run_length(following: Iterable[Set[str]]) -> int: return run -def _marker_run_pieces(seen: Sequence[int], pieces: Sequence[Sequence[int]], +def _marker_run_pieces(pieces: Sequence[Sequence[int]], tokens: Sequence[WorkToken], m: int) -> int: - """How many of `seen`'s pieces the marker run at seen[m] spans: 1 + """How many pieces the marker run at pieces[m] spans: 1 for a single-word marker, more for a phrase entry ('z domu'). classify already decided where the run ends and recorded it on the tokens -- "vocab:maiden-marker" on the head, "vocab:maiden-marker-cont" on the rest. - Each continuation is the NEXT piece and is always in `seen`, and + Each continuation is the NEXT piece, and that holds because classify REFUSES to tag a run whose tokens are not structurally contiguous. It is not a property of this walk, and - the reasons `seen` can skip a token are wider than they look. No - join has run yet, so a piece is one token. `seen` itself skips a - tail segment's delimiter cores, and a core between two marker words - would be a token between them, which classify would not have tagged - as a run. But `pieces` comes from a SEGMENT, and segment keeps only + the reasons a token can be missing are wider than they look. No + join has run yet, so a piece is one token. A tail segment's + delimiter cores are cut out before grouping (#549), and a core + between two marker words would be a token between them, which + classify would not have tagged as a run. And `pieces` comes from a + SEGMENT, and segment keeps only the tokens no stage has given a role, bucketed by the commas before them -- so a run half inside a bracketed clause, or split across a structure comma, is one no segment holds whole. Walking cont tags @@ -229,7 +230,7 @@ def _marker_run_pieces(seen: Sequence[int], pieces: Sequence[Sequence[int]], segment at all -- see _vocab.tag_marker_runs, which states the limit. """ return marker_run_length( - tokens[pieces[seen[k]][0]].tags for k in range(m + 1, len(seen))) + tokens[pieces[k][0]].tags for k in range(m + 1, len(pieces))) # rules.md#M2: "a word the clause gives up reads as a post-nominal or @@ -511,8 +512,7 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - beside: list[_Beside], - cores: Set[str]) -> bool: + beside: list[_Beside]) -> bool: """Whether the suffix piece at `k` is a connective PLACED TO JOIN between two name words of the clause `lo`..`hi`. @@ -542,14 +542,13 @@ def _link_joins_inside_the_clause(k: int, lo: int, hi: int, if not _is_conj_piece(pieces[k], ptags[k], tokens): return False if not beside: - beside.append(_run_neighbours(pieces, ptags, tokens, cores)) + beside.append(_run_neighbours(pieces, ptags, tokens)) return _between_name_words(k, lo, hi, pieces, ptags, tokens, beside[0]) def _maiden_take(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - cores: Set[str], one_case: bool | None, reader: TailReader, ambiguities: list[PendingAmbiguity], @@ -606,28 +605,18 @@ def _maiden_take(pieces: Sequence[Sequence[int]], fork and the reader decide it, and 'Doe, J. nee V' keeps maiden 'V' though 'Doe, J. V' reads suffix 'V'. - A tail segment's delimiter cores (`cores`, empty elsewhere) are - structure, not words, and group() drops them after the pass -- - before the pass moved ahead of the joins it dropped them first. - So the walk steps over them: a core is neither a word the marker - can take ('PhD née - Jones' read maiden '- Jones') nor the name - word M2 needs ahead of the marker ('- née Jones' took 'Jones', - leaving the core alone, which the drop then kept as the segment's - only piece). They stay in place for the drop, which still sees - the segment as written, which is why this returns indices rather - than a slice. + A tail segment's delimiter cores never reach this walk: group() + cuts the segment at them first and hands each part over on its + own, as the comma twin's segments would be (#549). So a core is + neither a word the marker can take ('PhD née - Jones' reads as + 'PhD née, Jones') nor the name word M2 needs ahead of the marker. """ - # the lone-core test; also in _run_neighbours and group()'s #206 - # drop, whose copy adds `len(pieces) > 1` -- keep in step - seen = [k for k in range(len(pieces)) - if not (len(pieces[k]) == 1 - and tokens[pieces[k][0]].text in cores)] - m = next((v for v in range(1, len(seen)) - if _is_maiden_marker_piece(pieces[seen[v]], tokens)), None) + m = next((v for v in range(1, len(pieces)) + if _is_maiden_marker_piece(pieces[v], tokens)), None) if m is None: return None # the marker may be a phrase, in which case it is several pieces - run = _marker_run_pieces(seen, pieces, tokens, m) + run = _marker_run_pieces(pieces, tokens, m) # "up to any trailing suffix": a suffix WORD anywhere after the # marker ends the maiden name, and so does the trailing numeral as # assign will read it, which the suffix-piece test does not see @@ -662,8 +651,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # pre-change 9,852 says so: dropping it moves 18 of them, on # 'John van der Berg Ma', 'John de Ma' and 'Freiherr von Berg MA' # under every one of the six policies. - skip = frozenset(range(len(pieces))) - frozenset(seen) - rest = peel_walk(seen[m], ptags, skip) + rest = peel_walk(m, ptags) # #535: where a trailing rule reads these words, the walk reads the # end of the name as that rule does -- the S2 peel and the H5 title # chain to their fixed point (`tail_reading`), so a title behind @@ -694,7 +682,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # kept 'ba' in the clause where 'Jane Doe nee Smith ba' gives # it up (#535 review). A marker run with nothing after it has # no floor to set; the walk below declines it anyway. - first = seen[m + run] if m + run < len(seen) else None + first = m + run if m + run < len(pieces) else None floor = rest.index(first) + 1 if first in rest else 1 rest, chained, peeled = tail_reading(rest, pieces, ptags, tokens, one_case, floor) @@ -713,7 +701,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # 'Dr. née Jones Smith V' leaves the numeral as assign's whole # rest and no fork fires at all. if trailing < len(pieces): - left = [i for i in seen if i < seen[m] or i >= trailing] + left = [i for i in range(len(pieces)) if i < m or i >= trailing] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] # The same peel pair as above, read over the view -- and, for @@ -771,7 +759,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # per member, and no re-entrancy -- the predicate never calls the # walk that calls it. if (reads is not None and peeled.names < len(rest) - and m + run + 1 < len(seen)): + and m + run + 1 < len(pieces)): # THE FIRST-WORD FLOOR: the stop never takes the FIRST word # after the marker -- a class member standing alone there # stays the maiden name. A clamp rather than a veto: where the @@ -781,7 +769,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # clamped piece may then be no member at all, and the test # below declines -- which changes nothing, the walk stopping # at that suffix word of its own accord. - stop = max(rest[peeled.names], seen[m + run + 1]) + stop = max(rest[peeled.names], m + run + 1) head = pieces[stop] # `len(head) == 1` is DEFENSIVE and measured inert # (2026-09-19) rather than unreachable: it asks a LONE piece's @@ -842,7 +830,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # class member. if (stop < trailing and len(head) == 1 and AMBIGUOUS_ACRONYM_TAG in tokens[head[0]].tags): - left = [i for i in seen if i < seen[m] or i >= stop] + left = [i for i in range(len(pieces)) if i < m or i >= stop] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] # where the member stands in that view: everything before @@ -909,7 +897,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], if chained and reads is not None: stop = min(chained) if stop < trailing: - left = [i for i in seen if i < seen[m] or i >= stop] + left = [i for i in range(len(pieces)) if i < m or i >= stop] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] at = left.index(stop) @@ -925,36 +913,16 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # bound it with, so the decline the `j <= m + run` test below # reaches is taken here instead -- `lo` would have no piece to name # ('Jane van der Berg née'). - if m + run >= len(seen): + if m + run >= len(pieces): return None # rules.md#M2: "a link inside the birth name does not end it" -- # the clause's OWN bounds for the link exception, which are not the # segment's. `lo` is the first piece after the marker run, so the - # marker is never the name word on a link's left, and a delimiter - # core between the marker and that first word is below `lo` by - # construction and cannot pass for one either. A core is the TAIL - # segment's alone (`extra_suffix_delimiters`, empty by default), - # and a dash standing where no tail segment can hold it is an - # ordinary word at EITHER policy: 'PhD née - i Jones' keeps maiden - # '- i Jones' configured and unconfigured alike, there being no - # comma to make a tail out of. It takes the tail a suffix comma - # builds for the dash to be a core at all, and then the two - # policies part company -- 'Smith, John, PhD née - i Jones' keeps - # maiden '- i Jones' by default and declines under a configured - # ' - ', which is the pair - # test_a_core_between_the_marker_and_the_first_word_is_below_lo - # holds (all four readings measured 2026-09-22). - # NOT theoretical, settled by measurement 2026-09-21 over corpus u - # cases.py u the property grids u a 50,925-name generated set with - # cores, under thirteen core-bearing policies: 25,536 of 596,392 - # maiden takes had a core standing exactly there, so the bound is - # load-bearing. PAST `lo` the bound no longer reaches a core, and - # `_run_neighbours` steps over it as it steps over a connective - # (#538): a core is structure, so the word on a link's side is the - # one past it, and the clause reads as the same text written - # without the core ('Smith, John, PhD née Puig Mr. - i Soler' ends - # at the link after the title, as '... Puig Mr. i Soler' does -- - # the title is the word on the link's left and refuses). + # marker is never the name word on a link's left. No delimiter + # core stands anywhere in the clause: a tail segment's cores are + # cut out before this walk (#549), and a dash where no tail + # segment holds it is an ordinary word at either policy ('PhD née + # - i Jones' keeps maiden '- i Jones' configured or not). # `peel_start` is where assign's trailing run begins -- over the # pieces as written for a NONE reader (`trailing_start`'s whole # answer, read off the peel pair above rather than re-running it), @@ -975,19 +943,19 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # visits. A dated snapshot from before #535, measured 2026-09-20 # with a probe here over the whole suite: 93,408 reaches of this # site, `peel_start > trailing` 0 of them. - lo = seen[m + run] + lo = m + run peel_start = (rest[peeled.names] if peeled.names < len(rest) else len(pieces)) j = m + run # The link exception's memo cell, filled inside the predicate on # the first connective it is asked about (see its docstring). beside: list[_Beside] = [] - while j < len(seen) and seen[j] < trailing: - k = seen[j] + while j < len(pieces) and j < trailing: + k = j if (not is_suffix_piece(pieces[k], ptags[k], tokens) or _link_joins_inside_the_clause(k, lo, peel_start, pieces, ptags, tokens, - beside, cores)): + beside)): j += 1 continue # rules.md#M2: "a word the clause gives up reads as a @@ -1017,8 +985,8 @@ def _maiden_take(pieces: Sequence[Sequence[int]], else len(pieces)) if _link_joins_inside_the_clause(k, lo, written_start, pieces, ptags, tokens, - beside, cores): - left = [i for i in seen if i < seen[m] or i >= k] + beside): + left = [i for i in range(len(pieces)) if i < m or i >= k] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] at = left.index(k) @@ -1064,7 +1032,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # 0 times over 1,760,904 parses, and deleting the test is # byte-identical (fields, ambiguities, token roles and tags) over # the 905,796-parse oracle at the same frame counts. - last = pieces[seen[j - 1]] + last = pieces[j - 1] word = tokens[last[0]] if (reads is not None and not word.tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS)): @@ -1074,7 +1042,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], f"a post-nominal; the maiden marker's clause keeps it " f"rather than reading it as one", tuple(last))) - return seen[m:m + run], seen[m + run:j] + return list(range(m, m + run)), list(range(m + run, j)) #: What `_between_name_words` reads instead of walking: two arrays over @@ -1088,17 +1056,9 @@ def _maiden_take(pieces: Sequence[Sequence[int]], def _run_neighbours(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], - tokens: Sequence[WorkToken], - cores: Set[str]) -> _Beside: - """The nearest piece that is neither a connective nor a delimiter - core, on each side of every index. - - A tail segment's delimiter CORE (`cores`, empty off a tail segment - and at the default policy) is stepped over exactly as a connective - is: it is structure the caller declared, the #206 drop takes a LONE - core out of the output, and a link read with it present must read as the - same text read without it (rules.md#M2, #538). Asked inline, beside - `_is_conj_piece`, so an empty set costs no frame. + tokens: Sequence[WorkToken]) -> _Beside: + """The nearest piece that is not a connective, on each side of + every index. EVERY MEMBER OF ONE RUN HAS THE SAME ANSWER, which is the whole of the fix: `_between_name_words` used to walk the run itself, so a @@ -1157,11 +1117,7 @@ def _run_neighbours(pieces: Sequence[Sequence[int]], prev = -1 for k in range(n): left[k] = prev - # the lone-core test; also in _maiden_take and group()'s #206 - # drop, whose copy adds `len(pieces) > 1` -- keep in step - is_conj = (_is_conj_piece(pieces[k], ptags[k], tokens) - or (len(pieces[k]) == 1 - and tokens[pieces[k][0]].text in cores)) + is_conj = _is_conj_piece(pieces[k], ptags[k], tokens) conj[k] = is_conj if not is_conj: prev = k @@ -1219,8 +1175,6 @@ def _is_rootname(piece: Sequence[int], ptags: Set[str], # is the CALLER's half: `_group_segment`'s `frozen` set is where a # connective this refuses is placed as the generation, and # `_link_joins_inside_the_clause` is the maiden walk's. -# rules.md#P3: "the search reads past it as it reads past a -# connective" (#538) -- `cores` in `beside`, below. # `Sequence[Sequence[int]]` rather than `Sequence[Piece]`, widened # when the maiden walk became a second caller: this reads a piece and # never edits one, and `_maiden_take` holds its pieces at the wider @@ -1259,8 +1213,7 @@ def _between_name_words(k: int, lo: int, hi: int, one ('Carod i y Rovira'), so the word this rule is about is the first one past the run -- and where the run runs out ('Juan i e') there is no name word on that side at all. `beside` is where that - stepping already happened, a lone delimiter core stepped over the - same way too (#538, rules.md#P3's separator sentence): + stepping already happened: `_run_neighbours` walked every run once for the whole segment, so this reads an index rather than walking to it. A SENTINEL OUT OF RANGE is how "the run ran out" arrives -- @@ -1308,7 +1261,6 @@ def _group_segment(seg: tuple[int, ...], additional: int, tokens: Sequence[WorkToken], bound_join: BoundJoin = BoundJoin.STRICT, ambiguities: list[PendingAmbiguity] | None = None, - cores: Set[str] = frozenset(), given_name_titles: Set[str] = frozenset(), opens_the_name: bool = False, *, @@ -1472,7 +1424,7 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # The tokens are not touched here: this function reads them and # returns what it took, and group() records the drop and the roles. taken: MaidenTake | None = None - take = _maiden_take(pieces, ptags, tokens, cores, one_case, + take = _maiden_take(pieces, ptags, tokens, one_case, reader, maiden_ambiguities, tail_follows=tail_follows) if take is not None: @@ -1573,11 +1525,8 @@ def merge(lo: int, hi: int, add: Set[str] = frozenset(), # of connectives has the same nearest name word on # each side, and asking per member walked the run once # per member -- quadratic in its length, 3.8x per - # doubling measured at `b9ed1429`. `cores` is passed - # deliberately here too: a link beside a declared - # delimiter reads as it would with the delimiter absent - # (#538). - beside = _run_neighbours(pieces, ptags, tokens, cores) + # doubling measured at `b9ed1429`. + beside = _run_neighbours(pieces, ptags, tokens) if not _between_name_words(k, lo, hi, pieces, ptags, tokens, beside): frozen.add(piece[0]) @@ -2040,9 +1989,8 @@ def group(state: ParseState) -> ParseState: # parts; the SUFFIX_COMMA pre-comma segment gets 0. additional = 1 if state.structure is Structure.FAMILY_COMMA else 0 # v1 expand_suffix_delimiter parity (#206): tail segments (wholly - # consumed as suffixes by assign) drop delimiter-core tokens, the - # same structural mechanism as the maiden marker (taken out in - # _group_segment, recorded in `dropped` just below) + # consumed as suffixes by assign) are cut at their delimiter-core + # tokens, which land in `dropped` as the maiden marker does cores = delimiter_cores(state.policy.extra_suffix_delimiters) tail_start = {Structure.SUFFIX_COMMA: 1, Structure.FAMILY_COMMA: 2}.get(state.structure) @@ -2065,7 +2013,6 @@ def group(state: ParseState) -> ParseState: # the parse consumes wholly as suffixes raises no report about # reading a word of it as a name" tail = tail_start is not None and seg_idx >= tail_start - seg_cores = cores if tail else frozenset() # #533: which rule reads what the maiden walk would leave, off # the three facts already in hand here. A tail segment is read # as credentials whole and segment 0 of a family comma is the @@ -2077,45 +2024,52 @@ def group(state: ParseState) -> ParseState: else TailReader.NONE) else: reader = TailReader.TRAILING - pieces, ptags, taken = _group_segment( - seg, additional, tokens, bound_join, - None if (family_comma or tail) else ambiguities, - seg_cores, - state.lexicon.given_name_titles, - opens_the_name=(seg_idx == 0 and not family_comma), - one_case=state.one_case, - reader=reader, - maiden_ambiguities=ambiguities, - tail_follows=(reader is TailReader.GIVEN_SLOT - and len(state.segments) > 2)) - # the marker is dropped and the maiden name's tokens become - # MAIDEN (#274); which pieces those are was settled in - # _group_segment, before the joins - if taken is not None: - marker_piece, maiden_pieces = taken - dropped.extend(marker_piece) - for piece in maiden_pieces: - for i in piece: - tokens[i] = copy_with( - tokens[i], role=Role.MAIDEN) - # rules.md#C1: "a part that is nothing but suffix words is the - # credential run and reads as suffixes, whole" -- WHOLE is this - # block's half of the rule, the routing being assign's. - # - # v1 expand_suffix_delimiter parity (#206): a delimiter core - # inside a segment separates suffix entries and is dropped, but - # a segment that IS only the core stays whole (v1 expand() - # splits within a part, never erases a lone part). Keyed on - # `tail` through `seg_cores`, which is empty off a tail - # segment, because the #206 parity is a TAIL rule. - # - # What this block decides is the #206 core DROP and nothing - # else: a delimiter core inside a tail segment leaves the - # pieces, and `dropped` is where that fact is recorded -- for - # the render, and for post_rules' entry pass, which reads it - # back as the one dropped token that separates two entries. - # - # It used to decide the ENTRY too, marking a continuation + # rules.md#C1: "a delimiter the policy declares parts a + # trailing suffix part as a comma would" (#549): the segment is + # grouped as the parts between its cores, each read exactly as + # the comma twin's own segment would be, and the cores are + # dropped (v1 expand_suffix_delimiter parity, #206), where + # post_rules' entry pass reads them back as entry boundaries. + # No join, maiden walk or link search reaches across a core, + # because none of them ever sees one. A segment that IS only + # its core keeps it (v1 expand() splits within a part, never + # erases a lone part). + parts: list[tuple[int, ...]] = [seg] + if tail and cores and len(seg) > 1: + parts = [()] + for i in seg: + if tokens[i].text in cores: + dropped.append(i) + parts.append(()) + else: + parts[-1] += (i,) + parts = [p for p in parts if p] + pieces = [] + ptags = [] + for part in parts: + part_pieces, part_ptags, taken = _group_segment( + part, additional, tokens, bound_join, + None if (family_comma or tail) else ambiguities, + state.lexicon.given_name_titles, + opens_the_name=(seg_idx == 0 and not family_comma), + one_case=state.one_case, + reader=reader, + maiden_ambiguities=ambiguities, + tail_follows=(reader is TailReader.GIVEN_SLOT + and len(state.segments) > 2)) + pieces.extend(part_pieces) + ptags.extend(part_ptags) + # the marker is dropped and the maiden name's tokens become + # MAIDEN (#274); which pieces those are was settled in + # _group_segment, before the joins + if taken is not None: + marker_piece, maiden_pieces = taken + dropped.extend(marker_piece) + for piece in maiden_pieces: + for i in piece: + tokens[i] = copy_with( + tokens[i], role=Role.MAIDEN) + # Group used to decide the suffix ENTRY, marking a continuation # token "joined" between pieces off segment SHAPE -- `tail` by # index, ORed since #429 with segment_suffix_reading's # per-piece content verdict, gated per piece and sticky across @@ -2127,23 +2081,6 @@ def group(state: ParseState) -> ParseState: # title between them renders elsewhere, so they join, while # 'Smith Jr., Mr. Jr.' has the writer's own comma and they do # not. - if seg_cores: - kept: list[int] = [] - for k in range(len(pieces)): - # the lone-core test; also in _maiden_take and - # _run_neighbours -- keep in step. This copy alone adds - # `len(pieces) > 1`: a segment that is nothing but its - # core keeps it - is_core = (len(pieces[k]) == 1 - and tokens[pieces[k][0]].text in seg_cores - and len(pieces) > 1) - if is_core: - dropped.extend(pieces[k]) - continue - kept.append(k) - if len(kept) != len(pieces): - pieces = [pieces[k] for k in kept] - ptags = [ptags[k] for k in kept] # continuation tokens of a suffix-merged piece (the ph-d merge) # carry the stable "joined" tag: the suffix string view joins # SUFFIX tokens with ", ", and the tag lets it heal the split diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index 91bc5631..24b0634f 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -624,24 +624,20 @@ class Peel(NamedTuple): # and does not. A BARE ambiguous acronym is consumed only when the name # has words to spare" # (v1's are_suffixes tail rule, with the roman-numeral special) -def peel_walk(start: int, ptags: Sequence[Set[str]], - skip: Set[int] = frozenset()) -> list[int]: +def peel_walk(start: int, ptags: Sequence[Set[str]]) -> list[int]: """The indices peel_trailing walks: `start` to the segment's end, minus the group-flagged credential pieces (the Ph. D. merge), - which assign reads as suffixes at any position, and minus `skip` - -- a tail segment's delimiter cores, which are structure rather - than words (the maiden walk's case, #424). Built here and nowhere + which assign reads as suffixes at any position. Built here and nowhere else, so the walk's input cannot drift between assign and the group sites that read it: the numeral fork is a last-piece test that reads the piece before as rest[k - 2], which holds only over this list.""" return [j for j in range(start, len(ptags)) - if j not in skip and "suffix" not in ptags[j]] + if "suffix" not in ptags[j]] def trailing_start(start: int, pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], - skip: Set[int] = frozenset(), *, one_case: bool | None) -> int: """Where assign's trailing suffix run begins, read over the pieces as they stand from `start`: the index of the first piece the S2 @@ -664,7 +660,7 @@ def trailing_start(start: int, pieces: Sequence[Sequence[int]], pair's peel and H5's title chain to a fixed point. So this function's answer is the whole answer only where no trailing title chain also reads the pieces.""" - rest = peel_walk(start, ptags, skip) + rest = peel_walk(start, ptags) peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) return rest[peeled.names] if peeled.names < len(rest) else len(pieces) diff --git a/nameparser/_pipeline/_post_rules.py b/nameparser/_pipeline/_post_rules.py index fa71a84f..a1334d75 100644 --- a/nameparser/_pipeline/_post_rules.py +++ b/nameparser/_pipeline/_post_rules.py @@ -122,7 +122,7 @@ def _mark_suffix_entries(tokens: list[WorkToken], state: ParseState) -> None: # Policy.extra_suffix_delimiters, imported rather than repeated # (mechanisms.md#ONE-PREDICATE-PER-QUESTION). The two sites read # ONE derivation and differ only in a gate: group drops a core - # only on a `tail` segment, through `seg_cores`, while this arm + # only on a `tail` segment, cutting the segment there, while this arm # reads `delimiter_cores` whole and asks by TEXT alone. So a # dropped token whose text the policy names as a delimiter parts # the run whatever dropped it. The gate is not needed here: a core diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 27ffacf2..5b1cbcec 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -6170,6 +6170,17 @@ def _check_cjk_shape_purity(self) -> None: "The core is gone before the render sees it, so the " "boundary is read off the dropped indices rather than " "off a surviving token. Parity at every baseline"), + Case("suffix_delimiter_core_parts_a_connective_join", + "Doe, John, PhD - and MD", + {"given": "John", "family": "Doe", "suffix": "PhD, and MD"}, + policy=_SD, + ambiguities=("comma-structure",), + notes="a declared core parts a tail as a comma would, so a " + "connective beside it joins nothing across it and the " + "core is dropped (rules.md C1, #549) -- 1.4.0's " + "expand_suffix_delimiter reading, where 2.0.0 through " + "2.3.0 merged the core into the joined piece and " + "rendered 'PhD - and MD'"), Case("suffix_delimiter_core_that_survives_is_a_boundary_too", "Smith, MD - PhD - FACS", {"title": "MD", "given": "-", "middle": "-", "family": "Smith", diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index ccfcfdc6..81f6d074 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -485,23 +485,24 @@ def test_the_connective_carveout_counts_the_surviving_name() -> None: def test_a_delimiter_core_in_a_suffix_tail_is_not_maiden_text() -> None: - """A tail segment drops its delimiter cores (#206) and the marker - takes what is left, in that order -- the order group() had before - the marker pass moved ahead of the joins. A core is not a word the - marker can take: it is structure, like the marker itself.""" + """A tail segment is cut at its delimiter cores before the marker + pass, and the cores are dropped (#206, #549). A core is not a word + the marker can take, and the marker cannot take across one: 'PhD + née - Jones' reads as 'PhD née, Jones' does, the marker ending its + part with nothing behind it, so there is no clause.""" out = _grouped("Smith, John, PhD née - Jones", policy=_DASH) core = next(i for i, t in enumerate(out.tokens) if t.text == "-") assert core in out.dropped - assert out.tokens[core].role is not Role.MAIDEN - assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == [ - "Jones"] + assert not any(t.role is Role.MAIDEN for t in out.tokens) + twin = _grouped("Smith, John, PhD née, Jones") + assert not any(t.role is Role.MAIDEN for t in twin.tokens) def test_the_walk_peels_past_a_trailing_core() -> None: - # The walk's peel skips the cores a tail segment drops, so a core - # standing last does not make the numeral "not last": the V is the - # suffix and the marker takes 'Jones' alone (#424; the test - # review's surviving mutant). + # A core standing last is cut off before the walk (#549), so it + # does not make the numeral "not last": the V is the suffix and the + # marker takes 'Jones' alone (#424; the test review's surviving + # mutant). out = _grouped("Smith, John, PhD née Jones V -", policy=_DASH) assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == [ "Jones"] @@ -510,8 +511,8 @@ def test_the_walk_peels_past_a_trailing_core() -> None: def test_a_core_is_screened_before_the_marker_looks_for_a_word_ahead() -> None: """M2 needs a name word BEFORE the marker. A core is not one, so a marker standing behind nothing but a core is a leading marker and - is not taken -- and the core, no longer the segment's only piece - once the tail is read as written, is dropped as usual.""" + is not taken -- the core, which is not the segment's only token, + parts the segment there and is dropped (#549).""" out = _grouped("Smith, John, - née Jones", policy=_DASH) assert not any(t.role is Role.MAIDEN for t in out.tokens) core = next(i for i, t in enumerate(out.tokens) if t.text == "-") @@ -1805,14 +1806,10 @@ def test_the_marker_is_not_the_name_word_on_the_links_left() -> None: def test_a_core_between_the_marker_and_the_first_word_is_below_lo( ) -> None: - # A delimiter core is TAIL-segment structure that group() drops - # after this pass, so it is never a word of the clause -- and - # between the marker and the first word it is below `lo`, which - # the bound refuses without the piece tests ever being asked. - # Reachable, not theoretical: measured 2026-09-21 over corpus u - # cases.py u the property grids u a 50,925-name generated set with - # cores, under thirteen core-bearing policies, 25,536 of 596,392 - # maiden takes had a core standing there. + # A delimiter core is TAIL-segment structure that group() cuts the + # segment at before this pass (#549), so it is never a word of the + # clause, and a marker with a core straight behind it ends its part + # with nothing to take, as 'PhD née, i Jones' does. out = _grouped("Smith, John, PhD née - i Jones", policy=_DASH, lexicon=_LINK_LEX) assert _maiden_texts(out) == [] @@ -1823,62 +1820,56 @@ def test_a_core_between_the_marker_and_the_first_word_is_below_lo( assert _maiden_texts(plain) == ["-", "i", "Jones"] -def test_a_core_beside_a_link_is_stepped_over_like_a_connective( -) -> None: - """rules.md#M2 gives the link exception a name word on each side, - and a delimiter core is structure the caller declared rather than - a name word -- so the word on a link's side is the one past the - core, and the clause reads as the same text written without it - (#538). - - Past the clause's first word the core used to be an ordinary index - to the neighbour walk, which stepped over connectives and nothing - else, so it stood in for the name word on the link's left and the - clause ran on past the link it otherwise ends at. The population is - decisions.md's 2026-09-22 #397 follow-up entry: none of it - reachable at the default policy, `extra_suffix_delimiters` being - empty there. +def test_a_core_ends_a_maiden_clause_as_a_comma_would() -> None: + """rules.md's C1: a declared delimiter in a trailing part "parts it as + a comma would" (#549), so a maiden clause ends at a core exactly as + it ends at the comma written in its place, and a link beside the + core has no name word on that side. + + This reverses #538, which read the clause as the text written + WITHOUT the core: 'Puig - i Soler' kept maiden 'Puig i Soler' + there. RECORDED NEGATIVE CONTROL: at 061f02da the third assertion + read ["Puig", "i", "Soler"]. """ out = _grouped("Smith, John, PhD née Puig Mr. - i Soler", policy=_DASH, lexicon=_LINK_LEX) assert _maiden_texts(out) == ["Puig", "Mr."] - # the same clause with the core taken out of it: the title IS the - # word on the link's left and refuses, so the clause ends there. - without = _grouped("Smith, John, PhD née Puig Mr. i Soler", - policy=_DASH, lexicon=_LINK_LEX) - assert _maiden_texts(without) == ["Puig", "Mr."] - # and between two NAME words the core is stepped over too, so the - # link joins exactly as it joins written without the core -- the - # skip reading, not a boundary one, which would have ended the - # clause at 'Puig' and pushed the link into the credentials. + twin = _grouped("Smith, John, PhD née Puig Mr., i Soler", + lexicon=_LINK_LEX) + assert _maiden_texts(twin) == ["Puig", "Mr."] + # between two NAME words the core ends the clause too, as the + # comma does: the link and the word behind it are the next part's. between = _grouped("Smith, John, PhD née Puig - i Soler", policy=_DASH, lexicon=_LINK_LEX) - assert _maiden_texts(between) == ["Puig", "i", "Soler"] - - -def test_a_core_beside_a_link_in_a_credential_tail_is_dropped() -> None: - """The same stepping applies outside a maiden clause, in the - `frozen` loop's own `_run_neighbours` call. In BOTH texts below, - what stands beyond the core is a credential or nothing -- never a - name word -- so the link's neighbour search, stepping past the - core, finds no name word there either and stays a lone suffix - word rather than joining. The core is then a lone piece with - nothing joined to it, which the #206 drop removes exactly as it - removes any lone core, and the entries it stood between separate - the way they already do in 'PhD - MD' -> 'PhD, MD' (#538). Where a - name word stands beyond the core instead, the link joins across it - and the core survives in the suffix text -- not this test's shape; - rules.md#P3's separator sentence states the join, decisions.md's - #538 entry the surviving core. - - RECORDED NEGATIVE CONTROL: at e0f1a2fa, before the frozen loop - stepped over a core, these read suffix 'PhD - i Soler' and - '- i Puig' -- the core kept inside the piece the link joined. + assert _maiden_texts(between) == ["Puig"] + twin = _grouped("Smith, John, PhD née Puig, i Soler", + lexicon=_LINK_LEX) + assert _maiden_texts(twin) == ["Puig"] + + +def test_a_core_beside_a_connective_in_a_credential_tail_is_dropped( +) -> None: + """No connective joins across a core, generational or ordinary, and + whatever stands beyond it: the core parts the tail as a comma would + and is dropped, so the entries it stood between separate the way + they do in 'PhD - MD' -> 'PhD, MD' (#206, #549). + + RECORDED NEGATIVE CONTROLS: at e0f1a2fa, before #538, the first two + read suffix 'PhD - i Soler' and '- i Puig'; at 061f02da, after #538 + and before #549, the last three read 'Puig - i Soler', 'Puig i - + Soler' and 'PhD - and MD' -- the core kept inside the piece a + connective joined across it. """ dash = Parser(policy=Policy(extra_suffix_delimiters=frozenset({" - "}))) assert str(dash.parse("Smith, John, PhD - i Soler").suffix) == \ "PhD, i Soler" assert str(dash.parse("Smith, John, - i Puig").suffix) == "i Puig" + assert str(dash.parse("Smith, John, Puig - i Soler").suffix) == \ + "Puig, i Soler" + assert str(dash.parse("Smith, John, Puig i - Soler").suffix) == \ + "Puig i, Soler" + assert str(dash.parse("Smith, John, PhD - and MD").suffix) == \ + "PhD, and MD" def test_a_marker_with_nothing_after_it_declines_before_the_bound( diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 3884212d..fce4c74e 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3863,7 +3863,7 @@ def _claim(rule: dict) -> _Claim: # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified # name by name. - _Claim(439, ('given', 'suffix', 'title'), "5a1100d3a0ac", None), + _Claim(441, ('given', 'suffix', 'title'), "9f312397b07e", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3974,7 +3974,7 @@ def _claim(rule: dict) -> _Claim: # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified # name by name. - _Claim(439, ('family', 'given'), "5a1100d3a0ac", None), + _Claim(441, ('family', 'given'), "9f312397b07e", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": @@ -4246,7 +4246,7 @@ def _claim(rule: dict) -> _Claim: # carries a shape tag. Reach again, verified name by name. # 2026-10-01, #575: 104 -> 105, 'Ortega y Gasset, Ed'. Reach. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": - _Claim(105, ('_initials',), "231bbda6768a", ('DEFAULT',)), + _Claim(106, ('_initials',), "8b55345f9ad4", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index da94becb..a3ef45c2 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -3218,44 +3218,52 @@ def test_the_decomposed_initials_walk_can_fail( for line in failures), failures -def test_a_delimiter_core_reads_as_if_it_were_not_written() -> None: - """#538: under a core-bearing policy, a maiden clause reads exactly - as the same text written without the core. The core is structure - the caller declared, the #206 drop takes a LONE core out of the - output, and - rules.md#M2's link exception asks for a NAME word on each side of - the link -- so the word on a link's side is the one past the core. - - Asked of the `maiden` field only, because that is the field #538 - is about: a core that lands inside a connective run the join - merges ('Puig Dr. i - y Soler') stays in the SUFFIX text under this - policy and the default alike, which is a separate gap. - - RECORDED NEGATIVE CONTROL: at e0f1a2fa, before the core was stepped - over, 6 of these 81 texts disagreed ('Puig Mr. - i Soler' read - maiden 'Puig Mr. i Soler' against 'Puig Mr.', and the same for - 'Puig Dr. - i y Soler', under each of the three heads). - """ - dash = Parser(policy=Policy(extra_suffix_delimiters=frozenset({" - "}))) - heads = ("Smith, John, PhD née", "Smith, John, MD née", - "Doe, Jane, PhD nee") - bodies = ("Puig Mr. i Soler", "Puig i Soler", "Puig Dr. i y Soler", - "Puig i i Soler", "Carod i Rovira Mr.", "Puig Mr. i Dr. Soler", - "Jones Smith i Soler", "Puig y Soler", "Puig Jr. i Soler") +_TWIN_HEADS = ("Smith, John,", "Smith, John, Jr.,", "Doe, Jane, PhD") +_TWIN_BODIES = ( + "PhD née Puig Mr. i Soler", "PhD née Puig i Soler", + "PhD née Puig Dr. i y Soler", "MD née Carod i Rovira Mr.", + "PhD née Jones Smith i Soler", "PhD née Puig y Soler", + "Puig i Soler", "Puig Dr. i y Soler", "PhD and MD", "PhD i MD", + "PhD MD FACS", "Puig y Soler") + + +def _comma_twin_findings(parser: Parser) -> tuple[list[str], int]: + """Each body with a ' - ' core in one interior gap, against the same + text with a typed comma there, all seven fields.""" failures = [] total = 0 - for head in heads: - for body in bodies: + for head in _TWIN_HEADS: + for body in _TWIN_BODIES: parts = body.split() - # a core in every gap PAST the first word: one between the - # marker and that word is below the clause's bound already - # (test_a_core_between_the_marker_and_the_first_word_is_below_lo) for gap in range(1, len(parts)): - written = " ".join(parts[:gap] + ["-"] + parts[gap:]) + left, right = " ".join(parts[:gap]), " ".join(parts[gap:]) total += 1 - got = str(dash.parse(f"{head} {written}").maiden) - want = str(dash.parse(f"{head} {body}").maiden) + got = parser.parse(f"{head} {left} - {right}").as_dict() + want = parser.parse(f"{head} {left}, {right}").as_dict() if got != want: - failures.append(f"{head} {written!r}: {got!r} != {want!r}") - assert total == 81 + failures.append(f"{head} {left} - {right}: " + f"{got} != {want}") + return failures, total + + +def test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin() -> None: + """rules.md's C1: "a delimiter the policy declares parts a trailing + suffix part as a comma would" (#549). Every field of a name with a + declared core standing in a part after the suffix comma equals the + field of the same text written with a comma in the core's place -- + the maiden clause, the connective join and the entry boundary + alike. The heads put the core in a TRAILING part only: before the + first comma, or in the given part of the listing form, the core is + a word (C1's Accepted entry), and its comma twin a different + structure. + + RECORDED NEGATIVE CONTROL: at 061f02da, where the join and the + maiden walk stepped over a core instead (#538), 102 of these 129 + texts disagreed -- 'Puig - i Soler' read suffix 'Puig - i Soler' + against 'Puig, i Soler', and 'PhD née Puig - i Soler' maiden + 'Puig i Soler' against 'Puig'. + """ + dash = Parser(policy=Policy(extra_suffix_delimiters=frozenset({" - "}))) + failures, total = _comma_twin_findings(dash) + assert total == 129 assert not failures, "\n".join(failures) diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 12184f4b..0c78811a 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -358,9 +358,11 @@ "Smith, John Prof." "Smith, John V" "Smith, John V." +"Smith, John, PhD - and MD" "Smith, John, PhD - i Soler" "Smith, John, PhD née Puig - i Soler" "Smith, John, PhD née Puig Mr. - i Soler" +"Smith, John, Puig - i Soler" "Smith, John, and" "Smith, Jr." "Smith, MA" From 9e784fe728db7854b140400b00a64c0425226b5f Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 17:25:18 -0700 Subject: [PATCH 2/4] Fix what the review of #549 found - group's tail split accumulated a tuple per token, quadratic in a part's length; build a list and freeze it once per part (4.7x for 4x the input, against 8.0x). - rules.md#M2's statement still described #538's stepping; it now defers to C1. P3's separator sentence is scoped to trailing parts, and both rules list C1 in interacts:. C1 states the lone-core exception and records, as Accepted, that a maiden marker beside a declared delimiter takes no clause -- the comma twin's reading, and the larger cost by count. - decisions.md#C1's #549 entry: the maiden-loss class in WHAT MOVES and in the weighing; the Jr., head figure (752 of 2,100, not 973 of 3,310); the ' / ' policy moving the same 2,624; the twin compared on role fields, comma-structure counts differing; the #418 repair added to SUPERSEDES. - release_log: the maiden-loss consequence, and the scope stated as the part read wholly as suffixes. - Stale prose in test_post_rules and a renamed test_group test. - The C1 example enters the rules corpus and the 1.4.0 ledger. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 12 ++++----- docs/design/rules.md | 28 +++++++++++++------- docs/release_log.rst | 2 +- nameparser/_pipeline/_group.py | 25 ++++++++++++----- tests/v2/pipeline/test_group.py | 2 +- tests/v2/pipeline/test_post_rules.py | 4 +-- tests/v2/test_ledger_guards.py | 25 +++++++++++++---- tools/differential/corpus_rules.jsonl | 1 + tools/differential/expected_since_1.4.0.toml | 8 +++++- 9 files changed, 76 insertions(+), 31 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 52464407..c2ae97d8 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -228,7 +228,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q - 2026-09-22 #397 follow-up — A DELIMITER CORE PAST THE CLAUSE'S FIRST WORD PASSES FOR THE NAME WORD BESIDE A LINK, RECORDED AS A DEVIATION RATHER THAN REPAIRED. The link exception above wants "a connective standing between two name words of the clause", and a separator the caller declared through `Policy.extra_suffix_delimiters` is structure, not a name word — so a link with one beside it joins nothing and should end the clause like any other suffix word. Between the marker and the clause's first word that already holds, the bound refusing the core before either piece test is asked. PAST that first word it does not: the core is an ordinary index to the run walk, which steps over connectives and nothing else, so it stands in for the name word on the link's left and the clause runs on past a title it would otherwise stop at. MEASURED 2026-09-22 under `extra_suffix_delimiters=(" - ",)`: `Smith, John, PhD née Puig Mr. - i Soler` reads maiden 'Puig Mr. i Soler' where its separator-less twin `Smith, John, PhD née Puig Mr. i Soler` stops at 'Puig Mr.'. THE POPULATION is the branch's own sweep, recorded with the code it describes (2026-09-21, corpus ∪ cases.py ∪ the property grids ∪ a 50,925-name generated set with cores, under thirteen core-bearing policies): the predicate is asked about a core in 51,072 of 900,023 calls, the answer differs from a core-skipping reading in 8,094 parses over 1,278 texts, and 1,824 of those move `maiden` on 288 texts — none of the 288 reachable at the default policy, `extra_suffix_delimiters` being empty there. NOT REPAIRED HERE: the fix threads the core set through three call sites into the run walk and moves the parent's reading as well, which makes it its own change rather than a rider on a review round. Open: #538. PINNED TWICE MEANWHILE. rules.md#M2 carries the shape as a `deviates: #538` example under an `extra_suffix_delimiters-dash` annotation — the first entry `tests/v2/rules_doc.py`'s registry has had for that field, named after the Policy field and carrying the delimiter in the suffix because the field's value is a set rather than a flag. And `tests/v2/pipeline/test_group.py::test_a_core_beside_a_link_wrongly_passes_for_a_word_until_538` holds the pair at the piece level, named so nobody reads it as the contract. ONE COST OF THE DOC EXAMPLE, worth knowing before the repair lands: its string enters `corpus_rules.jsonl`, where the differential gate parses it with the DEFAULT facade — no delimiter declared, so the dash is an ordinary name word and #538's reading is off the path entirely. It moves there for the 2026-09-20 link fix instead, and is classified as that at all five baselines: added to the `fix(#397) a link inside a maiden clause stays in the birth name` alternation at the four 2.x ledgers (suffix 'PhD i Soler' → 'PhD', maiden 'Puig Mr. -' → 'Puig Mr. - i Soler', identical at each), and given its own rule at 1.4.0, where v1 had no maiden markers and read the whole suffix-comma tail as one suffix. - 2026-09-26 #538 — A DELIMITER CORE IS STEPPED OVER BY THE LINK'S NEIGHBOUR SEARCH, AS A CONNECTIVE IS (Derek, 2026-09-26). The 2026-09-22 follow-up above recorded the deviation; this resolves it. `_run_neighbours` treats a lone core like a connective, so the word on a link's side is the one past the core and a maiden clause reads exactly as the same text written without the core: `Smith, John, PhD née Puig Mr. - i Soler` under `extra_suffix_delimiters=(" - ",)` now stops at maiden 'Puig Mr.' as its separator-less twin does. INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_reads_as_if_it_were_not_written` over 81 generated texts, maiden field: 6 disagreed at e0f1a2fa, 0 after. Declined: THE BOUNDARY READING, which the 2026-09-22 entry's wording implied ("a link with one beside it joins nothing") — a core ending the neighbour search on its side. It fixes the reported shape equally, and it moves `Smith, John, PhD née Puig - i Soler` and `… Puig i - Soler` from maiden 'Puig i Soler' to maiden 'Puig', suffix 'PhD i Soler' — a birth-name link pushed into the credentials, where the skip reading leaves both as they read. Default-unreachable either way. THE SAME STEPPING APPLIES OUTSIDE A CLAUSE: the `frozen` loop's own `_run_neighbours` call also takes `cores`, so a connective beside a core looks past it; where a credential or nothing stands beyond, it stays a lone suffix word rather than joining, and the core it stood beside -- now a lone piece with nothing joined to it -- is dropped as #206 drops any lone core, the same way it already drops one between two ordinary post-nominals: `Smith, John, PhD - i Soler` 'PhD - i Soler' -> 'PhD, i Soler', `Smith, John, PhD - i - MD` 'PhD - i - MD' -> 'PhD, i, MD', `Smith, John, - i Puig` '- i Puig' -> 'i Puig', `Smith, John, Puig i -` 'Puig i -' -> 'Puig i'. MEASURED 2026-09-26: 191 of 8,097 generated core-bearing tail texts (the three heads `Smith, John,`, `Smith, John, Jr.,` and `John Smith,`, each followed by every 2-4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr.} containing a core) move `suffix` and no other field, none reachable at the default policy; the comparison is each text parsed under `extra_suffix_delimiters=(" - ",)` with `_run_neighbours` handed its cores and with it handed none. Not repaired here: a core the join MERGES into a joined piece -- interior (`… Puig Dr. i - y Soler`, still 'PhD i - y Soler') or at the edge a link joins across, a name word standing beyond it rather than a credential or nothing (`Smith, John, Puig - i Soler` 'Puig - i Soler', its separator-less twin 'Puig i Soler'; `Smith, John, Puig i - Soler` 'Puig i - Soler', measured 2026-09-26) -- is not dropped from the suffix text; the join is right, the surviving separator is the gap. Also in that family: an ORDINARY connective's join still takes a core as its neighbour, the stepping above being for generational vocabulary alone (rules.md#P3) -- `Smith, John, PhD - and MD` keeps suffix 'PhD - and MD' (measured 2026-09-26), pre-existing and unchanged here. Open: #549, which carries both the surviving core and this one. -- 2026-10-03 #549 — SUPERSEDED: the skip reading the #538 entry above settles is reversed, and a declared delimiter in a trailing part now parts it as a comma would, so `Smith, John, PhD née Puig - i Soler` reads maiden 'Puig' with suffix 'PhD, i Soler', the reading #538 declined and the one 2.3.0 gave. The decision, its measurement and the invariant that replaces #538's are under C1's 2026-10-03 entry. +- 2026-10-03 #549 — SUPERSEDED: the skip reading the #538 entry above settles is reversed, and a declared delimiter in a trailing part now parts it as a comma would, so `Smith, John, PhD née Puig - i Soler` reads maiden 'Puig' with suffix 'PhD, i Soler', the reading #538 declined and the one 2.3.0 gave, and a marker with a core straight before or after it takes no clause (`Smith, John, PhD née - Jones` reads suffix 'PhD née, Jones', reversing the #418 entry's repair). The decision, its measurement and the invariant that replaces #538's are under C1's 2026-10-03 entry. - 2026-09-26 #535 — THE MAIDEN WALK READS THE TRAILING TITLE CHAIN, AND EVERY STOP ASKS ONE RELEASE QUESTION OF THE SPAN IT GIVES UP (Derek, 2026-09-26). This resolves the NOT TRANSPARENT HERE clause of the 2026-09-19 #533 entry above, which recorded `Jane Doe nee Smith MA Prof.` and `Jane Doe nee Smith Prof. MA` disagreeing and deferred the pair as one decision about what "trailing" means inside a clause. BOTH HALVES WERE TAKEN TOGETHER: a trailing title ends the clause (`Jane Doe nee Smith Prof.` reads title 'Prof.', maiden 'Smith', where 2.3.0 read maiden 'Smith Prof.'), and a credential or numeral standing in front of that title gets the stop it gets with the title absent (`… Smith MA Prof.` gives suffix 'MA', `… Smith V Prof.` suffix 'V'). Only the first half would have left the pair disagreeing in the other direction; H5's transparency is a statement about the peel and the chain read together to their fixed point, so the walk now reads the end of the clause the way assign reads the end of the name — `tail_reading`, the shared predicate mechanisms.md#ONE-PREDICATE-PER-QUESTION names — and over the forms the transparency property test below exercises, the two spellings give one answer. They split where something in or ahead of the clause stands to take the title — for example a particle chain (ahead of the clause or inside it), a bound-given join, a name left with no name word, or, after a family comma, a title word in front of a class member with a title behind it, which the given part's chain stops short of exactly as it does with no marker (`Doe, Jane nee Smith Rev. MA Prof.` reads maiden 'Smith Rev.', title 'Prof.', suffix 'MA', where `… Rev. Prof. MA` reads title 'Rev. Prof.', maiden 'Smith', suffix 'MA'; bare `Doe, Jane Rev. MA Prof.` reads middle 'Rev.') — and the Accepted pairs recorded here and in rules.md#M2 are examples of that, not a complete list. `tests/v2/test_properties.py::test_a_trailing_title_is_transparent_to_the_maiden_clause` holds that as an invariant over the two inputs, over heads and runs where the credential is given up and nothing ahead would take the title, with its negative control at e0f1a2fa in its docstring. ACCEPTED, THE TRANSPARENCY BOUNDARY: where the clause KEEPS the credential the spellings differ, because a clause is one contiguous run and a title inside the kept text cannot leave without the words behind it — `Doe nee Smith ba Prof.` reads title 'Prof.', maiden 'Smith ba', and `Doe nee Smith Prof. ba` maiden 'Smith Prof. ba'; `abdul nee Smith MA Prof.` reads title 'Prof.', family 'abdul', maiden 'Smith MA', and `abdul nee Smith Prof. MA` given 'abdul', maiden 'Smith Prof. MA' (2.3.0 kept every word in all four). The report differs with them: `Doe, Dr. nee Smith MA` and `… Prof. MA` report `suffix-or-name` on the kept MA, `… MA Prof.` reports nothing, the kept member no longer ENDING the clause, which is what M2's report is asked of. rules.md#M2 carries the first pair as an Accepted example. ACCEPTED, FURTHER SPLITS, measured 2026-09-26. Where something ahead of the clause would take the title, the spellings differ even with the credential given up in both, because the first-suffix-word stop ends the clause before the title check runs and the title check alone declines: `Jane van der Berg nee Smith PhD Prof.` reads title 'Prof.', maiden 'Smith', suffix 'PhD', and `… Prof. PhD` maiden 'Smith Prof.', suffix 'PhD' — H5's accepted particle-chain boundary inherited, the bare `Jane van der Berg Smith PhD Prof.` / `… Prof. PhD` splitting the same way — and `Berg, abdul nee Smith PhD Prof.` (bound-given join) and `Doe, Dr. nee Smith PhD Prof.` (no name word left) split likewise; the parent d9d80492 and 2.3.0 read all of these as the tree does, so only the claim is new. And a released particle with a title behind it is withdrawn: `Jane Doe nee Smith DO Prof.` reads title 'Prof.', maiden 'Smith DO', where `… Prof. DO` reads title 'Prof.', maiden 'Smith', suffix 'DO' (2.3.0 maiden 'Smith DO Prof.' and 'Smith Prof. DO'); `do` likewise. A particle INSIDE the clause does it from the other side, its chain able to take the credential behind it: `Jane Doe nee Smith do MA Prof.` reads title 'Prof.', maiden 'Smith do MA', and `… do Prof. MA` title 'Prof.', maiden 'Smith do', suffix 'MA' (2.3.0 kept every word in both). rules.md#M2 carries a pair of each shape as Accepted examples. ACCEPTED, H5'S REACH: `Jane Doe nee Smith King.` reads title 'King.', maiden 'Smith' (2.3.0 maiden 'Smith King.'), H5's accepted `Mary Jane King.` cost now reaching the last word of a birth name as it reaches the last word of a current one; rules.md#M2 carries it as an Accepted example. READER SCOPE IS UNCHANGED: the chain is consulted only where a trailing rule reads the clause's words (no comma, the part before a suffix comma, the given part after a family comma); before a family comma and in a tail segment the walk reads the peel alone and the clause keeps the title (`Doe nee Smith Prof., Jane` keeps maiden 'Smith Prof.'). THE FIRST-WORD FLOOR MOVED INTO THE CHAIN. `trailing_titles` and `tail_reading` take a `floor` (1 everywhere else), and the walk passes the position just past the marker's first word, so the chain never takes that word: `Jane Doe nee King.` keeps its maiden name (TITLES holds borne surnames and the marker announced one), and `Jane Doe nee Prof. Dr.` keeps 'Prof.' and gives up 'Dr.'. The first draft applied the floor as a clamp on the title stop after the chain had run, and that is measurably wrong rather than merely inelegant: a chain allowed to take the first word has already spliced it out of the count the re-peel reads, so `Jane Doe nee King. ba` kept 'ba' in the clause where `Jane Doe nee Smith ba` gives it up. `test_a_title_first_word_counts_as_a_word` is the invariant (a title-vocabulary first word against an ordinary one, the credential behind it read alike) and its docstring records the clamp's failure count. ONE RELEASE CHECK FOR THREE STOPS. The numeral, the credential and the title stop each ask `_release_reads_off` whether the name the take would leave reads what they give up as titles or suffixes with no join below the take absorbing it — rules.md#M2's invariant, "a word the clause gives up reads as a post-nominal or the clause keeps it", which the title stop now answers to as well. It is asked of the whole SPAN a stop gives up, not of the stop's own word: a first draft that checked the word alone broke the invariant on 513 of a 5,198-parse sweep (measured 2026-09-26 on that draft, which is not in the tree to recompute), because a stop gives up every word behind it — `DOE NEE SMITH PROF. MA` stopped at the title and handed the MA to the family, where the name left standing reads MA as a name; it keeps maiden 'SMITH PROF. MA' now. The question is asked per reader, the way that reader reads: the TRAILING reader is assign's own `tail_reading` over the view; the GIVEN_SLOT reader has its count settled by the comma, so it is the writing alone with a name word ahead. The name-word-ahead half is what keeps `Dr. nee Jones Smith Prof.` whole — the take would leave `Dr. Prof.`, whose Prof. would be the family name. THE JOINS. A released span holding a title with a particle ahead of it declines, because P2's chain runs on over a trailing title (rules.md#H5's Accepted `John van der Berg Prof.`): `Jane van der Berg nee Smith Prof.` keeps maiden 'Smith Prof.', and a released particle with a title behind it is withdrawn for the same reason (`Jane Doe nee Smith MA do Prof.` keeps maiden 'Smith MA do'). The bound-given (P5) half of `_join_takes_the_member` is asked for the GIVEN_SLOT reader alone: P5 is `BoundJoin.LENIENT` only after a family comma, and before one its STRICT reserve already refuses a join that would change a suffix reading, so asking there over-declined — `abdul nee Smith V` kept 'V' in the clause, where the scoped check lets it go to suffix as 2.3.0 did. The consequence is recorded by a case row rather than prevented: `abdul nee Smith Dr.` reads title 'Dr.', family 'abdul', maiden 'Smith', which is how `abdul Dr.` reads bare (2.3.0 read given 'abdul', maiden 'Smith Dr.'). THREE NUMERAL-STOP READINGS MOVE WITH IT, decided in rather than deferred, since each is the same invariant broken at the stop this change was already rewriting. The numeral stop never asked the join question: `Berg, abdul nee Smith V` read given 'abdul V' at 2.2.0 and 2.3.0 (given 'abdul nee', suffix 'V' at 2.0.0 and 2.1.0, #411's reserve differing between those pairs before the walk runs) and reads maiden 'Smith V' now. After a family comma the given slot reads a lone numeral as a suffix only where the given part is the LAST comma part — assign's own two-segment condition from #144 — and the walk now asks it too (`tail_follows`): `Doe, Jane nee Smith V, PhD` read middle 'V' at 2.2.0, 2.3.0 and e0f1a2fa, an M2 violation predating this change that its first draft had extended to `… V Prof., PhD`, and it reads maiden 'Smith V' now, as 2.0.0 read it. The withdrawal reaches through a title behind the numeral, so `Doe, Jane nee Smith Prof. V, PhD` keeps maiden 'Smith Prof. V' where `Doe, Jane nee Smith V Prof., PhD` gives up the title — the same asymmetry the bare given slot has with no marker in it (`Doe, Jane Prof. V, PhD` against `Doe, Jane V Prof., PhD`). And the numeral stop reads FROM the marker, unlike the other two, so a numeral straight after the marker is not held to the first-word floor, and it may decline the clause; examples, not a rule over every shape: `Dr. nee V` and `Doe, J. nee V` keep maiden 'V' as 2.3.0 did, though `Doe, J. V` reads suffix 'V', and `Doe nee V, Jane` and `Smith, John, PhD nee V` read no clause at all (family 'Doe nee V', suffix 'PhD nee V', as at 2.3.0) — `Jane Doe nee V Prof.` read maiden 'V Prof.' at 2.3.0 and reads family 'nee', suffix 'V', title 'Prof.' now, as `Jane Doe nee V` reads plus the title. After a family comma with another comma part behind the given one, the same #144 condition that keeps `Doe, Jane nee Smith V, PhD` whole now keeps a lone numeral in the clause as well: `Doe, Jane nee V, PhD` read given 'Jane', middle 'nee V', suffix 'PhD' at 2.3.0 and at the parent d9d80492 and reads maiden 'V', suffix 'PhD' now, as 2.0.0 and 2.1.0 read it, `Doe, Jane nee V, Jr.` likewise — the marker had been read as a name word there at 2.2.0 and 2.3.0, and a case row records it. A LINK THE WALK STOPS AT ASKS THE RELEASE QUESTION TOO. Reading the end of the clause through the title chain makes the link exception refuse a link it used to join — in `Jane Doe nee Smith i DO Prof.` the DO is the peel's once the title is chained — and a stop at a link gives up the words behind it. Unchecked, that put 'Doe i' in the middle name and 'DO Prof.' in the family (e0f1a2fa read maiden 'Smith i DO Prof.'), and `Berg, abdul nee Smith i V Prof., MD` read middle 'V'. So a link the exception refuses asks `_release_reads_off` of the run it would give up — only where the exception, bounded by the peel over the words as WRITTEN, would have joined it (the refusal is the title chain's), where a trailing rule reads the clause, and where the link is not the first word after the marker (a first-word stop declines the clause and gives nothing up); where that fails the clause keeps the link and walks on. Scoped that narrowly because a first version asked it of every refused link and kept links the name left standing reads off: `Doe, Jane nee Smith i V` kept maiden 'Smith i' where every release gives up suffix 'i V' (23,472 of a 411,936-parse link grid moved clean readings, measured 2026-09-26). With it, the given-part model reads the lenient numeral in both of assign's passes — the literal last piece first, then the last one standing once the chain has taken the titles behind it — and only where no comma part follows (`Doe, Jane nee Smith i V Prof.` gives up 'i V' and the title, as bare `Doe, Jane i V Prof.` reads them; `… i V Prof., PhD` keeps maiden 'Smith i V'). The same model moves the credential stop where a lenient numeral follows the credential: `Doe, Jane nee Smith MA V` now gives up 'MA V' (suffix 'MA V', maiden 'Smith') as bare `Doe, Jane MA V` reads it, where d9d80492 and 2.3.0 read maiden 'Smith MA', suffix 'V'; `… MA V, PhD` keeps maiden 'Smith MA V'. Over that link grid (links i, y, e, and, i y, y i, and i × 8 heads × bodies × `, MD` tail × 3 cases × 3 name orders), against the tree before the link check: 0 new violations, 6,195 fixed, and the clean readings that move now match the bare given part's (`DOE, JANE NEE SMITH MA I` gives up 'MA I', as `DOE, JANE MA I` reads suffix 'MA I'). `Jane Doe nee Smith i DO Prof.` now reads title 'Prof.', maiden 'Smith i DO' and reports the kept DO, and `Berg, abdul nee Smith i V Prof., MD` maiden 'Smith i V Prof.', suffix 'MD'. A RELEASED TITLE THAT IS ALSO A PARTICLE IS KEPT AFTER A FAMILY COMMA, because P6 attaches a particle trailing the given part to the family: released by the title stop, 'St.' in `Doe, Jane nee Smith St.` reads family 'St. Doe', so `_join_takes_the_member` declines it and the name reads maiden 'Smith St.', as e0f1a2fa and 2.3.0 read it, and so does `Doe, Jane nee Smith MA St.`. Titles only — a credential that is also a particle ('DO') is the given slot's own lean, which the credential stop has already asked (#533). A FOURTH MEASUREMENT, against e0f1a2fa rather than d9d80492: the M2 invariant over a fuzz of twelve heads ({`Jane Doe`, `Doe, Jane`, `Doe, Prof.`, `Jane Doe, PhD`, `Doe, J.`, `Berg, Jane van der`, `Jane van der Berg`, `Berg, abdul`, `abdul Berg`, `J. Doe`, `Prof. Jane Doe`, `Dr.`}), ` nee Smith ` and every one-to-three-word sequence over {MA, Ma, V, Prof., St., King., PhD, ba, do, DO, Jr., i, van, y} holding at least one of Prof., St., King., then nothing or `, MD`, as written, upper- and lower-cased (107,352 parses, measured 2026-09-26 on the narrowed tree; the intermediate tree above read 2,228): 588 texts that violated it at e0f1a2fa no longer do, and 12 newly do, all `Dr. nee Smith i MA|V` followed by `Prof.`, `King.` or `St.`, with or without `, MD`. Each is the title-carrying twin of `Dr. nee Smith i MA` / `Dr. nee Smith i V`, which read family 'i' at e0f1a2fa already: the bare-title head leaves no name word, the unguarded stop's class (#548), which the title now reaches because it leaves the clause. ACCEPTED: `Jane Doe (nee Smith Prof.)` reads title 'Prof.', maiden 'Smith' — bracket content ending in a period is suffix-shaped (S1), so the brackets are dropped and the clause is read bare, outside the reach of the delimiter precedence; rules.md#M2 carries it as an Accepted example. MEASURED 2026-09-26, the tree against its parent d9d80492, over three generated grids, and these are dated snapshots. Recipe: grid A is each head in {`Doe, Jane`, `Doe, J.`, `Doe, Prof.`, `Berg, abdul`, `Doe, Jane van`, `Jane Doe`, `J. Doe`, `Dr.`, `abdul`, `Jane van der Berg`, `John`}, then ` nee `, then every sequence of one to three words (repetition allowed) from {Smith, Jones, Prof., Dr., Sir, MA, DO, PhD, V, III, i, do, van, Ma, M.A., King., Rev., ba}, then either nothing or `, PhD`, each text as written, upper-cased and lower-cased, duplicates removed — 327,936 texts; grid B, aimed at the bound-given join, is the same construction over heads {`Dr. abdul`, `abdul rahman`, `Dr. abdul rahman`, `Berg, Dr. abdul`, `Berg, abdul rahman`, `abu`, `Berg, abu`, `Mr. abu`, `Dr.`, `Berg, Dr.`, `abdul`, `Berg, abdul`}, words {Smith, Prof., MA, V, do, van, PhD, Ma, ba, III, Jr., M.A.} and tails {nothing, `, PhD`, `, Jr., MD`}, plus these fourteen texts, as written only: `Jane Doe nee Ph. D. Prof.`, `Jane Doe nee Ph. D. Smith Prof.`, `Jane Doe z domu King. ba`, `Jane Doe z domu Smith Prof.`, `Jane Doe nee King. Prof. ba`, `Jane Doe nee King. Prof.`, `Doe, Jane nee Smith V,`, `Doe, Jane nee Smith V, ` (with the trailing space), `Doe, Jane, nee Smith V`, `Doe, Jane nee Smith V Prof., PhD`, `Doe, Jane nee Smith Prof. V, PhD`, `Doe, Jane nee Smith MA, PhD`, `Doe, Jane nee Smith V (Jr.)`, `Doe, Jane nee Smith V "Bo"` — 173,057 texts. The invariant tested is M2's: in a parse with a non-empty maiden field, every token written after the ` nee ` marker is roled maiden, title or suffix (so the two `z domu` texts are parsed and not checked). Grid C, a four-word probe of the title chain after a family comma, is each head in {`Doe, Jane`, `Jane Doe`, `Doe, J.`, `John Smith`, `Doe, Jane Mary`}, then ` nee `, then every sequence of four words (repetition allowed) from {Smith, Jones, Prof., Dr., King., MA, PhD, V, Jr., ba, do}, then either nothing or `, PhD`, as written only — 146,410 texts. Per grid, the counts are: texts that newly violate the invariant at the tree, texts that violated it at the parent and no longer do, texts that violate it at both. Grid A: 0, 2,012 and 17,854; grid C: 0, 1,368 and 15,534; grid B: 0, 1,222 and 11,149 (all re-measured 2026-09-26 on the tree with the link and particle-title checks below, as narrowed; an intermediate tree whose link check fired on every refused link read grid A 0, 3,772 and 16,094, its extra 1,760 fixes being links the as-written reading refuses too, which are #548's class), one of those last being the check itself counting the quoted nickname in `Doe, Jane nee Smith V "Bo"`. In grids A and B, every other remaining violation has the clause ending immediately before a word of the unambiguous suffix vocabulary written the way the plain suffix-word stop takes it (PhD, III, M.A., Jr., and a lower-case i or v; never a bare capital, which reads as an initial), which is where that stop ends a clause without asking a release question. DEFERRED: the walk's plain suffix-word stop — the one ending the clause at the first suffix word, as distinct from the trailing numeral, credential and title stops — asks no release question at all, and where the words behind it cannot read as post-nominals they land in a name part: on degenerate heads the stop word itself does (`Dr. nee Smith PhD Prof.` reads family 'PhD', `Doe nee Smith Jr. Prof., Jane` family 'Doe Prof.'), and with an ordinary head the words behind it do (`Doe, Jane nee Smith PhD Smith` reads middle 'Smith'). Unchanged here, and what "the clause keeps it" should mean where the kept words would follow a credential the clause itself ended at is its own question; rules.md#M2 carries `Doe nee Smith Jr. Prof., Jane` as an Accepted example meanwhile. The trailing numeral's stop has the same gap before a family comma, where it is made over the peel alone with no release question: `Doe nee Smith V, Jane` reads family 'Doe V', maiden 'Smith' (so did 2.2.0, 2.3.0 and the parent d9d80492; 2.0.0 and 2.1.0 read maiden 'Smith V'), and rules.md#M2 carries it beside the other. Open: #548. COST, measured 2026-09-26 on CPython 3.11.16 as profiler call events in one `Parser.parse` after a warm-up parse, parent → tree: `John Smith` 171 unchanged, as are `Jane Doe Prof.` and `John Smith MA`; `Jane Doe nee Smith` 244 → 246; `Jane Doe nee Smith MA` 377 → 388; `Jane Doe nee Smith Prof.` 276 → 372, the one shape that now runs the chain and a release check it never ran. Recompute: count `sys.setprofile` call events around the second of two `Parser().parse` calls, with the tree and d9d80492 each first on `sys.path`. - 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'. @@ -913,11 +913,11 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): COST, measured 2026-10-01 by `sys.setprofile` call counts (mean of 50, after one warm-up) against `git archive origin/master`: `Smith, John` 183 → 183, `John Smith, MA` 252 → 252, `John Smith, PhD` 209 → 217, `Dr. Juan de la Vega III` (the benchmark reference) 366 → 366. The +8 is the unambiguous credential path building its token list and asking each token whether it is a particle; a part with no particle never builds the units, and a frame-free approximation of `_normalize` was declined as a second spelling of it. - 2026-10-01 (Derek), #564 — AN UNLISTED ALL-CAPS WORD IS A CREDENTIAL IN THE PART AFTER A COMMA BEHIND TWO NAME WORDS, BY DEFAULT. The decision, its two-letter rule and the mixed-run membership are recorded under decisions.md#S2 (#564), where the caps shape's switch has always been decided; this bullet is C1's pointer to it. - 2026-10-03 (Derek), #549 — A DECLARED DELIMITER PARTS A TRAILING SUFFIX PART AS A COMMA WOULD, AND THIS REVERSES #538'S SKIP READING (M2, 2026-09-26). The question #549 asked was whether a core a connective join swallows should be skipped (the text read as if the core were absent) or should end the run it stands in; Derek's ground for the second was the one a caller has when declaring one: a delimiter set to ` - ` is meant to behave like a comma around suffixes, and the simplest parser that honours that is one where it IS a comma there. So `group` cuts a tail segment at its cores before grouping anything, groups each part as the comma twin's own segment would be grouped, and drops the cores, which post_rules' entry pass still reads back as boundaries. No join, maiden walk or link search can reach across a core because none of them ever sees one, and the machinery that let them step over it is gone: the `cores` parameter of `_group_segment`, `_maiden_take`, `_run_neighbours` and `_link_joins_inside_the_clause`, `_maiden_take`'s `seen` index list, the `skip` arguments of `peel_walk` and `trailing_start`, and the #206 drop block that ran after the joins. One call per parse fewer (`tools/perf/call_count.py`, py3.11, 2026-10-03: parse 369 against 370 at 061f02da). - WHAT MOVES, against 1.4.0 and against 2.3.0. A connective beside a core no longer joins across it: `HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix` is 'PhD, and MD', and `... Puig - y Soler` 'Puig, y Soler' -- 1.4.0's answers (its `expand_suffix_delimiter` split the part's text before anything joined), where 2.0.0, 2.1.0, 2.2.0 and 2.3.0 each gave 'PhD - and MD' and 'Puig - y Soler' (all measured 2026-10-03 with the released wheels). The generational link moves back too: 2.3.0 read `Smith, John, Puig - i Soler` as 'Puig, i Soler' and `Smith, John, PhD née Puig - i Soler` as maiden 'Puig', suffix 'PhD, i Soler', `i` being no connective there, and #397/#538 had moved them to 'Puig - i Soler' and maiden 'Puig i Soler' before either shipped; both read as 2.3.0 did again. An interior core still parts: `... PhD née Puig Dr. i - y Soler` gives suffix 'PhD i, y Soler'. At the default policy nothing moves, `extra_suffix_delimiters` being empty: the gate exits 0 at all five baselines. - THE DECLINED READING, TAKEN. #538 declined the boundary reading because it moves `Smith, John, PhD née Puig - i Soler` from maiden 'Puig i Soler' to maiden 'Puig' with suffix 'PhD, i Soler', "a birth-name link pushed into the credentials". That is what the same text with a comma typed in the core's place reads, and it is what 2.3.0 read; the comma reading is accepted as the cost of one rule for a declared separator, where the skip reading kept two (a lone core dropped like a comma, a core beside a connective stepped over like a word) and #549 was the seam between them. - INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin`: every field of a name with a ` - ` core in a part after the suffix comma equals that of the same text with a comma in its place, over 129 generated texts; 102 disagreed at 061f02da, 0 after. It replaces #538's `test_a_delimiter_core_reads_as_if_it_were_not_written`, whose invariant this decision reverses. Scoped to TRAILING parts on purpose: before the first comma, or in the given part of the listing form, the core is a word (C1's Accepted entry, v1 parity), and a comma typed there changes the name's structure rather than parting a suffix run, so the twin is a different parse. - MEASURED 2026-10-03, branch against 061f02da. POPULATION: each head in {`Smith, John,`, `Smith, John, Jr.,`, `John Smith,`, `John Smith`} followed by every 2-, 3- and 4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr., née, and} that contains `-`, 19,972 texts, plus the 1,505 distinct names of `tools/differential/corpus*.jsonl` at 061f02da. Under `extra_suffix_delimiters=(" - ",)` 2,624 of the generated parses move (suffix alone 2,384, suffix and maiden 240); under `(" / ",)`, with each ` - ` rewritten to ` / `, 2,204 (1,924 and 280); and under `(" and ",)` over the texts as written, where `and` is now the core and the dash a word, 2,528 (2,468 and 60). No other field and no ambiguity report moves under any of the three. In the corpus, one name moves under each of two policies: `Smith, John, PhD née Puig - i Soler` under ` - ` (above), and `Doe, Jane, and Jr.` under ` and `, suffix 'and Jr.' to 'Jr.', a leading core dropped as any lone core is. COMMA-TWIN AGREEMENT over the generated texts under ` - `, comparing each against the text with ` - ` replaced by `, ` under the default policy and keeping only texts whose core is neither first nor last after the head nor beside another core: `Smith, John,` 752 of 2,100 disagreed before and 0 after, `Smith, John, Jr.,` 973 of 3,310 and 0; `John Smith,` 1,956 of 2,100 both before and after, the head where the typed comma moves the name's structure. RECOMPUTE: parse each text under each policy at both trees (pin `nameparser.__file__` on each side, AGENTS.md's two-tree gotcha), compare the seven role fields and the sorted ambiguity kinds, and count the texts whose record differs. - SUPERSEDES the 2026-09-26 #538 entry under M2 (its stepping, its invariant and its Not-repaired list, which this closes) and the 2026-09-22 #397 follow-up's bound argument there (a core can no longer stand between a marker and the clause's first word, the segment being cut at it first). Not changed: a core OUTSIDE a tail segment is still a word (C1's Accepted entry), and a core inside a token (`RN/CRNA` under `/`) still reads by `splits_into_suffixes` and keeps the token whole. + WHAT MOVES, against 1.4.0 and against 2.3.0. A connective beside a core no longer joins across it: `HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix` is 'PhD, and MD', and `... Puig - y Soler` 'Puig, y Soler' -- 1.4.0's answers (its `expand_suffix_delimiter` split the part's text before anything joined), where 2.0.0, 2.1.0, 2.2.0 and 2.3.0 each gave 'PhD - and MD' and 'Puig - y Soler' (all measured 2026-10-03 with the released wheels). The generational link moves back too: 2.3.0 read `Smith, John, Puig - i Soler` as 'Puig, i Soler' and `Smith, John, PhD née Puig - i Soler` as maiden 'Puig', suffix 'PhD, i Soler', `i` being no connective there, and #397/#538 had moved them to 'Puig - i Soler' and maiden 'Puig i Soler' before either shipped; both read as 2.3.0 did again. An interior core still parts: `... PhD née Puig Dr. i - y Soler` gives suffix 'PhD i, y Soler'. A MAIDEN MARKER BESIDE A CORE TAKES NO CLAUSE, and this is the larger cost by count: with the core straight before the marker (`Smith, John, MD - née Jones Smith`) or straight after it (`Smith, John, PhD née - Jones`), the comma twin's marker has no name word ahead of it in its part, or nothing behind it, and M2 declines, so the parse reads suffix 'MD, née Jones Smith' and 'PhD née, Jones' with no maiden, where 2.0.0 through 2.3.0 read maiden 'Jones Smith' and 'Jones' (measured 2026-10-03; 1.4.0 had no maiden markers and gave the suffix reading). It is the one rules.md#C1 records as Accepted. At the default policy nothing moves, `extra_suffix_delimiters` being empty: the gate exits 0 at all five baselines. + THE DECLINED READING, TAKEN. #538 declined the boundary reading because it moves `Smith, John, PhD née Puig - i Soler` from maiden 'Puig i Soler' to maiden 'Puig' with suffix 'PhD, i Soler', "a birth-name link pushed into the credentials". That is what the same text with a comma typed in the core's place reads, and it is what 2.3.0 read. The marker beside a core (WHAT MOVES, above) is the same trade at a higher price, a maiden name lost rather than a link moved, and it is taken on the same ground; the review of this change found it, the first draft of this entry having weighed only the link shape. The comma reading is accepted as the cost of one rule for a declared separator, where the skip reading kept two (a lone core dropped like a comma, a core beside a connective stepped over like a word) and #549 was the seam between them. + INVARIANT, pinned by `tests/v2/test_properties.py::test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin`: every ROLE field of a name with a ` - ` core in a part after the suffix comma equals that of the same text with a comma in its place, over 129 generated texts; 102 disagreed at 061f02da, 0 after. It replaces #538's `test_a_delimiter_core_reads_as_if_it_were_not_written`, whose invariant this decision reverses. Scoped to TRAILING parts on purpose: before the first comma, or in the given part of the listing form, the core is a word (C1's Accepted entry, v1 parity), and a comma typed there changes the name's structure rather than parting a suffix run, so the twin is a different parse. The ambiguity reports are not part of it: C2 reports `comma-structure` once per extra comma part, so a typed comma can add a report the core does not (`Smith, John, Puig - Puig` against `Smith, John, Puig, Puig`). + MEASURED 2026-10-03, branch against 061f02da. POPULATION: each head in {`Smith, John,`, `Smith, John, Jr.,`, `John Smith,`, `John Smith`} followed by every 2-, 3- and 4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr., née, and} that contains `-`, 19,972 texts, plus the 1,505 distinct names of `tools/differential/corpus*.jsonl` at 061f02da. Under `extra_suffix_delimiters=(" - ",)` 2,624 of the generated parses move (suffix alone 2,384, suffix and maiden 240 -- every one of the 240 a maiden name lost to a marker beside a core: 100 with `née -`, 108 with `- née`, 32 with both); under `(" / ",)`, with every `-` word rewritten to `/`, the same 2,624 (2,384 and 240), the move being a property of the core's position rather than its spelling; and under `(" and ",)` over the texts as written, where `and` is now the core and the dash a word, 2,528 (2,468 and 60). No other field and no ambiguity report moves under any of the three. In the corpus, one name moves under each of two policies: `Smith, John, PhD née Puig - i Soler` under ` - ` (above), and `Doe, Jane, and Jr.` under ` and `, suffix 'and Jr.' to 'Jr.', a leading core dropped as any lone core is. COMMA-TWIN AGREEMENT over the generated texts under ` - `, comparing the seven role fields of each against the text with ` - ` replaced by `, ` under the default policy and keeping only texts whose core is neither first nor last after the head nor beside another core: `Smith, John,` 752 of 2,100 disagreed before and 0 after, `Smith, John, Jr.,` 752 of 2,100 and 0; `John Smith,` 1,956 of 2,100 both before and after, the head where the typed comma moves the name's structure. (A first draft counted 973 of 3,310 for the `Jr.,` head, its filter reading `Jr.,` as part of the body; the review of this change caught it.) RECOMPUTE: parse each text under each policy at both trees (pin `nameparser.__file__` on each side, AGENTS.md's two-tree gotcha), compare the seven role fields and the sorted ambiguity kinds, and count the texts whose record differs; for the twin, compare the role fields alone. + SUPERSEDES, under M2: the 2026-09-26 #538 entry (its stepping, its invariant and its Not-repaired list, which this closes); the 2026-09-22 #397 follow-up's bound argument (a core can no longer stand between a marker and the clause's first word, the segment being cut at it first); and the #418 reorder entry's closing repair, which screened a tail's cores out of the marker pass so that `Smith, John, PhD née - Jones` read maiden 'Jones' -- it reads suffix 'PhD née, Jones' now, as above. Entries there that describe the walk stepping over a tail's cores (the #424 and #533 bullets) describe machinery this change deletes; they are dated and left as written. Not changed: a core OUTSIDE a tail segment is still a word (C1's Accepted entry), and a core inside a token (`RN/CRNA` under `/`) still reads by `splits_into_suffixes` and keeps the token whole. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index ee14dda2..2033de63 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -547,8 +547,9 @@ P3. Rationale: connective words ("y", "of the") bind name words into nothing, and a word of that vocabulary ending a name, or standing before the credential a name ends with, is the generation it also spells. - A separator the caller declared ends the search as a comma would, - a connective never joining across one (C1). + A separator the caller declared, standing in a trailing suffix + part, ends the search there as a comma would, so a connective + never joins across it (C1); anywhere else it is a word. Both questions this rule asks of a name — how many words it has, and whether it is written in one case — are asked of the name's OWN words: a maiden marker taken as one, and the words it takes @@ -650,7 +651,7 @@ P3. Rationale: connective words ("y", "of the") bind name words into same two words unjoined are two name words and H1 does not fire. P1's leading run is the second (#395, landed): its run takes the "Vega y Santos" join whole or stops before it. - history: decisions.md#P3 · interacts: H1, P1, M2, R1, R3, R4, S2 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_post_rules.py + history: decisions.md#P3 · interacts: H1, P1, M2, R1, R3, R4, S2, C1 · implemented: nameparser/_pipeline/_classify.py, nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_post_rules.py P4. Rationale: a particle links forward from inside a name; at the very front there is no name yet to be inside. @@ -1413,10 +1414,10 @@ M2. Rationale: a maiden marker announces that what follows it is the reads them as post-nominals or titles; otherwise the clause keeps the link and runs on. The link first after the marker is not asked: stopping there declines the clause. - A separator the caller declared is structure rather than a name - word, and the link exception reads past it: the word on a link's - side is the one beyond the separator, so the clause reads as the - same clause written without it. + A separator the caller declared, standing in a trailing suffix + part, ends the clause as a comma would (C1): the words beyond it + are the next part's, and a marker with one straight before or + after it reads as it would with a comma typed there. The trailing numeral, credential and title stops are each asked TWICE for one reason: the count of words to spare includes the very words the marker @@ -1622,7 +1623,7 @@ M2. Rationale: a maiden marker announces that what follows it is the clause reads that member as the credential. "Jane Doe Jr. nee Smith Ma" → maiden="Smith Ma" "Jane Doe Jr. Ma" → suffix="Jr. Ma" · boundary - history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5 · implemented: nameparser/_pipeline/_group.py + history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5, C1 · implemented: nameparser/_pipeline/_group.py M3. Rationale: an enclosure says nothing about whether it means maiden, but a recognized marker word inside it does — the clause @@ -1859,7 +1860,8 @@ C1. Rationale: a credential run after the comma means the name is in comma would: the words on each side of it read exactly as the parts of the same text written with a comma in its place, so no join, maiden clause or connective reaches across it, and the - delimiter itself is dropped. Only a trailing part is parted: in + delimiter itself is dropped, unless it is the whole of its part. + Only a trailing part is parted: in the part before the first comma, or in the given part of the listing form, the delimiter is a word (the Accepted entry below). "Smith, John" → family="Smith" @@ -1981,6 +1983,14 @@ C1. Rationale: a credential run after the comma means the name is in here. It is also the one consequence of this question Latin script cannot witness, nothing there being glued to the end of a name — which is why this clause carries no example of its own. + Accepted: a maiden marker with a declared delimiter straight + before or after it takes no clause, the same words written with + a comma there taking none either — a marker needs a name word + ahead of it in its own part, and something to take behind it + (M2). Releases 2.0 through 2.3 read past the delimiter and took + the clause; the comma reading is accepted as the cost of one + rule for a declared separator. + "Smith, John, MD - née Jones Smith" extra_suffix_delimiters-dash → suffix="MD, née Jones Smith" Accepted: a delimiter core the policy names (T1) is a word here, not structure — v1 applied the delimiter to the suffix-comma form alone, and that limitation is kept as parity: "Smith, RN - diff --git a/docs/release_log.rst b/docs/release_log.rst index a325b4ed..b6f14fde 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -70,7 +70,7 @@ Release Log - **Fix a katakana name typed with a separate voicing mark being read family-first.** ``HumanName("ア゙イ タロウ")`` gives first ``ア゙イ``, last ``タロウ``, as ``アイ タロウ`` does, where 2.1 through 2.3 gave last ``ア゙イ``, first ``タロウ``. A dakuten or handakuten after a kana that has no precomposed voiced form (``ア゙``, ``ン゙``), or the spacing ``゛`` and ``゜``, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: ``マイケル ア゙イ`` gives first ``マイケル`` where it gave last ``マイケル``. Hiragana names and kanji-and-kana names with such a mark read as before. See the ``W4`` entry of ``docs/design/decisions.md`` (closes #596) - - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part exactly as a comma typed in its place would: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. Only a trailing suffix part -- one after the name's commas -- is affected; elsewhere the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) + - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. **A maiden marker with the delimiter straight before or after it no longer takes a maiden name**, as it takes none after a typed comma: ``Smith, John, MD - née Jones Smith`` gives suffix ``MD, née Jones Smith`` and no maiden, where 2.0 through 2.3 gave maiden ``Jones Smith``. Only a part read wholly as suffixes after the name's commas is affected; before the first comma, or in the given-name part of ``Family, Given``, the delimiter is still a word, as in 1.4.0 (``John Smith, PhD née Puig - i Soler`` keeps maiden ``Puig - i Soler``). The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) - **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597) diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index ff9738d0..12216112 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -9,8 +9,8 @@ Reads: token tags (from classify), Lexicon.given_name_titles (the P5 licence, #369) and Policy.extra_suffix_delimiters, whose delimiter-core tokens part a tail segment as a comma would and are -dropped (v1 suffix_delimiter parity, #549) -- no other Policy field. Policy.lenient_comma_suffixes left this list -with #436: it reached here only through segment_suffix_reading, whose +dropped (v1 suffix_delimiter parity, #549) -- no other Policy field. +Policy.lenient_comma_suffixes left this list with #436: it reached here only through segment_suffix_reading, whose render consumer was this stage's one-entry join and now lives in post_rules. The v1 "derived titles/prefixes" registration becomes piece_tags entries -- per-parse state that @@ -2034,16 +2034,29 @@ def group(state: ParseState) -> ParseState: # because none of them ever sees one. A segment that IS only # its core keeps it (v1 expand() splits within a part, never # erases a lone part). + # Accumulated in a list and frozen once per part: a tuple + # extended per token is quadratic in the part's length, C-level + # work no frame guard sees (#553's class; review of #549). + # Measured 2026-10-03, py3.11, 'Smith, John, ' + 'PhD ' * n + + # '- MD' under a ' - ' delimiter: 4.7x for n 4,000 -> 16,000 as + # written, the default policy's own 4.5x, where the tuple read + # 8.0x. No clock guard: _PREFIXED_SHAPES parses with the default + # policy and _POLICY_SHAPES repeats a unit with no prefix, and a + # tail needs both, so a row would mean a third table. parts: list[tuple[int, ...]] = [seg] if tail and cores and len(seg) > 1: - parts = [()] + parts = [] + current: list[int] = [] for i in seg: if tokens[i].text in cores: dropped.append(i) - parts.append(()) + if current: + parts.append(tuple(current)) + current = [] else: - parts[-1] += (i,) - parts = [p for p in parts if p] + current.append(i) + if current: + parts.append(tuple(current)) pieces = [] ptags = [] for part in parts: diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index 81f6d074..4c534a1c 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -1804,7 +1804,7 @@ def test_the_marker_is_not_the_name_word_on_the_links_left() -> None: assert _maiden_texts(plain) == ["i", "Soler"] -def test_a_core_between_the_marker_and_the_first_word_is_below_lo( +def test_a_core_straight_after_the_marker_leaves_it_nothing_to_take( ) -> None: # A delimiter core is TAIL-segment structure that group() cuts the # segment at before this pass (#549), so it is never a word of the diff --git a/tests/v2/pipeline/test_post_rules.py b/tests/v2/pipeline/test_post_rules.py index cb9ff6a2..83ad45f4 100644 --- a/tests/v2/pipeline/test_post_rules.py +++ b/tests/v2/pipeline/test_post_rules.py @@ -1013,8 +1013,8 @@ def test_a_bare_marker_maiden_clause_does_not_part_the_run() -> None: def test_a_marker_the_policy_names_as_a_delimiter_parts_the_run() -> None: """Where the two sites' gates differ, and the one case that - reaches it. group drops a core only on a `tail` segment, through - `seg_cores`; this pass reads `delimiter_cores` whole and asks by + reaches it. group drops a core only on a `tail` segment, cutting + the segment there (#549); this pass reads `delimiter_cores` whole and asks by TEXT, so a dropped token the policy names as a delimiter parts the run whatever dropped it. A maiden marker the policy ALSO lists is the reachable case, and it parts by the policy's own declaration diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index fce4c74e..5b105a6e 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -2320,12 +2320,14 @@ class _LatinCopy(NamedTuple): frozenset({"Jane Doe \\(nee Smith MA\\)", "Jane Doe \\(nee Smith Ma\\)", "Jane Doe \\(nee Smith\\) MA"}), # rules.md#M2's two live examples of the delimiter-core fix (#538), + # and rules.md#C1's Accepted example of a marker beside one (#549), # one alternative per corpus name. A list of names, not a copy of # any wordlist: what selects them is the doc's own choice of # examples for a policy this corpus does not configure, so there # is no vocabulary here for the alternation to drift from. frozenset({"Smith, John, PhD née Puig Mr\\. - i Soler", - "Smith, John, PhD née Puig - i Soler"}), + "Smith, John, PhD née Puig - i Soler", + "Smith, John, MD - née Jones Smith"}), # fix(#400)'s two openings: start-of-name or just after a family # comma. `abd` joins forward on the given side wherever that side # begins, and the alternation is over ANCHORS, not over words -- @@ -3692,8 +3694,11 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: 105 -> 108; gains 'Doe, Jane nee Smith PhD # MEng', 'Jane Doe Jr. nee Smith Ma', 'Jane Doe nee Smith PhD # MEng'. + # 2026-10-03, #549: 108 -> 109, 'Smith, John, MD - née Jones + # Smith', rules.md#C1's Accepted example. Reach -- it carries a + # marker. "fix(#274) maiden markers consumed": - _Claim(108, ('family', 'maiden', 'middle'), "19733e2db435", None), + _Claim(109, ('family', 'maiden', 'middle'), "87d3acaaea68", None), # 2026-09-19, #533: 5 -> 6, the same one new corpus name # '田中 太郎 旧姓 佐藤 MA' as the CJK rule above. "fix(cjk-maiden-marker) maiden marker consumed, compounding with the CJK order flip": @@ -3863,7 +3868,11 @@ def _claim(rule: dict) -> _Claim: # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified # name by name. - _Claim(441, ('given', 'suffix', 'title'), "9f312397b07e", None), + # 2026-10-03, #549: 439 -> 442, 'Smith, John, Puig - i Soler', + # 'Smith, John, PhD - and MD' and 'Smith, John, MD - née Jones + # Smith', #549's rules.md#C1 examples. Reach -- each is a + # comma name. + _Claim(442, ('given', 'suffix', 'title'), "bb89aae9aa98", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3974,7 +3983,9 @@ def _claim(rule: dict) -> _Claim: # 'García Márquez, MJ JK', 'John Smith, PhD XYZ' and 'MÜLLER # WEIß, HANS', #564's rules.md#C1 examples. Reach, verified # name by name. - _Claim(441, ('family', 'given'), "9f312397b07e", None), + # 2026-10-03, #549: 439 -> 442, the same three #549 examples. + # Reach. + _Claim(442, ('family', 'given'), "bb89aae9aa98", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": @@ -4245,6 +4256,8 @@ def _claim(rule: dict) -> _Claim: # unrelated 'y', which is in the corpus because its row # carries a shape tag. Reach again, verified name by name. # 2026-10-01, #575: 104 -> 105, 'Ortega y Gasset, Ed'. Reach. + # 2026-10-03, #549: 105 -> 106, 'Smith, John, PhD - and MD'. + # Reach. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": _Claim(106, ('_initials',), "8b55345f9ad4", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening @@ -4516,8 +4529,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-26, #538: 1 -> 2, rules.md#M2's second live example, # 'Smith, John, PhD née Puig - i Soler', joining the same # alternation for the same reason. + # 2026-10-03, #549: 2 -> 3, rules.md#C1's Accepted example, + # 'Smith, John, MD - née Jones Smith', joining it too. "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all": - _Claim(2, ('maiden', 'suffix'), '168b8bdef5db', None), + _Claim(3, ('maiden', 'suffix'), 'c1572309c3f2', None), # New rule (#274): one corpus name, 'Dr. nee Jones Smith # Prof.' -- only a title precedes the marker, so no name word # is left standing ahead of it. Its diff from this baseline is diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 0c78811a..364309f3 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -358,6 +358,7 @@ "Smith, John Prof." "Smith, John V" "Smith, John V." +"Smith, John, MD - née Jones Smith" "Smith, John, PhD - and MD" "Smith, John, PhD - i Soler" "Smith, John, PhD née Puig - i Soler" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 98d035d1..77282d3c 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4710,7 +4710,13 @@ issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the su # the same shape as the neighbour's own reading with 'Mr.' dropped # from it; the tree's reading (suffix 'PhD', maiden 'Puig - i Soler') # is unchanged by the fix, and only the corpus name is new. -name_regex = "^(?:Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler)$" +# 'Smith, John, MD - née Jones Smith' joined with #549 (2026-10-03): +# rules.md#C1's Accepted example of a marker beside a configured +# delimiter, reaching the DEFAULT facade with the dash an ordinary +# word again. Measured at this baseline: suffix 'MD - née Jones Smith' +# -> 'MD -', maiden '' -> 'Jones Smith' -- the marker half of this +# rule alone, there being no link in it. +name_regex = "^(?:Smith, John, PhD née Puig Mr\\. - i Soler|Smith, John, PhD née Puig - i Soler|Smith, John, MD - née Jones Smith)$" fields = ["maiden", "suffix"] [[change]] From 684afbc7660c153c1b16e23532f6c725696ee2b0 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 17:32:55 -0700 Subject: [PATCH 3/4] Fix what the second review of #549 found - release_log: the scope sentence names the comma forms it covers and drops an example whose reading moved between 2.3 and 2.4 for other reasons. - rules.md#C1: the Accepted entry's reason is the marker's position in its part (opening it, or ending it), not a name word ahead of it; M2 joins C1's interacts:. - expected_since_1.4.0.toml: drop a stale count of the alternation. Co-Authored-By: Claude Opus 5.5 --- docs/design/rules.md | 8 ++++---- docs/release_log.rst | 2 +- tools/differential/expected_since_1.4.0.toml | 2 +- 3 files changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/design/rules.md b/docs/design/rules.md index 2033de63..fe45bc2e 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1985,9 +1985,9 @@ C1. Rationale: a credential run after the comma means the name is in name — which is why this clause carries no example of its own. Accepted: a maiden marker with a declared delimiter straight before or after it takes no clause, the same words written with - a comma there taking none either — a marker needs a name word - ahead of it in its own part, and something to take behind it - (M2). Releases 2.0 through 2.3 read past the delimiter and took + a comma there taking none either: a marker opening a trailing + part has nothing ahead of it to follow, and one ending a part has + nothing behind it to take (M2). Releases 2.0 through 2.3 read past the delimiter and took the clause; the comma reading is accepted as the cost of one rule for a declared separator. "Smith, John, MD - née Jones Smith" extra_suffix_delimiters-dash → suffix="MD, née Jones Smith" @@ -2008,7 +2008,7 @@ C1. Rationale: a credential run after the comma means the name is in V` reads the suffix and `Smith, John PhD I.` continues the run, while adding a suffix comma after either turns that same letter into the middle initial. - history: decisions.md#C1 · interacts: H1, H2, P1, P2, P3, P5, P6, W3, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py + history: decisions.md#C1 · interacts: H1, H2, P1, P2, P3, P5, P6, W3, S2, S3, M2 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. diff --git a/docs/release_log.rst b/docs/release_log.rst index b6f14fde..59cdd1cf 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -70,7 +70,7 @@ Release Log - **Fix a katakana name typed with a separate voicing mark being read family-first.** ``HumanName("ア゙イ タロウ")`` gives first ``ア゙イ``, last ``タロウ``, as ``アイ タロウ`` does, where 2.1 through 2.3 gave last ``ア゙イ``, first ``タロウ``. A dakuten or handakuten after a kana that has no precomposed voiced form (``ア゙``, ``ン゙``), or the spacing ``゛`` and ``゜``, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: ``マイケル ア゙イ`` gives first ``マイケル`` where it gave last ``マイケル``. Hiragana names and kanji-and-kana names with such a mark read as before. See the ``W4`` entry of ``docs/design/decisions.md`` (closes #596) - - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. **A maiden marker with the delimiter straight before or after it no longer takes a maiden name**, as it takes none after a typed comma: ``Smith, John, MD - née Jones Smith`` gives suffix ``MD, née Jones Smith`` and no maiden, where 2.0 through 2.3 gave maiden ``Jones Smith``. Only a part read wholly as suffixes after the name's commas is affected; before the first comma, or in the given-name part of ``Family, Given``, the delimiter is still a word, as in 1.4.0 (``John Smith, PhD née Puig - i Soler`` keeps maiden ``Puig - i Soler``). The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) + - **Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.** ``HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix`` is ``PhD, and MD``, where 2.0 through 2.3 gave ``PhD - and MD``; ``Smith, John, Puig - y Soler`` gives ``Puig, y Soler`` where they gave ``Puig - y Soler``. Both are 1.4.0's answers again. A delimiter declared through ``suffix_delimiter`` or ``Policy(extra_suffix_delimiters=...)`` now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. **A maiden marker with the delimiter straight before or after it no longer takes a maiden name**, as it takes none after a typed comma: ``Smith, John, MD - née Jones Smith`` gives suffix ``MD, née Jones Smith`` and no maiden, where 2.0 through 2.3 gave maiden ``Jones Smith``. Only the parts after a suffix comma are affected -- the credentials of ``Name, PhD`` or of ``Family, Given, PhD``; in the name before the first comma, and in the given-name part of ``Family, Given``, the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #549) - **Fix halfwidth corner brackets not being read as a nickname.** ``HumanName("山田 「タロー」 タロウ")`` gives nickname ``タロー``, last ``山田``, first ``タロウ``, where every release gave middle ``「タロー」`` and first ``山田`` (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth ``「」`` are the corner brackets of legacy JIS X 0201 data, the same punctuation as ``「」``, and are now a default nickname pair in both APIs: ``DEFAULT_NICKNAME_DELIMITERS`` gains ``("「", "」")`` and the 1.x ``nickname_delimiters`` gains the key ``halfwidth_corner_brackets``. They are not limited to Japanese text: ``John 「Jack」 Smith`` gives nickname ``Jack`` where it gave middle ``「Jack」``. A ``Constants`` restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the ``N1`` entry of ``docs/design/decisions.md`` (closes #597) diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 77282d3c..13ea8ee2 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4699,7 +4699,7 @@ issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the su # about a FAMILY comma, where v1 left a middle name behind and the # field list says so. This tail leaves v1 a suffix and nothing else. # -# Literal, two names, for the reason those rules give: the shape would +# Literal names, for the reason those rules give: the shape would # be "a maiden marker and a bare letter", which reaches every clause # name in the corpora. 'Smith, John, PhD née Puig - i Soler' joined # this alternation with #538 (2026-09-26): it is rules.md#M2's second From d180516580d82ac07958400305b27b738f180f28 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 17:51:55 -0700 Subject: [PATCH 4/4] Fix what the PR-review-toolkit pass on #549 found MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - _maiden_take's walk loses the scaffolding the `seen` list left: `j < len(pieces)` was implied by `j < trailing`, and `k` aliased `j`. - Tests: a case row pins that a tail part holding only its core keeps it (the guard no test pinned before or after #549; recorded negative control 'PhD'); a case row pins 'Smith, John, PhD née - Jones' at parse level; the comma-twin invariant compares the reported kinds too, less comma-structure, with its negative control unchanged at 102 of 129. - Prose: the M2 delimiter examples are no longer called live examples of the #538 fix in the five ledgers and the guard test; a misquoted C1 excerpt; the lone-core comment states the several-cores case; a standing table count; an example fragment that read wrong beside the tail form; _marker_run_pieces' docstring after the rewrite; a renamed test whose mutant site #549 deleted; reflowed lines. Co-Authored-By: Claude Opus 5.5 --- docs/design/rules.md | 14 ++--- nameparser/_pipeline/_group.py | 58 ++++++++++---------- tests/v2/cases.py | 22 ++++++++ tests/v2/pipeline/test_group.py | 15 ++--- tests/v2/pipeline/test_post_rules.py | 4 +- tests/v2/test_ledger_guards.py | 5 +- tests/v2/test_properties.py | 24 +++++--- tools/differential/expected_since_1.4.0.toml | 4 +- tools/differential/expected_since_2.0.0.toml | 8 +-- tools/differential/expected_since_2.1.0.toml | 8 +-- tools/differential/expected_since_2.2.0.toml | 8 +-- tools/differential/expected_since_2.3.0.toml | 8 +-- 12 files changed, 107 insertions(+), 71 deletions(-) diff --git a/docs/design/rules.md b/docs/design/rules.md index fe45bc2e..f2fbd5c7 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1860,10 +1860,10 @@ C1. Rationale: a credential run after the comma means the name is in comma would: the words on each side of it read exactly as the parts of the same text written with a comma in its place, so no join, maiden clause or connective reaches across it, and the - delimiter itself is dropped, unless it is the whole of its part. - Only a trailing part is parted: in - the part before the first comma, or in the given part of the - listing form, the delimiter is a word (the Accepted entry below). + delimiter itself is dropped, unless it stands alone as the whole + of its part. Only a trailing part is parted: in the part before + the first comma, or in the given part of the listing form, the + delimiter is a word (the Accepted entry below). "Smith, John" → family="Smith" "سلمان، محمد" → family="سلمان" "田中、太郎" → family="" @@ -1987,9 +1987,9 @@ C1. Rationale: a credential run after the comma means the name is in before or after it takes no clause, the same words written with a comma there taking none either: a marker opening a trailing part has nothing ahead of it to follow, and one ending a part has - nothing behind it to take (M2). Releases 2.0 through 2.3 read past the delimiter and took - the clause; the comma reading is accepted as the cost of one - rule for a declared separator. + nothing behind it to take (M2). Releases 2.0 through 2.3 read + past the delimiter and took the clause; the comma reading is + accepted as the cost of one rule for a declared separator. "Smith, John, MD - née Jones Smith" extra_suffix_delimiters-dash → suffix="MD, née Jones Smith" Accepted: a delimiter core the policy names (T1) is a word here, not structure — v1 applied the delimiter to the suffix-comma diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 12216112..31535fe9 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -10,11 +10,12 @@ P5 licence, #369) and Policy.extra_suffix_delimiters, whose delimiter-core tokens part a tail segment as a comma would and are dropped (v1 suffix_delimiter parity, #549) -- no other Policy field. -Policy.lenient_comma_suffixes left this list with #436: it reached here only through segment_suffix_reading, whose -render consumer was this stage's one-entry join and now lives in -post_rules. The v1 "derived titles/prefixes" -registration becomes piece_tags entries -- per-parse state that -dissolves with the state (v1 kept per-parse sets for the same reason). +Policy.lenient_comma_suffixes left this list with #436: it reached +here only through segment_suffix_reading, whose render consumer was +this stage's one-entry join and now lives in post_rules. The v1 +"derived titles/prefixes" registration becomes piece_tags entries -- +per-parse state that dissolves with the state (v1 kept per-parse sets +for the same reason). Implements rules P2, P3, P4 and M2, and the group half of M1 (#329: the marker dropped inside EXTRACTED maiden @@ -209,14 +210,14 @@ def _marker_run_pieces(pieces: Sequence[Sequence[int]], Each continuation is the NEXT piece, and that holds because classify REFUSES to tag a run whose tokens are not structurally contiguous. It is not a property of this walk, and - the reasons a token can be missing are wider than they look. No - join has run yet, so a piece is one token. A tail segment's - delimiter cores are cut out before grouping (#549), and a core - between two marker words would be a token between them, which - classify would not have tagged as a run. And `pieces` comes from a - SEGMENT, and segment keeps only - the tokens no stage has given a role, bucketed by the commas before - them -- so a run half inside a bracketed clause, or split across a + the reasons a token can be absent between two pieces are wider + than they look. No join has run yet, so a piece is one token. A + tail segment is cut at its delimiter cores and grouped part by + part (#549), and a core between two marker words would be a token + between them, which classify would not have tagged as a run. And + `pieces` comes from a SEGMENT, or a part of one, and segment keeps + only the tokens no stage has given a role, bucketed by the commas + before them -- so a run half inside a bracketed clause, or split across a structure comma, is one no segment holds whole. Walking cont tags without classify's refusal read a proper PREFIX of the phrase as the whole marker, and 'Anna z (domu) Nowak' lost its given name to @@ -921,8 +922,9 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # marker is never the name word on a link's left. No delimiter # core stands anywhere in the clause: a tail segment's cores are # cut out before this walk (#549), and a dash where no tail - # segment holds it is an ordinary word at either policy ('PhD née - # - i Jones' keeps maiden '- i Jones' configured or not). + # segment holds it is an ordinary word at either policy ('John PhD + # née - i Jones' keeps maiden '- i Jones' configured or not, where + # 'Smith, John, PhD née - i Jones' configured takes no clause). # `peel_start` is where assign's trailing run begins -- over the # pieces as written for a NONE reader (`trailing_start`'s whole # answer, read off the peel pair above rather than re-running it), @@ -946,14 +948,13 @@ def _maiden_take(pieces: Sequence[Sequence[int]], lo = m + run peel_start = (rest[peeled.names] if peeled.names < len(rest) else len(pieces)) - j = m + run + j = lo # The link exception's memo cell, filled inside the predicate on # the first connective it is asked about (see its docstring). beside: list[_Beside] = [] - while j < len(pieces) and j < trailing: - k = j - if (not is_suffix_piece(pieces[k], ptags[k], tokens) - or _link_joins_inside_the_clause(k, lo, peel_start, + while j < trailing: + if (not is_suffix_piece(pieces[j], ptags[j], tokens) + or _link_joins_inside_the_clause(j, lo, peel_start, pieces, ptags, tokens, beside)): j += 1 @@ -977,19 +978,19 @@ def _maiden_take(pieces: Sequence[Sequence[int]], # declines the clause outright, which gives nothing up. The # as-written peel is read only here, on a refused link. if (j > m + run and reads is not None - and _is_conj_piece(pieces[k], ptags[k], tokens)): + and _is_conj_piece(pieces[j], ptags[j], tokens)): as_written = peel_trailing(written, pieces, ptags, tokens, one_case) written_start = (written[as_written.names] if as_written.names < len(written) else len(pieces)) - if _link_joins_inside_the_clause(k, lo, written_start, + if _link_joins_inside_the_clause(j, lo, written_start, pieces, ptags, tokens, beside): - left = [i for i in range(len(pieces)) if i < m or i >= k] + left = [i for i in range(len(pieces)) if i < m or i >= j] view = [pieces[i] for i in left] view_tags = [ptags[i] for i in left] - at = left.index(k) + at = left.index(j) if not _release_reads_off(view, view_tags, tokens, at, len(view), at, reads, one_case, tail_follows=tail_follows): @@ -2031,9 +2032,10 @@ def group(state: ParseState) -> ParseState: # dropped (v1 expand_suffix_delimiter parity, #206), where # post_rules' entry pass reads them back as entry boundaries. # No join, maiden walk or link search reaches across a core, - # because none of them ever sees one. A segment that IS only - # its core keeps it (v1 expand() splits within a part, never - # erases a lone part). + # because none of them ever sees one. A one-token segment that + # is its core keeps it (v1 expand() splits within a part, never + # erases a lone part); a segment of several cores and nothing + # else drops them all, as the #206 drop did. # Accumulated in a list and frozen once per part: a tuple # extended per token is quadratic in the part's length, C-level # work no frame guard sees (#553's class; review of #549). @@ -2042,7 +2044,7 @@ def group(state: ParseState) -> ParseState: # written, the default policy's own 4.5x, where the tuple read # 8.0x. No clock guard: _PREFIXED_SHAPES parses with the default # policy and _POLICY_SHAPES repeats a unit with no prefix, and a - # tail needs both, so a row would mean a third table. + # tail needs both, so a row would mean a table of its own. parts: list[tuple[int, ...]] = [seg] if tail and cores and len(seg) > 1: parts = [] diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 5b1cbcec..05919e80 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -6181,6 +6181,28 @@ def _check_cjk_shape_purity(self) -> None: "expand_suffix_delimiter reading, where 2.0.0 through " "2.3.0 merged the core into the joined piece and " "rendered 'PhD - and MD'"), + Case("suffix_delimiter_lone_core_part_is_kept", + "Smith, John, PhD, -", + {"given": "John", "family": "Smith", "suffix": "PhD, -"}, + policy=_SD, + notes="a tail part that is nothing but its one core keeps it " + "(v1 expand() split within a part and never erased a " + "lone part, #206): the core is cut and dropped only " + "where the part holds something else. Pairs with " + "suffix_delimiter_core_parts_a_connective_join. " + "RECORDED NEGATIVE CONTROL: with the one-token guard " + "removed this reads suffix 'PhD' (review of #549, " + "which found no test pinned the guard before or after " + "the change)"), + Case("suffix_delimiter_core_after_a_marker_takes_no_clause", + "Smith, John, PhD née - Jones", + {"given": "John", "family": "Smith", "suffix": "PhD née, Jones"}, + policy=_SD, + ambiguities=("comma-structure",), + notes="the marker ends its part with nothing behind it, as it " + "does before a typed comma, so it takes no clause " + "(rules.md C1's Accepted entry, #549). 2.0.0 through " + "2.3.0 read maiden 'Jones', stepping past the core"), Case("suffix_delimiter_core_that_survives_is_a_boundary_too", "Smith, MD - PhD - FACS", {"title": "MD", "given": "-", "middle": "-", "family": "Smith", diff --git a/tests/v2/pipeline/test_group.py b/tests/v2/pipeline/test_group.py index 4c534a1c..dc22180d 100644 --- a/tests/v2/pipeline/test_group.py +++ b/tests/v2/pipeline/test_group.py @@ -498,11 +498,12 @@ def test_a_delimiter_core_in_a_suffix_tail_is_not_maiden_text() -> None: assert not any(t.role is Role.MAIDEN for t in twin.tokens) -def test_the_walk_peels_past_a_trailing_core() -> None: +def test_a_trailing_core_is_cut_before_the_walk() -> None: # A core standing last is cut off before the walk (#549), so it # does not make the numeral "not last": the V is the suffix and the - # marker takes 'Jones' alone (#424; the test review's surviving - # mutant). + # marker takes 'Jones' alone. (#424's test review found a surviving + # mutant in the walk's core skip; #549 deleted the skip, so the + # site that mutant lived in is gone.) out = _grouped("Smith, John, PhD née Jones V -", policy=_DASH) assert [t.text for t in out.tokens if t.role is Role.MAIDEN] == [ "Jones"] @@ -1821,10 +1822,10 @@ def test_a_core_straight_after_the_marker_leaves_it_nothing_to_take( def test_a_core_ends_a_maiden_clause_as_a_comma_would() -> None: - """rules.md's C1: a declared delimiter in a trailing part "parts it as - a comma would" (#549), so a maiden clause ends at a core exactly as - it ends at the comma written in its place, and a link beside the - core has no name word on that side. + """rules.md's C1: "a delimiter the policy declares parts a trailing + suffix part as a comma would" (#549), so a maiden clause ends at a + core exactly as it ends at the comma written in its place, and a + link beside the core has no name word on that side. This reverses #538, which read the clause as the text written WITHOUT the core: 'Puig - i Soler' kept maiden 'Puig i Soler' diff --git a/tests/v2/pipeline/test_post_rules.py b/tests/v2/pipeline/test_post_rules.py index 83ad45f4..d3abeb34 100644 --- a/tests/v2/pipeline/test_post_rules.py +++ b/tests/v2/pipeline/test_post_rules.py @@ -1014,8 +1014,8 @@ def test_a_bare_marker_maiden_clause_does_not_part_the_run() -> None: def test_a_marker_the_policy_names_as_a_delimiter_parts_the_run() -> None: """Where the two sites' gates differ, and the one case that reaches it. group drops a core only on a `tail` segment, cutting - the segment there (#549); this pass reads `delimiter_cores` whole and asks by - TEXT, so a dropped token the policy names as a delimiter parts the + the segment there (#549); this pass reads `delimiter_cores` whole + and asks by TEXT, so a dropped token the policy names as a delimiter parts the run whatever dropped it. A maiden marker the policy ALSO lists is the reachable case, and it parts by the policy's own declaration (measured 2026-09-06).""" diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 5b105a6e..0de2f12b 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -2319,7 +2319,8 @@ class _LatinCopy(NamedTuple): "Jane Doe nee MA PhD"}), frozenset({"Jane Doe \\(nee Smith MA\\)", "Jane Doe \\(nee Smith Ma\\)", "Jane Doe \\(nee Smith\\) MA"}), - # rules.md#M2's two live examples of the delimiter-core fix (#538), + # rules.md#M2's two delimiter examples (added with #538, read by + # the comma reading since #549), # and rules.md#C1's Accepted example of a marker beside one (#549), # one alternative per corpus name. A list of names, not a copy of # any wordlist: what selects them is the doc's own choice of @@ -2999,7 +3000,7 @@ class _LatinCopy(NamedTuple): # vocabulary decides; a member spelled as the shape (a bare letter # after a maiden marker) would reach every clause name in the # corpora and pre-excuse the readings the walk must refuse. The - # last two are rules.md#M2's two live examples of the #538 fix, + # last two are rules.md#M2's two delimiter examples (added with #538), # which the corpus parses at the DEFAULT policy and so for this # rule's own sentence rather than for #538's. frozenset({"Doe, Jane nee Puig i Soler", "Jane Doe nee Puig i Soler", diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index a3ef45c2..b9dfb775 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -3227,6 +3227,15 @@ def test_the_decomposed_initials_walk_can_fail( "PhD MD FACS", "Puig y Soler") +def _twin_record(name: ParsedName) -> tuple[object, ...]: + """The role fields and the reported kinds, less `comma-structure`: + C2 reports it once per extra comma part, so the typed comma adds + one the core does not, by construction rather than by reading.""" + kinds = sorted(a.kind for a in name.ambiguities + if a.kind is not AmbiguityKind.COMMA_STRUCTURE) + return (name.as_dict(), kinds) + + def _comma_twin_findings(parser: Parser) -> tuple[list[str], int]: """Each body with a ' - ' core in one interior gap, against the same text with a typed comma there, all seven fields.""" @@ -3238,8 +3247,8 @@ def _comma_twin_findings(parser: Parser) -> tuple[list[str], int]: for gap in range(1, len(parts)): left, right = " ".join(parts[:gap]), " ".join(parts[gap:]) total += 1 - got = parser.parse(f"{head} {left} - {right}").as_dict() - want = parser.parse(f"{head} {left}, {right}").as_dict() + got = _twin_record(parser.parse(f"{head} {left} - {right}")) + want = _twin_record(parser.parse(f"{head} {left}, {right}")) if got != want: failures.append(f"{head} {left} - {right}: " f"{got} != {want}") @@ -3248,11 +3257,12 @@ def _comma_twin_findings(parser: Parser) -> tuple[list[str], int]: def test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin() -> None: """rules.md's C1: "a delimiter the policy declares parts a trailing - suffix part as a comma would" (#549). Every field of a name with a - declared core standing in a part after the suffix comma equals the - field of the same text written with a comma in the core's place -- - the maiden clause, the connective join and the entry boundary - alike. The heads put the core in a TRAILING part only: before the + suffix part as a comma would" (#549). Every role field of a name + with a declared core standing in a part after the suffix comma, and + every kind it reports but `comma-structure` (`_twin_record`), + equals that of the same text written with a comma in the core's + place -- the maiden clause, the connective join and the entry + boundary alike. The heads put the core in a TRAILING part only: before the first comma, or in the given part of the listing form, the core is a word (C1's Accepted entry), and its comma twin a different structure. diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 13ea8ee2..7fd0bf33 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4682,8 +4682,8 @@ fields = ["family", "maiden", "middle", "suffix"] issue = "fix(#274/#397) a maiden clause inside a suffix-comma tail leaves the suffix field, link and all" # 'Smith, John, PhD née Puig Mr. - i Soler', which arrives from # rules.md#M2 rather than from a report: it is one of rules.md#M2's -# live examples of the #538 fix (under a configured ' - ' it reads -# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# delimiter examples, added with #538 (under a configured ' - ' it +# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, # so what the gate sees is the DEFAULT facade reading, where the dash # is an ordinary name word like any other. # diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 3587c28b..71fbfc50 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3582,15 +3582,15 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # # The THIRD name arrives from rules.md#M2 rather than from a report: # 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's -# live examples of the #538 fix (under a configured ' - ' it reads -# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# delimiter examples, added with #538 (under a configured ' - ' it +# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, # so the gate parses it with the DEFAULT facade, where the dash is an # ordinary name word: the link has one on each side and the clause # keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). -# That is this rule's own sentence and no part of #538, whose reading -# needs the delimiter declared; {maiden, suffix} is a subset of the +# That is this rule's own sentence and no part of #538 or #549, whose +# readings need the delimiter declared; {maiden, suffix} is a subset of the # union above. # # Literal-anchored to the four. The shape is "a one-letter connective diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 753faa83..5bb58411 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3494,15 +3494,15 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # # The THIRD name arrives from rules.md#M2 rather than from a report: # 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's -# live examples of the #538 fix (under a configured ' - ' it reads -# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# delimiter examples, added with #538 (under a configured ' - ' it +# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, # so the gate parses it with the DEFAULT facade, where the dash is an # ordinary name word: the link has one on each side and the clause # keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). -# That is this rule's own sentence and no part of #538, whose reading -# needs the delimiter declared; {maiden, suffix} is a subset of the +# That is this rule's own sentence and no part of #538 or #549, whose +# readings need the delimiter declared; {maiden, suffix} is a subset of the # union above. # # Literal-anchored to the four. The shape is "a one-letter connective diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index da87f178..41ff5d07 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1892,15 +1892,15 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # # The THIRD name arrives from rules.md#M2 rather than from a report: # 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's -# live examples of the #538 fix (under a configured ' - ' it reads -# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# delimiter examples, added with #538 (under a configured ' - ' it +# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, # so the gate parses it with the DEFAULT facade, where the dash is an # ordinary name word: the link has one on each side and the clause # keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). -# That is this rule's own sentence and no part of #538, whose reading -# needs the delimiter declared; {maiden, suffix} is a subset of the +# That is this rule's own sentence and no part of #538 or #549, whose +# readings need the delimiter declared; {maiden, suffix} is a subset of the # union above. # # Literal-anchored to the four. The shape is "a one-letter connective diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 7ba1a21f..11ad2ab6 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1142,15 +1142,15 @@ issue = "fix(#397) a link inside a maiden clause stays in the birth name" # # The THIRD name arrives from rules.md#M2 rather than from a report: # 'Smith, John, PhD née Puig Mr. - i Soler' is one of rules.md#M2's -# live examples of the #538 fix (under a configured ' - ' it reads -# maiden 'Puig Mr.'). This corpus does not configure that delimiter, +# delimiter examples, added with #538 (under a configured ' - ' it +# reads maiden 'Puig Mr.', by the comma reading since #549). This corpus does not configure that delimiter, # so the gate parses it with the DEFAULT facade, where the dash is an # ordinary name word: the link has one on each side and the clause # keeps 'i Soler' where this baseline left it in the suffix: suffix # 'PhD i Soler' -> 'PhD', maiden 'Puig Mr. -' -> 'Puig Mr. - i Soler' # (measured 2026-09-22 at all four 2.x baselines, identical at each). -# That is this rule's own sentence and no part of #538, whose reading -# needs the delimiter declared; {maiden, suffix} is a subset of the +# That is this rule's own sentence and no part of #538 or #549, whose +# readings need the delimiter declared; {maiden, suffix} is a subset of the # union above. # # Literal-anchored to the four. The shape is "a one-letter connective