Part a trailing suffix part at a declared suffix delimiter as a comma would (#549) - #600
Merged
Merged
Conversation
…549) A delimiter declared through extra_suffix_delimiters / suffix_delimiter now separates a tail segment exactly as a comma typed in its place: group cuts the segment at its cores before grouping, groups each part as the comma twin's segment would be grouped, and drops the cores. No join, maiden walk or link search can reach across a core, so the #538 stepping machinery goes: the cores parameter on four functions, _maiden_take's seen index list, the skip arguments of peel_walk and trailing_start, and the post-join #206 drop block. 'Smith, John, PhD - and MD' reads suffix 'PhD, and MD' (1.4.0's answer; 2.0-2.3 gave 'PhD - and MD'), and 'Smith, John, PhD née Puig - i Soler' reads maiden 'Puig', suffix 'PhD, i Soler' again, as 2.3.0 did -- the boundary reading #538 declined, taken now. The default policy declares no delimiter; the differential gate exits 0 at all five baselines. One call per parse fewer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- group's tail split accumulated a tuple per token, quadratic in a part's length; build a list and freeze it once per part (4.7x for 4x the input, against 8.0x). - rules.md#M2's statement still described #538's stepping; it now defers to C1. P3's separator sentence is scoped to trailing parts, and both rules list C1 in interacts:. C1 states the lone-core exception and records, as Accepted, that a maiden marker beside a declared delimiter takes no clause -- the comma twin's reading, and the larger cost by count. - decisions.md#C1's #549 entry: the maiden-loss class in WHAT MOVES and in the weighing; the Jr., head figure (752 of 2,100, not 973 of 3,310); the ' / ' policy moving the same 2,624; the twin compared on role fields, comma-structure counts differing; the #418 repair added to SUPERSEDES. - release_log: the maiden-loss consequence, and the scope stated as the part read wholly as suffixes. - Stale prose in test_post_rules and a renamed test_group test. - The C1 example enters the rules corpus and the 1.4.0 ledger. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- release_log: the scope sentence names the comma forms it covers and drops an example whose reading moved between 2.3 and 2.4 for other reasons. - rules.md#C1: the Accepted entry's reason is the marker's position in its part (opening it, or ending it), not a name word ahead of it; M2 joins C1's interacts:. - expected_since_1.4.0.toml: drop a stale count of the alternation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #600 +/- ##
=======================================
Coverage 98.98% 98.98%
=======================================
Files 45 45
Lines 4235 4238 +3
=======================================
+ Hits 4192 4195 +3
Misses 43 43 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
- _maiden_take's walk loses the scaffolding the `seen` list left: `j < len(pieces)` was implied by `j < trailing`, and `k` aliased `j`. - Tests: a case row pins that a tail part holding only its core keeps it (the guard no test pinned before or after #549; recorded negative control 'PhD'); a case row pins 'Smith, John, PhD née - Jones' at parse level; the comma-twin invariant compares the reported kinds too, less comma-structure, with its negative control unchanged at 102 of 129. - Prose: the M2 delimiter examples are no longer called live examples of the #538 fix in the five ledgers and the guard test; a misquoted C1 excerpt; the lone-core comment states the several-cores case; a standing table count; an example fragment that read wrong beside the tail form; _marker_run_pieces' docstring after the rewrite; a renamed test whose mutant site #549 deleted; reflowed lines. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #549.
What changes
A delimiter declared through
Policy(extra_suffix_delimiters=...)orHumanName(suffix_delimiter=...)now splits a trailing suffix part exactly where a typed comma would.group()cuts a tail segment at each delimiter before any grouping, groups each piece the way the comma version's own segment would be grouped, and drops the delimiter. No join, maiden walk or link search can reach across a delimiter, because none of them ever sees one.This reverses #538's "skip" reading, where the connective search stepped over a delimiter as if it weren't written. The skip reading was the source of #549: a delimiter standing alone was dropped like a comma, while one next to a connective was merged into the joined phrase like a word.
Because delimiters never reach grouping, the machinery that stepped over them is deleted:
coresparameter on four functions;_maiden_take'sseenindex list;skiparguments ofpeel_walkandtrailing_start;The result is one call per parse fewer (369 vs 370,
tools/perf/call_count.py, py3.11).Behavior, with
suffix_delimiter=" - "(each checked against the released wheels)Smith, John, PhD - and MDPhD, and MDPhD - and MDPhD, and MDSmith, John, Puig - y SolerPuig, y SolerPuig - y SolerPuig, y SolerSmith, John, PhD née Puig - i SolerPuig(2.3.0)Puig, suffixPhD, i SolerSmith, John, MD - née Jones SmithMD - née Jones SmithJones SmithMD, née Jones Smith, no maidenThe last row is an accepted cost (decided 2026-10-03). A maiden marker with the delimiter straight before or after it takes no clause, as it takes none after a typed comma. By count it is the larger move: every one of the 240 maiden movers in the generated set loses the maiden. Recorded as Accepted in
rules.md#C1, in bold in the release note, and pinned by the case rowsuffix_delimiter_core_after_a_marker_takes_no_clause.At the default policy nothing changes, since no delimiter is declared. Outside a trailing suffix part (before the first comma, or in the given part of
Family, Given), the delimiter is still a word, as in 1.4.0.Measurements (2026-10-03, branch vs 061f02d)
The full figures and the recompute recipe are in the new
decisions.md#C1entry.-delimiter: 2,624 parses move (suffix only 2,384; suffix + maiden 240). No other field moves and no ambiguity report moves./delimiter: the same 2,624 parses.anddelimiter: 2,528 parses.Smith, John,andSmith, John, Jr.,heads: 752 of 2,100 disagree each → 0.John Smith,head: 1,956 of 2,100, unchanged. That head is out of scope, because there a typed comma changes how the name is split into parts.Tests and docs
test_a_delimiter_core_in_a_tail_reads_as_its_comma_twin(129 texts). Its recorded negative control is 102 disagreements at 061f02d. It replaces Should a suffix delimiter inside a maiden clause count as a name word beside a link?Smith, John, PhD née Puig Mr. - i Solerkeepsi Solerin the maiden name #538's "reads as if not written" invariant.Smith, John, PhD née Puig Mr. - i Solerkeepsi Solerin the maiden name #538 tests intest_group.pyare rewritten for the comma reading, with both earlier answers recorded as negative controls. There's a new case row,Doe, John, PhD - and MD.rules.md:decisions.md: a C1 entry, plus pointers marking Should a suffix delimiter inside a maiden clause count as a name word beside a link?Smith, John, PhD née Puig Mr. - i Solerkeepsi Solerin the maiden name #538's entry andjuan y garcia nee jonesloses the family name — P3's word count includes the words the maiden name takes away #418's repair as superseded.release_log.rst: a new Should a declared suffix delimiter inside a joined connective run be dropped?Smith, John, Puig - i Solerkeeps the-in the suffix #549 bullet, and theJosep Carod i Rovirareads familyRovira— Catalan and Polish link surnames withi, which is not a conjunction #397/Should a suffix delimiter inside a maiden clause count as a name word beside a link?Smith, John, PhD née Puig Mr. - i Solerkeepsi Solerin the maiden name #538 bullet corrected.Review
Four passes: design docs, code, a second round on the first fix commit, then
/pr-review-toolkit:review-pr(code, tests, comments). None found a correctness bug. Over 517k parses on both trees, every field that moved equals what the typed-comma version gives, and every recorded negative control reproduced._maiden_take, and doc errors (stale rule statements, a wrong figure, misquoted excerpts, "Should a suffix delimiter inside a maiden clause count as a name word beside a link?Smith, John, PhD née Puig Mr. - i Solerkeepsi Solerin the maiden name #538 fix" wording in five ledgers).PhD née - Jones. The comma-comparison grid now also compares reported ambiguity kinds; its negative control is unchanged at 102 of 129.🤖 Generated with Claude Code