Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,7 +104,7 @@ jobs:
with:
ref: ${{ github.event.pull_request.base.sha }} # trusted base; PR head is fetched as review data
fetch-depth: 0
- uses: 0xPolygon/codegenie@v0.6.1
- uses: 0xPolygon/codegenie@v0.6.3
with:
# Works with any model!
model: "openrouter/deepseek/deepseek-v4.1-flash:max"
Expand Down Expand Up @@ -272,10 +272,18 @@ The planning check rejects degraded plans even when every hunk was reviewed. The

SVG files are skipped by default. Set `[review] skipSvgReview = false` in `codegenie.toml`, or run `codegenie review --no-skip-svg-review`, to include them subject to other exclusion rules. `--skip-svg-review` enables the skip explicitly. This controls changed-file review; repository evidence searches can still find SVG content and disclose oversized matches they omit.

Local investigation budgets are soft targets, with hard ceilings of **2×** the target for aggregate tool calls, investigation rounds, and result characters. This applies to packet review, system review, and verification. At normal depth with `budgetBoost = 1`, investigate packets target 20 calls / 6 rounds / 48,000 characters for deep coverage, 6 / 2 / 12,000 for normal coverage, and 4 / 2 / 4,000 for light coverage. System review targets 6 / 2 / 12,000; verification targets 8 / 3 / 32,000. Depth and budget scaling apply before deriving the ceilings. Zero budgets remain disabled. Existing source extensions are subsumed by this allowance and cannot add capacity beyond it. Per-result caps, explicit spending limits, provider context limits, stage deadlines, and repair budgets are unchanged.

Before the first request and after each repository-tool batch, models receive remaining soft targets and a nudge to finish or resolve concrete remaining questions when a target is reached. Traces record soft targets, hard ceilings, and remaining capacity in `tool_budget_initial` / `tool_budget_remaining`, plus a single `tool_budget_soft_target_reached` event per investigation. Crossing a target alone is not a refusal or an incomplete-review diagnostic. Executed calls and cache hits each count; invalid arguments and pre-execution refusals do not consume those calls, but a separate refusal limit and the round ceiling bound repeated invalid requests. Source allocations are guidance within the shared character allowance. Discovery results retain their per-result cap of 4,000 characters (2,000 for light investigations), subject to existing budget scaling. Searches and file lists keep whole entries and disclose omissions. Invalid raw tool arguments receive correction feedback before SDK coercion. `read_range` requires both inclusive 1-based bounds; requests entirely beyond EOF return empty text, not a different line.

For formats without a syntax adapter, outlines explicitly mark symbol extraction as unavailable. Small files include complete source text; larger files provide a `read_range` hint. Empty symbol arrays do not establish that definitions are absent. Complete source supplied through an outline can also serve as evidence during attention reconciliation.

Composition uses the next lower supported reasoning level by default, including retries: for a model supporting `low`, `high`, and `max`, `max` becomes `high`. Set `[review] compositionReasoningStepDown = false` in `codegenie.toml` to keep the configured review reasoning level for composition. The lowest supported level stays unchanged; models without advertised reasoning levels retain the configured behavior. Override this per run with `codegenie review --composition-reasoning-step-down` or `--no-composition-reasoning-step-down`. Omitting both flags preserves the configuration, which defaults to `true`. Investigation and verification keep their configured reasoning; traces record configured and selected levels. Structured-output repairs continue to use the model’s lowest supported reasoning level. Each composition attempt has a 300-second deadline, with at most one retry. The outer composition deadline is 780 seconds (two attempts plus the shared 180-second repair allowance); overall review cancellation still takes precedence. Repair attempts share that 180-second allowance, rather than receiving 180 seconds each.

Composition validates source references before acceptance. It locally removes repeated known references and misplaced references already correctly accounted for in the same finding, records those removals, and validates the whole result. Remaining attribution errors receive bounded repairs in a fresh context with exact field paths and source inventories. Attribution patches replace only permitted reference lists; finding order and prose stay intact, and the assembled report must pass full validation. If a recommendation lacks support, a bounded composition repair may instead omit or rewrite that advice section while preserving the diagnosis and retaining its original sources. Reports consolidate identical evidence and keep additional verbatim evidence and caveats in expandable sections. If synthesis fails, the report identifies its source-based presentation and retains distinct contributions. `stages/10-composition/composition-sources.json` records all inputs and dispositions; references establish attribution, not proof of semantic equivalence. Verification distinguishes essential missing proof from secondary uncertainty: unresolved hypotheses remain visible under human attention, while established defects may still have uncertainty about severity.

Verification receives up to 6,000 characters of relevant source already collected by completed packets, with file, revision and tool provenance. Attention reconciliation also matches concern file/symbol metadata when selecting source evidence; selection does not establish that a question is answered.

Verification assesses proposed fixes and tests independently from the defect, using the existing optional assessment fields and investigation budget. The verifier selects a concrete remedy that preserves the original caller requirement and checks its test against both the defect and a weakened guarantee. In verifier submissions, `suggestionText` may be omitted: the harness binds the assessment to the named final suggestion before retaining a repair draft. Explicit text mismatches still lose support; changing a suggestion alone cannot transfer a retained assessment to it. Support requires evidence for the observable requirement; tests should reject weakened guarantees without excluding other valid implementations. Prominent fix/test sections and structured recommendation fields contain only supported current proposals. Unverified, incompatible and replaced proposals remain in expandable provenance. The composer is instructed to keep summaries diagnosis-focused and advice out of impact/verification prose; source validation checks attribution and eligibility, not the semantic correctness of arbitrary prose.

## Development
Expand Down
10 changes: 10 additions & 0 deletions evals/synthetics/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,16 @@ findings in a document remain separate. The latter deliberately reverses PR #28'
file-wide merging tradeoff. Existing code fingerprint tests, including executable
examples under `docs/`, retain symbol-based identity.

The text-tool investigation suite uses temporary committed Git repositories and
production tool schemas, wrappers and Git grep. It covers `.ridl` and arbitrary
unsupported formats: small outlines, locating a late section in a large file,
exact bounded reads, base/head isolation from dirty files, POSIX regexes, shared
glob semantics, invalid-query correction, text-only mentions, and bounded result
packing followed by exact source reads. Boundary cases (missing/fractional bounds,
EOF, empty/missing files, UTF-8, CRLF and truncation) live in
`tests/text-tool-contracts.test.ts`; runner tests verify invalid arguments do not
consume source-call slots or masquerade as provider failures.

These evaluate harness behavior, not whether a model identifies or phrases a
finding consistently. Use `codegenie eval --eval-dir ...` for live model comparisons.
Add additional `*.test.ts` scenarios here as harness failure patterns are found.
Loading
Loading