Skip to content

v1.87.6.0 Release gstack 1.87.6 - #2898

Open
garrytan wants to merge 417 commits into
mainfrom
nouakchott
Open

garrytan wants to merge 417 commits into
mainfrom
nouakchott

Conversation

@garrytan

@garrytan garrytan commented Sep 18, 2026

Copy link
Copy Markdown
Owner

Summary

  • Release gstack 1.87.6.0 / package 1.87.6.
  • Tighten plan-review gates so saved decisions, report read-backs, and publication checks remain ordered across permissions and host routing.
  • Make coverage audits read concrete source/test files before diagramming, and clarify plan-eng targeted audit timing/report ordering.
  • Refresh generated skills and Codex/Factory ship goldens.

Validation

  • bun run test via bin/gstack-evidence: all 6 shards passed, evidence /home/vercel-sandbox/.gstack/projects/garrytan-gstack/logs/2026-09-18T00-31-12-399Z-tests-3495771-d2d532ba.log.
  • bun run build: passed.
  • bun run gen:skill-docs --dry-run: fresh.
  • bun run gen:skill-docs --host codex --dry-run: fresh.
  • Adjacent free sweep: 867 passed across 7 files.
  • Paid gate plan-eng-coverage-audit: passed, eval /home/vercel-sandbox/.gstack/projects/garrytan-gstack/evals/1.87.6-nouakchott-e2e-2026-09-18-0022.json.
  • Paid judge plan-eng-review/SKILL.md sections: passed, eval /home/vercel-sandbox/.gstack/projects/garrytan-gstack/evals/1.87.6-nouakchott-llm-judge-2026-09-18-0023.json.

Known warning: generated ship skills still exceed the 160KB token-ceiling warning on several host renders; parity/golden tests pass with the refreshed output.


Open workspace in Conductor

garrytan and others added 30 commits September 10, 2026 21:48
Keep CEO approach, engineering scope and outside-review choices from approving independent remedies together. Carry declared contracts through DX polish and resolve new gaps before editing the plan. Regenerate every host and retain existing stop boundaries.

Validation: 654 focused tests passed across nine files; all-host generation passed. Full free and periodic validation pending.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
Align the Design review philosophy and rating recipe with its section protocol: resolve one proposed fix, then apply only that approved decision and retain honest scores for declined fixes.

Validation: 469 focused tests passed across four files; all-host generation passed.

Co-Authored-By: OpenAI Codex <noreply@openai.com>
Match owned completion hooks to submitted choices, reject conflicting or late answers, and retain bounded failure evidence.
Require the selected mode acknowledgement, a completed follow-up question, and its current decoded display while preserving existing posture assertions.
Preserve explicit AUTO decisions, apply native single-select defaults, and bind complete cropped questions and report permissions to their owned requests. Require seeded review findings instead of crediting setup menus.

Keep captured failure controls and additive selection dependencies. The integrated candidate passed 3,099 focused tests across 65 files; affected paid validation remains required before publication.
Preserve complete option rows when native descriptions wrap, retain current
owned Write arguments before journal publication, and keep engineering and
DX answers within their declared fixture interfaces. Add captured free
regressions without increasing model budgets or relaxing completion checks.
Guard owned methodology reads and reviewer dispatches, detect complete driver
loads through Bash, and distinguish report-only edits from implementation
changes. Follow authenticated native UUID ancestry when journal writes arrive
out of order and verify earlier native content for cached phase reads.

Keep current close acknowledgment and parent publication in order, require CEO
entry before later phases, and register captured failure regressions.
Match complete native Edit panes and truncated question borders, reject stderr close before EOF, and stop the CEO split fixture once its acknowledged scope decisions are collected. Keep semantic validation, process failures, report requirements, and absolute deadlines authoritative.

Add captured-event and real-process regressions with selection dependencies. Focused checks pass; final integrated paid and full-suite acceptance remain pending.
Recover missed native UUID ancestry through the existing strict graph while preserving ordinary event order and legacy scoping. Bind publication hooks to Claude's original project directory while retaining current cwd for requested file paths.

Captured public-event regressions, existing caller checks, and a pinned native CLI loopback verify both fixes. Preserve failed attempts and require fresh paid and final full-suite acceptance.
@trunk-io

trunk-io Bot commented Sep 18, 2026

Copy link
Copy Markdown

Merging to main in this repository is managed by Trunk.

  • To merge this pull request, check the box to the left or comment /trunk merge below.

After your PR is submitted to the merge queue, this comment will be automatically updated with its status. If the PR fails, failure details will also be posted here

@github-actions github-actions Bot changed the title Release gstack 1.87.6 v1.87.6.0 Release gstack 1.87.6 Sep 18, 2026
@github-actions

Copy link
Copy Markdown

E2E Evals: ❌ FAIL

97/99 tests passed | $21.96 total cost | reconcile exit: 1 | ⚠ 4 flaky pass(es) — recorded, not blocking

Shard Result Status Cost
e2e-browse/skill-e2e-bws 7/7 $0.57
e2e/skill-e2e-coverage-audit 2/2 $0.81
e2e-cso/skill-e2e-cso 1/1 $0.9
e2e-deploy/skill-e2e-deploy 6/6 $3.1
e2e-design/skill-e2e-design 3/3 $0.88
e2e-docsync-spawned/skill-e2e-docsync-spawned 1/1 $0.21
e2e-hermetic/skill-e2e-hermetic-canary 2/2 $0.03
e2e-learnings/skill-e2e-learnings 1/1 $0.22
e2e-opus-47/skill-e2e-opus-47 6/6 $0.67
e2e-plan-tune-cathedral/skill-e2e-plan-tune-cathedral 5/5 $0
e2e-plan-tune/skill-e2e-plan-tune 1/1 $0.41
e2e/skill-e2e-plan 7/7 ✅⚠ $2.84
e2e-qa-workflow/skill-e2e-qa-workflow 3/3 $2.2
e2e-retro/skill-e2e-retro 1/1 $0.6
e2e-review-army/skill-e2e-review-army 5/5 $1.89
e2e-review-attribution/skill-e2e-review-attribution 3/3 $0.5
e2e-review/skill-e2e-review 2/2 $0.57
e2e-session-intelligence/skill-e2e-session-intelligence 4/4 $0.48
e2e-ship-docsync/skill-e2e-ship-docsync 1/1 $0.56
e2e-skillify/skill-e2e-skillify 3/3 $1.51
e2e-third-party-actions/skill-e2e-third-party-actions 5/5 $0.52
e2e-triage/skill-e2e-triage 1/1 $0.47
e2e/skill-e2e-workflow 4/4 $1.46
llm-judge/skill-llm-eval 23/25 $0.56
Fail-closed reconciliation
[test:paid] report: 6/6 slices, 52 planned shards, tier=gate
  slice 1  passed              0s  test/skill-e2e-memory-pipeline.test.ts
  slice 1  passed             89s  test/skill-e2e-plan-ceo-plan-mode.test.ts
  slice 1  passed            227s  test/skill-e2e-plan-mode-no-op.test.ts
  slice 1  passed             54s  test/skill-e2e-plan-tune.test.ts
  slice 1  passed            145s  test/skill-e2e-review.test.ts
  slice 1  passed            299s  test/skill-e2e-workflow.test.ts
  slice 2  passed              5s  test/skill-e2e-hermetic-canary.test.ts
  slice 2  passed             67s  test/skill-e2e-office-hours-auto-mode.test.ts
  slice 2  passed              0s  test/skill-e2e-plan-decision-classification.test.ts
  slice 2  passed            429s  test/skill-e2e-plan.test.ts
  slice 2  passed              0s  test/skill-e2e-qa-bugs.test.ts
  slice 2  passed             50s  test/skill-e2e-session-intelligence.test.ts
  slice 2  passed              0s  test/skill-llm-eval-spec.test.ts
  slice 3  passed              0s  test/skill-e2e-autoplan-dual-voice.test.ts
  slice 3  passed            250s  test/skill-e2e-cso.test.ts
  slice 3  passed             64s  test/skill-e2e-docsync-spawned.test.ts
  slice 3  passed              0s  test/skill-e2e-ios-swift-build.test.ts
  slice 3  passed              0s  test/skill-e2e-office-hours-phase4.test.ts
  slice 3  passed              0s  test/skill-e2e-plan-devex-peer-comparison-classification.test.ts
  slice 3  passed              0s  test/skill-e2e-plan-format.test.ts
  slice 3  passed            153s  test/skill-e2e-retro.test.ts
  slice 3  passed            160s  test/skill-e2e-skillify.test.ts
  slice 3  passed              0s  test/skill-routing-e2e.test.ts
  slice 4  passed             98s  test/skill-e2e-bws.test.ts
  slice 4  passed            493s  test/skill-e2e-deploy.test.ts
  slice 4  passed              0s  test/skill-e2e-first-task-scaffold.test.ts
  slice 4  passed              0s  test/skill-e2e-ios.test.ts
  slice 4  passed              0s  test/skill-e2e-office-hours.test.ts
  slice 4  passed            138s  test/skill-e2e-plan-devex-plan-mode.test.ts
  slice 4  passed              0s  test/skill-e2e-plan-prosons.test.ts
  slice 4  passed            395s  test/skill-e2e-review-army.test.ts
  slice 4  passed             70s  test/skill-e2e-third-party-actions.test.ts
  slice 5  passed             56s  test/skill-e2e-coverage-audit.test.ts
  slice 5  passed             39s  test/skill-e2e-diagram.test.ts
  slice 5  passed              0s  test/skill-e2e-ios-device.test.ts
  slice 5  passed              0s  test/skill-e2e-office-hours-brain-writeback.test.ts
  slice 5  passed            270s  test/skill-e2e-plan-ceo-finding-floor.test.ts
  slice 5  passed            469s  test/skill-e2e-plan-design-with-ui.test.ts
  slice 5  passed            179s  test/skill-e2e-plan-devex-finding-floor.test.ts
  slice 5  passed            395s  test/skill-e2e-qa-workflow.test.ts
  slice 5  passed            115s  test/skill-e2e-ship-docsync.test.ts
  slice 5  failed            366s  test/skill-llm-eval.test.ts
  slice 6  passed              0s  test/llm-judge-recommendation.test.ts
  slice 6  passed             48s  test/skill-e2e-ask-user-question-format-compliance.test.ts
  slice 6  passed              0s  test/skill-e2e-context-skills.test.ts
  slice 6  passed            146s  test/skill-e2e-design.test.ts
  slice 6  passed              0s  test/skill-e2e-gbrain-roundtrip-local.test.ts
  slice 6  passed             28s  test/skill-e2e-learnings.test.ts
  slice 6  passed             15s  test/skill-e2e-opus-47.test.ts
  slice 6  passed              0s  test/skill-e2e-plan-tune-cathedral.test.ts
  slice 6  passed            121s  test/skill-e2e-review-attribution.test.ts
  slice 6  passed             73s  test/skill-e2e-triage.test.ts

Sliced lane: diff-selected gate census via scripts/test-paid-shards.ts (planner → 6 executors → fail-closed report)

Failures

  • ❌ plan-ceo-review/SKILL.md modes: unknown
  • ❌ plan-eng-review/SKILL.md sections: unknown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant