From 5cfd9100bbe3cb25db26a9a6cec98d61abebedf8 Mon Sep 17 00:00:00 2001 From: masnwilliams <43387599+masnwilliams@users.noreply.github.com> Date: Thu, 1 Oct 2026 17:30:30 +0000 Subject: [PATCH] Make ClawBench Braintrust publishing opt-in --- .github/workflows/benchmark-clawbench.yml | 5 ++++- benchmarks/harbor/README.md | 4 ++-- 2 files changed, 6 insertions(+), 3 deletions(-) diff --git a/.github/workflows/benchmark-clawbench.yml b/.github/workflows/benchmark-clawbench.yml index a701e6df..c63b4238 100644 --- a/.github/workflows/benchmark-clawbench.yml +++ b/.github/workflows/benchmark-clawbench.yml @@ -379,8 +379,11 @@ jobs: if ((candidate_status != 0)); then tail -100 "$RUNNER_TEMP/candidate.log" >&2; fi if ((baseline_status != 0)); then tail -100 "$RUNNER_TEMP/baseline.log" >&2; fi + # Opt-in: the report below renders from the Harbor job directories, so + # an unpublished run still gets its job summary and PR comment. - name: Publish Braintrust experiment id: publish + if: vars.BENCHMARK_PUBLISH_BRAINTRUST == 'true' continue-on-error: true shell: bash env: @@ -478,6 +481,6 @@ jobs: run: | [[ "$CANDIDATE_STATUS" == "0" ]] [[ "$BASELINE_STATUS" == "0" ]] - [[ "$PUBLISH_OUTCOME" == "success" ]] + [[ "$PUBLISH_OUTCOME" == "success" || "$PUBLISH_OUTCOME" == "skipped" ]] [[ "$REPORT_OUTCOME" == "success" ]] jq -e 'all(.arms[]; .complete == true)' "$RUNNER_TEMP/benchmark-summary.json" >/dev/null diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index d52eff34..5cfec21b 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -23,7 +23,7 @@ The image records the current Git SHA, and the generated task records the ClawBe - `PURELY_MAIL_API_KEY` and `PURELY_MAIL_DOMAIN` for ClawBench account tasks - `OPENAI_API_KEY` for Codex, or Anthropic credentials for Claude Code - the ClawBench judge variables when using a hosted judge: `CLAWBENCH_JUDGE_BASE_URL`, `CLAWBENCH_JUDGE_API_KEY`, `CLAWBENCH_JUDGE_MODEL`, and `CLAWBENCH_JUDGE_API_TYPE` -- `BRAINTRUST_API_KEY` and `BRAINTRUST_PROJECT` when publishing results +- `BRAINTRUST_API_KEY` and `BRAINTRUST_PROJECT` when publishing results (in CI, also the `BENCHMARK_PUBLISH_BRAINTRUST=true` repository variable) The runner checks `/auth/context` before generating trials and stops unless the benchmark credential and effective connection resolve to the same non-empty project scope. @@ -71,7 +71,7 @@ An organization member or repository collaborator can also start the full PR com /benchmark clawbench ``` -The command parser does not execute comment text. It accepts only the exact command, rejects fork pull requests and untrusted commenters, and resolves the candidate and merge-base SHAs through GitHub's API. The workflow uses the `benchmarks` environment for credentials, updates one benchmark comment on the pull request, and publishes the same results to Braintrust. +The command parser does not execute comment text. It accepts only the exact command, rejects fork pull requests and untrusted commenters, and resolves the candidate and merge-base SHAs through GitHub's API. The workflow uses the `benchmarks` environment for credentials, updates one benchmark comment on the pull request, and publishes the same results to Braintrust when the `BENCHMARK_PUBLISH_BRAINTRUST` repository variable is `true`. ## Results