Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .github/workflows/benchmark-clawbench.yml
Original file line number Diff line number Diff line change
Expand Up @@ -379,8 +379,11 @@ jobs:
if ((candidate_status != 0)); then tail -100 "$RUNNER_TEMP/candidate.log" >&2; fi
if ((baseline_status != 0)); then tail -100 "$RUNNER_TEMP/baseline.log" >&2; fi

# Opt-in: the report below renders from the Harbor job directories, so
# an unpublished run still gets its job summary and PR comment.
- name: Publish Braintrust experiment
id: publish
if: vars.BENCHMARK_PUBLISH_BRAINTRUST == 'true'
continue-on-error: true
shell: bash
env:
Expand Down Expand Up @@ -478,6 +481,6 @@ jobs:
run: |
[[ "$CANDIDATE_STATUS" == "0" ]]
[[ "$BASELINE_STATUS" == "0" ]]
[[ "$PUBLISH_OUTCOME" == "success" ]]
[[ "$PUBLISH_OUTCOME" == "success" || "$PUBLISH_OUTCOME" == "skipped" ]]
[[ "$REPORT_OUTCOME" == "success" ]]
jq -e 'all(.arms[]; .complete == true)' "$RUNNER_TEMP/benchmark-summary.json" >/dev/null
4 changes: 2 additions & 2 deletions benchmarks/harbor/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ The image records the current Git SHA, and the generated task records the ClawBe
- `PURELY_MAIL_API_KEY` and `PURELY_MAIL_DOMAIN` for ClawBench account tasks
- `OPENAI_API_KEY` for Codex, or Anthropic credentials for Claude Code
- the ClawBench judge variables when using a hosted judge: `CLAWBENCH_JUDGE_BASE_URL`, `CLAWBENCH_JUDGE_API_KEY`, `CLAWBENCH_JUDGE_MODEL`, and `CLAWBENCH_JUDGE_API_TYPE`
- `BRAINTRUST_API_KEY` and `BRAINTRUST_PROJECT` when publishing results
- `BRAINTRUST_API_KEY` and `BRAINTRUST_PROJECT` when publishing results (in CI, also the `BENCHMARK_PUBLISH_BRAINTRUST=true` repository variable)

The runner checks `/auth/context` before generating trials and stops unless the benchmark credential and effective connection resolve to the same non-empty project scope.

Expand Down Expand Up @@ -71,7 +71,7 @@ An organization member or repository collaborator can also start the full PR com
/benchmark clawbench
```

The command parser does not execute comment text. It accepts only the exact command, rejects fork pull requests and untrusted commenters, and resolves the candidate and merge-base SHAs through GitHub's API. The workflow uses the `benchmarks` environment for credentials, updates one benchmark comment on the pull request, and publishes the same results to Braintrust.
The command parser does not execute comment text. It accepts only the exact command, rejects fork pull requests and untrusted commenters, and resolves the candidate and merge-base SHAs through GitHub's API. The workflow uses the `benchmarks` environment for credentials, updates one benchmark comment on the pull request, and publishes the same results to Braintrust when the `BENCHMARK_PUBLISH_BRAINTRUST` repository variable is `true`.

## Results

Expand Down
Loading